For systematic reviews · meta-analyses · ML datasets

An ocean of knowledge within 200+ million papers, turn it into a discovery.

No credit card · free tier forever · exports to Excel, JSON

100+ languages in → one verified database out

Pricing

One credit = one paper. That's the whole formula.

However many fields you extract. 100 papers ≈ 1,000–2,000 clean rows.

960papers

→ the Researcher Pro ($79 for 2,000 papers) plan covers this comfortably.

Early Access Special

Free

$0

150 papers · Early Access · no card

  • 150 free credits (~150 papers)
  • 3-pass AI verification engine
  • AI Judge™ agreement scoring
  • Excel / CSV / JSON export
Claim 150 Papers
Best price in market

Starter

$29

500 papers · one-time package

  • 500 credits (~500 papers)
  • Parallel batch queueing
  • Discrepancy arbitrator
  • DOI / PMID full-text resolver
Get Starter
Most Popular

Researcher Pro

$79

2,000 papers · one-time package

  • 2,000 credits (~2,000 papers)
  • Up to 5x extraction speed
  • Discrepancy arbitrator & AI Judge™
  • Priority researcher support
Get Researcher Pro
Best Ever Value

Lab & Dataset

$299

10,000 papers · one-time package

  • 10,000 credits (~10,000 papers)
  • Up to 10x Turbo speed
  • Multi-seat pooled lab access
  • Dedicated AI scientist support
Get Lab Package
Enterprise

Institution

Custom

universities · departments · pharma

  • Unlimited / pooled credit quota
  • On-premise / private cloud
  • DPA & security compliance
  • Dedicated account manager & SLA
Contact us

1 credit = 1 paper, any number of fields. Academic discount available for students and postdocs. Write to us from your university address.

Used by researchers at

How it works

Three steps between a pile of PDFs and a publishable dataset.

STEP 01

Search & add papers

Use our automated search engine across 200M+ papers (no need to manually regroup papers), or drag in PDFs and paste DOIs.

STEP 02

Describe your schema

Say what you need in plain language. Our AI Expert proposes the fields; you edit anything.

sample sizeodds ratioSeebeck (µV/K)growth method+ your own
STEP 03

Review & export

The AI Judge™ flags disagreement, never silently resolved. Then export to Excel, CSV, or JSON.

The trust loop

Kateeb doesn't just extract. It converges.

The AI Judge™ compares model outputs against the sources and helps investigate disagreements. Agreement is an automated signal, not independently measured accuracy. The scores shown here illustrate the workflow.

01Download the papers

Paste DOIs; full texts fetched, every row DOI-linked.

02Generate the schema

Plain language in; typed fields out.

03Build the database

Every value pinned to its page coordinates.

04The AI Judge™

Independent models re-read the sources and score agreement.

Refine & re-run

Tighten the schema where the Judge disagreed; run again.

Pass 1 · AI Judge™

0%

90% illustrative agreement target ▲

running cross-model verification…

Trusted database · DOI-linked, value-traced

Start free · 150 papersNo credit card · 150 papers free

Evidence-linked data

Every number knows where it came from.

Every value is pinned to its exact coordinates on the source page. If it isn't in the paper, the cell stays empty with N/A.

Try it: click any value in the spreadsheet

extraction · hypertension_rct_007.pdf
FieldValue
1Sample size (n)418
2Mean difference (SBP)−5.3 mmHg
395% CI−7.1 to −3.5

click a value → see its source ↗

source · page 6, results

3.2 Primary outcome

Of the 418 randomised participants, 402 completed follow-up. At 12 weeks, the intervention group showed a mean reduction in systolic blood pressure of −5.3 mmHg versus control (95% CI −7.1 to −3.5; p < 0.001), consistent with the pre-registered analysis plan.


cross-model agreement confirmed
Cross-model agreementInspired by duplicate extraction workflows, providing an automated quality-control signal.
Coordinate-anchoredThe audit trail attaches straight to your PRISMA supplementary materials.
Abstention by designEmpty stays empty. No generative fill or plausible inventions.

After extraction

The spreadsheet is just the beginning.

Trends for your field. Figures for your paper. Training data your models can trust.

0.40.81.2246thermal conductivity κ (W/mK)ZTκ 5.0 · ZT 0.61κ 4.2 · ZT 0.86κ 2.1 · ZT 0.95κ 1.4 · ZT 1.12κ 3.2 · ZT 0.78κ 2.8 · ZT 0.90κ 1.8 · ZT 1.05κ 4.6 · ZT 0.70κ 3.6 · ZT 0.82κ 2.4 · ZT 1.00κ 1.1 · ZT 1.20κ 5.4 · ZT 0.55Bi₂Te₃/Sb₂Te₃ · 1.12
From 62 papers to one trend, one click.Chart any field against any other, across your whole corpus, in one click.illustrative data · export PNG / SVG
FAIR-Aligned ExportExport your database with persistent identifiers (DOI), reference files (RIS / BibTeX), and structured schema definitions (CSV / JSON) conforming to FAIR data principles.
Open Standard FormatsDirect export to CSV and JSON formatted for immediate analysis in Python, R (meta), Stata, or spreadsheet tools with zero lock-in.
Source ProvenanceEvery extracted data point preserves source paper titles, DOIs, and extraction metadata for complete auditability and replication.

Start free · 150 papersNo credit card · trends, figures, ML exports included

Scientific Data Marketplace · Reusable Schemas · Verified Datasets

Explore, reuse, and publish research data with complete provenance.

Discover ready-to-plot scientific databases extracted from papers with AI Judge cross-validation, reusable domain schemas, and open datasets from global research repositories (DataCite, Zenodo, Dataverse).

FREE VERIFIED DATABASEAI JUDGE™ CHECKED
Version 1.0 · Static Release

marketplace · generative-ai-evaluation-benchmarks

Generative AI evaluation and safety

Fast-moving evaluation, safety, security, hardware, and autonomous-system research.

DOI Coverage100% (73 of 73)
AI Cross-Validation27% (20 papers)
AI Agreement93.4% (8 Resolved)
LicenceCC BY 4.0
7 Research Opportunities:
Generative AI evaluation and safetyAI-agent reliability and tool useAI for scientific discoveryCybersecurity and IoT vulnerabilitiesQuantum computing, sensing and materialsRobotics and autonomous systemsSemiconductors and edge AI
model_namestringmodel_versionstringbenchmark_namestringtask_typeenumevaluation_scorenumbersafety_dimensionenumevaluation_languagestringdoistring
Sample Extracted Evidence Row (Provenance-Grounded)
model_namemodel_versionbenchmark_nametask_typeevaluation_scoresafety_dimensionevaluation_languagedoi
Llama-3-70B-Instructv1.0MMLU-ProReasoning68.4TruthfulnessEnglish10.48550/arXiv.2407.21783
Every dataset version includes: source_doi, validation_summary.csv, data_dictionary.csv, manifest.json (SHA-256), README.md

Built with Sci-database

The Islamic Ink Atlas: 1,690 historical recipes, turned into a living research portal.

Centuries of recipes extracted into a searchable database with an interactive map and knowledge graph. Every recipe keeps its citation and a certainty level.

Our own team's project at QEERI-Materials: a demonstration, not a customer story. Judge it live.

Explore the Atlas ↗
The Islamic Ink Atlas: Geographical & Temporal Visualization of ink and pigment recipes across time and space

Questions researchers actually ask

Before you trust us with your review

Will journal reviewers accept AI-extracted data?
Acceptance depends on the journal, the research question and the authors' checks. Sci-database supports source-linked extraction and automated second-pass review on selected papers. Keep source DOIs, available page references, validation coverage and corrections with the data, and document human checks. An AI agreement score does not guarantee correctness or publication acceptance. Read the methodology and limitations.
What happens when the AI gets a number wrong?
You catch it before it reaches your analysis. When the independent models disagree, the value is flagged by the AI Judge for your review, never silently resolved. You see both candidates next to the source region and decide. The fields most prone to disagreement are the ones you'd expect: nested tables, figure-only data, and unit conversions.
How do I know my database is accurate?
Inspect the cross-model agreement score alongside validation coverage, then audit the underlying values. Models can agree and still be wrong. Improving a schema can help, but repeated tuning on the same papers is not a held-out accuracy test. Retain source DOIs and available evidence locators, check units and missing values, and record a human-reviewed sample before drawing scientific conclusions.
Do you train models on my uploads?
No. Your PDFs and your extracted data are never used to train models. Uploads are encrypted in transit and at rest, and you can delete a paper or an entire project at any time.
I'm not in medicine. Is this for me?
Very much so. Our most active users are in materials science, physics, and astronomy, extracting thermoelectric parameters, DFT results, and supernova observations. If you can name the fields, the engine can extract them. Browse the public schema gallery for real examples.
Can I import papers without uploading PDFs?
Yes. Paste DOIs or PMIDs and we fetch the records directly, with PubMed integration built in. You can mix imported and uploaded papers in the same project.
How do credits work?
1 credit = 1 paper, regardless of how many fields you extract from it. The free tier includes 150 credits, enough to pilot a real review before paying anything. We issue invoices institutions can process for grant purchases.
What formats can I export, and does it support FAIR data principles?
Database and citation exports include CSV, JSON, RIS and BibTeX. The Export tab can package schema, provenance and validation information. These features support FAIR practices, but do not certify FAIR compliance or accuracy. Supply reuse rights, creator information, limitations and a durable versioned release when publishing. Inspect a worked export example.

Unifying human knowledge,
one database at a time.

Every language, every field, every era: your literature is already a dataset. Start with 150 papers free and judge the evidence trail with your own eyes.

Start free · 150 papers

Developed at Hamad Bin Khalifa University · Patent QF Ref: 2024-066