What Sci-database is, who it's for, and the 4 academic trust principles.
1 · Introduction
What is Sci-database?
Sci-database turns piles of research PDFs into one structured, verified database. Instead of spending months manually copying values from hundreds of papers into a spreadsheet, you:
- Search 200M+ scientific papers through public repositories (OpenAlex, CORE, Semantic Scholar), or upload your own PDFs.
- Describe the data you need in plain English — the AI proposes a structured schema (the columns of your future database), which you can edit field by field.
- Extract — the Kateeb engine reads each full paper and fills your schema. One paper can yield several records (for example, one row per material or per experimental condition).
- Validate — a second, independent AI model re-extracts a sample of your papers, and the platform shows a field-by-field comparison. Disagreements are flagged, never silently resolved.
- Export your database to Excel-ready CSV or JSON — or share it publicly to contribute to open science.

Who is it for?
| You are… | You use Sci-database to… |
|---|---|
| A systematic reviewer / meta-analyst | Extract outcomes, sample sizes, effect sizes and study characteristics from hundreds of trials, then feed R (meta), Stata, or RevMan |
| A materials scientist / physicist / chemist | Build property databases (e.g. Seebeck coefficients, ZT values, band gaps, synthesis methods) for review papers |
| An ML researcher | Turn literature into labeled training data at scale |
| Any researcher writing a review | Replace months of manual data entry with an afternoon of supervised extraction |
If your field publishes papers, you can build a database from them — the platform is domain-agnostic. Public schemas already exist for thermoelectrics, dental composites, hypertension RCTs, IoT security, supernova observations, climate projections, and more.

The trust principles
Sci-database is built for academic use, where a wrong number in a meta-analysis is worse than a missing one. Four principles run through the whole product:
1. Zero synthetic data
The extraction model operates under a strict instruction: if a value is missing from the paper, or the model is not highly confident, it must return exactly N/A — never a guess, never a plausible-sounding fabrication.
2. Nothing is silently resolved
When the validator model disagrees with the original extraction, the disagreement is shown to you in a red-highlighted comparison row. You (optionally helped by a third AI arbiter) decide what the truth is.

3. Converge, don't hope
Extraction is not one shot. The intended loop is: extract → validate against a second model → investigate disagreements → let the AI improve your schema's field descriptions → re-run. Each pass raises the cross-model agreement; the platform's trust threshold is 90%+ agreement before a database is considered "trusted".
4. Your data stays yours
- Your papers and extracted data are private to your account by default.
- Your schema (the field definitions only — no data) may be shared in the public gallery to help other researchers; this is stated in the Terms you accept at sign-up.
- Uploads are never used to train models.
- You can delete your files, jobs, and databases at any time.
The Kateeb engine
The extraction engine is branded Kateeb (كاتب, Arabic for "scribe"). In validation views, the original extraction column is always labeled "Original (Kateeb)" so you can tell at a glance which model produced which value. Sci-database was developed at Hamad Bin Khalifa University (Doha, Qatar) and is patent-registered (QF Ref: 2024-066).
Next: Getting started →