Sci-database Documentation
Turn research papers into verified, structured databases with every value traceable to its source. Learn how to search 200M+ papers, design extraction schemas, run the Kateeb engine, and validate findings with the dual-model AI Judge.
Table of Contents
12 Chapters · 45 min totalGetting Started
2 chaptersCore Workflow
4 chapters3 · Searching Papers
4 min readSearch 200M+ papers across OpenAlex, CORE, and Semantic Scholar with focused keywords.
4 · Files & Uploads
4 min readAutomatic PDF download verification, drag-and-drop uploads, and fixing unverified files.
5 · Building Your Schema
5 min readAI-assisted parameter generation, controlled vocabularies, and unit pairing.
6 · Running Extractions
4 min readThe Kateeb engine, multi-record extraction, zero-hallucination rules, and resume.
Quality & Trust
3 chapters7 · Validation & AI Judge
6 min readCross-model validation, investigating disagreements, and converging above 90% agreement.
8 · Database & Exports
3 min readExplore extracted records, inspect individual papers, and export to CSV, JSON, or Excel.
9 · Sharing & Marketplace
3 min readPublic schema gallery, community data sharing, and listing golden databases.
Reference
3 chapters10 · Credits & Pricing
3 min read1 credit = 1 paper. Transparent pricing, earning free credits, and subscription plans.
11 · FAQ
4 min readAnswers to questions researchers, peer reviewers, and institutional boards ask.
12 · Troubleshooting
5 min readEvery error message, why it happens, and step-by-step resolution guides.
Sci-database Documentation
Turn research papers into verified, structured databases — with every value traceable to its source.
Sci-database is an AI platform for researchers who need structured data out of the scientific literature: systematic reviews, meta-analyses, review papers, and machine-learning datasets. You search 200M+ papers (or upload your own PDFs), describe the data you need in plain English, and the Kateeb extraction engine (كاتب, "scribe") reads every paper and fills your database — then a second, independent AI model cross-checks the results so you can see exactly where the models agree and where they don't.

The workflow at a glance
┌─────────────┐ ┌──────────────┐ ┌─────────────┐ ┌──────────────┐ ┌─────────────┐
│ 1. SEARCH │ → │ 2. VERIFY │ → │ 3. SCHEMA │ → │ 4. EXTRACT │ → │ 5. VALIDATE │
│ 200M+ papers│ │ PDFs fetched │ │ AI suggests │ │ Kateeb reads │ │ AI Judge │
│ or upload │ │ & checked │ │ you edit │ │ every paper │ │ cross-checks│
└─────────────┘ └──────────────┘ └─────────────┘ └──────────────┘ └─────────────┘
│
┌────────────────────────────────────┘
▼
┌─────────────────────┐
│ 6. EXPORT & SHARE │
│ Excel · CSV · JSON │
│ or publish/sell it │
└─────────────────────┘
Papers in. One verified spreadsheet out. Zero synthetic data — if a value is not in the paper, the cell says N/A.
Documentation contents
- Introduction — what Sci-database is, who it's for, and the design principles that make its output trustworthy.
- Getting started — sign in, get your free credits, complete your profile.
- Searching for papers — OpenAlex, CORE, and Semantic Scholar; filters, API keys, and the queue.
- Files & uploads — automatic PDF verification, uploading your own PDFs, fixing papers that can't be fetched.
- Building your schema — describe your database in plain English, let the AI draft the fields, then edit and save.
- Running an extraction — processing jobs, statuses, stop & resume, result emails.
- Validation & the AI Judge — cross-model checking, investigating disagreements, and converging above 90% agreement.
- Your database & exports — viewing results and exporting to CSV / JSON / Excel.
- Sharing & the marketplace — public schemas, free sharing, and listing a "golden" database for sale.
- Credits & pricing — every credit cost in one table.
- FAQ — the questions researchers actually ask.
- Troubleshooting — every error message and how to fix it.
Developed by Dr. El Tayeb Bentria at Hamad Bin Khalifa University (HBKU), Doha, Qatar · Patent QF Ref: 2024-066.