Home/About
Academic Research & Innovation

Turning 200 Million Research Papers into Verified Scientific Databases

Sci-database was created to solve the most tedious bottleneck in modern science: researchers spending months manually reading PDF tables, copying numbers into spreadsheets, and struggling with irreproducible systematic reviews.

Patent Protection

QF Patent Ref: 2024-066

Proprietary coordinate-grounded multi-pass extraction algorithms developed and patented under Qatar Foundation.

Institution Origin

HBKU · Doha, Qatar

Incubated at Hamad Bin Khalifa University and Qatar Energy & Environment Research Institute (QEERI).

Global Literature

200M+ Papers Connected

Direct indexing and semantic search across OpenAlex, PubMed, Crossref, and arXiv in 100+ languages.

The Kateeb (كتيب) Engine

Named after the historical Arabic word for scribe (كتيب), the engine was built with one core scientific principle: strict evidence grounding and auditable provenance.

General-purpose LLMs can misread tables, associate incorrect units, or synthesize plausible-sounding values. Sci-database solves this by anchoring every single extracted cell to exact PDF page coordinates and abstaining (N/A) when data is absent, using an auditable 3-stage pipeline:

STAGE 01

Coordinate Grounding

Every single extracted cell is pinned to its exact bounding box coordinates and page index in the source PDF.

STAGE 02

Dual-Model Extraction

Inspired by duplicate extraction workflows, independent models read the source text to cross-check numbers, units, and conditions.

STAGE 03

AI Judge Verification & Arbitration

The AI Judge measures cross-model agreement and flags ambiguous values with full provenance—never silently resolving discrepancies.

Live Showcase · 3,300+ Verified Records

Tested on Complex Historical & Scientific Data

Our team deployed Kateeb to build the Islamic Ink Atlas—extracting over 3,300 historical chemistry recipes, materials, and pigment compositions across medieval Arabic and Persian manuscripts.

Our Commitments to Researchers

Your Data is Never Used for Training

Uploaded preprints, unpublished results, and custom schemas are strictly quarantined. We never train commercial models on your research.

Full Provenance & Audit Trails

Export your database to Excel, CSV, or JSON with every single data point hyperlinked to its DOI, author, and page quote.

Institutional Collaboration & Support

We support academic labs, universities, and health institutes with custom on-premise deployments, DPA reviews, and high-volume pipelines.

Start Extracting Free (150 Papers)

No credit card required · Instant access · Free tier forever