Extraction, AI validation & scientific reuse: methodology

Product methodology · Updated 2 September 2026 · Maintained by Sci-database

Can AI-extracted data be used in a scientific paper?

It can support research when the authors verify fitness for their question, document extraction and corrections, and meet the journal’s requirements. A DOI, a peer-reviewed source, a high AI agreement score or a checked-export label is not a guarantee that an extracted value is correct or that reviewers will accept the analysis.

What the workflow does

  1. Define the schema: record field definitions, units, missing-value rules and the relevant population or experimental conditions.
  2. Extract: produce structured rows from the supplied literature. Retain the source DOI and inspect page references or coordinates where available; missing locators remain missing.
  3. Validate a sample: run the second-pass model on selected papers and compare outputs. A completed comparison is automated evidence, not a human audit.
  4. Investigate disagreements: AI Judge arbitration can propose corrections. Compare important or disputed values against the source yourself and keep unresolved cases visible.
  5. Export and archive: keep the data, schema, source citations, validation evidence, export settings and limitations together. Publish only material you have permission to distribute.

Validate tab instructions · Download a reproducible worked example

What do the percentages mean?

DOI coverage
Extracted records with a recognized source DOI divided by all extracted records. This checks identifier presence, not DOI resolution or whether the paper supports the value.
Validation coverage
Papers with extracted records and an accepted validation result divided by papers with extracted records. The export implementation accepts completed results, legacy results without a status, or results containing validated data. Always inspect failed and incomplete runs separately.
Export summary AI agreement
The arithmetic mean of available numeric per-paper agreement scores, with the number of scored papers reported. It is not a pooled field-level accuracy estimate. A missing score is unknown, not zero or 100%.
Per-record agreement
Agreement among comparable field outputs under the current comparison rules. Schema changes, numeric normalization and missing values can affect this measure. Keep the comparison version with the report.

Why 20% validation or 90% agreement is not certification

A sample may miss rare errors or difficult tables, and two models can make the same mistake. The adequacy of a 20% sample depends on how papers were selected and the consequences of error. A 90% agreement target is a workflow target, not a demonstrated 90% accuracy rate.

For a defensible accuracy claim, define the evaluation unit and tolerances in advance, have qualified researchers establish reference answers, evaluate held-out papers that were not used to improve the prompt, and report sample size, errors, missing outputs and uncertainty. Check both disagreements and a sample of model-agreed values.

A practical minimum before plotting or citing

  • Confirm units, populations, time points, duplicates and missing values for the variables used in the analysis.
  • Keep source DOI, stable row ID, schema definitions, validation coverage, corrections and unresolved cases with the exported rows.
  • Record the extraction and validation model names, dates, prompt/schema version and software version where available. Disclose missing execution metadata.
  • Document your human checks and exclusions. Archive the exact analysis dataset and code; cite that release separately from the source papers.

FAIR support and its limits

FAIR concerns findability, access conditions, interoperability and reuse. Machine-readable files and source links help, but do not themselves establish FAIR compliance or scientific validity. A reusable release also needs descriptive metadata, clear rights and durable access. The original framework is described by Wilkinson et al., Scientific Data (2016).

Marketplace URLs describe available listings; they are not registered dataset DOIs or immutable archives. Current listings do not collect a named dataset creator. Confirm authorship and licensing with the publisher, and deposit a versioned release in an appropriate repository when the research requires persistent citation.

What evidence is published here?

The worked example uses one curated bibliographic row and the actual plotting-export code. It demonstrates traceable file structure, not an AI extraction experiment. No held-out accuracy benchmark, human-adjudicated reference dataset or independent replication is presented on this page. Workflow illustrations and affiliation or institutional-reference statements should not be treated as performance evidence.