Extraction, AI validation & scientific reuse: methodology
Product methodology · Updated 2 September 2026 · Maintained by Sci-database
Can AI-extracted data be used in a scientific paper?
It can support research when the authors verify fitness for their question, document extraction and corrections, and meet the journal’s requirements. A DOI, a peer-reviewed source, a high AI agreement score or a checked-export label is not a guarantee that an extracted value is correct or that reviewers will accept the analysis.
What the workflow does
- Define the schema: record field definitions, units, missing-value rules and the relevant population or experimental conditions.
- Extract: produce structured rows from the supplied literature. Retain the source DOI and inspect page references or coordinates where available; missing locators remain missing.
- Validate a sample: run the second-pass model on selected papers and compare outputs. A completed comparison is automated evidence, not a human audit.
- Investigate disagreements: AI Judge arbitration can propose corrections. Compare important or disputed values against the source yourself and keep unresolved cases visible.
- Export and archive: keep the data, schema, source citations, validation evidence, export settings and limitations together. Publish only material you have permission to distribute.
Validate tab instructions · Download a reproducible worked example
What do the percentages mean?
- DOI coverage
- Extracted records with a recognized source DOI divided by all extracted records. This checks identifier presence, not DOI resolution or whether the paper supports the value.
- Validation coverage
- Papers with extracted records and an accepted validation result divided by papers with extracted records. The export implementation accepts completed results, legacy results without a status, or results containing validated data. Always inspect failed and incomplete runs separately.
- Export summary AI agreement
- The arithmetic mean of available numeric per-paper agreement scores, with the number of scored papers reported. It is not a pooled field-level accuracy estimate. A missing score is unknown, not zero or 100%.
- Per-record agreement
- Agreement among comparable field outputs under the current comparison rules. Schema changes, numeric normalization and missing values can affect this measure. Keep the comparison version with the report.
Why 20% validation or 90% agreement is not certification
A sample may miss rare errors or difficult tables, and two models can make the same mistake. The adequacy of a 20% sample depends on how papers were selected and the consequences of error. A 90% agreement target is a workflow target, not a demonstrated 90% accuracy rate.
For a defensible accuracy claim, define the evaluation unit and tolerances in advance, have qualified researchers establish reference answers, evaluate held-out papers that were not used to improve the prompt, and report sample size, errors, missing outputs and uncertainty. Check both disagreements and a sample of model-agreed values.
A practical minimum before plotting or citing
- Confirm units, populations, time points, duplicates and missing values for the variables used in the analysis.
- Keep source DOI, stable row ID, schema definitions, validation coverage, corrections and unresolved cases with the exported rows.
- Record the extraction and validation model names, dates, prompt/schema version and software version where available. Disclose missing execution metadata.
- Document your human checks and exclusions. Archive the exact analysis dataset and code; cite that release separately from the source papers.
FAIR support and its limits
FAIR concerns findability, access conditions, interoperability and reuse. Machine-readable files and source links help, but do not themselves establish FAIR compliance or scientific validity. A reusable release also needs descriptive metadata, clear rights and durable access. The original framework is described by Wilkinson et al., Scientific Data (2016).
Marketplace URLs describe available listings; they are not registered dataset DOIs or immutable archives. Current listings do not collect a named dataset creator. Confirm authorship and licensing with the publisher, and deposit a versioned release in an appropriate repository when the research requires persistent citation.
What evidence is published here?
The worked example uses one curated bibliographic row and the actual plotting-export code. It demonstrates traceable file structure, not an AI extraction experiment. No held-out accuracy benchmark, human-adjudicated reference dataset or independent replication is presented on this page. Workflow illustrations and affiliation or institutional-reference statements should not be treated as performance evidence.
