Skip to main content

Extract structured data from academic papers, faster.

Turn collections of academic PDFs into structured datasets using the extraction questions that matter to your review.

Provide your papers and extraction framework, and receive a structured Excel workbook with an answer and reasoning note for every question — ready for your review.

Academic PDFs

Extraction questions

Structured extraction

Excel output

Researcher review

Spend less time searching through papers by hand.

Systematic review data extraction often means locating the same types of information — study design, sample size, outcomes, model performance — across dozens or hundreds of papers, one at a time.

This service lets you define your extraction framework once in a spreadsheet and apply it consistently across your full set of academic papers, producing structured results you can review before analysis.

How it works

Your extraction questions, applied consistently across your papers.

  1. 01

    Define what you need

    Provide a spreadsheet (CSV or Excel) listing your extraction questions. Each row defines one question with a short name, the question text, an answer format (text, number, category, or list), and optional allowed categories for categorical answers.

  2. 02

    Provide your papers

    Supply your academic papers as PDF files. The extraction process reads document text and processes tables and figures alongside the main content, looking only for information explicitly stated in each paper.

  3. 03

    Receive structured results

    You receive a single Excel workbook with one row per paper — or one row per extracted item when a paper reports multiple entities. Each question appears as a column paired with a reasoning column explaining where the answer was found.

  4. 04

    Review and verify

    Use the cross-checking application to review results one paper at a time. Edit extracted values, view reasoning notes, optionally open the source PDF side by side, and mark each field as correct or incorrect before export.

From extraction questions to a usable dataset.

Each question becomes a column in your workbook, paired with a reasoning note.

Illustrative example
pdf_namestudy_designstudy_design_reasoningvariable_namepopulation_sizeaurocauroc_reasoning
Smith_2021Retrospective cohortMethods section, paragraph 2Logistic regression model4,8320.78Table 3, performance metrics
Smith_2021Retrospective cohortMethods section, paragraph 2Random forest model4,8320.81Table 3, performance metrics
Jones_2019Prospective cohortAbstract and MethodsCox proportional hazards1,205NOT_FOUNDAUROC not reported for this model

When a paper reports multiple items (for example, separate prediction models), the workbook expands to one row per item with paper-level answers repeated.

Designed to be checked, not blindly trusted.

AI-assisted extraction supports your judgement — it does not replace it.

After delivery, you review results in a local browser application. Load the Excel workbook, work through one paper at a time, and optionally open the matching source PDF side by side. Each extracted field shows the value alongside a read-only reasoning note from the extraction.

Edit any value, mark fields as correct or incorrect, and export your reviewed data with a parallel correctness grid. All processing stays in your browser — nothing is uploaded during verification.

Paper extraction review

Review one paper at a time, edit cells, mark each field correct or incorrect, then export.

Smith_2021

Correct: 1 · Incorrect: 0 · Pending: 1 · Rows complete: 1/2

1 / 2 variables reviewed

← Previous paperPaper 1 of 3 · Variable 1 of 2Next paper →

Viewing: Smith_2021.pdf

Source PDF displayed here. Highlight text in a field to search within the document.

Variable-level data

auroc

0.78

Reasoning

Table 3, performance metrics for logistic regression model
Agree (save as-is)CorrectIncorrect

population_size

4,832

Reasoning

Methods section: development cohort sample size
Agree (save as-is)CorrectIncorrect

Reduce the repetitive work while keeping the researcher in control.

Built for systematic review workflows

Consistent

Apply the same extraction framework across every paper in your review, with the same question definitions and answer formats throughout.

Flexible

Define questions relevant to your specific review — text, numbers, categories from your own list, or lists of items extracted per paper.

Reviewable

Inspect and amend every extracted value before analysis. Reasoning notes help you locate answers in the source paper.

Practical

Receive a structured Excel workbook that fits existing evidence-synthesis workflows, with an export path for reviewed results and correctness labels.

Choose the size of your extraction

Select the package that best matches your review. After payment, you will be directed to submit your papers and extraction framework.

Pilot

Up to 10 papers

£X

For small reviews, pilot projects or testing the workflow.

  • Up to 10 academic PDFs
  • Researcher-defined extraction questions
  • Structured Excel workbook with reasoning notes
  • Cross-checking application for review
Choose Pilot

Standard

Up to 50 papers

£Y

For typical systematic review extraction projects.

  • Up to 50 academic PDFs
  • Paper-level and variable-level extraction
  • Tables and figures read alongside text
  • Export reviewed results with correctness labels
Choose Standard

Large

Up to 150 papers

£Z

For larger systematic reviews and evidence-synthesis projects.

  • Up to 150 academic PDFs
  • Flexible extraction question formats
  • One row per paper or per extracted item
  • Local browser-based verification workflow
Choose Large

Larger or more complex review?

Contact us to discuss projects that fall outside the standard packages.

Contact Us

What happens next?

  1. 1

    Pay securely

    Complete checkout through Stripe. You will receive a confirmation with next steps.

  2. 2

    Submit your project

    After payment, follow the submission link to provide your PDF papers and extraction question spreadsheet (CSV or Excel).

  3. 3

    Receive and review

    We process your papers and deliver a structured Excel workbook. Use the cross-checking application to review, edit, and export your verified results.

Frequently asked questions

What do I provide?
You provide your academic papers as PDF files and a spreadsheet (CSV or Excel) defining your extraction questions. Each question needs a short name, the question text, and an answer format (text, number, category from a list, or semicolon-separated list).
Can I define my own extraction questions?
Yes. You define every extraction question in your own spreadsheet. Each row is one question with a name, the question text, an output type, and optional allowed categories for categorical answers. The same framework is applied consistently across all papers in your project.
Can information be extracted from tables and figures?
Yes, where the content is visible in the paper. Tables and figures are processed alongside the document text. Extraction is limited to what is explicitly stated in the paper — the system does not infer values that are not reported.
What do I receive?
You receive a single Excel workbook (.xlsx) with one sheet. Each extraction question appears as a column, paired with a reasoning column that briefly explains where the answer was found. If your review extracts multiple items per paper, you receive one row per paper-item pair with a variable_name column.
Can one paper contain multiple things I need extracted?
Yes. If your extraction framework includes a question that lists distinct items (for example, separate prediction models in a study), the workbook expands to one row per item. Paper-level answers are repeated on each row so you can review variable-level data alongside the source paper.
How do I cross-check results?
After delivery, you use the cross-checking application in your browser. Load the Excel workbook, review one paper at a time, and optionally open the matching source PDF side by side. Each extracted field shows the value and a read-only reasoning note. You can edit values and mark each field as correct or incorrect.
Can I edit extracted results?
Yes. In the cross-checking application you can edit any extracted value directly. You can also add new data points, add or remove variable rows, and copy values between rows. Reasoning columns are read-only reference material from the extraction.
Can I save or download reviewed results?
Yes. The cross-checking application autosaves your progress in the browser. You can also download a combined Excel workbook with two sheets: Data (your edited values) and Labels (correct/incorrect marks for each field). Session files can be downloaded and reloaded for backup.
Should I check the results?
Yes. AI-assisted extraction should be reviewed before being relied upon in analysis, systematic review conclusions, or publication. The cross-checking workflow is designed to support your judgement, not replace it.
How long will it take?
Typical turnaround is 5–10 working days depending on the number of papers and complexity of your extraction framework. We will confirm an estimated delivery date when you submit your project.
What happens to my papers?
Uploaded materials are handled according to our privacy and data-retention policies. The cross-checking application runs entirely in your browser — your reviewed data does not leave your machine during verification.

Ready to test it on your review?

Choose a package and turn your extraction framework into structured, reviewable results.

View Pricing