Can AI reliably extract breast-cancer trial data from health records?
For many fields, but not all. In a peer-reviewed single-hospital study of 113 patients in 11 breast-cancer trials, an AI pipeline matched manually curated trial forms well on several variables, yet accuracy fell to 46% for treatment start dates and 23% for end dates. The system remains an assistive extraction tool requiring human verification.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1MIRROR retrospectively compared AI-extracted data with routinely curated electronic case-report forms for 113 patients enrolled in 11 breast-cancer trials at one Seville hospital.
- 2The pipeline processed 138,454 structured records and about 1.85 million words of unstructured text across 18 data domains, using rules, named-entity recognition, a RoBERTa model and GPT-4 according to the field.
- 3Several fields reached 90%–100% accuracy, but study-treatment start and end dates reached only 46% and 23%; the study did not measure saved staff time, cost, patient outcomes or performance at another hospital.
Research topic
Whether a multi-method AI pipeline can reproduce breast-cancer clinical-trial data captured manually in electronic case-report forms

The direct answer: useful for defined fields, unreliable as a complete replacement
The MIRROR study shows that an AI pipeline can reproduce many data fields used in breast-cancer research with high agreement against manually curated trial forms. Birth year and administrative sex matched at 100% in this dataset; menopausal status reached 91% accuracy, drug-use status 96%, current cancer 90%, receptor information 88% and study treatment 99%. The system also attached references back to the source text, giving a reviewer a way to inspect why each value was extracted.
Reliability was not uniform. Accuracy was 83% for smoking and previous cancer, 81% for previous non-cancer disease and 87% for laboratory tests. Most importantly, extracting the start date of study treatment achieved 46% accuracy and the end date 23%. Those are not marginal fields in a clinical trial. The authors therefore position the system as an assistive tool that can accelerate well-defined extraction while leaving complex and temporally sensitive variables under human review.[1]
What the study compared
MIRROR was a retrospective longitudinal study at Virgen del Rocío University Hospital in Seville. It included 113 adults with early or advanced breast cancer who had participated in 11 clinical trials between January 2012 and December 2021. The hospital was one recruiting centre for each trial. The comparison used electronic case-report forms, or eCRFs, that clinical-research staff had populated during the original trials as the operational reference standard.
The AI system extracted the same kinds of information directly from the hospital’s electronic records. The evaluation covered 18 domains: birth date, sex, menopausal status, drug use, alcohol use, smoking, previous cancer, previous non-cancer disease, family history, current cancer, receptor status, study treatment, concomitant medication, laboratory tests, radiological response, overall survival, and study-treatment start and end dates. Family history could not be scored because the eCRFs did not contain it, even though 11 patients had family-history information in their records.[1]
The pipeline combined several techniques rather than one model
The paper does not describe a single all-purpose language model reading every chart. Its CapTrial pipeline applied different methods to different data types. Structured hospital tables were ingested and normalized through extract-transform-load routines. Unstructured documents were selected and preprocessed, then analysed with regular expressions, medical named-entity recognition, a masked RoBERTa model and, for fields requiring context across notes, a GPT-4-based aggregation step. Extracted entities were normalized into predefined formats before patient-level comparison with the eCRF.
No task-specific model was trained on the complete MIRROR cohort. For selected tasks, a limited development subset of at most 10% of available data was used to refine extraction pipelines. Reported performance was then assessed on the remaining 90%–100%, depending on the field. That separation reduces direct reuse of evaluation examples, although the paper does not provide external testing from another hospital, record system or language environment.[1]
The denominator was substantial, but it still represents one local system
Across the 113 patients, the system processed 138,454 structured records and approximately 1.85 million words of unstructured text. Those volumes are more informative than quoting patient count alone because each person’s longitudinal record can contain many laboratory results, medication entries, pathology reports and clinical notes. The system’s ability to retain a reference to the original text is especially valuable when thousands of candidate entries must be checked.
Volume does not create geographic or technical generalisability. All records came from one hospital, and the proprietary extraction software was not integrated directly with the hospital’s main EHR. The researchers had to work with an intermediate data system, introducing interoperability and data-loss constraints. The results therefore show what this pipeline could retrieve from this local record environment, not what any hospital could expect from a plug-and-play product.[1]
Where performance was strongest—and where it broke down
The paper reports 100% accuracy, precision, recall and F1 for birth year and sex. Menopausal status achieved 91% accuracy and 90% F1; drug-use status 96% accuracy and 94% F1; alcohol use 87% accuracy and 93% F1; smoking 83% accuracy and 81% F1; previous cancer 83% accuracy and 88% F1; previous non-cancer disease 81% accuracy and 90% F1; receptor status 88% across the four reported measures; and radiological response 83% accuracy with a 91% F1. Some fields were evaluated as direct matches rather than binary classifications, so precision and recall were not always applicable.
High headline accuracy sometimes depended on restricting the comparison to dates or entries that could be aligned between systems. Concomitant medication and overall survival were reported at 100% accuracy under the study’s matching rules, while underlying record counts differed between the eCRF and EHR-derived data. By contrast, study-treatment start and end dates performed poorly because the reference date and the most recent relevant clinical report did not consistently correspond. This is exactly the kind of temporal ambiguity that makes automated clinical abstraction difficult.[1]
Why the manual trial form is an imperfect reference standard
MIRROR treated the trial eCRF as operational ground truth because trained research staff had completed it during regulated studies. That is pragmatic, but it does not prove the manual value was always correct. The researchers did not generate a second independent human annotation set and did not calculate inter-annotator agreement. Manual transcription errors, incomplete forms and discrepancies between the trial snapshot and later health-record information can all look like AI errors when concordance is the metric.
The source-linking feature partly addresses this problem. When the AI and eCRF disagreed, reviewers could inspect the exact EHR fragment and determine whether the pipeline misread the record, the information was missing, or the reference form differed. The paper says this helped identify discrepancies, but it does not report a separate adjudicated accuracy estimate after resolving every disagreement. The safest interpretation is therefore agreement with routine manual capture, not absolute clinical truth.[1]
What this could change for trial teams and patients
Clinical-trial teams spend substantial time transferring information from health records into structured research forms. Automating dependable fields could reduce repetitive work and let staff focus on ambiguous cases, data quality and patient-facing activity. Source references can also make review faster than a system that returns an answer without showing where it came from. More complete extraction might support real-world evidence studies and reduce duplicate entry across research systems.
Those benefits were not measured here. The study did not compare abstraction time, staff workload, cost, query rates, regulatory inspection findings or downstream trial decisions. It did not alter patient care or test whether better data capture improves recruitment, safety reporting or treatment outcomes. Patient records also raise privacy and governance issues: any deployment would need strict access control, audit trails, local validation and a defined person accountable for approving extracted data.[1]
Funding, commercial interests and what would change the assessment
The authors report no specific grant funding for MIRROR. Several authors were full-time employees of MEDSIR, which sponsored the included trials, or Science4Tech Solutions, the company that developed CapTrial; other authors disclosed research funding, advisory work, travel support, patents or ownership interests involving healthcare companies. Those relationships do not invalidate the results, but they make independent replication especially important.
Confidence would rise with prospective testing across hospitals, EHR vendors, languages and independently adjudicated records. A stronger study would lock the pipeline before evaluation, use a second human abstraction team, report agreement and adjudication by field, and measure time, cost and correction workload. It should also test whether date extraction and temporally complex medication histories improve without losing traceability. Until then, MIRROR supports targeted, human-supervised automation—not autonomous population of clinical-trial databases.[1]
What this means for people
- Research staff could spend less time copying well-defined fields and more time reviewing complex records and supporting participants.
- Traceable source references can make errors easier to find, but only if a qualified person actually reviews consequential fields.
- Patients should not assume this evidence improves treatment or trial access: the study evaluated data capture, not clinical decisions or outcomes.
Global context
Clinical-trial data are recorded differently across hospitals, countries and vendors. MIRROR demonstrates a useful Spanish single-site workflow, but its performance depends on local terminology, documentation habits and system interfaces. International adoption would require local privacy controls, field-by-field validation and evidence that human-supervised automation remains reliable under different record structures.
What the evidence does not yet show
- All 113 patients came from one Spanish hospital and 11 trials run through the same research network; there was no external hospital or EHR validation.
- The routinely curated eCRF was used as the operational reference standard without a parallel independent annotation set or formal inter-annotator agreement.
- For selected fields, up to 10% of available data helped refine extraction pipelines; evaluation used the remaining 90%–100%, but no fully independent external dataset.
- Accuracy varied sharply by field, falling to 46% for study-treatment start dates and 23% for end dates.
- The study did not measure staff time, cost, regulatory quality, patient outcomes or whether automated extraction changed a trial decision.
- The proprietary system was not directly integrated into the hospital’s main EHR, creating an additional interoperability constraint.
What to watch next
- Prospective multi-hospital validation across EHR vendors and languages.
- Independent dual-human abstraction and adjudication of AI–eCRF disagreements.
- Measured changes in staff time, correction workload, cost and regulatory audit findings.
- Improved temporal reasoning for treatment dates, medication changes and longitudinal response.
- Clear governance for privacy, source traceability, human approval and vendor accountability.
Living evidence record
Impact record IAI-0Y25SZ2
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
7 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what npj Breast Cancer published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
Evidence trail
Sources used for this report
Links checked 7 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Health & Life Sciences
Can clinical AI recognise when the patient record does not support an answer?
A new clinical-agent benchmark raises a practical question for health systems: can an assistant explain what the record cannot establish? Our analysis examines evidence, local testing and the burden on staff.
6 min · 2 sources
Health & Life Sciences
Did longer AI-assisted Parkinson’s rehab improve movement?
No clear motor advantage emerged between one, two and three months of home training. The peer-reviewed Chinese trial randomized 120 people, analysed 71 and had no usual-care group, so its exploratory cognitive signal cannot establish benefit.
7 min · 1 source
Health & Life Sciences
Can AI support breast-ultrasound decisions across countries?
A peer-reviewed South Korean-led study validated an interpretable retrieval-augmented system across 8,311 images from 11 cohorts in seven countries and tested assistance with four readers. The retrospective evidence is encouraging, but it is not a prospective screening trial or proof of better patient outcomes.
8 min · 3 sources
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.