Can AI make tumour mitosis counts more consistent?
A preprint study paired 13 pathologists' unaided and AI-assisted reviews of 385 tumour slides from three European centres. Agreement rose and counting time fell, but the unreviewed study did not establish which counts were correct, whether diagnoses improved or whether patients benefited.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
Research topic
AI-assisted hotspot selection and mitotic-figure counting across multiple tumour types

At a glance
- 1Thirteen pathologists assessed 385 digitised tumour slides without assistance and again with MitPro after at least a two-week washout; each slide was read by three randomly assigned pathologists in both stages.
- 2Combined inter-reader agreement rose from an ICC of 0.589 to 0.949, while mean pathologist-level median assessment time fell from 286.4 to 127.8 seconds. Agreement remained much weaker for thyroid cases than for the combined dataset.
- 3The preprint tested reproducibility and speed, not accuracy against an independent ground truth, final diagnosis, treatment decisions or patient outcomes. It was supported by the product developer and Innovate UK, and three authors disclosed company equity or directorships.
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Living evidence record
Impact record IAI-0IKZSQ0
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent support
Present
Record status
Monitoring
Last checked
3 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what arXiv published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
Mitotic counting is important, but it is not perfectly reproducible
Pathologists count cells undergoing division because mitotic activity can contribute to tumour grading, prognosis and treatment planning. The task is deceptively difficult. A reader first has to find a small hotspot within a much larger tissue section, then distinguish genuine mitotic figures from lookalikes and count them within a defined area. Differences in hotspot selection or recognition can produce different counts from the same slide.
The MitPro study asked a practical question: can software that proposes hotspots and highlights candidate mitoses make different pathologists more consistent and faster while leaving the final decision with the human reader? That is narrower than asking whether an AI can diagnose cancer. The system was used as decision support. Pathologists could reject its hotspot, correct false positives and false negatives, and confirm the final count themselves.
The evidence is timely but preliminary. The manuscript was submitted to arXiv on 1 October 2026 and has not been peer reviewed. Its strongest claim concerns reproducibility and assessment time within a structured reader study, not clinical accuracy or improved patient care.[1]
The paired study used 385 independent patient slides
The retrospective, non-interventional study assembled one whole-slide haematoxylin-and-eosin image from each of 385 patients at centres in Lithuania, the United Kingdom and Spain. The primary dataset covered breast resection and biopsy specimens, melanoma, thyroid tumours, meningioma, gastrointestinal neuroendocrine tumours, sarcoma and leiomyoma. Slides came from three scanner models and were independent of the data used to train or analytically validate the model.
Thirteen pathologists contributed to the primary analysis. Each slide was independently reviewed by three randomly selected readers. In stage one, pathologists used the relevant standard counting method without AI assistance. After a washout of at least two weeks, the same reader-slide pair was assessed again with the software. Counting areas were one or two square millimetres depending on the applicable tumour guideline. Time ran from opening the slide to recording the final decision.
The paired design is useful because it reduces variation caused by comparing one group of readers with another. It does not remove every source of bias. All readers completed the unaided stage first and the assisted stage second, so practice, familiarity with a case set or a general learning effect could contribute to the later-stage advantage despite the washout. A counterbalanced order would have tested that possibility more directly.[1]
The model was trained separately from the reader-study cases
MitPro used a U-Net-based architecture trained on 28,222 annotated mitotic figures. Before the clinical reader study, the authors analytically validated detection on 463 one-square-millimetre regions from 215 whole-slide images containing 1,827 consensus mitoses. They report an F1 score of 0.803, precision of 0.728 and recall of 0.894 in that separate detection test. Those metrics describe candidate detection in annotated regions, not the accuracy of complete tumour grades or patient diagnoses.
In the assisted workflow, the software proposed three hotspots and highlighted candidate mitotic figures. Readers could choose a proposed region or draw their own, then review and amend the suggestions. The primary workflow ran inside the HALO AP image-management platform; supporting analyses used Sectra. The study also reports a separate 76-slide supporting set covering glioblastoma and non-small-cell lung cancer.
Separating training data from the reader-study slides is an important safeguard against evaluating on memorised examples. However, a single product evaluated with training supplied by its developer is still not the same as an independent head-to-head comparison across tools. Laboratories also vary in staining, scanners, specimen mix and quality-control practice beyond the combinations represented here.[1]
Agreement improved sharply in the combined analysis
The primary endpoint was a one-way random-effects intraclass correlation coefficient, or ICC, for agreement among readers. The combined ICC rose from 0.589 without assistance to 0.949 with assistance. The estimated difference was 0.360, with a 95% confidence interval from 0.260 to 0.449, calculated with a paired cluster bootstrap using 20,000 samples. The manuscript also reports increased agreement in the categorical mitotic scores used for grading.
The headline combined number should not conceal variation by tumour type. Thyroid cases improved from an ICC of 0.001 to 0.405, but the assisted value remained far below the overall 0.949. A system can therefore improve reproducibility without making agreement uniformly high. The primary study was powered at the main tumour-group level, not for every rarer histological subtype, so subtype results remain less certain.
Greater agreement does not prove that readers converged on the correct count. The study intentionally had no independent ground truth for the 385 clinical slides. The authors state that the observed shift toward higher counts should not be interpreted as evidence of greater accuracy. Consensus can be clinically useful, but people can agree with one another and still share the same systematic error.[1]
Counting was faster, although stage order matters
Mean pathologist-level median assessment time fell from 286.4 seconds in the unaided stage to 127.8 seconds in the assisted stage. The estimated paired saving was 151.8 seconds, with a reported 95% confidence interval from 144.1 to 161.8 seconds—about a 55% reduction. Time savings appeared across readers, tumour groups, patient age and sex groups, scanners and centres, although the authors caution that centre results were not designed as direct performance comparisons.
For a laboratory processing many relevant cases, a saving of roughly two and a half minutes per assessment could accumulate. Yet the timer covered the structured counting task, not the complete pathology workflow. It did not measure slide preparation, accessioning, quality control, software integration, report authorisation, downstream discussion or time spent managing system failures. Nor did the study test whether faster counts reduce total diagnostic turnaround.
The fixed sequence is particularly relevant to time. Readers already knew the task and had encountered the same cases when they reached stage two. The washout reduces immediate recall but cannot guarantee that all familiarity effects disappeared. A randomised crossover or counterbalanced design would give stronger evidence that the tool, rather than stage order, caused the full measured saving.[1]
Readers usually accepted a proposed hotspot but retained an override
Pathologists used an AI-proposed hotspot in 87.2% of assisted assessments. They selected a manually defined region in the remaining 12.8%, and every pathologist used the manual option at least once. That result supports an assistive design rather than an autonomous one: proposed regions were often useful, but readers sometimes judged another area more appropriate.
Counts shifted slightly upward with assistance. Across assessments, categorical scores were unchanged 64.9% of the time, higher in 29.1% and lower in 5.9%. Complete agreement among all three pathologists rose from 50% of slides in stage one to 75% in stage two. Within-stage consensus matching rose from 82% to 92%.
The upward movement deserves scrutiny because a higher mitotic score can alter a grade in some contexts. The authors found that score changes were not statistically significant and occurred about as frequently as routine between-reader differences in stage one. Without an independent reference, however, the study cannot establish whether a higher assisted score corrected an initial miss or introduced overcounting. Monitoring for automation bias and known difficult mimics would be essential in service.[1]
Commercial involvement is material to interpretation
The work was supported by Innovate UK grant 10040491 and Histofy Ltd, the company developing MitPro. The manuscript says one author led development of the tool and the study, two authors are directors and shareholders of Histofy, and another is a shareholder. The corresponding affiliation is also connected to the company. These disclosures do not make the results false, but they increase the importance of independent replication, preregistered endpoints and access to sufficiently detailed data and code.
Ethical approval was granted by the London–Fulham Research Ethics Committee under reference 25/LO/0158 and IRAS project 342246. The international sample is a strength relative to a single-centre reader study, yet all primary sites were European. The dataset cannot represent every staining protocol, scanner, population, laboratory process or rare tumour subtype encountered worldwide.
The preprint does not report a prospective clinical rollout comparing laboratories with and without the system. Interface findings also cannot simply be transferred between vendors: supporting results appeared similar in HALO AP and Sectra, but there was no reader-level crossover between the two systems. Implementation costs, licensing, training, cybersecurity, uptime and integration work were outside the evidence presented.[1]
What this could mean for patients and pathology teams
For pathologists, the immediate promise is a second set of computational eyes for two repetitive parts of the task: finding an active region and flagging candidate figures. More consistent counts could reduce avoidable disagreement and give specialists more time for interpretation. Preserving an override is important because the reader remains responsible for deciding what is biologically plausible and clinically relevant.
For patients, the evidence is indirect. The study did not measure whether final tumour grades became more accurate, whether disagreements requiring review fell, whether treatment recommendations changed, or whether outcomes improved. It also did not quantify harmful false positives, false negatives or delays caused by technical problems. Patients should not read the time reduction as proof of faster or safer care.
A cautious operational use would treat the model as auditable decision support. Laboratories would validate it locally, define when manual review or a second reader is required, track overrides and discordant cases, and regularly test slides containing difficult mitotic mimics. Performance should be monitored by tumour type rather than inferred from the combined result.[1]
What evidence would change the assessment
Confidence would rise with a peer-reviewed, preregistered replication led independently of the developer. A randomised and counterbalanced reader design should separate assistance from learning and order effects. Test cases should come from more countries, laboratories, scanners and rare subtypes, with results reported separately where biology or workflow differs.
The decisive accuracy study needs a defensible independent reference, perhaps adjudication by an expert panel using additional stains or follow-up information where appropriate. Researchers should measure not only cell counts but final grade, diagnostic decisions, treatment implications, false-positive and false-negative patterns, and reader behaviour when the model is deliberately wrong. Prospective trials should include full workflow time, implementation cost, downtime and quality-control burden.
For now, the bounded conclusion is that AI assistance made 13 pathologists more consistent with one another and faster on a repeated set of 385 digitised slides. That is a useful signal for a laborious pathology task. It is not evidence that the assisted counts were more accurate, that a diagnosis improved or that any patient experienced a better outcome.[1]
What this means for people
- Pathologists may spend less time searching and counting, but remain responsible for hotspot choice, correction and the final assessment.
- Laboratories could gain more reproducible counts, while taking on validation, integration, training, licensing and ongoing quality-control work.
- Patients could benefit only if better consistency translates into more accurate and timely diagnoses; this study did not test that link.
Global context
The reader study spans Lithuania, the United Kingdom and Spain, but digital-pathology infrastructure, specialist availability and tumour workloads differ substantially worldwide. Benefits in well-equipped European centres cannot be assumed in laboratories with other scanners, stains, staffing, connectivity or quality systems. Independent tests in Asia, the Middle East, Africa, Latin America and other European and North American settings are needed before the combined result can be treated as global evidence.
What the evidence does not yet show
- The paper is a preprint submitted on 1 October 2026 and has not been peer reviewed.
- The clinical reader study measured reproducibility and time, not accuracy against an independent ground truth, final diagnosis, treatment choice or patient outcome.
- All readers completed the unaided stage before the assisted stage, so learning, case familiarity or order effects may contribute despite the two-week washout.
- Three European centres, three scanners and seven primary tumour types cannot represent every population, staining process, laboratory workflow or rare subtype.
- The developer supported the work; one author led product development and the study, and three authors disclosed company shareholdings or directorships.
- Task timing did not measure the complete laboratory workflow, integration cost, downtime, cybersecurity, quality-control burden or diagnostic turnaround.
What to watch next
- Peer review and any changes to the methods, confidence intervals or interpretation in a journal version.
- Independent, preregistered replications with counterbalanced reader order and tumour-specific accuracy references.
- Prospective studies of final grade, diagnosis, treatment implications, full workflow time and patient-relevant errors.
- Post-deployment monitoring for automation bias, overcounting, difficult mitotic mimics, overrides and performance drift by tumour type.
Evidence trail
Sources used for this report
Links checked 3 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Health & Life Sciences
Can a medical AI support a clinical board without becoming the decision-maker?
A peer-reviewed German study put OpenEvidence into discussions of 100 real inflammatory-disease cases. Specialists found the answers useful, but complete agreement with the board was limited, prompt tuning was inconclusive and outputs changed over time. The study tested workflow feasibility—not patient benefit or safety.
9 min · 1 source
Health & Life Sciences
Can hospitals share AI insight for less?
A peer-reviewed benchmark across seven medical datasets found that consensus-based learning matched federated-learning accuracy overall while cutting measured training time and data transfer. The result is promising engineering evidence, not proof of clinical benefit or privacy.
8 min · 2 sources
Health & Life Sciences
Can brain MRI reveal more than BMI?
A peer-reviewed study trained a deep-learning model on 45,702 MRI scans from six cohorts. Brain-derived features tracked BMI and separated several disease groups better than BMI alone—but the disease analysis stayed inside UK Biobank and cannot establish cause, diagnosis or clinical benefit.
9 min · 2 sources
Reader discussion
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.