Can AI identify leukaemia cells reliably across patients?
A hybrid image model correctly classified 1,583 of 1,599 held-out cell images, but the split was image-level rather than patient-disjoint. The 99% result is an internal benchmark, not evidence that the system works in a new clinic.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1The model used 10,661 labelled cell images from the training component of C-NMC 2019, divided into 7,462 training, 1,600 validation and 1,599 held-out test images.
- 2It correctly classified 1,583 of 1,599 test images, reporting 99.00% accuracy and approximately 0.990 precision, recall and F1, with an AUC of 0.99.
- 3The split was stratified by image rather than by patient, so visually related cells from the same person may occur across partitions and inflate apparent generalisation.
Research topic
Whether a MobileNet and vision-transformer hybrid can distinguish acute lymphoblastic leukaemia blasts from healthy haematogones in a public blood-smear image benchmark
The answer: it classified images accurately, but cross-patient reliability is unproven
A peer-reviewed study reports that a hybrid MobileNet and vision-transformer system correctly classified 1,583 of 1,599 held-out microscope images as acute lymphoblastic leukaemia blasts or healthy haematogones. That corresponds to 99.00% accuracy; precision, recall and F1 were each about 0.990, and area under the receiver-operating-characteristic curve was 0.99. On this internal benchmark, the model separated the two cell-image categories very well.
The result does not show 99% diagnostic accuracy in patients. The researchers split images, not people, and used a single public dataset. Cells from one person can share staining, scanner, preparation and biological characteristics. If related images appear in both training and test partitions, a model may recognise those shared patterns without learning features that travel reliably to a new patient, laboratory or microscope. The paper itself frames the result as internal benchmark performance rather than multicentre clinical generalisation.[1]
The Impact Brief · Free
Follow the evidence in health & life sciences.
Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.
What the researchers actually tested
The experiment used 10,661 labelled cell images from the training component of the C-NMC 2019 challenge dataset. A stratified image-level partition assigned 7,462 images to training, 1,600 to validation and 1,599 to the final test. Data augmentation was restricted to the training partition. The task was binary image classification: malignant acute lymphoblastic leukaemia blasts versus healthy haematogones, which can look similar under a microscope.
The proposed MoViTHAFX-Net has two main feature branches. MobileNet extracts local visual details such as cell shape and texture, while a vision transformer models relationships across more distant parts of an image. A hierarchical attention-fusion component refines and combines three learnable morphology-guided projections before classification. Grad-CAM++ produces post-hoc visualisations intended to show which image regions influenced a decision. This is a technically specific pipeline, not a clinical workflow evaluation.[1]
Why 1,583 correct images still leave an important uncertainty
Sixteen test images were misclassified. That is a small image-level error count, but it does not reveal how errors cluster by patient, slide, laboratory or difficult cell subtype. If several mistakes came from one person, the patient-level implication could be larger than the image average suggests. Conversely, a clinician normally interprets multiple cells together with counts, history and other tests; an isolated cell prediction is not itself a diagnosis. The unit of analysis therefore matters as much as the percentage.
A patient-disjoint split would place every image from one person in only one partition. External validation would go further by testing specimens collected and processed elsewhere, ideally with different microscopes, staining protocols and patient populations. Neither condition is established here. The benchmark's held-out label is genuine at image level, but it should not be translated into a claim about unseen patients or routine services.[1]
What attention fusion and heatmaps do—and do not—show
Ablation experiments in the same setting indicate that hierarchical attention fusion adds an incremental advantage over simply concatenating the convolutional and transformer features. That supports the authors' architectural claim within the benchmark. It does not establish that the added component is necessary in other datasets, that the gain survives a patient-disjoint split or that a smaller model could not achieve a similar result with less computation.
Grad-CAM++ can help reviewers see whether highlighted areas fall on the cell rather than the background, but a plausible heatmap is not a causal explanation. It can be stable while the model relies on a staining artefact, and it can vary without changing the clinically relevant evidence. The article reports qualitative visualisation, not a blinded study in which haematologists judged whether explanations were faithful, useful or improved decisions.[1]
What this means for laboratory staff and patients
For laboratory teams, a well-validated image model could prioritise suspicious fields, support quality checks or provide a second reading cue when workloads are high. This study does not test any of those uses. It does not measure time saved, cases escalated appropriately, false reassurance, staff confidence or how the tool behaves when image quality is poor. A screening aid also needs a defined response to uncertainty rather than forcing every image into one of two categories.
For patients, the central risk is that an impressive image score is mistaken for diagnostic proof. Acute lymphoblastic leukaemia diagnosis depends on a broader clinical and laboratory assessment. A missed malignant cell could delay review, while a false positive could trigger anxiety and unnecessary investigation. The appropriate near-term interpretation is that the architecture deserves patient-separated external testing, not that it is ready to replace microscopy expertise or established diagnostic pathways.[1]
Limits and what would change the assessment
The authors are based at Vellore Institute of Technology in India, report no external funding and declare no competing interests. The article is a citable accepted version that may receive editorial corrections before the final Version of Record. Its evidence comes from one public benchmark's training component, one custom image-level partition and one binary task. It does not report prospective use, multicentre variation, patient-disjoint performance or clinical outcomes.
Confidence would rise with a preregistered patient-level split followed by a frozen-model evaluation in independent hospitals. Results should include sensitivity and specificity by patient, laboratory, microscope and clinically important subgroup; calibration and abstention; reader studies with haematologists; and workload or turnaround-time effects. The most decisive evidence would be a prospective comparison showing that the tool improves an actual diagnostic workflow without increasing missed cases or unnecessary referrals.[1]
What this means for people
- Laboratory staff could eventually use a validated model to prioritise images, but this study did not test work in a laboratory.
- Patients could be harmed if an internal image benchmark is presented as a clinical diagnostic guarantee.
- Health services need patient-level sensitivity, subgroup performance and workload evidence before considering deployment.
Global context
Public image datasets let teams in many countries compare methods, but staining, microscopes, disease prevalence, referral pathways and staffing differ across health systems. This India-based study offers a clear benchmark result and an equally important warning about image-level splitting. Cross-border clinical use would require local patient-separated validation, quality controls and regulation appropriate to a diagnostic support tool.
What the evidence does not yet show
- The study split images rather than patients, allowing patient-specific visual patterns to cross partitions.
- All results come from one public benchmark and a binary cell-image task.
- Grad-CAM++ visualisations were qualitative and do not prove faithful clinical reasoning.
- No clinician-reader, workflow, prospective or multicentre evaluation was reported.
- Image-level accuracy cannot be interpreted as patient-level diagnostic accuracy.
What to watch next
- Patient-disjoint replication
- Independent multicentre testing
- Calibration and abstention
- Prospective laboratory workflow studies
Living evidence record
Impact record IAI-0S6DQ9Y
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
11 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what Scientific Reports published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
Evidence trail
Sources used for this report
Links checked 11 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Health & Life Sciences
Can AI segment and grade lung tumours from CT images?
A combined segmentation and classification framework reported 94.78% accuracy on a 264-image held-out partition. The evidence comes from one processed repository dataset, with no independent hospital or prospective validation.
6 min · 1 source
Health & Life Sciences
Can a high-AUC diabetes model still be unsafe?
New analysis today of a peer-reviewed 2 October audit of 12 machine-learning approaches on two public diabetes datasets. Similarly ranked models can differ materially in calibration, uncertainty and safe deferral, but this is a benchmark study—not a clinical trial, diagnostic approval or evidence of improved patient outcomes.
7 min · 3 sources
Health & Life Sciences
Can this MRI model draw brain-tumour boundaries reliably?
A peer-reviewed multimodal segmentation model was developed on 2,422 public MRI volumes and externally tested on 125 cases. Accuracy remained useful but fell outside the development data; no prospective clinical workflow, radiologist comparison or patient-outcome test was performed.
8 min · 2 sources
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.