Can AI segment and grade lung tumours from CT images?
A combined segmentation and classification framework reported 94.78% accuracy on a 264-image held-out partition. The evidence comes from one processed repository dataset, with no independent hospital or prospective validation.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1The study used 972 CT images: 708 in the training partition and 264 in the repository-defined held-out test partition.
- 2The complete framework reported a Dice score of 0.8563 for segmentation and 94.78% classification accuracy, F1 of 0.9476, AUC of 0.9872 and MCC of 0.9215.
- 3The test comes from one processed repository dataset and consolidates Lung-RADS labels into three categories, so independent patient-level and multi-institutional performance remains unknown.
Research topic
Whether boundary-aware segmentation, radiomic and deep-feature fusion, and peritumour context can jointly segment lung tumours and classify consolidated Lung-RADS categories
The answer: it performed well on one held-out image set, not across hospitals
A peer-reviewed study reports that a combined lung-tumour segmentation and classification framework achieved 94.78% accuracy on a repository-defined test partition of 264 computed-tomography images. Its classification F1 was 0.9476, area under the receiver-operating-characteristic curve was 0.9872 and Matthews correlation coefficient was 0.9215. The segmentation component achieved a Dice score of 0.8563, indicating substantial but incomplete overlap between predicted and reference tumour regions.
These numbers show that the components worked together on the selected benchmark. They do not establish reliable diagnosis in a new hospital. The full dataset contained 972 images from one processed repository resource, and the accessible report does not establish an external institution, prospective patient series or live clinical workflow. Differences in scanners, reconstruction, contrast, tumour prevalence and annotation practice can all reduce performance after transfer.[1]
The Impact Brief · Free
Follow the evidence in health & life sciences.
Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.
How the three-part system was evaluated
The processed Lung Cancer Segmentation Dataset with Lung-RADS classes supplied 708 training images and 264 held-out test images. The source labels were consolidated into three groups: Mild for Lung-RADS 2, Moderate for Lung-RADS 3 and Severe for Lung-RADS 4A to 4B. That makes the experiment a three-category image-classification task alongside tumour segmentation, not a direct measurement of cancer mortality, treatment response or the full set of decisions a radiologist makes.
BAMR-Net first segments a tumour using multi-scale residual learning and boundary refinement. TRFF-Net then fuses engineered radiomic measurements with learned deep features using adjustable weights. MTCA-Net adds information from the tumour and its surrounding tissue. The complete BAMR-HF pipeline combines those stages. This architecture is designed around a sensible clinical intuition: boundaries, internal texture and peritumour context may each carry different information.[1]
The component comparisons support the design, within one setting
The radiomic-deep fusion component reported 90.38% test accuracy, while the context-aware component reached 93.27%. Combining the full sequence raised accuracy to 94.78%. Those comparisons suggest that segmentation, fused features and surrounding context each contributed under the same experimental conditions. Because all comparisons use the same underlying dataset, they do not show which component will remain valuable when a scanner, hospital or patient population changes.
A Dice score of 0.8563 is not perfect boundary agreement. Remaining differences could matter if a tool is used for radiation planning, volumetric follow-up or extracting radiomic features from a predicted contour. The study's classification result should also not obscure the segmentation uncertainty: later stages may inherit errors from the first stage. A clinically useful system needs case-level analysis of where boundaries fail, not only an average overlap score.[1]
Lung-RADS labels are useful targets but not a final diagnosis
Lung-RADS supports structured assessment of findings in lung-cancer screening, but consolidating 4A and 4B into one severe group removes distinctions that may affect follow-up. The three study categories simplify a harder clinical continuum. A model can score highly after consolidation while making errors near boundaries that matter to surveillance or referral. Class prevalence also affects how often a positive prediction will be correct in practice. Its output does not replace pathology, longitudinal imaging, symptoms or the radiologist's assessment of the whole examination.
The denominator is images, not a reported number of independently tested patients or institutions. Multiple images can come from related examinations, and model performance can be inflated when correlated records are divided across development stages. The repository-defined held-out partition is better than evaluating on training data, but confidence depends on how patients and acquisitions were separated. External patient-level testing is the appropriate next standard.[1]
What this could mean for radiologists and patients
A validated system could outline suspected tumours, extract reproducible features and direct attention to higher-risk scans. That might reduce repetitive contouring or help services triage growing volumes. This study does not measure reporting time, disagreement between readers, changed referrals or cancer outcomes. It also does not show how the model handles no tumour, multiple nodules, postoperative anatomy, motion artefact or an image from a scanner outside the repository.
For patients, false reassurance and unnecessary escalation are both consequential. A missed or undergraded finding can delay assessment; an overgraded finding can lead to anxiety, repeat imaging or invasive procedures. Automation should expose uncertainty and make it easy for clinicians to inspect and correct the contour. A high benchmark AUC does not determine the safest threshold for a population with a different disease prevalence or available follow-up capacity.[1]
Limits and what would change the assessment
The two authors are affiliated with engineering colleges in Tamil Nadu, India, and declare no competing interests. The article is a citable accepted version that may receive editorial corrections before the final Version of Record. Its main limitation is concentration in one processed 972-image resource. The evidence does not cover independent hospitals, prospective workflow, clinical reader comparison, subgroup performance or outcomes after an AI-supported decision.
Confidence would rise with a frozen model tested on patient-disjoint datasets from several institutions, scanner vendors and countries. Reports should give class counts, confidence intervals, calibration, case-level segmentation failures and sensitivity at clinically chosen thresholds. Reader studies should compare unaided and AI-assisted radiologists, recording time, corrections and downstream decisions. Prospective evidence should then test whether the tool improves care without increasing missed cancers, unnecessary investigations or workload.[1]
What this means for people
- Radiologists could eventually gain contouring and triage support, but the study did not test clinical work.
- Patients face harm from both missed suspicious findings and unnecessary escalation if thresholds do not transfer.
- Health services need calibrated case-level performance and downstream workload evidence before deployment.
Global context
Lung-screening programmes differ in eligibility, prevalence, scanner fleets and follow-up capacity. A model trained on a processed public resource may not retain the same calibration in another country or health system. This India-based engineering study supports further external validation; it does not support importing one benchmark threshold or accuracy figure directly into clinical practice.
What the evidence does not yet show
- The evidence comes from one processed repository dataset with 972 images.
- The accessible report does not establish independent hospital or prospective validation.
- Lung-RADS labels were consolidated into three categories, reducing clinical granularity.
- Average Dice and classification metrics do not reveal every clinically important boundary or class error.
- No radiologist-reader study, workflow outcome or patient outcome was measured.
What to watch next
- Patient-disjoint multicentre validation
- Scanner and subgroup performance
- Reader studies
- Prospective effects on referrals and outcomes
Living evidence record
Impact record IAI-133YANV
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
11 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what Scientific Reports published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
Evidence trail
Sources used for this report
Links checked 11 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Health & Life Sciences
Can AI identify leukaemia cells reliably across patients?
A hybrid image model correctly classified 1,583 of 1,599 held-out cell images, but the split was image-level rather than patient-disjoint. The 99% result is an internal benchmark, not evidence that the system works in a new clinic.
6 min · 1 source
Health & Life Sciences
Can a high-AUC diabetes model still be unsafe?
New analysis today of a peer-reviewed 2 October audit of 12 machine-learning approaches on two public diabetes datasets. Similarly ranked models can differ materially in calibration, uncertainty and safe deferral, but this is a benchmark study—not a clinical trial, diagnostic approval or evidence of improved patient outcomes.
7 min · 3 sources
Health & Life Sciences
Can this MRI model draw brain-tumour boundaries reliably?
A peer-reviewed multimodal segmentation model was developed on 2,422 public MRI volumes and externally tested on 125 cases. Accuracy remained useful but fell outside the development data; no prospective clinical workflow, radiologist comparison or patient-outcome test was performed.
8 min · 2 sources
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.