Back to the news portal
Health & Life SciencesResearch paperResearchSource analysisIndiaSouth AsiaGlobal radiology

Can AI segment and grade lung tumours from CT images?

A combined segmentation and classification framework reported 94.78% accuracy on a 264-image held-out partition. The evidence comes from one processed repository dataset, with no independent hospital or prospective validation.

By The Impact of AI Editorial DeskReleased 11 October 2026 at 17:57 BST6 min read1 source

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

Share
Social links
LinkedInXBlueskyRedditEmail

At a glance

  • 1The study used 972 CT images: 708 in the training partition and 264 in the repository-defined held-out test partition.
  • 2The complete framework reported a Dice score of 0.8563 for segmentation and 94.78% classification accuracy, F1 of 0.9476, AUC of 0.9872 and MCC of 0.9215.
  • 3The test comes from one processed repository dataset and consolidates Lung-RADS labels into three categories, so independent patient-level and multi-institutional performance remains unknown.
Key themesClinical AILung cancerCT imagingRadiomicsDiagnostic validation

Research topic

Whether boundary-aware segmentation, radiomic and deep-feature fusion, and peritumour context can jointly segment lung tumours and classify consolidated Lung-RADS categories

The answer: it performed well on one held-out image set, not across hospitals

A peer-reviewed study reports that a combined lung-tumour segmentation and classification framework achieved 94.78% accuracy on a repository-defined test partition of 264 computed-tomography images. Its classification F1 was 0.9476, area under the receiver-operating-characteristic curve was 0.9872 and Matthews correlation coefficient was 0.9215. The segmentation component achieved a Dice score of 0.8563, indicating substantial but incomplete overlap between predicted and reference tumour regions.

These numbers show that the components worked together on the selected benchmark. They do not establish reliable diagnosis in a new hospital. The full dataset contained 972 images from one processed repository resource, and the accessible report does not establish an external institution, prospective patient series or live clinical workflow. Differences in scanners, reconstruction, contrast, tumour prevalence and annotation practice can all reduce performance after transfer.[1]

The Impact Brief · Free

Follow the evidence in health & life sciences.

Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.

Choose your topics (optional)

One concise, source-linked briefing. Unsubscribe at any time.

How the three-part system was evaluated

The processed Lung Cancer Segmentation Dataset with Lung-RADS classes supplied 708 training images and 264 held-out test images. The source labels were consolidated into three groups: Mild for Lung-RADS 2, Moderate for Lung-RADS 3 and Severe for Lung-RADS 4A to 4B. That makes the experiment a three-category image-classification task alongside tumour segmentation, not a direct measurement of cancer mortality, treatment response or the full set of decisions a radiologist makes.

BAMR-Net first segments a tumour using multi-scale residual learning and boundary refinement. TRFF-Net then fuses engineered radiomic measurements with learned deep features using adjustable weights. MTCA-Net adds information from the tumour and its surrounding tissue. The complete BAMR-HF pipeline combines those stages. This architecture is designed around a sensible clinical intuition: boundaries, internal texture and peritumour context may each carry different information.[1]

The component comparisons support the design, within one setting

The radiomic-deep fusion component reported 90.38% test accuracy, while the context-aware component reached 93.27%. Combining the full sequence raised accuracy to 94.78%. Those comparisons suggest that segmentation, fused features and surrounding context each contributed under the same experimental conditions. Because all comparisons use the same underlying dataset, they do not show which component will remain valuable when a scanner, hospital or patient population changes.

A Dice score of 0.8563 is not perfect boundary agreement. Remaining differences could matter if a tool is used for radiation planning, volumetric follow-up or extracting radiomic features from a predicted contour. The study's classification result should also not obscure the segmentation uncertainty: later stages may inherit errors from the first stage. A clinically useful system needs case-level analysis of where boundaries fail, not only an average overlap score.[1]

Lung-RADS labels are useful targets but not a final diagnosis

Lung-RADS supports structured assessment of findings in lung-cancer screening, but consolidating 4A and 4B into one severe group removes distinctions that may affect follow-up. The three study categories simplify a harder clinical continuum. A model can score highly after consolidation while making errors near boundaries that matter to surveillance or referral. Class prevalence also affects how often a positive prediction will be correct in practice. Its output does not replace pathology, longitudinal imaging, symptoms or the radiologist's assessment of the whole examination.

The denominator is images, not a reported number of independently tested patients or institutions. Multiple images can come from related examinations, and model performance can be inflated when correlated records are divided across development stages. The repository-defined held-out partition is better than evaluating on training data, but confidence depends on how patients and acquisitions were separated. External patient-level testing is the appropriate next standard.[1]

What this could mean for radiologists and patients

A validated system could outline suspected tumours, extract reproducible features and direct attention to higher-risk scans. That might reduce repetitive contouring or help services triage growing volumes. This study does not measure reporting time, disagreement between readers, changed referrals or cancer outcomes. It also does not show how the model handles no tumour, multiple nodules, postoperative anatomy, motion artefact or an image from a scanner outside the repository.

For patients, false reassurance and unnecessary escalation are both consequential. A missed or undergraded finding can delay assessment; an overgraded finding can lead to anxiety, repeat imaging or invasive procedures. Automation should expose uncertainty and make it easy for clinicians to inspect and correct the contour. A high benchmark AUC does not determine the safest threshold for a population with a different disease prevalence or available follow-up capacity.[1]

Limits and what would change the assessment

The two authors are affiliated with engineering colleges in Tamil Nadu, India, and declare no competing interests. The article is a citable accepted version that may receive editorial corrections before the final Version of Record. Its main limitation is concentration in one processed 972-image resource. The evidence does not cover independent hospitals, prospective workflow, clinical reader comparison, subgroup performance or outcomes after an AI-supported decision.

Confidence would rise with a frozen model tested on patient-disjoint datasets from several institutions, scanner vendors and countries. Reports should give class counts, confidence intervals, calibration, case-level segmentation failures and sensitivity at clinically chosen thresholds. Reader studies should compare unaided and AI-assisted radiologists, recording time, corrections and downstream decisions. Prospective evidence should then test whether the tool improves care without increasing missed cancers, unnecessary investigations or workload.[1]

What this means for people

  • Radiologists could eventually gain contouring and triage support, but the study did not test clinical work.
  • Patients face harm from both missed suspicious findings and unnecessary escalation if thresholds do not transfer.
  • Health services need calibrated case-level performance and downstream workload evidence before deployment.

Global context

Lung-screening programmes differ in eligibility, prevalence, scanner fleets and follow-up capacity. A model trained on a processed public resource may not retain the same calibration in another country or health system. This India-based engineering study supports further external validation; it does not support importing one benchmark threshold or accuracy figure directly into clinical practice.

What the evidence does not yet show

  • The evidence comes from one processed repository dataset with 972 images.
  • The accessible report does not establish independent hospital or prospective validation.
  • Lung-RADS labels were consolidated into three categories, reducing clinical granularity.
  • Average Dice and classification metrics do not reveal every clinically important boundary or class error.
  • No radiologist-reader study, workflow outcome or patient outcome was measured.

What to watch next

  • Patient-disjoint multicentre validation
  • Scanner and subgroup performance
  • Reader studies
  • Prospective effects on referrals and outcomes

Living evidence record

Impact record IAI-133YANV

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent or research support

Present

Record status

Monitoring

Last checked

11 October 2026

Source trail

1 direct source across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Single-source reporting disclosure

This record analyses one direct source. It can establish what Scientific Reports published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.

Evidence trail

Sources used for this report

Links checked 11 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Health & Life Sciences

Can AI identify leukaemia cells reliably across patients?

A hybrid image model correctly classified 1,583 of 1,599 held-out cell images, but the split was image-level rather than patient-disjoint. The 99% result is an internal benchmark, not evidence that the system works in a new clinic.

6 min · 1 source

Health & Life Sciences

Can a high-AUC diabetes model still be unsafe?

New analysis today of a peer-reviewed 2 October audit of 12 machine-learning approaches on two public diabetes datasets. Similarly ranked models can differ materially in calibration, uncertainty and safe deferral, but this is a benchmark study—not a clinical trial, diagnostic approval or evidence of improved patient outcomes.

7 min · 3 sources

Health & Life Sciences

Can this MRI model draw brain-tumour boundaries reliably?

A peer-reviewed multimodal segmentation model was developed on 2,422 public MRI volumes and externally tested on 125 cases. Accuracy remained useful but fell outside the development data; no prospective clinical workflow, radiologist comparison or patient-outcome test was performed.

8 min · 2 sources

Reader commentary

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Explore commentary across the portal →

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.