Back to the news portal
Health & Life SciencesResearch paperResearchSource analysisSpainEuropeInternational

Can mouse MRI predict glioblastoma treatment response?

A deep-learning model separated cure from relapse across 298 MRI examinations, but the effective test cohort was only 10 treated mice from one laboratory. It is preclinical evidence, not a patient-ready predictor.

By The Impact of AI Health & Life Sciences DeskReleased 4 October 2026 at 16:57 BST7 min read1 source

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesGlioblastomaMedical imagingDeep learningRadiomicsPreclinical researchTreatment response

Research topic

Whether radiomics and transfer learning can distinguish eventual cure from relapse using longitudinal T2-weighted MRI in a treated mouse glioblastoma model

The Impact of AI research cover asking whether mouse MRI can predict glioblastoma treatment response, with a conceptual mouse-brain scan and two uncertain outcome paths.
AI-generated editorial illustration. The mouse-brain scan, model layers and outcome paths are conceptual; they do not depict a patient, an actual laboratory animal, measured MRI, treatment decision or validated clinical system.

At a glance

  • 1The study began with 18 mice: eight untreated controls and 10 treated animals, evenly divided between eventual cure and relapse. The response classifiers used the 10 treated animals.
  • 2A fine-tuned EfficientNetB0 reached an examination-level AUC of 0.868 and sensitivity of 0.818, compared with an AUC of 0.770 and sensitivity of 0.545 for radiomics plus XGBoost.
  • 3The effective independent sample was 10 mice from one GL261 model and one institution. Repeated examinations are not 298 independent animals, and no patient data or external cohort were tested.

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Living evidence record

Impact record IAI-1D5KIOT

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent support

Present

Record status

Monitoring

Last checked

4 October 2026

Source trail

1 direct source across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Single-source reporting disclosure

This record analyses one direct source. It can establish what Scientific Reports published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.

The question is response monitoring, not diagnosis

Glioblastoma treatment can produce imaging changes that are difficult to interpret. Apparent progression may reflect treatment effects, while a temporary reduction in enhancement may not represent lasting control. Researchers at the Universitat Politècnica de Catalunya, Universitat Autònoma de Barcelona and CIBER-BBN asked a narrower preclinical question: could ordinary T2-weighted MRI contain patterns that separate mice eventually cured by temozolomide from mice whose tumours later relapsed?

The team compared two approaches. A radiomics pipeline converted manually segmented tumours into engineered shape, intensity and texture features, then classified them with XGBoost. A deep-learning pipeline adapted EfficientNetB0, a convolutional network originally pretrained on natural images, to the MRI slices. This was not a test of detecting glioblastoma, choosing a drug or predicting survival in people. It was a retrospective discrimination task inside one controlled mouse model.[1]

Ten treated mice supplied the decisive comparison

The underlying cohort contained 18 C57BL/6 mice implanted with GL261 glioblastoma cells. Eight untreated controls contributed 36 MRI examinations. Five temozolomide-treated mice had an initial response followed by relapse and contributed 84 examinations; five treated mice had complete regression and long-term survival and contributed 49. The classification analysis excluded the controls and compared the five relapsing with the five cured animals.

The paper reports 298 examination-level samples for its modelling analysis because individual imaging sessions contained multiple images. That number must not be mistaken for the number of independent subjects. Images from one mouse are correlated. The authors explicitly identify 10 treated animals as the effective independent sample and use grouped leave-one-animal-out cross-validation so that all examinations from the held-out mouse remain outside a training fold.

For radiomics, 1,218 candidate features were extracted from manually segmented regions. Imputation, variance and correlation filtering, statistical selection, scaling, oversampling and model fitting were performed within each training fold. An average of about 22 features survived. This design addresses an important leakage risk, although feature selection and hyperparameter choices in such a tiny dataset can still produce unstable estimates that require external replication.[1]

Deep learning performed better on examinations

At the examination level, radiomics plus XGBoost reached an area under the receiver operating characteristic curve of 0.770, with a 95% bootstrap confidence interval from 0.703 to 0.832. Its sensitivity for cured examinations was 0.545 and its positive predictive value was 0.444. A frozen EfficientNet feature extractor followed by XGBoost reached an AUC of 0.801. Fine-tuning the final 50 network layers produced the strongest examination-level AUC, 0.868, with a reported 95% interval from 0.810 to 0.918, sensitivity of 0.818 and positive predictive value of 0.491.

At the animal level, majority voting produced AUCs of 0.92 for radiomics, 0.92 for the frozen feature extractor and 0.88 for the fine-tuned model. All three correctly identified all five relapsing mice. The fine-tuned model identified four of five cured mice; the other approaches identified three. With only five animals in either outcome group, one changed classification moves sensitivity by 20 percentage points. Those animal-level figures are therefore descriptive rather than a stable clinical estimate.

The authors also divided scans into early, middle and late windows. Fine-tuned EfficientNet AUC rose from 0.788 in the early window to 1.000 in the late window. That trend is plausible because treatment effects become more visible with time, but the windows reuse data from the same small group. A perfect late-window result does not establish perfect prediction in a new mouse, let alone an early warning tool for patients.[1]

Subject-aware validation helps, but does not create an external test

Keeping each mouse wholly inside either training or testing is the correct safeguard against an easy form of leakage. Training augmentation and oversampling were also confined to training partitions, and the authors used a fixed seed and deterministic computation. However, the held-out mouse was supplied to the neural-network training process as validation data for monitoring, although the paper says it did not influence early stopping or checkpoint selection because every fold ran for 20 epochs and used final weights. Future work should keep an untouched external cohort entirely outside model development and monitoring.

The network processed individual time points, not each animal’s sequence as a temporal model. Its attention maps often included tumour and surrounding tissue, but Grad-CAM is a visualisation of model sensitivity rather than proof of a biological mechanism. The radiomic features that differed between outcome groups are similarly hypothesis-generating. Several illustrated trajectories rest on five cured and five relapsing animals and were not established as repeatable biomarkers.[1]

What this means for patients now

The immediate effect on patient care is none. The mice carried one implanted tumour model, received a specific temozolomide schedule and were scanned on a 7-tesla preclinical instrument at one centre. Human glioblastoma varies across molecular subtypes, treatments, scanners and clinical histories. People also undergo surgery, radiotherapy and combined regimens that can change images in ways this experiment did not represent.

The result is useful as a methods signal: standard T2-weighted imaging may contain response information worth testing, and learned image features may capture more of it than one engineered-feature pipeline. A credible translational path would require preregistered studies in independent animal models, then multi-centre patient cohorts with prospectively defined time points, clinically relevant comparators, calibration, subgroup analysis and decision-impact testing. Any clinical system would also need to show whether it changes treatment safely rather than merely ranking outcomes after the fact.[1]

Funding and what would change the assessment

The work was supported by Spanish national research grants, the Generalitat de Catalunya and Instituto de Salud Carlos III. The authors declared no competing interests and say data may be requested from the corresponding author. Animal procedures were approved by the Universitat Autònoma de Barcelona ethics committee under two named protocols and followed Spanish and European animal-protection rules.

Confidence would rise if an independent team reproduced the result on more animals, different tumour models, scanners and treatment regimens, with a locked model and untouched external test set. For relevance to people, the essential evidence is prospective multi-centre validation against expert assessment and patient outcomes. Until then, the careful conclusion is that deep learning separated two outcomes in a very small preclinical cohort—not that AI can predict a person’s glioblastoma response.[1]

What this means for people

  • Patients should not interpret these mouse results as a test available for their own MRI or a reason to change treatment.
  • Researchers gain a leakage-aware preclinical comparison that can inform larger response-monitoring studies.
  • Clinicians would need external validation, calibration and decision-impact evidence before relying on a similar system.

Global context

The study comes from Spanish research institutions and a shared preclinical MRI facility. Its methods address a global clinical problem, but scanner access, treatment protocols and patient populations vary widely. Translation would require data from multiple health systems and resource settings, not only high-field research centres, with attention to molecular diversity and the uneven availability of advanced imaging.

What the evidence does not yet show

  • The decisive cohort contained only 10 treated mice: five eventually cured and five that relapsed.
  • All data came from one GL261 mouse model, one institution, one scanner setting and one treatment schedule; there was no external validation.
  • The 298 examination-level observations are clustered within animals and are not 298 independent subjects.
  • Manual tumour segmentation may introduce observer variability, and neither pipeline explicitly modelled full longitudinal sequences.
  • No patients, clinical decisions or patient-centred outcomes were tested.

What to watch next

  • Locked-model replication in independent animal cohorts and alternative glioblastoma models.
  • Multi-centre patient studies that preserve subject-level separation and report calibration as well as discrimination.
  • Prospective comparisons with radiologist assessment and standard response criteria at clinically useful early time points.
  • Evidence that model-informed decisions improve outcomes without unnecessary treatment changes.

Evidence trail

Sources used for this report

Links checked 4 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Health & Life Sciences

Can brain MRI reveal more than BMI?

A peer-reviewed study trained a deep-learning model on 45,702 MRI scans from six cohorts. Brain-derived features tracked BMI and separated several disease groups better than BMI alone—but the disease analysis stayed inside UK Biobank and cannot establish cause, diagnosis or clinical benefit.

9 min · 2 sources

Health & Life Sciences

Can AI support breast-ultrasound decisions across countries?

A peer-reviewed South Korean-led study validated an interpretable retrieval-augmented system across 8,311 images from 11 cohorts in seven countries and tested assistance with four readers. The retrospective evidence is encouraging, but it is not a prospective screening trial or proof of better patient outcomes.

8 min · 3 sources

Health & Life Sciences

Can this MRI model draw brain-tumour boundaries reliably?

A peer-reviewed multimodal segmentation model was developed on 2,422 public MRI volumes and externally tested on 125 cases. Accuracy remained useful but fell outside the development data; no prospective clinical workflow, radiologist comparison or patient-outcome test was performed.

8 min · 2 sources

Reader commentary

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Explore commentary across the portal →

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.