Back to the news portal
Health & Life SciencesNew analysis today · source 10 October 2026Research paperResearchSource analysisIraqAsia & Middle EastEuropeGlobal health

Can prostate-pathology AI generalise to Iraqi patients?

Three research models trained outside Iraq showed strong agreement with pathologists on 339 biopsy slides from 185 patients in Erbil. The result supports local validation, not autonomous diagnosis or immediate clinical deployment.

By The Impact of AI Editorial DeskReleased 11 October 2026 at 06:56 BST8 min read3 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

Share
Social links
LinkedInXBlueskyRedditEmail

At a glance

  • 1The retrospective consecutive series contained 339 slides from 185 patients treated in Erbil from 2013 to 2024; each slide was scanned on three devices, creating 1,017 whole-slide images.
  • 2Cancer-detection AUC ranged from 0.916 to 0.966, while the task-specific model was more specific than the two foundation-model pipelines and required substantially less computing energy in prior measurements.
  • 3Agreement with pathologists does not establish clinical benefit, and the study was not designed as an equivalence or non-inferiority trial; the algorithms remain research systems without diagnostic approval.
Key themesProstate cancerDigital pathologyExternal validationHealth equityClinical AIDiagnostic safety

Research topic

Whether prostate-biopsy AI developed on Scandinavian and North American data retains diagnostic and grading performance in an Iraqi cohort and across three slide scanners

The answer: the models travelled better than many clinical AI systems do, but they are not ready to diagnose patients

Three pathology AI systems trained outside Iraq produced strong cancer-detection and grading results on prostate biopsies from the Kurdistan Region. The best task-specific model reached an area under the receiver operating characteristic curve of 0.966 for separating malignant from benign slides. Across the three models, sensitivity was 97.1% to 98.9%. Their average agreement with pathologists on ISUP grade was close to the agreement among pathologists in the study's 59-slide, three-reader subset.

That is meaningful external-validation evidence because geographic, laboratory, patient and scanner data were separate from the systems' development material. It is not evidence that a laboratory can remove human review, that treatment decisions improve, or that errors are acceptable in practice. The authors explicitly state that the systems are research algorithms without CE-IVD or US FDA approval. The safest interpretation is that local prospective evaluation is now more credible—not that deployment has been proved safe.[1]

The Impact Brief · Free

Follow the evidence in health & life sciences.

Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.

Choose your topics (optional)

One concise, source-linked briefing. Unsubscribe at any time.

What was tested and who was represented

Researchers retrieved a consecutive series of prostate core-needle biopsies from PAR Private Hospital in Erbil. The archive covered patients investigated for suspected adenocarcinoma between mid-2013 and mid-2024. After excluding prostatectomy and transurethral-resection material, the final cohort included 339 glass slides from 185 patients. The mean age was 68 years, with a range from 33 to 95. One pathologist graded all 339 slides, a second graded 337, and a uropathologist independently reviewed a stratified 59-slide subset plus difficult discrepant cases.

Each physical slide was digitised on two high-throughput scanners and a compact Grundium device, yielding 1,017 whole-slide images. The three AI pipelines used the same attention-based multiple-instance framework but different feature encoders: a model trained specifically on prostate biopsies, plus the UNI and Virchow2 foundation models. Development used 51,247 whole-slide images from Sweden and Norway and earlier validation used 25,601 images from 15 European and Australian clinical sites. The Iraqi material was not used to tune or recalibrate them.[1][2][3]

The headline numbers—and the errors behind them

Against the primary pathologist, cancer-detection AUC ranged from 0.916 for UNI to 0.966 for the task-specific system. Sensitivity was high for all three, but specificity separated them: 94.5% for the task-specific model, 85.4% for UNI and 88.4% for Virchow2. That difference matters in daily work because false-positive calls can force pathologists to re-examine benign glands, adding workload rather than relieving it. Grading agreement also favoured the task-specific system, with quadratically weighted kappa of 0.869 for ISUP grade compared with 0.788 for UNI.

The prespecified marked-error review is more informative than a single average. Across 337 slides read by the first two pathologists, the task-specific, UNI and Virchow2 models made five, 16 and seven marked errors respectively under a definition that included calling a benign slide ISUP grade 2 or higher, or missing a grade-2-or-higher cancer as benign. When a third pathologist re-read slides on which at least two models made marked errors, many apparent discrepancies changed under that reference. All nine difficult slides would have required immunohistochemistry for a definitive diagnosis, showing how an uncertain human reference constrains claims about machine accuracy.[1]

A compact scanner may lower the cost of local validation

Predictions were consistent when the same slides were scanned on three devices. Average cross-scanner quadratically weighted kappa for ISUP grade was 0.956 for the task-specific model, 0.905 for UNI and 0.929 for Virchow2. Most remaining differences were one-step grade shifts. This result is useful for hospitals that cannot justify a full digital-pathology estate before seeing whether AI performs on their own tissue preparation and staining workflow.

A compact scanner could support a bounded validation project without implying that the rest of the operational stack is solved. Slides in this study were still scanned at one Swedish site, so differences in local staff, temperature, maintenance and laboratory procedures were not tested. The paper also reports prior energy measurements of 0.63 watt-hours per biopsy for the task-specific model, 6.74 for UNI and 22.09 for Virchow2. Those comparisons reinforce that a larger foundation model is not automatically the most practical or accurate option for a resource-constrained laboratory.[1]

What this changes for pathologists and patients

For pathologists, the study supports a staged approach: digitise a representative local set, compare model outputs with blinded readers, inspect clinically serious disagreements and measure the added review burden before any live use. A system with high sensitivity but lower specificity may create queues of false alarms. A smaller model may be easier to run on premises, which can reduce cloud cost and support local data control, but it still needs version control, calibration checks, downtime procedures and a named clinician accountable for the final report.

For patients, the potential benefit is more consistent access to specialist review where pathology capacity is limited. The study does not show faster diagnosis, fewer repeat biopsies, better treatment selection or improved survival. It also cannot separate biological population differences from slide preparation and laboratory practice. Presenting the result as proof of clinical benefit would skip the outcomes that matter most to patients and could encourage use before errors, escalation routes and responsibility are understood.[1]

Limits, disclosures and evidence that would change the assessment

The cohort came from one private hospital and contained relatively few low-grade ISUP group 1 slides—seven of 339. Sample size reflected the eligible archive rather than a formal power calculation. No molecular characterisation was available, no immunohistochemistry resolved the nine most difficult cases, and no patient outcome served as an objective reference. The comparison with pathologist agreement was not designed to demonstrate statistical equivalence or non-inferiority. The multi-scanner test reused the same physical slides and one scanning site rather than independent routine laboratories.

Funding included Swedish research foundations, cancer funding and national computing infrastructure; the paper reports no product deployment trial. The openly shared image set and outputs make replication possible. Confidence would rise with prospective multi-hospital validation in Iraq and neighbouring countries, locked thresholds, commercial comparator systems, workload and turnaround-time measures, and follow-up against treatment or outcome data. A controlled clinical study showing fewer consequential errors without extra review burden would change this from promising technical generalisation into evidence for patient benefit.[1][2][3]

What this means for people

  • Pathologists may gain a practical route to test AI locally before committing to full laboratory digitisation.
  • False-positive calls from the foundation-model pipelines could increase review work in already constrained services.
  • Patients could benefit only if technical agreement translates into faster, safer and more consistent diagnosis under human responsibility.

Global context

The study addresses a persistent geographic gap: pathology AI has been developed mainly on North American and European material, while many hospitals in the Middle East remain under-digitised. It provides unusually direct evidence from Iraq and releases the image set for reuse. Generalisation still needs testing across laboratories, health systems and populations; a successful result in one Erbil hospital does not establish regional or global performance.

What the evidence does not yet show

  • This was a retrospective single-hospital cohort of 185 patients and 339 slides, not a prospective clinical deployment.
  • The pathologist comparison used 59 slides and was not designed to establish equivalence or non-inferiority.
  • Nine highly discrepant slides lacked the immunohistochemistry needed for an unequivocal reference diagnosis.
  • Only seven slides were ISUP grade group 1, limiting assessment of low-grade cancer versus benign mimics.
  • All three scanner datasets came from the same physical slides scanned at one site, so laboratory-to-laboratory variation remains untested.
  • The systems are research algorithms without regulatory approval for clinical diagnosis.

What to watch next

  • Prospective validation across Iraqi and neighbouring laboratories using locked models and locally scanned routine slides.
  • Comparisons with regulatory-authorised commercial systems and consensus or outcome-based reference standards.
  • Turnaround time, false-positive review burden, missed cancers, calibration and clinician override patterns in live workflow.
  • Patient outcomes, treatment decisions and equity across grade, age, laboratory and population subgroups.

Living evidence record

Impact record IAI-1B3XAOW

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent or research support

Present

Record status

Monitoring

Last checked

11 October 2026

Source trail

3 direct sources across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Related-source reporting disclosure

This record analyses 3 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.

Evidence trail

Sources used for this report

Links checked 11 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Health & Life Sciences

Can a burn-risk model help before the outcome is known?

A peer-reviewed Iranian study reports very high internal discrimination in 2,266 burn admissions. But the model uses hospital stay, ICU stay and infection information accumulated during care, overpredicted mortality in parts of the test set and has not been validated outside one hospital, so it is not an admission score or a deployable clinical tool.

10 min · 1 source

Health & Life Sciences

Can this MRI model draw brain-tumour boundaries reliably?

A peer-reviewed multimodal segmentation model was developed on 2,422 public MRI volumes and externally tested on 125 cases. Accuracy remained useful but fell outside the development data; no prospective clinical workflow, radiologist comparison or patient-outcome test was performed.

8 min · 2 sources

Health & Life Sciences

Can AI prioritise cognitive assessment in sleep apnoea?

A random-forest model separated concurrent mild cognitive impairment in an external Chinese cohort, but overestimated probabilities and was evaluated almost entirely in men. It is not a diagnostic or prognostic tool.

6 min · 1 source

Reader commentary

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Explore commentary across the portal →

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.