Back to the news portal
Health & Life SciencesResearch paperResearchSource analysisGermanyEuropeGlobal oncology

Can consistent AI replace sarcoma tumour-board judgement?

No. A structured Claude workflow reproduced 94% of 1,071 decision-code instances across repeated runs, but the study did not test whether those recommendations were clinically correct or improved care.

By The Impact of AI Editorial DeskReleased 11 October 2026 at 08:03 BST7 min read1 source

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

Share
Social links
LinkedInXBlueskyRedditEmail

At a glance

  • 1The pilot processed 51 de-identified sarcoma cases in 154 runs and compared 21 decision codes per case, producing 1,071 case-code instances.
  • 2Overall exact reproducibility was 94.0%; only 17 of 51 cases were identical across all 21 codes, and reconstruction, staging text and radiotherapy fields varied most.
  • 3The study did not compare recommendations with an expert tumour board or patient outcomes, so reproducibility cannot be read as clinical correctness.
Key themesSarcomaTumour boardsLarge language modelsClinical decision supportReproducibilityStructured output

Research topic

Whether a schema-enforced large-language-model workflow returns stable multidisciplinary sarcoma decision codes when the same de-identified case is processed repeatedly

The answer: structured output made the model steadier, but steadiness is not medical validity

A Claude Opus 4.5 workflow produced the same value across repeated runs for 1,007 of 1,071 case-level decision codes, an overall reproducibility estimate of 94.0% with a cluster-bootstrap 95% confidence interval of 92.0% to 95.8%. That is evidence that a constrained schema, explicit semantics and conditional rules can reduce run-to-run variation in a clinical simulation. It does not establish that the recommendations were appropriate.

A system can consistently return the same wrong answer. The researchers deliberately framed this as a methodological pilot and made concordance with historical multidisciplinary-team decisions follow-up work. No clinician panel scored the recommendations for correctness, no patient entered an AI-guided pathway, and no outcome improved. The immediate result is therefore about building an auditable experimental instrument for later comparison—not about replacing surgeons, oncologists, radiologists, pathologists or specialist nurses.[1]

The Impact Brief · Free

Follow the evidence in health & life sciences.

Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.

Choose your topics (optional)

One concise, source-linked briefing. Unsubscribe at any time.

How the 51-case experiment was organised

The dataset contained 51 fully de-identified sarcoma cases spanning 17 histological subtypes from 2020 to 2025. Ten cases were used during three framework-development phases. The final schema and prompt were then frozen before the remaining 41 cases were processed, although development was not prospectively logged and the authors appropriately avoid calling those 41 a formal held-out test set. All cases were later run under the same final version.

Fifty cases had three completed runs and one case had four, giving 154 analysed outputs. Claude Opus 4.5 was accessed through the official SDK at temperature zero with a maximum of 4,000 output tokens. Tool-use output was validated against a Pydantic schema; all saved runs returned the required block and passed validation. The prompt encoded meanings and conditional relationships across surgical, reconstructive, radiotherapy, imaging, staging and systemic-treatment fields.[1]

What 94% reproducibility means—and what it hides

The denominator was 51 cases multiplied by 21 decision codes. A code counted as stable only when every available run for that case returned exactly the same value. There was no case folding, whitespace harmonisation or semantic matching, so equivalent free-text wording could count as unstable. Reproducibility was 94.2% among the 41 cases not used to specify the framework and 93.3% among the ten development cases. Seventeen cases, one third of the total, were stable on every code.

Variation was not evenly distributed. Reconstruction decisions were stable in 72.5% of case-code comparisons, free-text staging recommendations in 76.5%, and radiotherapy dose specifications in 86.3%. The lowest-stability case reached 61.9% and involved an older patient for whom outputs alternated among observation, wide excision and different reconstruction options. These are not trivial formatting differences; they sit close to choices that would require specialist deliberation, patient preferences and knowledge of local capability.[1]

The ablation shows where rules helped

A post-hoc ablation reran 15 cases with and without the semantic and conditional-logic rules. Reproducibility was 94.6% with the full framework and 82.5% without it, a 12.1-percentage-point difference with a 95% confidence interval of 7.6 to 17.1 points. Most of the improvement came from exact wording in three free-text fields, which rose by 60 points. Across the 18 categorical and numeric fields, the difference was 4.1 points and its confidence interval crossed zero.

That nuance prevents an exaggerated conclusion. Schema enforcement clearly standardised the container and language, while the evidence for better stability in coded clinical choices was smaller. Structured output still has practical value: it makes missing fields, impossible combinations and version changes easier to audit. But schema conformity is a software property, not a guarantee that a treatment plan matches guidelines, multidisciplinary judgement or an individual patient's goals.[1]

What this means for clinicians and patients

For researchers, the framework offers a way to measure a model's behaviour repeatedly before asking experts to judge it. Future evaluations can lock the model and prompt, compare outputs with documented tumour-board decisions, inspect disagreement by speciality and test multiple providers. Hospitals considering decision-support research should also log prompt versions, model identifiers, API failures, schema validation, manual overrides and the provenance of every case field.

For clinicians and patients, nothing in this pilot supports using the output as an independent recommendation. Sarcoma care is uncommon, multidisciplinary and preference-sensitive; reconstruction and radiotherapy—the areas with greatest remaining variation—depend on anatomy, function, timing, prior treatment and local expertise. A safe study would present outputs to a specialist team as research material, preserve ordinary care, and measure whether the system adds information or merely creates extra review work.[1]

Limits, disclosures and evidence that would change the assessment

The study used one proprietary model, one institution's de-identified cases and one final prompt. Temperature zero reduces sampling variation but does not guarantee determinism. Failed schema-validation attempts were not logged because they produced no saved file, so the analysed denominator describes successful outputs. Hallucination screening was performed by a single author, and absence of detected fabricated content under five categories does not prove that every coded recommendation was factually or clinically correct.

The model was not compared with a multidisciplinary panel, guidelines, alternative models or patient outcomes. Source cases remain at the institution, limiting independent case-level replication, although aggregated results and framework materials are reported as deposited. Confidence would rise with prospectively logged, multi-centre, multi-model evaluation; blinded expert adjudication; predefined safety-critical disagreement rules; and workflow studies. Evidence of improved plan quality without added delay or automation bias would be required before considering clinical benefit.[1]

What this means for people

  • Specialists may gain a more auditable research interface for testing AI recommendations before any clinical trial.
  • Patients should not interpret repeatable output as a second opinion or treatment recommendation.
  • Variable reconstruction and radiotherapy fields reinforce the need for multidisciplinary and patient-specific judgement.

Global context

Structured-output research can make clinical AI easier to audit across health systems, but the required tumour-board roles, guidelines and treatment options vary internationally. A schema designed in one German institution may not represent resources or pathways elsewhere. Cross-country validation must therefore examine both model behaviour and whether the coded choices match local clinical responsibility, access and patient preferences.

What the evidence does not yet show

  • The primary endpoint was exact repeated-output stability, not clinical correctness or guideline concordance.
  • Only Claude Opus 4.5 and one final prompt/schema were evaluated on 51 cases from one institution.
  • Ten cases informed framework development, and the remaining 41 were not a prospectively declared held-out set.
  • A single reviewer screened outputs for hallucination categories; multi-rater clinical adjudication was not performed.
  • Unsuccessful schema-validation attempts were not logged, so reliability of all attempted invocations cannot be calculated.
  • No patient outcome, workflow burden, treatment decision or specialist-panel comparison was measured.

What to watch next

  • Blinded concordance against documented sarcoma-board decisions and current clinical guidance.
  • Independent multi-centre tests across models, languages, institutions and less structured case records.
  • Specialist adjudication of reconstruction, staging and radiotherapy disagreements.
  • Prospective workflow evidence on review burden, automation bias, delays, overrides and patient outcomes.

Living evidence record

Impact record IAI-1KI0FNY

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent or research support

Present

Record status

Monitoring

Last checked

11 October 2026

Source trail

1 direct source across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Single-source reporting disclosure

This record analyses one direct source. It can establish what Scientific Reports published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.

Evidence trail

Sources used for this report

Links checked 11 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Health & Life Sciences

Can a medical AI support a clinical board without becoming the decision-maker?

A peer-reviewed German study put OpenEvidence into discussions of 100 real inflammatory-disease cases. Specialists found the answers useful, but complete agreement with the board was limited, prompt tuning was inconclusive and outputs changed over time. The study tested workflow feasibility—not patient benefit or safety.

9 min · 1 source

Health & Life Sciences

Can AI make tumour mitosis counts more consistent?

A preprint study paired 13 pathologists' unaided and AI-assisted reviews of 385 tumour slides from three European centres. Agreement rose and counting time fell, but the unreviewed study did not establish which counts were correct, whether diagnoses improved or whether patients benefited.

11 min · 1 source

Health & Life Sciences

Does primary-care AI improve outcomes?

Not yet on the available evidence. A peer-reviewed review found 10 real-world studies of clinician-facing AI in primary care: some improved detection or care processes, but neither trial measuring patient-important outcomes demonstrated benefit, and the evidence was low or very low certainty.

10 min · 2 sources

Reader commentary

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Explore commentary across the portal →

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.