Back to the news portal
AI Risks & SafetyResearch paperResearchSource analysisUnited States guidelinesChinaAustriaAustraliaSingaporeGlobal clinical AI

Do AI models favour one cancer guideline?

Four models produced 6,000 choices across 15 conflicts between US immunotherapy guidelines and usually selected NCCN. But 10 conflicts involved one lung-cancer setting, and removing them weakened or reversed the pattern for most models.

By The Impact of AI Editorial DeskReleased 8 October 2026 at 14:33 BST7 min read2 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

Share
Social links
LinkedInXBlueskyRedditEmail

At a glance

  • 1Researchers extracted 15 disagreements from ASCO and NCCN immune-checkpoint-inhibitor guidelines dated October 2023 to October 2024.
  • 2GPT-4o, Claude 3.7 Sonnet, Gemini 1.5 Pro and DeepSeek-V3 each answered every case 100 times: 6,000 model outputs in total.
  • 3NCCN was selected in 71.3% to 100% of outputs, depending on the model, and an order-reversal control did not remove the direction of preference.
Key themesClinical AILarge language modelsCancer guidelinesImmunotherapyModel preferenceDecision support

Research topic

Repeated large-language-model choices between discordant ASCO and NCCN immune-checkpoint-inhibitor recommendations

The Impact of AI research cover showing four generic AI systems comparing two conceptual clinical-guideline documents under a caution marker and expert-review panel, with the 15-case denominator visible.
AI-generated editorial illustration. The AI systems, guideline documents and review panel are conceptual; no real provider logo, clinician, patient record or treatment decision is depicted.

The direct answer: the models showed a preference, but the small case mix shaped it

All four tested language models more often chose National Comprehensive Cancer Network recommendations when asked to select between NCCN and American Society of Clinical Oncology guidance that differed on immune-checkpoint-inhibitor treatment. Gemini 1.5 Pro selected NCCN in every run; Claude 3.7 Sonnet, DeepSeek-V3 and GPT-4o selected it in roughly seven out of 10. Reversing the order of the two guideline options did not remove the direction of the result, making a simple first-choice position effect less likely.

The experiment is not evidence that NCCN is better, that the models made the clinically correct decision or that one can safely resolve a guideline conflict. It measured preference across only 15 disagreements. Ten concerned non-small-cell lung cancer, and the sensitivity analysis changed the picture substantially when those cases were removed. The practical finding is that a general-purpose model may add a repeatable, model-specific tilt to an already uncertain clinical question unless its sources and reasoning are made explicit.[1][2]

The study generated 6,000 choices from 15 guideline disagreements

The researchers searched ASCO and NCCN immune-checkpoint-inhibitor guidelines published or updated between October 2023 and October 2024. Two investigators independently extracted discordant recommendations for comparable tumour types, patient characteristics and treatment contexts, then resolved disagreements by consensus or senior adjudication. The 15 final cases were classified by cancer type, line of therapy, clinical stage, monotherapy or combination immunotherapy, concurrent chemotherapy and checkpoint-inhibitor class.

Four model versions were accessed through their official web interfaces in April 2025: GPT-4o, Claude 3.7 Sonnet, Gemini 1.5 Pro and DeepSeek-V3. Researchers manually entered a standard prompt that included each recommendation, evidence level, recommendation grade and treatment context. Each model answered each of the 15 cases 100 times in separate conversations without prior context. That produces 1,500 choices per model and 6,000 overall, although repeated generations do not create 6,000 independent clinical scenarios.[1]

The headline preference ranged from 71.3% to 100%

Gemini selected NCCN in 100% of runs. Claude selected it in 73.7%, DeepSeek in 72.0% and GPT-4o in 71.3%. The paper reports uncertainty intervals clustered around the 15 questions, rather than treating all repeated outputs as if they were unrelated cases. Claude's reported 95% interval was 62.0% to 85.3%, DeepSeek's 61.3% to 84.5%, and GPT-4o's 53.3% to 88.4%. Those broad intervals are an important reminder that the clinical-case denominator is 15, not 6,000.

A control reversed the presentation order of the options and the models continued to favour NCCN, according to the paper. That addresses one familiar source of model bias, but not others. Differences in wording, evidence summaries, publication prominence or training-corpus availability could influence a model without reflecting a considered clinical judgment. The study did not inspect proprietary training data or establish why the preference arose, so explanations about corpus distribution remain hypotheses.[1]

Removing 10 lung-cancer cases altered three of the four conclusions

Ten of the 15 discordant recommendations concerned non-small-cell lung cancer. When the authors repeated the analysis on the remaining five questions, GPT-4o selected NCCN in 39.4% of trials, reversing the direction of its full-set preference. Claude's NCCN share was 56.6% and was not statistically significant. DeepSeek retained a significant but weaker preference at 70.0%. Gemini remained at 100%. This is the study's most consequential qualification because it shows that one disease area drove much of the pooled pattern.

Five remaining cases are too few for broad claims across oncology. The result also cannot be assumed to extend to European, Asian or national guidelines written for different health systems, drug approvals and resource constraints. A useful replication would pre-register a larger, balanced set of disagreements across tumour types and guideline producers, freeze model versions and prompts, and report each clinical question separately before pooling. Otherwise, repeated sampling can make a narrow case set appear more general than it is.[1]

Three oncology experts scored the answers, but this was not a care study

Three oncology specialists who were not involved in generating the questions independently rated responses for accuracy, comprehensiveness, readability and potential harm on five-point scales, resolving initial rating differences by consensus. Mean accuracy scores ranged from 3.51 for Gemini to 3.89 for DeepSeek, and readability from 3.91 to 4.29. Those descriptive scores suggest the outputs often appeared serviceable to experts, but they do not show that clinicians or patients would make better decisions with them.

The study did not randomise doctors to use or avoid an assistant, compare treatment plans with a multidisciplinary tumour board, or follow any patient outcome. It also did not test whether a model appropriately abstained when guidelines reflected different values or evidence thresholds. Potential harm was a reviewer rating of text, not an observed adverse event. Before clinical use, a system would need current source retrieval, provenance, calibrated uncertainty, escalation rules and prospective workflow testing under governance appropriate to a medical device or decision-support tool.[1]

Model settings and time make this a snapshot, not a permanent ranking

The models were tested through consumer web interfaces rather than a reproducible API, and the authors manually entered prompts. Default generation settings differed: GPT-4o and DeepSeek used temperature 1.0, Claude 0.7 and Gemini 0.2. No system prompt was used and no safety-filter refusal occurred. Those choices reflect the available interfaces but complicate direct comparison, because lower temperature can reduce variation and may help explain Gemini's rigid 100% pattern.

The underlying guidelines ended in October 2024, model access occurred in April 2025 and publication followed in October 2026. Both clinical recommendations and model behaviour can change. The next assessment should repeat the protocol with current and archived versions, equalise controllable settings, publish complete prompts and raw outputs, and include retrieval from the latest authoritative guideline text. For patients and clinicians, the safe lesson is not to select a preferred chatbot; it is to verify the source, date and rationale whenever a model turns legitimate guideline disagreement into one confident answer.[1]

What this means for people

  • A consistent hidden preference could steer clinicians or patients toward one recommendation without explaining why guidelines differ.
  • Transparent source dates and citations could make disagreement visible and prompt specialist review rather than false certainty.
  • Models trained on prominent US materials may transfer poorly to countries with different approvals, resources and treatment pathways.

Global context

The researchers were affiliated with institutions in China, Austria, Australia and Singapore, but the evaluated documents were two US guideline families. Cancer treatment availability, reimbursement, regulatory approval and clinical capacity vary sharply across countries. A responsible global evaluation must test local guidance and make resource assumptions visible rather than treating US recommendations as a universal reference.

What the evidence does not yet show

  • The clinical denominator is 15 guideline disagreements; 10 concern non-small-cell lung cancer.
  • Repeated outputs increase precision about model variability but do not create new clinical cases.
  • The two compared guideline families are US based, limiting geographic transfer.
  • Different web-interface temperatures and inaccessible training data limit causal interpretation and model comparison.
  • No clinician decision, patient outcome, adverse event or prospective deployment was measured.

What to watch next

  • A larger pre-registered and tumour-balanced replication using current guideline versions.
  • Equal model settings, published prompts, raw outputs and versioned interfaces or APIs.
  • Tests against guidelines from additional countries and health-system contexts.
  • Prospective evidence that source-grounded decision support improves clinician reasoning without suppressing legitimate uncertainty.

Living evidence record

Impact record IAI-1NNQK74

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent or research support

Present

Record status

Monitoring

Last checked

8 October 2026

Source trail

2 direct sources across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Related-source reporting disclosure

This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.

Evidence trail

Sources used for this report

Links checked 8 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

AI Risks & Safety

Do AI explanations prevent over-reliance?

Not in this unreviewed experiment. Explanations reduced raw acceptance but did not improve reliance calibration among novices doing a clinical-text annotation task; over-reliance tended to increase across the session.

7 min · 1 source

AI Risks & Safety

Can hidden image text mislead dental AI?

A peer-reviewed German stress test found that adversarial text placed inside 270 dental radiographs could flip four vision-language models from an abnormal to a normal finding. OCR sanitisation sharply reduced the measured attacks, but the experiment used a permissive prompt, a pathology-heavy benchmark and no live clinical system.

7 min · 3 sources

AI Risks & Safety

Can hashes make AI conversations auditable without publishing them?

A peer-reviewed experiment converted nearly four million public chatbot interactions into cryptographic commitments and detected eight induced ledger manipulations. It is a promising integrity mechanism, not proof that a conversation is true, authorised or safe.

9 min · 1 source

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Reader commentary

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Explore commentary across the portal →

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.