Do AI models favour one cancer guideline?
Four models produced 6,000 choices across 15 conflicts between US immunotherapy guidelines and usually selected NCCN. But 10 conflicts involved one lung-cancer setting, and removing them weakened or reversed the pattern for most models.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1Researchers extracted 15 disagreements from ASCO and NCCN immune-checkpoint-inhibitor guidelines dated October 2023 to October 2024.
- 2GPT-4o, Claude 3.7 Sonnet, Gemini 1.5 Pro and DeepSeek-V3 each answered every case 100 times: 6,000 model outputs in total.
- 3NCCN was selected in 71.3% to 100% of outputs, depending on the model, and an order-reversal control did not remove the direction of preference.
Research topic
Repeated large-language-model choices between discordant ASCO and NCCN immune-checkpoint-inhibitor recommendations

The direct answer: the models showed a preference, but the small case mix shaped it
All four tested language models more often chose National Comprehensive Cancer Network recommendations when asked to select between NCCN and American Society of Clinical Oncology guidance that differed on immune-checkpoint-inhibitor treatment. Gemini 1.5 Pro selected NCCN in every run; Claude 3.7 Sonnet, DeepSeek-V3 and GPT-4o selected it in roughly seven out of 10. Reversing the order of the two guideline options did not remove the direction of the result, making a simple first-choice position effect less likely.
The experiment is not evidence that NCCN is better, that the models made the clinically correct decision or that one can safely resolve a guideline conflict. It measured preference across only 15 disagreements. Ten concerned non-small-cell lung cancer, and the sensitivity analysis changed the picture substantially when those cases were removed. The practical finding is that a general-purpose model may add a repeatable, model-specific tilt to an already uncertain clinical question unless its sources and reasoning are made explicit.[1][2]
The study generated 6,000 choices from 15 guideline disagreements
The researchers searched ASCO and NCCN immune-checkpoint-inhibitor guidelines published or updated between October 2023 and October 2024. Two investigators independently extracted discordant recommendations for comparable tumour types, patient characteristics and treatment contexts, then resolved disagreements by consensus or senior adjudication. The 15 final cases were classified by cancer type, line of therapy, clinical stage, monotherapy or combination immunotherapy, concurrent chemotherapy and checkpoint-inhibitor class.
Four model versions were accessed through their official web interfaces in April 2025: GPT-4o, Claude 3.7 Sonnet, Gemini 1.5 Pro and DeepSeek-V3. Researchers manually entered a standard prompt that included each recommendation, evidence level, recommendation grade and treatment context. Each model answered each of the 15 cases 100 times in separate conversations without prior context. That produces 1,500 choices per model and 6,000 overall, although repeated generations do not create 6,000 independent clinical scenarios.[1]
The headline preference ranged from 71.3% to 100%
Gemini selected NCCN in 100% of runs. Claude selected it in 73.7%, DeepSeek in 72.0% and GPT-4o in 71.3%. The paper reports uncertainty intervals clustered around the 15 questions, rather than treating all repeated outputs as if they were unrelated cases. Claude's reported 95% interval was 62.0% to 85.3%, DeepSeek's 61.3% to 84.5%, and GPT-4o's 53.3% to 88.4%. Those broad intervals are an important reminder that the clinical-case denominator is 15, not 6,000.
A control reversed the presentation order of the options and the models continued to favour NCCN, according to the paper. That addresses one familiar source of model bias, but not others. Differences in wording, evidence summaries, publication prominence or training-corpus availability could influence a model without reflecting a considered clinical judgment. The study did not inspect proprietary training data or establish why the preference arose, so explanations about corpus distribution remain hypotheses.[1]
Removing 10 lung-cancer cases altered three of the four conclusions
Ten of the 15 discordant recommendations concerned non-small-cell lung cancer. When the authors repeated the analysis on the remaining five questions, GPT-4o selected NCCN in 39.4% of trials, reversing the direction of its full-set preference. Claude's NCCN share was 56.6% and was not statistically significant. DeepSeek retained a significant but weaker preference at 70.0%. Gemini remained at 100%. This is the study's most consequential qualification because it shows that one disease area drove much of the pooled pattern.
Five remaining cases are too few for broad claims across oncology. The result also cannot be assumed to extend to European, Asian or national guidelines written for different health systems, drug approvals and resource constraints. A useful replication would pre-register a larger, balanced set of disagreements across tumour types and guideline producers, freeze model versions and prompts, and report each clinical question separately before pooling. Otherwise, repeated sampling can make a narrow case set appear more general than it is.[1]
Three oncology experts scored the answers, but this was not a care study
Three oncology specialists who were not involved in generating the questions independently rated responses for accuracy, comprehensiveness, readability and potential harm on five-point scales, resolving initial rating differences by consensus. Mean accuracy scores ranged from 3.51 for Gemini to 3.89 for DeepSeek, and readability from 3.91 to 4.29. Those descriptive scores suggest the outputs often appeared serviceable to experts, but they do not show that clinicians or patients would make better decisions with them.
The study did not randomise doctors to use or avoid an assistant, compare treatment plans with a multidisciplinary tumour board, or follow any patient outcome. It also did not test whether a model appropriately abstained when guidelines reflected different values or evidence thresholds. Potential harm was a reviewer rating of text, not an observed adverse event. Before clinical use, a system would need current source retrieval, provenance, calibrated uncertainty, escalation rules and prospective workflow testing under governance appropriate to a medical device or decision-support tool.[1]
Model settings and time make this a snapshot, not a permanent ranking
The models were tested through consumer web interfaces rather than a reproducible API, and the authors manually entered prompts. Default generation settings differed: GPT-4o and DeepSeek used temperature 1.0, Claude 0.7 and Gemini 0.2. No system prompt was used and no safety-filter refusal occurred. Those choices reflect the available interfaces but complicate direct comparison, because lower temperature can reduce variation and may help explain Gemini's rigid 100% pattern.
The underlying guidelines ended in October 2024, model access occurred in April 2025 and publication followed in October 2026. Both clinical recommendations and model behaviour can change. The next assessment should repeat the protocol with current and archived versions, equalise controllable settings, publish complete prompts and raw outputs, and include retrieval from the latest authoritative guideline text. For patients and clinicians, the safe lesson is not to select a preferred chatbot; it is to verify the source, date and rationale whenever a model turns legitimate guideline disagreement into one confident answer.[1]
What this means for people
- A consistent hidden preference could steer clinicians or patients toward one recommendation without explaining why guidelines differ.
- Transparent source dates and citations could make disagreement visible and prompt specialist review rather than false certainty.
- Models trained on prominent US materials may transfer poorly to countries with different approvals, resources and treatment pathways.
Global context
The researchers were affiliated with institutions in China, Austria, Australia and Singapore, but the evaluated documents were two US guideline families. Cancer treatment availability, reimbursement, regulatory approval and clinical capacity vary sharply across countries. A responsible global evaluation must test local guidance and make resource assumptions visible rather than treating US recommendations as a universal reference.
What the evidence does not yet show
- The clinical denominator is 15 guideline disagreements; 10 concern non-small-cell lung cancer.
- Repeated outputs increase precision about model variability but do not create new clinical cases.
- The two compared guideline families are US based, limiting geographic transfer.
- Different web-interface temperatures and inaccessible training data limit causal interpretation and model comparison.
- No clinician decision, patient outcome, adverse event or prospective deployment was measured.
What to watch next
- A larger pre-registered and tumour-balanced replication using current guideline versions.
- Equal model settings, published prompts, raw outputs and versioned interfaces or APIs.
- Tests against guidelines from additional countries and health-system contexts.
- Prospective evidence that source-grounded decision support improves clinician reasoning without suppressing legitimate uncertainty.
Living evidence record
Impact record IAI-1NNQK74
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
8 October 2026
Source trail
2 direct sources across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
Evidence trail
Sources used for this report
Links checked 8 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
AI Risks & Safety
Do AI explanations prevent over-reliance?
Not in this unreviewed experiment. Explanations reduced raw acceptance but did not improve reliance calibration among novices doing a clinical-text annotation task; over-reliance tended to increase across the session.
7 min · 1 source
AI Risks & Safety
Can hidden image text mislead dental AI?
A peer-reviewed German stress test found that adversarial text placed inside 270 dental radiographs could flip four vision-language models from an abnormal to a normal finding. OCR sanitisation sharply reduced the measured attacks, but the experiment used a permissive prompt, a pathology-heavy benchmark and no live clinical system.
7 min · 3 sources
AI Risks & Safety
Can hashes make AI conversations auditable without publishing them?
A peer-reviewed experiment converted nearly four million public chatbot interactions into cryptographic commitments and detected eight induced ledger manipulations. It is a promising integrity mechanism, not proof that a conversation is true, authorised or safe.
9 min · 1 source
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.