Can chatbots match urologists on cancer decisions?
In 100 synthetic cancer cases, three leading chatbots were less often fully successful than two early-career urologists and made safety-critical errors far more often. Repeated answers also changed substantially.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1Two source-masked uro-oncology experts scored 500 initial responses to 100 synthetic cases across five cancers and five decision categories using a prespecified 2025 European Association of Urology-anchored rubric.
- 2Composite success was 60% to 70% for the three chatbots versus 85% and 89% for the two urologists; safety-critical errors occurred in 25% to 26% of chatbot answers versus 1% and 2% of clinician answers.
- 3Across 60 vignette-platform combinations repeated three times, only 45% kept the same success classification and 48.3% kept the same safety-critical status.
Research topic
Whether three consumer large language models match two early-career urologists on guideline-based urological oncology decisions and repeat the same answer across runs
The answer: not in this controlled test
Three leading chatbots did not match two early-career urologists on the full combination of guideline concordance, safety, completeness and correct risk or stage assessment in a peer-reviewed benchmark of 100 synthetic cancer cases. GPT-5.2 Instant met the study's composite success standard in 70 cases, Claude Sonnet 4.5 in 64 and Gemini 3 Pro Preview in 60. The two clinicians succeeded in 85 and 89 cases. The overall difference was statistically significant, although the 15-percentage-point gap between the stronger chatbot and the lower-scoring urologist did not remain significant after correction for multiple comparisons.
The safety result is more concerning than the ranking. Reviewers identified a safety-critical error in 25%, 25% and 26% of the three chatbots' initial answers, compared with 2% and 1% for the clinicians. All six clinician-versus-chatbot safety comparisons were significant after adjustment. These were written answers to invented vignettes, not treatment delivered to patients, but the gap argues against using an unverified consumer chatbot as an independent cancer decision-maker.[1]
The Impact Brief · Free
Follow the evidence in health & life sciences.
Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.
What the researchers actually measured
The Türkiye-based team constructed 100 synthetic vignettes spanning five urological cancers and five types of decision. Each source received the same standardised prompt and could not use external resources. Two uro-oncology experts, blinded to whether an answer came from a model or a clinician, independently scored all 500 initial responses. The rubric was prespecified and anchored to 2025 European Association of Urology guidance.
Success required a score of at least two in each of four domains: guideline concordance, clinical safety, decision completeness and risk or stage accuracy. A disqualifying clinical error also made an answer unsuccessful. This is a demanding composite, which is appropriate for consequential decisions because an otherwise polished recommendation can still be unsafe if one crucial component is wrong. The source-masked design reduces the risk that reviewers rewarded or penalised an answer because they recognised its author.[1]
Repeatability was weak even when the interface looked deterministic
The researchers also ran repeated queries for selected cases rather than assuming that one response represented a model. Across 60 vignette-and-platform combinations assessed over three runs, only 45% kept exactly the same primary success classification, 48.3% kept the same safety-critical status and 28.3% retained the complete four-domain score profile. Stable output was not necessarily correct: some combinations repeatedly succeeded, while others repeatedly failed or reproduced an error.
That distinction matters for clinicians and procurement teams. A system can give an acceptable answer during evaluation and a materially different answer to the same case later. Conversely, repeated wording can conceal a reliably unsafe decision. Quality assurance therefore needs repeated testing, version records and monitoring of the clinical classification, not merely a demonstration that the chatbot can produce fluent prose once.[1]
What the comparison can and cannot tell clinicians
The human comparison is useful but narrow. The study included two early-career urologists, both members of the research team, rather than a large, independent sample across training levels and health systems. The paper therefore does not establish an average human benchmark or show how senior multidisciplinary teams would perform. It also does not measure whether a clinician using an AI assistant would outperform either the clinician or model alone.
For a practising urologist, the most defensible use is as evidence that generic model competence should not be inferred from confident language or isolated cases. If a hospital considers decision support, it should test the exact model version, prompts and workflow against its own protocols; define when staff must disregard or escalate advice; and measure harmful omissions as well as headline accuracy. The study supports evaluation, not replacement of specialist judgement.[1]
The study does not establish patient benefit or harm
No patient records were used, no clinician acted on a model answer and no treatment, delay, complication or survival outcome was measured. Synthetic cases allow controlled coverage of rare or risky decisions and avoid exposing patients, but they remove incomplete notes, conflicting priorities, comorbidities, local capacity and the conversation needed for shared decision-making. A chatbot could perform differently when information arrives across a real record rather than a compact vignette.
The tested systems were consumer model snapshots. Their providers can change behaviour, safety layers and available knowledge, so these percentages should not be projected to another version without new testing. The full narrative responses were not publicly released, although the authors made structured data and ratings available. That constrains outside scrutiny of how the models reasoned, phrased uncertainty or produced the errors.[1]
What would change the assessment
Confidence would rise with preregistered replications using larger, independent clinician samples, multiple languages and health systems, and cases drawn prospectively from real clinical work under appropriate privacy and ethics controls. Studies should freeze model versions, repeat queries, publish de-identified outputs where possible and report results by cancer, decision type and severity of harm. External reviewers should test whether the rubric captures the decisions that matter locally.
The more important next question is whether carefully designed assistance improves a clinician's decision without increasing automation bias, workload or delay. That requires randomised or prospective workflow studies with patient-relevant outcomes and clear human override. The article reports no funding and no competing interests. It also states that OpenAI Codex was used only for manuscript editing, organisation and computational verification, not to generate the evaluated answers or assign outcomes. For now, the evidence supports caution: a chatbot may help structure information, but it has not earned autonomous authority over cancer care.[1]
What this means for people
- Patients should not treat a consumer chatbot answer as an independent cancer diagnosis or treatment plan.
- Clinicians need repeat testing and explicit escalation rules because the same model can change its safety classification across runs.
- Health services can use the benchmark as a starting point for local assurance, not as evidence that a tool is ready for autonomous use.
Global context
Urological cancer guidelines, treatment access and specialist staffing differ across health systems. A controlled Türkiye-based benchmark provides a useful warning about current consumer models, but local validation is essential before adoption elsewhere. Regions with fewer specialists may feel the strongest pressure to rely on chatbots and may also have the least capacity to detect a plausible but unsafe recommendation.
What the evidence does not yet show
- The benchmark used synthetic vignettes rather than patients, clinical records or real treatment decisions.
- Only two early-career urologists were included, both from the study team, so the human comparator is not representative of all clinicians.
- Consumer model snapshots can change and the results should not be transferred to later versions without retesting.
- Full narrative responses were not released, limiting independent review beyond the structured ratings.
- The study did not evaluate clinician-plus-AI performance, patient outcomes, workflow burden or automation bias.
What to watch next
- Independent multi-centre replication
- Clinician-plus-AI trials
- Frozen-version safety monitoring
- Patient-relevant outcomes
Living evidence record
Impact record IAI-1PZ419K
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
11 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what BMC Urology published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
Evidence trail
Sources used for this report
Links checked 11 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Health & Life Sciences
Can consistent AI replace sarcoma tumour-board judgement?
No. A structured Claude workflow reproduced 94% of 1,071 decision-code instances across repeated runs, but the study did not test whether those recommendations were clinically correct or improved care.
7 min · 1 source
Health & Life Sciences
Did clinicians prefer AI discharge summaries after long hospital stays?
In a retrospective 60-case comparison, 12 physicians usually preferred GPT-5.2 summaries and annotated fewer omissions. Reviewers knew which summary was AI-written, one hospital supplied the records, and no patient outcome or time saving was tested.
7 min · 3 sources
Health & Life Sciences
Can a medical AI support a clinical board without becoming the decision-maker?
A peer-reviewed German study put OpenEvidence into discussions of 100 real inflammatory-disease cases. Specialists found the answers useful, but complete agreement with the board was limited, prompt tuning was inconclusive and outputs changed over time. The study tested workflow feasibility—not patient benefit or safety.
9 min · 1 source
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.