Can a medical AI support a clinical board without becoming the decision-maker?
A peer-reviewed German study put OpenEvidence into discussions of 100 real inflammatory-disease cases. Specialists found the answers useful, but complete agreement with the board was limited, prompt tuning was inconclusive and outputs changed over time. The study tested workflow feasibility—not patient benefit or safety.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
Research topic
Prospective evaluation of a medical large language model in a multidisciplinary inflammatory-disease clinical board

At a glance
- 1Twenty-one specialists assessed OpenEvidence during discussion of 100 consecutive real cases, generating 759 clinician-case ratings; the design tested whether the tool could fit into a board workflow, not whether it improved patient outcomes.
- 2Two physician raters judged the model fully aligned with the final board recommendation in 40.2% and 50.5% of 97 evaluable cases; around one fifth to one quarter were rated as having no agreement.
- 3A structured prompt produced a numerical improvement that did not reach the study's significance threshold, while reruns months later changed answer length and most cited references—evidence that expert verification cannot be treated as optional.
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Living evidence record
Impact record IAI-1L57FOY
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent support
Present
Record status
Monitoring
Last checked
3 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what Scientific Reports published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
The study tested a real workflow, not an exam benchmark
Researchers at University Hospital Erlangen and Friedrich-Alexander-Universität Erlangen-Nürnberg prospectively observed 100 consecutive cases discussed by an interdisciplinary board between July and November 2025. The cases involved chronic inflammatory disease and commonly crossed specialty boundaries. Dermatology appeared in 82% of cases, gastroenterology in 52% and rheumatology in 36%, with smaller contributions from haematology-oncology, neurology, paediatrics and nephrology. On average, 2.3 specialties were directly involved in a case.
That makes the paper more informative than a question-answer benchmark built from examination items. These were real clinical problems seen by a functioning board, and 21 specialists took part. Yet it remains a prospective, observational feasibility study at one German centre. It did not randomise cases to AI-supported and unsupported care, compare decisions against later diagnoses, or measure treatment response, complications, time to care, costs or patient experience. The strongest justified conclusion is therefore about whether an AI evidence tool could be inserted into this particular meeting process.[1]
Human discussion came first, but the final reference was not independent of the AI
For each case, clinicians first discussed the question without the model and entered a preliminary recommendation in the hospital documentation system. A prepared, de-identified case prompt was then submitted to the commercial OpenEvidence platform. The board saw the answer, could add or amend points, and only then confirmed its final recommendation. No personal patient data was entered into the service, according to the paper.
This sequence protects an important human-first step, but it complicates interpretation of agreement. The final board recommendation used as the clinical reference could itself incorporate something raised by the model. Agreement with that final decision is therefore not the same as an independent test against an untouched expert answer, and still less a test against patient truth. The authors' design is useful for studying integration; it cannot establish that the model matched an independent gold standard or caused a better decision.[1]
Specialists liked the answers more than the agreement scores imply
The board members completed 759 clinician-case evaluations using eight five-point rating items and a free-text field. The highest mean score was 4.06 for providing clinically relevant information. Recommendation for future use averaged 3.95. Time saved was the lowest-rated item at 3.50 and also varied most. All eight medians were four, suggesting that specialists generally regarded the tool as useful and workable during discussion.
Perceptions were not uniform. Nephrologists gave the most conservative overall ratings, while some other specialties were more positive; the paper warns that only two to six raters represented each specialty and case exposure was uneven, so those differences should not be treated as comparative evidence. Only seven of the 759 evaluations included a free-text comment. Positive scores show acceptability among participating clinicians, but repeated ratings from the same small group are not equivalent to 759 independent users or 759 independent cases.[1]
Full agreement was a minority-to-bare-majority result
Two physician authors independently reviewed outputs for 97 evaluable cases and classified agreement with the final board recommendation as none, partial or complete. One rater judged complete agreement in 40.2% of cases and the other in 50.5%. They judged no agreement in 21.6% and 24.7%, respectively. The raters assigned the same category in 62 of 97 cases; within that consensus subset, 51.6% were complete, 22.6% partial and 25.8% no agreement.
Those figures resist both optimistic and dismissive readings. Partial or complete agreement in a majority suggests the system often surfaced material compatible with expert thinking. But a tool that is not fully aligned in roughly half the cases, and is judged not aligned in around a quarter, cannot safely replace the specialists capable of spotting what is missing. Inter-rater agreement was only moderate by the paper's reported statistics, which also shows that even physicians on the same board did not always agree about how closely an answer matched the recommendation.[1]
Prompt optimisation was not a proven fix
The researchers later reran all cases with a structured prompt asking the model to assess whether information was sufficient, request missing details and then give a justified recommendation. The model asked follow-up questions in every case, but the researchers supplied no additional information before it generated the recommendation. Complete agreement increased numerically to 50.5% and 51.5% for the two raters; in the consensus subset it reached 56.5%.
Across the paired comparison, 29 cases improved, 48 were unchanged and 20 worsened. The mean score rose from 1.222 to 1.361 on the study's zero-to-two scale, but the reported p-value was 0.0642 and the 95% confidence interval for the effect crossed zero. That does not prove the structured prompt had no value; it means this study did not establish a reliable improvement at its stated threshold. Prompt templates may help organise questions, but they are not a substitute for missing clinical information, external validation or outcome testing.[1]
The same cases did not produce stable evidence summaries
When the researchers resubmitted the cases months later, the outputs changed substantially. For the original prompts, mean response length rose by 91.1%, the number of cited references by 93.0%, and only 20.4% of the originally cited references were retained. For the optimised prompts, responses grew by 33.5%, citation counts by 22.1%, and 23.6% of references were retained. The platform is proprietary and continuously updated; the provider did not disclose a model or product version for the first two rounds, so the study cannot separate model updates from other sources of variation.
A same-day comparison was more reassuring but still not identical: changing the prompt left the core clinical recommendation unchanged in 87 of 100 cases. The practical lesson is not that later answers were necessarily worse. Longer responses and newer sources could reflect improvement. It is that a clinician cannot assume yesterday's answer, evidence set or rationale will be reproducible today. Governance needs dated outputs, recorded prompts, accessible citations, version information where available and a clear responsibility for re-checking the underlying evidence.[1]
What this means for patients and clinical teams
For clinicians, an evidence-synthesis model may be most useful as a second-pass literature assistant after experts have framed the case and identified missing facts. It could surface therapeutic options or papers during a meeting, especially when knowledge spans several disciplines. The workflow should preserve the preliminary human recommendation, require a qualified reviewer to open the cited evidence, and document why any model suggestion changed—or did not change—the final plan.
For patients, the paper offers no evidence of better outcomes, fewer errors or safer care. It did not measure hallucinations or clinical harm. People should not infer that use of the model improved their diagnosis or treatment simply because clinicians rated some answers as relevant. Deployment would also require local rules for de-identification, data processing, audit trails, conflicts, consent or notice where applicable, and a route to challenge a decision. The study's ethics approval and avoidance of personal data are useful design choices, not a complete governance framework for routine use.[1]
What evidence would change the assessment
Confidence would increase with preregistered, multicentre trials that keep the comparator independent: teams could be randomised to standard review or model-assisted review, with predefined measures of decision quality, time, resource use, errors, equity and patient outcomes. Cases should span institutions, languages, specialties and levels of complexity. Independent adjudicators should be blinded to the system used, and subgroup results should show where the tool helps, where it adds noise and whether benefits persist after clinicians become familiar with it.
Future evaluations also need fixed model identifiers or archived outputs, reproducibility checks, verified citation accuracy, explicit hallucination and harm review, and disclosure of commercial relationships. This study reported German Research Foundation funding for two authors, open-access support through Projekt DEAL and no competing interests. Its result is valuable precisely because it is bounded: OpenEvidence could be fitted into one expert board and was often perceived as useful, but the evidence does not support autonomous use, a safety claim or a promise of improved care.[1]
What this means for people
- Patients should not treat a model-supported discussion as evidence that care is safer or more accurate; this study did not measure outcomes or harm.
- Clinicians may gain faster access to literature across specialties, but the burden of checking sources and judging missing context remains with qualified professionals.
- Health systems need audit trails, privacy controls and routes for review so an unstable proprietary output does not become an unchallengeable part of care.
Global context
Health systems worldwide are testing language models as evidence assistants while regulators and hospitals still lack consistent ways to evaluate changing proprietary systems. This German single-centre study adds rare real-workflow evidence, but its design and instability results strengthen the case for local validation, recorded versions and human accountability before wider clinical reliance.
What the evidence does not yet show
- Single-centre observational study of 100 inflammatory-disease cases and 21 specialists; there was no randomised control group or external validation.
- The final board recommendation was confirmed after clinicians saw the model output, so it was not a fully independent reference standard.
- Two physician authors rated concordance; one had helped prepare 10 prompts, and blinding to the evaluation round was not possible.
- The proprietary platform did not disclose a model or product version for the first two rounds, limiting reproducibility and attribution of changes over time.
- Patient outcomes, clinical harm, hallucinations, citation accuracy and economic effects were not assessed.
What to watch next
- Independent multicentre trials comparing model-assisted boards with standard care on predefined clinical and operational outcomes.
- Prospective auditing of citation accuracy, omitted evidence, hallucinations, unsafe recommendations and subgroup performance.
- Versioned, reproducible outputs and governance that records prompts, sources, reviewer decisions and later model changes.
- Evidence that any time saving or broader evidence access translates into safer, fairer or more effective care.
Evidence trail
Sources used for this report
Links checked 3 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Health & Life Sciences
Can AI make tumour mitosis counts more consistent?
A preprint study paired 13 pathologists' unaided and AI-assisted reviews of 385 tumour slides from three European centres. Agreement rose and counting time fell, but the unreviewed study did not establish which counts were correct, whether diagnoses improved or whether patients benefited.
11 min · 1 source
Health & Life Sciences
Did clinicians prefer AI discharge summaries after long hospital stays?
In a retrospective 60-case comparison, 12 physicians usually preferred GPT-5.2 summaries and annotated fewer omissions. Reviewers knew which summary was AI-written, one hospital supplied the records, and no patient outcome or time saving was tested.
7 min · 3 sources
Health & Life Sciences
What should Europe require before hospital AI becomes routine?
A same-day report from Europe’s science and medical academies argues that health AI needs evidence suited to changing systems, interoperable data, accountable regulation and a workforce able to challenge the technology—not simply more pilots.
6 min · 2 sources
Reader discussion
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.