Back to the news portal
Health & Life SciencesResearch paperResearchSource analysisGermanyEurope

Can a medical AI support a clinical board without becoming the decision-maker?

A peer-reviewed German study put OpenEvidence into discussions of 100 real inflammatory-disease cases. Specialists found the answers useful, but complete agreement with the board was limited, prompt tuning was inconclusive and outputs changed over time. The study tested workflow feasibility—not patient benefit or safety.

By The Impact of AI Health DeskReleased 3 October 2026 at 02:05 BST9 min read1 source

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesclinical decision supportmedical AIlarge language modelsmultidisciplinary caremodel driftpatient safety

Research topic

Prospective evaluation of a medical large language model in a multidisciplinary inflammatory-disease clinical board

The Impact of AI research cover asking whether a medical AI can support a clinical board, with a conceptual round table of specialist chairs linked to an evidence document and an AI reasoning layer.
AI-generated editorial illustration. The clinical board, specialist chairs and evidence display are conceptual; they do not depict a study participant, patient record, provider interface, measured result or actual board meeting.

At a glance

  • 1Twenty-one specialists assessed OpenEvidence during discussion of 100 consecutive real cases, generating 759 clinician-case ratings; the design tested whether the tool could fit into a board workflow, not whether it improved patient outcomes.
  • 2Two physician raters judged the model fully aligned with the final board recommendation in 40.2% and 50.5% of 97 evaluable cases; around one fifth to one quarter were rated as having no agreement.
  • 3A structured prompt produced a numerical improvement that did not reach the study's significance threshold, while reruns months later changed answer length and most cited references—evidence that expert verification cannot be treated as optional.

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Living evidence record

Impact record IAI-1L57FOY

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent support

Present

Record status

Monitoring

Last checked

3 October 2026

Source trail

1 direct source across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Single-source reporting disclosure

This record analyses one direct source. It can establish what Scientific Reports published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.

The study tested a real workflow, not an exam benchmark

Researchers at University Hospital Erlangen and Friedrich-Alexander-Universität Erlangen-Nürnberg prospectively observed 100 consecutive cases discussed by an interdisciplinary board between July and November 2025. The cases involved chronic inflammatory disease and commonly crossed specialty boundaries. Dermatology appeared in 82% of cases, gastroenterology in 52% and rheumatology in 36%, with smaller contributions from haematology-oncology, neurology, paediatrics and nephrology. On average, 2.3 specialties were directly involved in a case.

That makes the paper more informative than a question-answer benchmark built from examination items. These were real clinical problems seen by a functioning board, and 21 specialists took part. Yet it remains a prospective, observational feasibility study at one German centre. It did not randomise cases to AI-supported and unsupported care, compare decisions against later diagnoses, or measure treatment response, complications, time to care, costs or patient experience. The strongest justified conclusion is therefore about whether an AI evidence tool could be inserted into this particular meeting process.[1]

Human discussion came first, but the final reference was not independent of the AI

For each case, clinicians first discussed the question without the model and entered a preliminary recommendation in the hospital documentation system. A prepared, de-identified case prompt was then submitted to the commercial OpenEvidence platform. The board saw the answer, could add or amend points, and only then confirmed its final recommendation. No personal patient data was entered into the service, according to the paper.

This sequence protects an important human-first step, but it complicates interpretation of agreement. The final board recommendation used as the clinical reference could itself incorporate something raised by the model. Agreement with that final decision is therefore not the same as an independent test against an untouched expert answer, and still less a test against patient truth. The authors' design is useful for studying integration; it cannot establish that the model matched an independent gold standard or caused a better decision.[1]

Specialists liked the answers more than the agreement scores imply

The board members completed 759 clinician-case evaluations using eight five-point rating items and a free-text field. The highest mean score was 4.06 for providing clinically relevant information. Recommendation for future use averaged 3.95. Time saved was the lowest-rated item at 3.50 and also varied most. All eight medians were four, suggesting that specialists generally regarded the tool as useful and workable during discussion.

Perceptions were not uniform. Nephrologists gave the most conservative overall ratings, while some other specialties were more positive; the paper warns that only two to six raters represented each specialty and case exposure was uneven, so those differences should not be treated as comparative evidence. Only seven of the 759 evaluations included a free-text comment. Positive scores show acceptability among participating clinicians, but repeated ratings from the same small group are not equivalent to 759 independent users or 759 independent cases.[1]

Full agreement was a minority-to-bare-majority result

Two physician authors independently reviewed outputs for 97 evaluable cases and classified agreement with the final board recommendation as none, partial or complete. One rater judged complete agreement in 40.2% of cases and the other in 50.5%. They judged no agreement in 21.6% and 24.7%, respectively. The raters assigned the same category in 62 of 97 cases; within that consensus subset, 51.6% were complete, 22.6% partial and 25.8% no agreement.

Those figures resist both optimistic and dismissive readings. Partial or complete agreement in a majority suggests the system often surfaced material compatible with expert thinking. But a tool that is not fully aligned in roughly half the cases, and is judged not aligned in around a quarter, cannot safely replace the specialists capable of spotting what is missing. Inter-rater agreement was only moderate by the paper's reported statistics, which also shows that even physicians on the same board did not always agree about how closely an answer matched the recommendation.[1]

Prompt optimisation was not a proven fix

The researchers later reran all cases with a structured prompt asking the model to assess whether information was sufficient, request missing details and then give a justified recommendation. The model asked follow-up questions in every case, but the researchers supplied no additional information before it generated the recommendation. Complete agreement increased numerically to 50.5% and 51.5% for the two raters; in the consensus subset it reached 56.5%.

Across the paired comparison, 29 cases improved, 48 were unchanged and 20 worsened. The mean score rose from 1.222 to 1.361 on the study's zero-to-two scale, but the reported p-value was 0.0642 and the 95% confidence interval for the effect crossed zero. That does not prove the structured prompt had no value; it means this study did not establish a reliable improvement at its stated threshold. Prompt templates may help organise questions, but they are not a substitute for missing clinical information, external validation or outcome testing.[1]

The same cases did not produce stable evidence summaries

When the researchers resubmitted the cases months later, the outputs changed substantially. For the original prompts, mean response length rose by 91.1%, the number of cited references by 93.0%, and only 20.4% of the originally cited references were retained. For the optimised prompts, responses grew by 33.5%, citation counts by 22.1%, and 23.6% of references were retained. The platform is proprietary and continuously updated; the provider did not disclose a model or product version for the first two rounds, so the study cannot separate model updates from other sources of variation.

A same-day comparison was more reassuring but still not identical: changing the prompt left the core clinical recommendation unchanged in 87 of 100 cases. The practical lesson is not that later answers were necessarily worse. Longer responses and newer sources could reflect improvement. It is that a clinician cannot assume yesterday's answer, evidence set or rationale will be reproducible today. Governance needs dated outputs, recorded prompts, accessible citations, version information where available and a clear responsibility for re-checking the underlying evidence.[1]

What this means for patients and clinical teams

For clinicians, an evidence-synthesis model may be most useful as a second-pass literature assistant after experts have framed the case and identified missing facts. It could surface therapeutic options or papers during a meeting, especially when knowledge spans several disciplines. The workflow should preserve the preliminary human recommendation, require a qualified reviewer to open the cited evidence, and document why any model suggestion changed—or did not change—the final plan.

For patients, the paper offers no evidence of better outcomes, fewer errors or safer care. It did not measure hallucinations or clinical harm. People should not infer that use of the model improved their diagnosis or treatment simply because clinicians rated some answers as relevant. Deployment would also require local rules for de-identification, data processing, audit trails, conflicts, consent or notice where applicable, and a route to challenge a decision. The study's ethics approval and avoidance of personal data are useful design choices, not a complete governance framework for routine use.[1]

What evidence would change the assessment

Confidence would increase with preregistered, multicentre trials that keep the comparator independent: teams could be randomised to standard review or model-assisted review, with predefined measures of decision quality, time, resource use, errors, equity and patient outcomes. Cases should span institutions, languages, specialties and levels of complexity. Independent adjudicators should be blinded to the system used, and subgroup results should show where the tool helps, where it adds noise and whether benefits persist after clinicians become familiar with it.

Future evaluations also need fixed model identifiers or archived outputs, reproducibility checks, verified citation accuracy, explicit hallucination and harm review, and disclosure of commercial relationships. This study reported German Research Foundation funding for two authors, open-access support through Projekt DEAL and no competing interests. Its result is valuable precisely because it is bounded: OpenEvidence could be fitted into one expert board and was often perceived as useful, but the evidence does not support autonomous use, a safety claim or a promise of improved care.[1]

What this means for people

  • Patients should not treat a model-supported discussion as evidence that care is safer or more accurate; this study did not measure outcomes or harm.
  • Clinicians may gain faster access to literature across specialties, but the burden of checking sources and judging missing context remains with qualified professionals.
  • Health systems need audit trails, privacy controls and routes for review so an unstable proprietary output does not become an unchallengeable part of care.

Global context

Health systems worldwide are testing language models as evidence assistants while regulators and hospitals still lack consistent ways to evaluate changing proprietary systems. This German single-centre study adds rare real-workflow evidence, but its design and instability results strengthen the case for local validation, recorded versions and human accountability before wider clinical reliance.

What the evidence does not yet show

  • Single-centre observational study of 100 inflammatory-disease cases and 21 specialists; there was no randomised control group or external validation.
  • The final board recommendation was confirmed after clinicians saw the model output, so it was not a fully independent reference standard.
  • Two physician authors rated concordance; one had helped prepare 10 prompts, and blinding to the evaluation round was not possible.
  • The proprietary platform did not disclose a model or product version for the first two rounds, limiting reproducibility and attribution of changes over time.
  • Patient outcomes, clinical harm, hallucinations, citation accuracy and economic effects were not assessed.

What to watch next

  • Independent multicentre trials comparing model-assisted boards with standard care on predefined clinical and operational outcomes.
  • Prospective auditing of citation accuracy, omitted evidence, hallucinations, unsafe recommendations and subgroup performance.
  • Versioned, reproducible outputs and governance that records prompts, sources, reviewer decisions and later model changes.
  • Evidence that any time saving or broader evidence access translates into safer, fairer or more effective care.

Evidence trail

Sources used for this report

Links checked 3 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Health & Life Sciences

Can AI make tumour mitosis counts more consistent?

A preprint study paired 13 pathologists' unaided and AI-assisted reviews of 385 tumour slides from three European centres. Agreement rose and counting time fell, but the unreviewed study did not establish which counts were correct, whether diagnoses improved or whether patients benefited.

11 min · 1 source

Health & Life Sciences

Did clinicians prefer AI discharge summaries after long hospital stays?

In a retrospective 60-case comparison, 12 physicians usually preferred GPT-5.2 summaries and annotated fewer omissions. Reviewers knew which summary was AI-written, one hospital supplied the records, and no patient outcome or time saving was tested.

7 min · 3 sources

Health & Life Sciences

What should Europe require before hospital AI becomes routine?

A same-day report from Europe’s science and medical academies argues that health AI needs evidence suited to changing systems, interoperable data, accountable regulation and a workforce able to challenge the technology—not simply more pilots.

6 min · 2 sources

Reader discussion

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.