Back to the news portal
Government & PolicyResearch paperResearchSource analysisCanadaFranceGermanySouth KoreaUnited KingdomUnited States

Can health systems evaluate generative AI?

A peer-reviewed review screened more than 43,000 policy and academic records and compared 56 guidance documents across six health systems. It found established routes for trials, economics and real-world evidence, but no specific framework for generative AI's variable outputs and changing performance.

By The Impact of AI Editorial DeskReleased 8 October 2026 at 20:03 BST10 min read1 source

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

Share
Social links
LinkedInXBlueskyRedditEmail

At a glance

  • 1The authors screened 22,772 policy records and 20,571 academic records, then included 56 policy documents from Canada, France, Germany, South Korea, the UK and the US.
  • 2All six systems had routes for efficacy, effectiveness or economic evidence, but the review found no specific guidance for generative AI's stochastic outputs, prompt sensitivity or changing model behaviour.
  • 3This is a qualitative policy map, not a trial of a product or proof that current assessment has caused harm. The six systems were purposively selected, document quality was not graded, and the work was funded by Roche Diagnostics.
Key themesHealth technology assessmentGenerative AIEvidence standardsReimbursementHealth policyReal-world evidence

Research topic

Whether health-technology-assessment guidance can evaluate the variable outputs and changing performance of generative AI

The Impact of AI research cover asking whether health systems can evaluate generative AI, with six conceptual guidance documents surrounding an AI evidence core and a qualification that this is a 56-document policy review with no clinical outcomes.
AI-generated editorial illustration. The six document panels, evidence scale, checklist and AI core are conceptual; they do not depict an official framework, regulator seal, clinical system, patient record or evidence that a generative-AI product is safe or effective.

The direct answer: the machinery exists, but it was built for more stable technologies

Health systems already know how to ask whether a technology works, costs less than alternatives and delivers value in routine care. The new comparative review finds that those systems do not yet give assessors a specific method for generative AI's defining problem: the same model can produce different answers to the same task, respond sharply to small prompt changes and change after an update. That gap matters because a favourable evaluation can influence reimbursement, coverage and procurement, while an incomplete one can mistake a strong average score for dependable performance in real care.

The paper does not show that every current decision is wrong or that a particular AI tool is unsafe. It maps written guidance, not decisions or patient outcomes. Its narrower conclusion is important enough: across six deliberately different health systems, the authors found established clinical and economic evidence requirements but no dedicated guidance for sampling variable model outputs, calculating how many outputs must be tested, or reassessing a system as its behaviour changes.[1]

What the researchers reviewed

The researchers used a policy-mapping method based on scoping-review principles and reported the search with PRISMA-ScR. They chose Canada, France, Germany, South Korea, the United Kingdom and the United States to cover different health-technology-assessment structures and because the team could work in English, French, German and Korean. The set spans a single-payer threshold model, comparative-benefit assessment, medico-economic evaluation, a federated system, conditional reimbursement and the US mixture of public coverage bodies and a private value-assessment organisation.

Repository searches identified 22,772 policy records: 15,504 from Canada, 241 from France, 3,410 from Germany, 39 from South Korea, 851 from the UK and 2,727 from the US. A supplementary PubMed search found 20,571 academic records. After screening and eligibility checks, 56 policy documents remained for qualitative synthesis. The searches were first run in early 2025 and updated on 1 November 2025, so the paper describes guidance available at that cut-off rather than every policy change since.

The team grouped relevant passages into four domains: efficacy and effectiveness, health economics, social and equity considerations, and uncertainty management. One author extracted themes and clustered them; three co-authors reviewed the clustering. The study therefore offers a structured cross-system comparison, but it is not a double-coded systematic review of every document and it did not grade the quality or legal force of the included guidance.[1]

Where the six systems already agree

All six systems distinguish technical or clinical efficacy from effectiveness in real practice. Randomised controlled trials remain preferred for higher-risk technologies that can change diagnosis or treatment, while several systems allow alternatives when software evolves too quickly or a randomised design is impractical. France explicitly permits pragmatic trials and target-trial emulation in relevant circumstances; the UK and US recognise observational and real-world evidence; Germany, South Korea and the US have conditional pathways that can tie access to further evidence collection.

Economic evaluation also has a familiar core. Cost-effectiveness and cost-utility analysis remain central, often using quality-adjusted life years. Some frameworks allow cost-consequence approaches or adaptive inputs when one summary measure misses workflow or caregiver effects. That flexibility is useful for software that saves time or changes how care is organised, but the review found little operational direction for modelling a model whose performance distribution can itself move.

Social and equity assessment was less consistent. Some systems ask developers to test performance across demographic groups or use qualitative evidence, while South Korea had no specific social or equity guidance in the reviewed set. A generic requirement to consider bias is not the same as a protocol for testing prompt, user, language and update effects across groups. The practical task is to specify which populations, tasks and deployment settings must be represented—and what happens when performance falls outside an agreed range.[1]

The generative-AI mismatch is about variability, not just opacity

Conventional assessment often assumes a technology has a bounded, reasonably stable set of outputs. Generative models sample from a distribution. Repeating an apparently identical prompt can yield different text, and rephrasing the prompt or changing the user can shift the result again. A single test answer is therefore not a sufficient unit of evidence. Assessors need rules for repeated sampling, clustered observations and uncertainty intervals that reflect both patients and model outputs.

Updates create a second problem. A clinical evaluation may describe one model version while the deployed service changes through fine-tuning, vendor updates or altered system prompts. The authors point to predetermined change-control plans in medical-device regulation as a possible template, but assessment bodies still need triggers for re-testing, fresh economic modelling and withdrawal or restriction. Conditional reimbursement can help only if the required real-world evidence, review period and action thresholds are specified in advance.

The review proposes directions rather than a finished checklist: sample multiple outputs; test sensitivity to prompts and users; use updating methods when performance evolves; define change controls; and make post-market evidence obligations explicit. Those suggestions are plausible, but this study did not compare them experimentally. A framework that adds repeated tests without measuring clinically important outcomes could create paperwork without resolving the decision problem.[1]

What this means for patients, clinicians, payers and developers

For patients, the gap can work in both directions. A health system may over-accept an impressive validation result that does not survive ordinary prompts, languages or model updates. It may also under-value a useful tool when existing measures overlook time saved, access improved or caregiver burden reduced. Neither risk justifies bypassing clinical evidence. Workflow benefit belongs alongside, not in place of, accuracy, safety, equity and patient outcomes.

Clinicians need to know which version was evaluated, what inputs were tested and when the system should not be trusted. Payers need contracts that preserve audit access and connect material changes to reassessment. Developers should expect a stronger evidence package than a single benchmark: prespecified tasks, repeated outputs, relevant comparators, subgroup reporting, calibrated thresholds and monitoring after deployment. Procurement teams should not treat a regulator's market authorisation, a technology-assessment recommendation and a local implementation decision as interchangeable approvals.[1]

The evidence limits and the commercial context

The six jurisdictions were purposively selected, not sampled to represent the world. Australia and the Netherlands were omitted, and many lower- and middle-income health systems have different capacity, financing and data constraints. The corpus mixed binding rules, methods manuals and non-binding frameworks. Treating them together helps map what is written, but a recommendation in a policy paper does not carry the same force as statute or a reimbursement rule.

The authors did not formally assess document quality. One researcher conducted the supplementary academic search, and one extracted and clustered themes before post-hoc review by three colleagues. The public documents and extraction sheet support scrutiny, but independent coding and a prospective update would strengthen confidence. The policy cut-off also means the study cannot establish whether agencies have since adopted unpublished practices or newer guidance.

Roche Diagnostics funded the underlying project. Four co-authors were Roche Diagnostics employees, and LSE employed six authors while receiving Roche funding for the analysis. The paper says the academic team retained editorial control and that no specific funding was received for writing this article. Those disclosures do not invalidate the results, but they matter because diagnostics companies have a commercial interest in how evidence and reimbursement frameworks evolve.[1]

What would change the assessment

The next step is not another high-level principle. Assessment bodies should publish operational protocols and test them on real products: how many outputs to sample, how to vary prompts and users, how to define a material model change, which subgroup and setting checks are mandatory, and which failures pause or reverse coverage. Public case studies should show how those rules change a decision compared with current practice.

Confidence would increase through a broader, independently coded update covering more regions and every new policy after November 2025. Comparative evaluations could then examine whether generative-AI-specific guidance improves the quality, speed and consistency of decisions without creating avoidable barriers. The decisive evidence is prospective: fewer unsafe deployments and missed benefits, better-calibrated reimbursement, transparent reassessment after updates and measurable patient outcomes. Until then, the paper identifies a credible gap, not proof that one proposed remedy has solved it.[1]

What this means for people

  • Patients need evidence that performance holds across ordinary prompts, users, languages and model versions—not only a favourable average benchmark.
  • Clinicians need version-specific limits, escalation rules and a clear human route when a model is uncertain or wrong.
  • Payers and procurement teams need change-control and audit terms so a materially altered product is not treated as the one originally evaluated.
  • Developers benefit from predictable evidence standards, but flexible pathways must not trade away clinical outcomes or equity checks for speed.

Global context

The six systems cover North America, Europe and East Asia and span materially different financing and assessment structures, which makes the shared gap notable. They do not cover Latin America, Africa, South Asia, the Middle East or most lower-resource systems. Those settings may face additional limits in local validation data, language coverage, regulator capacity and bargaining power with vendors. International convergence on minimum evidence fields could help, but thresholds and post-market duties must still reflect local populations, care pathways and the ability to act on a warning.

What the evidence does not yet show

  • The six health systems were purposively selected and do not represent all global assessment architectures or resource settings.
  • The corpus mixed documents with different legal and normative force, and the researchers did not formally grade document quality.
  • One author performed the supplementary literature search and one extracted and clustered themes before post-hoc review, rather than dual independent coding throughout.
  • The search was updated to 1 November 2025, so later policy changes and unpublished operating practice are outside the review.
  • The study mapped guidance rather than assessment decisions, deployed products, safety events or patient outcomes.
  • Roche Diagnostics funded the project, employed four co-authors and funded the LSE analysis that employed six other authors.

What to watch next

  • Operational HTA protocols for repeated output sampling, prompt sensitivity and model-update reassessment.
  • Public case studies showing whether generative-AI-specific rules change reimbursement or coverage decisions.
  • Independent updates covering more regions and guidance published after November 2025.
  • Prospective evidence linking assessment reforms to safer deployment, equitable access and patient outcomes.

Living evidence record

Impact record IAI-0I7DCLM

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent or research support

Present

Record status

Monitoring

Last checked

8 October 2026

Source trail

1 direct source across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Single-source reporting disclosure

This record analyses one direct source. It can establish what npj Digital Medicine published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.

Evidence trail

Sources used for this report

Links checked 8 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Government & Policy

Should children’s clinical AI follow different rules?

The American Academy of Pediatrics says generative-AI tools used in children’s care need pediatric-specific evidence, family-centred design, disclosure and continuous monitoring. The policy is influential guidance—not a trial, regulation or finding that any tool improves outcomes.

10 min · 1 source

Government & Policy

Can HHS cut clinical trials below four years without weakening the evidence?

ARPA-H wants predictive models, common controls, continuous statistics and agentic operations to replace fragmented trial phases. Three linked projects carry awards of up to $100.03 million, but SURPASS is still an open solicitation and no shortened trial or patient benefit has been demonstrated.

6 min · 5 sources

Government & Policy

What does NSF's new AI-science package actually fund?

NSF announced an AI-native laboratory-instrument programme of up to $75 million, more than $300 million in partner commitments for X-Labs, two research prizes and accelerated STEM training. Several are plans rather than open awards, and key budgets, rules and timelines remain unpublished.

9 min · 2 sources

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Reader commentary

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Explore commentary across the portal →

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.