Back to the news portal
Health & Life SciencesNew analysis today · source 8 October 2026Research paperResearchSource analysisChinaUnited StatesEuropeAustraliaGlobal primary care

Does primary-care AI improve outcomes?

Not yet on the available evidence. A peer-reviewed review found 10 real-world studies of clinician-facing AI in primary care: some improved detection or care processes, but neither trial measuring patient-important outcomes demonstrated benefit, and the evidence was low or very low certainty.

By The Impact of AI Editorial DeskReleased 9 October 2026 at 13:00 BST10 min read2 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

Share
Social links
LinkedInXBlueskyRedditEmail

At a glance

  • 1The review assessed 66 full texts and included 24 studies; only 10 contained a learned or case-based AI component and formed the appraised AI evidence base.
  • 2Five cluster-randomised trials, one individual randomised trial, two nonrandomised evaluations and two prospective diagnostic studies produced mixed results across different diseases and outcomes.
  • 3Neither trial measuring patient-important outcomes demonstrated benefit; certainty was low for clinical and diagnostic outcomes and very low for clinician efficiency.
Key themesPrimary careClinical decision supportPatient outcomesImplementation scienceEvidence qualityHealth AI

Research topic

Whether clinician-facing AI decision support used in routine primary or ambulatory care improves diagnosis, clinical decisions, efficiency or patient-important outcomes

The Impact of AI research cover asking whether primary-care AI improves outcomes, with a conceptual clinician, patient and transparent decision-support pathway; it states that 10 real-world studies were reviewed and patient benefit was not demonstrated.
AI-generated editorial illustration. The consultation, model panel and care pathway are conceptual; they do not depict a real patient, clinician, medical record or successful deployment.

The direct answer: real-world use has not yet shown consistent patient benefit

The strongest answer from this review is cautious: clinician-facing AI has sometimes changed detection or care processes in primary care, but the evidence does not establish that it consistently improves outcomes that patients feel or experience. The authors identified 10 eligible real-world AI studies across cardiovascular case-finding, dementia, melanoma, glaucoma, diabetic retinopathy, urinary tract infection, HIV prevention, falls, childhood asthma and musculoskeletal pain. The interventions and outcomes were too different for a single pooled effect, and the review found no reliable chain from a model output to a better clinical decision and then to a better patient outcome.

That distinction matters because primary care is where broad populations first meet a health system. A tool can classify an image well, flag a risk or increase a recorded diagnosis without improving health. It can also create extra referrals, false reassurance, alert burden or work that an accuracy score does not capture. The review therefore treats diagnostic performance, clinician decisions, efficiency and patient-important outcomes as separate evidential levels rather than assuming success at one level proves success at the next.[1]

The Impact Brief · Free

Follow the evidence in health & life sciences.

Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.

Choose your topics (optional)

One concise, source-linked briefing. Unsubscribe at any time.

What the reviewers actually included

The team searched five principal bibliographic databases, a supplementary EBSCOhost platform search and two trial registries for work published from 2011 through 2025. Two reviewers independently screened records. Of 66 reports assessed in full, 24 met the broad setting and outcome criteria. Ten used a learned or case-based AI component and formed the appraised population. Fourteen used deterministic alerts, fixed rules or risk equations and were retained only as context, not counted as evidence that AI itself was effective.

The 10 AI studies comprised five cluster-randomised trials, one individually randomised trial, two nonrandomised intervention evaluations and two prospective diagnostic-accuracy studies. That is more relevant to service delivery than a retrospective benchmark because the systems were used by clinicians in routine primary or ambulatory care with real patients. It is still a small and heterogeneous evidence base. Different conditions, comparators, workflows, adoption levels and endpoints mean a favourable result in one setting cannot be carried across the group.[1][2]

Detection improved in one trial and did not in another

In a pragmatic trial of AI-enabled electrocardiograms, new diagnoses of low ejection fraction rose from 1.6% to 2.1%, an odds ratio of 1.32 with a 95% confidence interval from 1.01 to 1.61. The review counted 22,641 participants for that evidence body. The result shows that a deployed case-finding system can change who is identified, but it does not by itself show that treatment, symptoms, hospital admissions or survival improved because of the additional diagnoses.

A separate cluster trial tested an electronic-health-record machine-learning marker for dementia among 5,325 participants. Used alone, it did not increase new diagnosis: 10.3% were diagnosed in the algorithm arm versus 12.4% in the comparator, with an adjusted odds ratio of 0.84 and a confidence interval from 0.63 to 1.11. A combined arm pairing the marker with a patient-reported instrument reached 15.4%. Because the combined intervention has more than one component, the extra diagnoses cannot be attributed to the algorithm alone.[1]

Good discrimination did not establish better care

The two prospective diagnostic studies produced very different forms of evidence. A Swedish primary-care dermoscopy application reached an area under the receiver operating characteristic curve of 0.960 for melanoma among 253 lesions. All lesions still went through the standard diagnostic work-up, so the study did not isolate whether the AI changed management or outcomes. In Australian general practice, a glaucoma classifier achieved an area under the curve of 0.80, sensitivity of 65% and specificity of 94.6% among 277 analysable participants.

The glaucoma denominator also exposes an implementation problem: only 277 of 414 recruited participants produced analysable images. A model may perform acceptably on usable inputs while the service fails to obtain those inputs for a substantial share of patients. The reviewers rated both diagnostic studies at high overall risk of bias under QUADAS-2, for clinician-selected sampling and differential verification in the melanoma study and for flow and image-acquisition exclusions in the glaucoma study.[1]

Care processes moved, but attribution was difficult

Some process results looked promising. An interpretable urinary-tract-infection system was associated with treatment success rising from 75% to 80% in a controlled before-after comparison, and to 83% among confirmed users. Yet the lack of concurrent randomisation leaves room for changes in practice, staffing, patient mix or other interventions to explain some of the difference. A diabetic-retinopathy system affected screening criteria for three of four general practitioners, but referral agreement varied and the before-after design could not isolate the AI's contribution.

The HIV pre-exposure-prophylaxis trial was null overall: initiation was 6% with machine-learning prompts and 4.5% in the comparator, with a hazard ratio of 1.32 and a 95% confidence interval from 0.84 to 2.10. A fall-prevention intervention improved shared decision-making but had inconclusive medication effects. These examples show why adoption, fidelity and the other parts of a multicomponent intervention need measurement alongside the model. A technically sound prediction can have little effect if clinicians do not see, trust or act on it, while an apparent effect may come from accompanying workflow changes.[1]

Patient outcomes were the weakest part of the case

Only two trials contributed patient-important outcome evidence. One involved 184 children in an asthma decision-support trial; the reported odds ratio for asthma exacerbation was 0.82, but its wide confidence interval from 0.37 to 1.96 included substantial benefit, no effect and harm. The second involved 724 adults receiving physiotherapy for musculoskeletal pain. Its global perceived-effect estimate was similarly imprecise, while one functional-improvement result favoured usual care rather than the AI-supported intervention.

The reviewers rated this patient-outcome body as low certainty. They also rated diagnostic case-finding and prospective diagnostic evidence as low certainty, and clinician-efficiency evidence as very low certainty. The efficiency estimate came from 28 of 42 clinicians who reported that record review took 3.5 minutes with the asthma tool versus an estimated 11.3 minutes without it. That was unblinded self-report rather than observed, randomised time measurement, so it cannot support a firm productivity claim.[1]

How the review handled bias and uncertainty

The authors used design-specific appraisal tools: standard and cluster versions of Cochrane's RoB 2 for randomised trials, ROBINS-I for nonrandomised interventions and QUADAS-2 for diagnostic studies. All six randomised trials had some concerns overall. The two nonrandomised studies were judged at serious risk of bias, largely because confounding and changes over time could explain the results. The review applied GRADE to four outcome bodies and did not manufacture a pooled estimate where the studies were clinically too different.

The authors did not formally assess publication bias because every rated outcome body contained at most two studies. That is a real limit: positive deployments may be more likely to be written up than failed or abandoned ones. Harms were also rarely prespecified or quantified. The paper names overdiagnosis, false positives, automation bias, inappropriate reassurance and alert burden as plausible consequences that the evidence base did not measure well enough to assess.[1][2]

What this means for patients, clinicians and purchasers

For patients, the evidence does not justify assuming that an AI flag improves health simply because it detects more cases or has a high discrimination score. Services should explain whether a tool changes a referral, test or treatment; what happens after a false positive or false negative; and whether outcomes are monitored for the people actually served. Human review remains essential, but it is not a complete safeguard if clinicians face opaque recommendations, excessive alerts or no practical route to challenge the system.

For clinicians and health-system purchasers, the review argues for trials that measure the whole pathway. A useful evaluation should record technical performance, who receives and acts on an output, changes in decisions, observed workload, downstream resource use, harms and patient outcomes. Local validation and calibration still matter, but they are only part of the question. Procurement claims should be tied to the endpoint actually measured rather than translating an accuracy statistic into a promise of safer, faster or cheaper care.[1]

Global context, funding and what would change the assessment

The reviewers were based at Tsinghua Medicine in Beijing, while the underlying deployments spanned several health systems and conditions. The paper reports no financial support and no conflicts of interest. Its conclusions are global in relevance but not a universal effect estimate: primary-care access, clinician roles, referral pathways, disease prevalence, digital records and liability rules differ sharply across countries. A tool that fits one service can create different costs or errors elsewhere.

Confidence would change with larger preregistered trials of locked systems in multiple primary-care settings, using contemporaneous comparators and adequate follow-up. They should measure patient-important outcomes and harms, not just detection; record uptake and overrides; observe rather than estimate time; and report subgroup performance and image or data acquisition failures. Until that evidence exists, the fairest conclusion is that some real-world systems alter parts of care, but consistent patient benefit remains unproven.[1][2]

What this means for people

  • Patients should not be told that a model improves health when the evaluation measured only detection or diagnostic discrimination.
  • Clinicians need visible uncertainty, workable override routes and evidence about workload as well as accuracy.
  • Purchasers should require outcome and harm monitoring in the population and workflow where a system will be used.

Global context

The review was conducted in China and synthesised deployments from several countries, but primary care is organised differently around the world. Referral capacity, access to follow-up tests, digital-record quality, clinician responsibilities and disease prevalence can change both benefit and harm. Cross-country transfer therefore needs local prospective evaluation rather than assuming a published effect will travel.

What the evidence does not yet show

  • Only 10 real-world AI-enabled studies met the review criteria, spread across different conditions, systems and outcomes.
  • No meta-analysis was appropriate because the interventions and endpoints were too heterogeneous.
  • All six randomised trials had some risk-of-bias concerns; the two nonrandomised studies were judged at serious risk.
  • Harms, implementation fidelity and clinician workload were incompletely or inconsistently measured.
  • The review searched evidence through December 2025, so it cannot assess deployments reported later.

What to watch next

  • Multisite randomised trials that connect model output, clinician action and patient-important outcomes.
  • Observed workload and resource-use measures rather than clinician estimates.
  • Prespecified monitoring for false reassurance, overdiagnosis, alert burden and unequal error rates.
  • Reporting of acquisition failures, nonuse, overrides and local calibration after deployment.
  • New evidence published after the review's December 2025 search cut-off.

Living evidence record

Impact record IAI-0UD7QRI

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent or research support

Present

Record status

Monitoring

Last checked

9 October 2026

Source trail

2 direct sources across 2 source types.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Related-source reporting disclosure

This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.

Evidence trail

Sources used for this report

Links checked 9 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Health & Life Sciences

Is operating-room AI ready for clinical use?

Not on the published evidence yet. A peer-reviewed scoping review screened 3,020 records but found only one completed feasibility study with five analysed patients; four larger prospective studies had no results posted.

8 min · 1 source

Health & Life Sciences

Can interpretable AI distinguish ovarian tumours before surgery?

A peer-reviewed retrospective study trained five classifiers on laboratory data from 349 patients at one Chinese hospital. Its best model reached 93.7% leave-one-out accuracy, but the small reused cohort and absence of external or prospective validation rule out clinical deployment.

7 min · 2 sources

Reader commentary

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Explore commentary across the portal →

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.