Back to the news portal
Health & Life SciencesResearch paperResearchSource analysisSouth KoreaEast AsiaGlobal health

How accurate must diagnostic AI be before people accept it?

A Korean survey of 1,328 people found mean minimum acceptable sensitivity and specificity of 90% for AI disease detection. That is a demanding stated threshold—not a clinical performance standard, a treatment decision or proof of what people would accept when facing a real diagnosis.

By The Impact of AI Editorial DeskReleased 7 October 2026 at 15:59 BST9 min read1 source

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

Share
Social links
LinkedInXBlueskyRedditEmail

At a glance

  • 1The nationwide Korean survey included 1,222 adults from the general population and 106 long-term lung-cancer survivors, for 1,328 respondents in total.
  • 2Mean minimum acceptable sensitivity and specificity were both 90%; 81.6% required sensitivity of at least 88%, and the same share required specificity of at least 90%.
  • 3The study measures stated preferences in hypothetical questions. It does not establish a regulatory threshold, evaluate a device or show how people trade accuracy against access, cost, speed or clinician oversight in real care.
Key themesDiagnostic AIPatient preferencesSensitivitySpecificityShared decision-makingEvidence thresholds

Research topic

What minimum sensitivity and specificity patients and the public say they would accept from AI-based disease detection

The Impact of AI research cover asking how accurate diagnostic AI must be before people accept it, with a conceptual clinician, diverse public group and diagnostic threshold interface.
AI-generated editorial illustration. The clinician, public group and interface are conceptual; they do not depict study participants, patient records, a tested device or actual clinical measurements.

The direct answer: respondents wanted about 90% on both measures

People in this Korean survey set a high bar for AI used to detect disease. Across 1,328 respondents, the mean minimum acceptable sensitivity was 90% and the mean minimum acceptable specificity was also 90%. Sensitivity describes how often a test identifies people who truly have the disease; specificity describes how often it correctly clears people who do not. A system can improve one while worsening the other, so asking about both is essential.

The distribution was demanding as well as the average. The researchers report that 81.6% of participants required sensitivity of at least 88%, meaning they would tolerate no more than 12 missed cases among 100 people who actually had the disease. The same proportion required specificity of at least 90%, equivalent to accepting no more than 10 false alarms among 100 people without the disease under the survey framing.

Those numbers answer a question about stated acceptability, not clinical safety. Respondents did not use an AI device, receive a diagnosis or choose between an AI system and a clinician. The result should therefore inform design, consent and public communication, while regulators and health systems still need disease-specific evidence about harms, benefits and the consequences of mistakes.[1]

Who was asked and how the question worked

The researchers combined two Korean groups. The main nationwide general-population sample contained 1,222 adults. A second group contained 106 people who had survived lung cancer for the long term. Together they produced the 1,328-person denominator. Keeping the two groups visible matters because personal experience of cancer could change how someone values a missed diagnosis relative to a false alarm.

Rather than asking whether AI should simply be ‘accurate’, the questionnaire translated performance into counts. Participants considered a hypothetical group of 100 people who had a disease and stated the maximum number the AI could miss; that response was converted into a minimum acceptable sensitivity. They separately considered 100 people without the disease and stated the maximum number who could be wrongly flagged; that became minimum acceptable specificity.

This count-based framing makes the two error types easier to picture than abstract percentages, but it also simplifies medicine. The seriousness of missing an aggressive cancer is not the same as missing a mild, self-limiting condition. A false positive can range from a repeat scan to an invasive biopsy. The paper estimates a general threshold across a hypothetical disease-detection setting rather than one universal value that can safely govern every test.[1]

Cancer experience was associated with stricter expectations

The analysis found that a history of cancer was associated with higher required performance. That is plausible: someone who has experienced diagnosis and treatment may have a sharper understanding of the cost of a delayed finding or an unnecessary alarm. But an association in this survey cannot show that cancer experience caused the difference. Age, healthcare contact, education, risk perception and other characteristics may also shape the answer.

The lung-cancer survivor group was much smaller than the public sample—106 people compared with 1,222. Its value is therefore in adding an experienced patient perspective, not in producing a precise national estimate for all cancer survivors. The study does not justify treating survivors as one uniform group or assuming that people with other diseases would make the same trade-off.

For developers, the practical lesson is that one acceptance threshold may conceal clinically important variation. Patient and public involvement should happen at the level of the intended disease, care pathway and consequence of error. A screening tool for a low-prevalence condition, for example, may create a large number of false positives even when its specificity sounds high, while a triage system may be designed to favour sensitivity and send more people for human review.[1]

Why 90% is not a universal pass mark

A 90% sensitivity sounds intuitive, but its meaning depends on what happens after the test. If AI is the only route to diagnosis, missing one in ten affected people could be unacceptable. If it is a preliminary flag checked by a clinician alongside other evidence, the system-level miss rate could be lower. Specificity has the same dependence on context: the burden of ten false alarms per 100 unaffected people changes with disease prevalence and the follow-up procedure.

The authors compare respondents' expectations with published performance for IDx-DR, an autonomous system for diabetic-retinopathy detection, citing observed sensitivity of 87.4% and specificity of 89.5%. That comparison illustrates a potential expectation gap, but it is not a head-to-head evaluation. The survey concerned hypothetical disease detection, while the device operates in a defined population, workflow and regulatory context with its own endpoints and confidence intervals.

Regulation also cannot be reduced to a single public-preference number. Authorities consider intended use, the severity of false negatives and false positives, human oversight, subgroup performance, calibration, cybersecurity and the quality of the reference standard. Public acceptability is one input into that judgement. It does not replace clinical validation or automatically determine whether a benefit–risk balance is favourable.[1]

What the survey means for patients and clinicians

Patients should receive more than a headline accuracy claim. Meaningful information would separate sensitivity from specificity, state the population and disease prevalence in which the system was tested, explain what happens after a positive or negative output and disclose whether a clinician can override the result. The survey suggests that people care about both missed disease and false alarms; consent material should make both visible.

Clinicians may also face a communication problem when a useful system falls short of a round 90% threshold on one metric. The answer is not to obscure the number, but to explain the entire pathway: whether the tool expands access, whether a second test catches misses, whether subgroup results differ and what alternatives are available. A lower-performing component can still improve care if embedded in a reliable safety net; a higher-performing model can still harm people if used outside its validated setting.

Health systems could use preference research to decide where human review is most important and what performance information should appear in patient-facing materials. They should not infer that every respondent would reject a system below 90%. Survey averages compress individual choices, and real decisions can change when waiting time, price, geography or the availability of a specialist is introduced.[1]

The main uncertainty is the gap between a hypothetical answer and real care

The study's strongest contribution is a concrete, patient-facing way to elicit acceptable error. Its central limitation is that intentions stated in a survey may not predict behaviour under uncertainty, symptoms or time pressure. Participants also judged a generic AI detector. They did not compare named systems, review evidence about trade-offs or experience consequences after a result.

Cultural and health-system context matter. The participants were in South Korea, where access pathways, trust in institutions, digital-health experience and attitudes toward cancer screening may differ from other countries. A sample designed to represent a national public cannot establish an international preference. Replication should test different diseases, languages, health systems and groups with varied experience of disability and diagnostic error.

Funding came from a Korean National Research Foundation grant identified by the paper as RS-2024-00440881; the authors state that the funder had no role in the research and declare no competing interests. Those disclosures reduce one obvious concern but do not remove the usual limits of self-reported, cross-sectional survey evidence.[1]

What evidence would change the assessment

Confidence would increase if researchers repeated the questions prospectively at the point where patients actually choose a diagnostic pathway, then compared stated thresholds with decisions and outcomes. Experiments could vary disease severity, prevalence, clinician involvement, waiting time, cost and the follow-up burden. Qualitative interviews would help explain why a respondent accepts one error but not another.

For deployed systems, the decisive evidence is still clinical: prospective accuracy with confidence intervals, subgroup performance, calibration, referral patterns, downstream tests, delayed diagnoses and patient-reported outcomes. Studies should compare AI-assisted pathways with the current standard of care rather than testing a model in isolation. They should also report whether better access offsets any additional false alarms or missed cases.

The 90% average is therefore best read as a warning against vague claims that the public will accept whatever accuracy is technically achievable. People expressed demanding expectations for both kinds of error. The next task is to connect those expectations to disease-specific choices and to show, with real-world evidence, whether an AI-supported pathway meets them.[1]

What this means for people

  • Patients may gain clearer explanations of missed-disease and false-alarm risks instead of a single marketing accuracy figure.
  • Clinicians need evidence and time to explain how an AI result fits into the wider diagnostic pathway.
  • Developers and health systems receive a prompt to involve patients before fixing thresholds, but not a universal 90% target.

Global context

The study provides current evidence from South Korea, where one national public sample was supplemented with a smaller group of long-term lung-cancer survivors. Its measurement approach can travel, but the numerical thresholds should not be exported unchanged. Disease burden, access, trust, follow-up care and the role of clinicians differ across health systems, so internationally useful guidance requires replication in specific pathways and populations.

What the evidence does not yet show

  • The outcomes are stated minimum thresholds in hypothetical scenarios, not observed acceptance, adherence or clinical outcomes.
  • The study does not evaluate a diagnostic system or establish a regulatory performance standard.
  • A generic disease-detection question cannot capture the different consequences of errors across conditions and care pathways.
  • All participants were in South Korea; preferences may differ across countries, cultures and health systems.
  • The long-term lung-cancer survivor subgroup contained 106 people and should not be treated as representative of every patient group.

What to watch next

  • Disease-specific preference studies that vary severity, prevalence, treatment options and follow-up burden.
  • Prospective comparisons of stated thresholds with choices made during real diagnostic care.
  • Performance reporting that separates sensitivity, specificity, subgroup results and the full pathway's outcomes.
  • Patient-facing consent and transparency standards for autonomous and clinician-assisted diagnostic AI.
  • Whether access gains or human review change what performance patients consider acceptable.

Living evidence record

Impact record IAI-1I2SFXJ

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent or research support

Present

Record status

Monitoring

Last checked

7 October 2026

Source trail

1 direct source across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Single-source reporting disclosure

This record analyses one direct source. It can establish what npj Digital Medicine published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.

Evidence trail

Sources used for this report

Links checked 7 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Health & Life Sciences

Can AI spot liposarcoma on ultrasound?

A peer-reviewed Chinese study reports strong results in a 95-patient internal test, including 0.97 accuracy. But all 317 patients came from one hospital and there was no external validation, so this is a proof of concept—not a clinically cleared diagnostic system.

9 min · 1 source

Health & Life Sciences

Can AI improve GLP-1 and weight-loss treatment?

A new analysis of research from China and South Korea: predicting treatment response and designing a different drug candidate are promising ideas, but neither establishes that using AI improves weight-loss outcomes.

6 min · 4 sources

Health & Life Sciences

What does it take to scale healthcare AI safely?

An OECD conference report turns input from 142 experts across 34 countries into ten proposed actions for health systems. It is a practical consensus framework for trust, infrastructure and financing—not evidence that scaled AI has improved patient outcomes.

10 min · 2 sources

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Reader commentary

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Explore commentary across the portal →

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.