Back to the news portal
AI Risks & SafetyNew analysis today · source 6 October 2026Research paperResearchSource analysisCanadaNorth AmericaGlobal
Source record 1. arXiv

Do AI explanations prevent over-reliance?

Not in this unreviewed experiment. Explanations reduced raw acceptance but did not improve reliance calibration among novices doing a clinical-text annotation task; over-reliance tended to increase across the session.

By The Impact of AI Editorial DeskReleased 7 October 2026 at 06:03 BST7 min read1 source

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

Share
Social links
LinkedInXBlueskyRedditEmail

At a glance

  • 1The final sample of 110 online participants completed 26,460 clinical-text annotation decisions across four conditions with or without AI, explanations or metacognitive prompts.
  • 2AI assistance improved task accuracy, but users accepted suggestions 75% of the time overall; explanations lowered acceptance from 79% to 72% without reliably improving calibration.
  • 3Over-reliance tended to grow across the session and varied greatly by person. The laboratory task and crowd-worker sample do not establish effects in clinical practice.
Key themesHuman-AI interactionExplanationsAutomation biasReliance calibrationDecision support

Research topic

Whether showing natural-language explanations changes how well novice users accept correct AI suggestions and resist incorrect ones when performance feedback is limited

The Impact of AI research cover asking whether AI explanations prevent over-reliance, with a conceptual accept-or-reject decision path and a visible unreviewed-preprint label.
AI-generated editorial illustration. The explanation panel, decision path, hand and balance are conceptual and do not reproduce clinical records, participant screens or measured outcomes.

Explanations did not calibrate novice trust

The study's answer is no: adding an explanation to an AI suggestion did not reliably improve whether novices accepted correct advice and resisted wrong advice. Explanations reduced the overall acceptance rate, but lower acceptance is not automatically better calibration. A cautious user can reject useful advice as easily as a credulous user can accept a mistake. The researchers instead measured alignment between reliance and AI correctness item by item.

Across the session, participants tended to drift toward over-reliance rather than learn a stable boundary for when to trust the system. That result matters for products used where ground truth arrives late or not at all: explanations can feel informative without giving users the domain knowledge or feedback needed to evaluate them. It does not mean explanations are useless in every setting, and it does not show that clinicians behave like online novices.[1]

Four conditions separated assistance and reflection

Participants annotated clinical concepts in text by deciding whether passages matched categories. The experiment used four conditions: no AI; no AI plus a metacognitive intervention; AI suggestions without explanations; and AI suggestions with explanations. The task was chosen because categories could be defined and AI correctness could be evaluated, while participants generally lacked specialist clinical-annotation experience. Feedback about performance was deliberately limited, reproducing situations in which users cannot easily learn from known outcomes.

The AI interface allowed a participant to accept, modify or reject suggested annotations. In the explanation condition, suggestions arrived with generated rationales. The study asked not only whether AI improved accuracy, but how reliance changed with sequential tasks, whether explanations improved calibration and whether self-reported confidence or understanding predicted behaviour. This is a more demanding question than whether users liked explanations or said they trusted them.[1]

The analysed sample was 110, not the 26,460 decisions

The team recruited 121 adults through Prolific, requiring high self-reported English fluency and paying £9 per hour. Eleven were excluded for not completing the task or making only one action, leaving 110: 31 with no AI, 28 with no AI plus metacognitive prompts, 23 with AI and 28 with AI plus explanations. They made 26,460 text-level decisions, but repeated decisions from one person are not independent participants.

Participants reported little prior annotation experience and little healthcare knowledge, although average self-rated task understanding and confidence were high before the task. The number of tasks completed varied widely, from one to 102, with a mean of 13.84. Mixed-effects models help account for repeated observations, but large variation in exposure means some participants contribute much more evidence about session trajectories than others.[1]

AI improved accuracy while reliance remained poorly calibrated

Mean accuracy was 0.429 in the AI-only group and 0.454 with AI plus explanations, compared with 0.211 and 0.180 in the two non-AI groups. Assistance therefore made participants better at this constructed annotation task. Yet in the two AI conditions, they accepted suggestions 75.0% of the time, modified 20.2% and rejected only 4.8%. Acceptance was 79.0% without explanations and 72.0% with them.

The authors adapted a Brier-score logic into a reliance calibration score: lower values mean behaviour is better aligned with whether the AI is correct. Average over-reliance was far larger than under-reliance. A Bayesian mixed-effects analysis estimated increasing over-reliance across task sequence, with a 95% credible interval spanning a small negative value but a 0.94 posterior probability of a positive slope; under-reliance decreased with a 0.97 posterior probability. These are probabilistic trends, not deterministic laws.[1]

An explanation can persuade without teaching verification

Natural-language rationales add content, but users still need a way to judge that content. An explanation may restate a prediction fluently, highlight plausible evidence or increase the feeling that the system has reasoned. Without immediate ground truth, novices cannot easily learn which features distinguish a sound rationale from a post-hoc story. In this experiment, explanation exposure did not reliably repair that information gap.

The distinction between confidence and understanding is also important. A person may feel capable while lacking the conceptual model needed to scrutinise suggestions. The paper reports that higher self-reported task understanding was associated with more selective reliance, whereas confidence alone was not a dependable guide. The authors propose prompting users to record why they accepted or rejected advice, then reflecting those reasons back at natural breakpoints. That intervention remains a design proposal, not a tested safety control.[1]

The task is not a clinical deployment

Although the text concerned clinical entities, participants were crowd workers doing annotation, not clinicians diagnosing or treating patients. The outcome was label accuracy, not health, workflow quality or safety. The authors chose categories intended to be distinguishable; if categories are ambiguous or instructions poor, misunderstanding could change reliance in ways this design does not capture. Prolific recruitment also does not represent the full demographic, educational or professional range of future users.

Motivation came partly from payment, and the study did not model cognitive load or intrinsic motivation. No funding or competing-interest statement was visible in the accessible preprint text, so readers cannot yet assess those disclosures. The paper is under review and may change. These constraints do not erase the observed behaviour; they limit how far it can travel—from a controlled novice annotation setting to medicine, finance, public administration or other high-stakes work.[1]

What would change the assessment

Confidence would rise with preregistered replications that manipulate explanation quality, AI accuracy and feedback timing independently. Researchers should test whether explanations help when users can inspect sources, compare counter-evidence or receive delayed outcomes, and whether requiring a brief rationale improves calibration rather than merely slowing decisions. Outcomes should include appropriate acceptance, appropriate rejection, task accuracy, time, cognitive load and learning across sessions.

High-stakes claims require domain professionals working on representative cases with real consequences and supervision. Studies should compare explanations with uncertainty displays, source citations, independent verification and deliberate-friction designs, including combinations. Until then, product teams should not count an explanation box as proof of safe human oversight. Explanations may support comprehension, contestability or audit, but the user also needs domain knowledge, usable evidence and feedback to decide when the machine is wrong.[1]

What this means for people

  • Users may become more accurate with AI while still developing an unhealthy pattern of deference.
  • A persuasive explanation can raise confidence without supplying the evidence needed to check a prediction.
  • Organisations remain responsible for monitoring, escalation and verification; an explanation interface is not a substitute for those controls.

Global context

The experiment was run online and the authors are affiliated with a Canadian research group, but explanation interfaces are deployed globally. Language, expertise, institutional authority and access to feedback can all alter reliance. In settings where users cannot challenge a system or inspect original evidence, explanation text may carry more persuasive weight and less corrective value, making local evaluation and governance essential.

What the evidence does not yet show

  • The paper is an unreviewed preprint and may change after peer review.
  • The 110 analysed participants were Prolific crowd workers with little healthcare or annotation experience, not clinicians.
  • The clinical-entity annotation task was controlled and low-stakes; results do not establish effects on real-world decisions or outcomes.
  • Task exposure varied greatly across participants, and eleven recruits were excluded for minimal or incomplete participation.
  • The study did not test source-linked explanations, calibrated uncertainty, delayed feedback or longer-term learning.

What to watch next

  • Independent replications that vary explanation quality separately from AI accuracy.
  • Tests with professionals, representative cases and delayed real-world feedback.
  • Comparisons among explanations, uncertainty, citations, independent checks and deliberate friction.
  • Evidence that metacognitive prompts improve calibration rather than just reduce acceptance or speed.

Living evidence record

Impact record IAI-0M41LPI

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent or research support

Present

Record status

Monitoring

Last checked

7 October 2026

Source trail

1 direct source across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Single-source reporting disclosure

This record analyses one direct source. It can establish what arXiv published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.

Evidence trail

Sources used for this report

Links checked 7 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

AI Risks & Safety

Can AI agents learn what is appropriate?

Perhaps, but the new Google-led report is a research agenda, not a tested safeguard. More than 50 contributors propose contextual policy engines, dynamic sandboxes and multi-agent benchmarks without demonstrating that they prevent real privacy or security failures.

8 min · 2 sources

AI Risks & Safety

What does a safety resignation reveal about OpenAI?

David Robinson, who says he led safety-report writing for 12 frontier launches and helped draft OpenAI’s current Preparedness Framework, has resigned and called for safety practices closer to aviation or nuclear power. His essay is consequential first-person testimony, not an independent audit or proof of imminent harm.

10 min · 4 sources

AI Risks & Safety

Can clinicians see the evidence behind approved diagnostic AI?

A peer-reviewed audit found public, device-specific performance evidence for 30 of 77 approved pathology and haematology AI products. The result is a transparency finding—not proof that the other 47 lack regulatory evidence or that the documented products are clinically equivalent.

9 min · 3 sources

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Reader commentary

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Explore commentary across the portal →

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.