What did a real-clinic test of Google’s medical chatbot actually show?
In a supervised, single-centre feasibility study, 100 adults completed an AMIE text chat before an urgent primary-care visit and no conversation met the prespecified threshold for a safety stop. The Lancet paper does not show that the chatbot can diagnose, triage or improve care without clinicians.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1The study enrolled 114 adults; 100 completed an AMIE text chat and 98 also attended the scheduled primary-care visit at one Boston clinic.
- 2No conversation met the prespecified threshold for a safety stop, but every interaction was watched live by a board-certified physician; one hallucination and five requests for clinical clarification were recorded.
- 3AMIE’s differential included the final diagnosis in 90% of 98 cases and in its top three in 74.5%, but clinicians produced more practical and cost-effective management plans.
Research topic
Whether a physician-supervised conversational AI can collect histories and discuss possible diagnoses with real patients before urgent primary-care visits without triggering prespecified safety interventions
The answer: promising preparation under intensive supervision
The strongest finding is narrow but useful. One hundred adults used Google’s Articulate Medical Intelligence Explorer, or AMIE, through secure text chat before a booked urgent primary-care appointment. A board-certified internal-medicine physician watched every interaction live and could stop it for risk of harm, serious distress, a clinical safety concern or a patient request. None crossed that prespecified stop threshold.
That is evidence that the supervised workflow was feasible in this clinic. It is not evidence that an unsupervised chatbot is safe for triage, diagnosis or treatment. The monitor remained part of the intervention, the appointment was already scheduled, emergencies and some higher-risk groups were outside the study, and the researchers did not compare patient outcomes against usual care.
The chatbot asked about symptoms and medical history, offered possible diagnoses for the patient to discuss with the doctor, and generated a transcript and summary. Of the 100 people who completed the AI conversation, 98 attended the subsequent visit. The peer-reviewed paper therefore moves this work beyond simulated patients, while leaving the central effectiveness question unanswered: did the additional conversation improve care enough to justify its cost, supervision and risks?[1][2][4]
The Impact Brief · Free
Follow the evidence in health & life sciences.
Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.
Who took part and what was measured
The prospective study ran between April and November 2025 at Beth Israel Deaconess Medical Center’s ambulatory primary-care practice in Boston. It enrolled 114 adults seeking care for a new, non-emergency episodic complaint. One hundred completed the pre-visit chat and 98 completed both that interaction and their scheduled appointment. Participation was voluntary and the AI conversation could occur up to five days before the visit.
The design was single-arm: everyone in the analysed workflow used AMIE. There was no contemporaneous group receiving the same care without the chatbot, so changes in consultation time, diagnostic accuracy, workload or patient outcomes cannot be attributed to the system. The main outcomes instead concerned feasibility, prespecified safety stops, ratings of conversational quality, patient and clinician experience, and retrospective comparison of diagnostic and management outputs.
The 98 participants represented a range of ages, racial and ethnic groups, health literacy and technical familiarity, but were younger than the clinic’s wider urgent-care population. More than half of the 1,452 urgent-care visits during the study period involved people older than 60, while the study skewed younger. That recruitment difference matters because comfort with text chat, symptom complexity and the need for assistance can vary with age and disability.[1][2][3][4]
Zero stops did not mean zero errors
No AI conversation required termination under the four registered safety criteria. Supervisors nevertheless identified one hallucination and added clinical clarification in five cases. Those events did not meet the stop threshold, but they show why ‘zero safety stops’ should not be translated into ‘zero safety problems’. A larger sample or less intensive oversight could expose different failure rates.
The safety endpoint also concerned the conversation itself. It did not establish that suggested diagnoses were correct, that management advice improved treatment, or that delayed or unnecessary care was avoided. One clinician later described an interaction as somewhat harmful because including lymphoma among the possible diagnoses may have caused anxiety. A system can complete a conversation without an emergency interruption and still create downstream confusion or worry.
For patients, the practical safeguard was human review: a monitor could intervene during the chat, and the treating clinician remained responsible for the visit. Removing either layer would create a different intervention and would need a new evaluation. Health services should therefore treat the staffing model—not only the language model—as part of the evidence.[2][4]
Diagnostic matches were encouraging, but context was unequal
Eight weeks after each encounter, reviewers determined a final diagnosis from the medical record. AMIE’s differential-diagnosis list contained that diagnosis somewhere in 88 of 98 cases, or 89.8%, and within its first three suggestions in 73 cases, or 74.5%. For the 46 cases confirmed by laboratory, microbiology, pathology or imaging, performance remained high, although the subset was small.
Blinded clinical evaluators found no statistically significant difference between AMIE and primary-care providers in overall differential-diagnosis quality, or in the appropriateness and safety of management plans. But clinicians’ plans were rated significantly better for practicality and cost effectiveness. The AI had no access to the electronic record, could not examine the patient and received text only; the clinicians had richer information and responsibility for actual care. That makes the comparison informative about generated lists, not a head-to-head trial of independent clinical performance.
A final diagnosis appearing somewhere in a long list is also not the same as choosing the correct action. The potential value is preparatory: identifying history to verify and possibilities to discuss. Evidence of benefit would require showing that this preparation reduces missed information, shortens work without shifting burden to monitors, or improves decisions and outcomes compared with a well-defined usual-care workflow.[1][2]
Clinicians saw preparation value; the denominator was smaller
Treating clinicians completed 60 post-visit surveys, but 16 reported that they had not reviewed the AMIE transcript or summary before the appointment. The relevant denominator for questions about preparation was therefore 44. In 33 of those 44 cases—75%—clinicians said the material helped them prepare, and in 25—57%—they said it may have changed their approach.
Those are self-reported impressions, not measured time savings or demonstrated improvements in care. The study did not quantify how long the AI chat took, how much supervision cost, whether clinicians spent less time documenting, or whether reviewing the transcript added work. It also did not measure follow-up, adverse events, referrals, testing, prescribing or symptom resolution.
Patients generally rated listening and explanation favourably, and attitudes towards AI became more positive after the chat. Yet trust was not complete. Participants expressed particular concern about confidentiality and about the system’s honesty or trustworthiness. Those concerns are central when symptoms, diagnoses and personal history are handled by a company-built model, and they cannot be answered by conversational quality scores alone.[2][4]
What would change the assessment
The next step should be a randomised, multi-centre comparison against a clearly described pre-visit workflow. It should include older adults, people using interpreters, disabled patients, lower-literacy groups and a wider range of complaints. Outcomes should cover clinically important omissions and false statements, patient anxiety, urgent escalation, diagnostic changes, tests and referrals, visit duration, after-hours work, total staff time and costs.
Researchers should also compare realistic levels of oversight. A dedicated board-certified physician watching every chat may be appropriate for an early study but could erase the workload benefit at scale. Trials need prespecified rules for when automated alerts, asynchronous review or immediate human takeover are required, and they should report both missed interventions and unnecessary escalations.
The study was funded by Alphabet, Google’s parent company. The authors disclosed that one senior investigator served as a visiting researcher at Google during part of the work, with full disclosures in the paper. Commercial sponsorship does not invalidate the result, but it increases the importance of independent replication, transparent error reporting and access to protocols. Until those tests exist, the evidence supports supervised research and carefully bounded pilots—not autonomous patient-facing diagnosis.[1][2][3][4]
What this means for people
- Patients may arrive with a more structured history and clearer questions, but should know that the chatbot is experimental and does not replace examination or urgent-care advice.
- Clinicians may gain a useful pre-visit summary, although review and correction can create work and responsibility that must be measured rather than assumed away.
- Health-service leaders need to cost the human oversight, privacy controls and escalation pathway as part of the intervention, not as optional extras.
Global context
The evidence comes from one well-resourced academic clinic in Boston using English-language text chat and live physician oversight. Primary-care access, digital literacy, liability, privacy law, staffing and emergency pathways differ internationally. A workflow that is feasible in this setting may be unaffordable or unsafe where clinicians cannot monitor conversations or promptly examine referred patients.
What the evidence does not yet show
- This was a prospective but single-arm feasibility study at one academic primary-care clinic; it cannot show that AMIE improved outcomes compared with usual care.
- Every chat was watched live by a board-certified physician, so the result does not apply to unsupervised use or prove that the model itself was safe.
- The study enrolled 114 adults, 100 completed the AI interaction and 98 completed the visit; the sample skewed younger than the clinic’s wider urgent-care population.
- Diagnostic matching was retrospective, and having the final diagnosis somewhere in a list is not evidence that the system selected safe or useful treatment.
- The research was funded by Alphabet and involved Google researchers; independent replication and full cost measurement remain necessary.
What to watch next
- Randomised, multi-centre trials comparing the full AI-plus-supervision workflow with usual pre-visit history taking.
- Patient outcomes, clinically weighted errors, anxiety, escalations, testing, referrals and follow-up—not only ratings and diagnostic-list overlap.
- Total clinician and supervisor time, documentation work, costs and whether any savings persist outside an intensively monitored study.
- Performance for older adults, interpreters, lower-literacy users, disabled patients and complaints that require examination or urgent escalation.
- Independent replication, privacy audits and clear rules for data retention, human takeover and responsibility when the chatbot is wrong.
Living evidence record
Impact record IAI-0UNEHAH
Evidence stage
Studied
Confidence
Corroborated
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
10 October 2026
Source trail
4 direct sources across 2 source types.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 4 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
Evidence trail
Sources used for this report
Links checked 10 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Health & Life Sciences
Did clinicians prefer AI discharge summaries after long hospital stays?
In a retrospective 60-case comparison, 12 physicians usually preferred GPT-5.2 summaries and annotated fewer omissions. Reviewers knew which summary was AI-written, one hospital supplied the records, and no patient outcome or time saving was tested.
7 min · 3 sources
Health & Life Sciences
Can a sleep-study ECG predict future heart risk?
A US study of 38,195 sleep-clinic patients found that a neural-network score added useful risk information for atrial fibrillation, heart failure and death. The gains were smaller for stroke and heart attack, and the retrospective study did not test whether using the score improves care.
8 min · 1 source
Health & Life Sciences
Does primary-care AI improve outcomes?
Not yet on the available evidence. A peer-reviewed review found 10 real-world studies of clinician-facing AI in primary care: some improved detection or care processes, but neither trial measuring patient-important outcomes demonstrated benefit, and the evidence was low or very low certainty.
10 min · 2 sources
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.