When should AI autism-screening results reach families?
Twenty-six US participants favoured earlier support but warned against opaque EHR predictions arriving before useful next steps. The study maps implementation concerns; it does not validate a screening model.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1The US qualitative study used semi-structured interviews and focus groups with 8 caregivers of young autistic children, 10 clinicians and 8 autistic adults: 26 participants in total.
- 2Participants saw possible value in earlier identification and services but raised concerns about incomplete EHR data, black-box models, emotional burden and screening that bypasses caregiver involvement.
- 3The study elicited perspectives on a proposed implementation context; it did not evaluate a model's sensitivity, specificity, fairness, clinical utility or effect on waiting times and outcomes.
Research topic
How caregivers, autistic adults and clinicians think AI-based electronic-health-record prediction should be communicated and governed in early autism screening
The answer: early enough to help, but not before people can explain what happens next
Across interviews and focus groups, caregivers, clinicians and autistic adults expressed cautious optimism about electronic-health-record models that might flag a child's elevated likelihood of autism. They associated earlier identification with clearer understanding of developmental concerns and earlier access to services. But they did not treat 'earlier' as automatically better. Participants wanted a result to arrive when it could support action, explanation and shared decisions rather than leave a family with an opaque warning and no route forward.
That is an implementation finding, not evidence that a particular AI model works. The study recruited 26 people in the United States—eight caregivers of young autistic children, ten clinicians involved in screening or early intervention and eight autistic adults. It used thematic analysis of semi-structured interviews and focus groups. No prediction model was evaluated, no child was screened by the research and no diagnostic or service outcome was measured.[1]
The Impact Brief · Free
Follow the evidence in health & life sciences.
Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.
Three groups contributed different forms of expertise
The study deliberately included people who would encounter a prediction from different positions. Caregivers could speak to uncertainty and navigation of services; clinicians could describe workflow and communication pressures; autistic adults could assess the meaning and possible consequences of being identified through an automated process. Combining those perspectives helps move an ethics discussion beyond accuracy alone, although 8, 10 and 8 participants remain small purposive groups rather than a representative survey.
Researchers analysed the discussions thematically to identify concerns, sources of trust and preferences for communicating likelihood within shared decision-making. The paper was supported by US National Institutes of Health grants. It also discloses commercial and professional relationships for some authors, and notes that the academic medical organisation employing most authors was planning implementation of an AI prediction model considered in the study, with some authors involved. That context makes the qualitative evidence especially relevant but requires transparency about potential implementation interests.[1]
Incomplete records can turn a technical score into an equity problem
Participants worried that electronic records are not neutral inventories of a child's development. They reflect who has access to care, which concerns a clinician records, how consistently information moves between services and whether previous encounters captured relevant behaviour. A model trained on those records may therefore learn gaps as well as patterns. Families with fragmented access or clinicians working in under-resourced settings could receive less reliable scores even if average performance appears strong.
A credible deployment would need subgroup evaluation, documentation of missingness and a way for families and clinicians to correct or supplement the source information. A likelihood score should not silently become a diagnosis, nor should a low score block referral when a caregiver or clinician remains concerned. Equity monitoring must measure what happens after the alert—who receives an evaluation, how long they wait and whether services are available—not only whether the model's retrospective predictions match a recorded outcome.[1]
Transparency has to be useful in the consultation
Participants raised the limits of black-box systems and the risk that an automated flag could reduce caregiver involvement. A generic statement that 'AI found a pattern' is not enough for shared decision-making. Clinicians need to know the intended population, validation setting, timing, uncertainty and major data limitations. Families need plain language about what the score does and does not mean, why further assessment is being suggested, what choices remain open and how to ask for review.
Timing is part of that explanation. Delivering a probability through a portal without context could increase anxiety, especially if specialist assessment is months away. Holding every result until certainty would erase the potential advantage of earlier screening. The interviews point towards a supported disclosure process: a trained professional discusses the signal, considers the child's broader context, invites the caregiver's observations and connects the result to a concrete next step. The model should inform that conversation rather than substitute for it.[1]
What health services should test before implementation
A health service would first need external validation on records that resemble its own population and data quality. Predefined measures should include sensitivity, specificity, calibration and positive predictive value at the intended threshold, with confidence intervals and subgroup analysis. Because autism prevalence and referral pathways affect predictive value, results from one health system cannot be copied directly into another. Prospective silent testing can reveal data failures before scores reach clinicians or families.
The workflow then needs its own evaluation. Services should track additional assessments, false reassurance, clinician overrides, caregiver understanding, distress, time to diagnosis, time to support and unequal access. They should specify who is accountable for reviewing an alert and what happens if no appointment is available. Consent and notice require careful design because EHR-based prediction may run before a family actively requests an autism assessment. Governance should include autistic people and caregivers beyond a single consultation exercise.[1]
Limits and what would change the assessment
Qualitative work can explain concerns and preferences in depth, but it does not estimate how prevalent a view is. Twenty-six US participants cannot represent the diversity of autistic people, families, languages, health systems or views about early identification. Participants responded to the concept and implementation of AI clinical prediction; their views may change after encountering a real score, a false positive, a delayed referral or a service that offers meaningful support. The study also cannot compare alternative consent or communication designs quantitatively.
Confidence in an implementation pathway would rise if diverse sites co-designed and prospectively compared communication approaches alongside model validation. Researchers should include minimally speaking autistic people through accessible methods, families who face barriers to care and communities underrepresented in EHR datasets. Evidence that supported disclosure improves understanding and timely access without increasing inequity or avoidable distress would strengthen the case. Evidence of poor calibration, service bottlenecks or reduced caregiver voice would argue for redesign or stopping deployment.[1]
What this means for people
- Families need a concrete next step and human explanation, not an unexplained probability in a patient portal.
- Clinicians need authority to incorporate caregiver observations and override a score when the record is incomplete.
- Autistic people should help govern how models frame early identification, support and the meaning of uncertainty.
Global context
The study reflects US electronic records and service pathways. Countries differ in routine developmental screening, diagnostic waiting times, privacy law, insurance, school support and the availability of early services. A technically identical model could therefore produce different consequences. International transfer requires local validation and a service-capacity assessment, not only translation of the interface.
What the evidence does not yet show
- The qualitative sample was 26 people in the United States and was not designed to estimate population prevalence of any view.
- No AI model, threshold, predictive-performance measure or clinical outcome was tested in this study.
- Participants discussed implementation concepts rather than receiving a real prediction in routine care.
- The institution was planning related AI implementation, and some authors were involved; the paper discloses this context and other interests.
- Findings may not transfer to health systems with different records, diagnostic pathways, languages or service capacity.
What to watch next
- External validation with calibration, predictive value and subgroup performance in the intended health system.
- Prospective evidence on caregiver understanding, distress, referral times, service access and clinician workload.
- Accessible co-design with broader autistic communities and families underrepresented in electronic records.
- Clear accountability, correction and appeal processes when an EHR-derived score conflicts with lived observations.
Living evidence record
Impact record IAI-1LGVF6A
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
11 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what BMC Medical Ethics published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
Evidence trail
Sources used for this report
Links checked 11 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Health & Life Sciences
Can AI prioritise cognitive assessment in sleep apnoea?
A random-forest model separated concurrent mild cognitive impairment in an external Chinese cohort, but overestimated probabilities and was evaluated almost entirely in men. It is not a diagnostic or prognostic tool.
6 min · 1 source
Health & Life Sciences
Can EHR-note AI help distinguish epilepsy from PNES?
EpiScreen separated epilepsy from psychogenic non-epileptic seizures across 13,633 records from two US datasets and improved accuracy in an eight-clinician simulation. It remains retrospective decision support, not a substitute for neurological assessment or video-EEG.
8 min · 2 sources
Health & Life Sciences
Can AI flag a longer ICU stay?
A peer-reviewed model separated longer-stay risk across 5,281 retrospective sepsis admissions in one US development cohort and two Korean external cohorts. Its AUROC fell from 0.848 internally to 0.781 at the smaller Korean site, and no prospective study tested whether alerts improve care or capacity.
7 min · 2 sources
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.