Back to the news portal
Health & Life SciencesResearch paperResearchSource analysisUnited StatesNorth AmericaGlobal health

Can EHR-note AI help distinguish epilepsy from PNES?

EpiScreen separated epilepsy from psychogenic non-epileptic seizures across 13,633 records from two US datasets and improved accuracy in an eight-clinician simulation. It remains retrospective decision support, not a substitute for neurological assessment or video-EEG.

By The Impact of AI Editorial DeskReleased 8 October 2026 at 18:57 BST8 min read2 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

Share
Social links
LinkedInXBlueskyRedditEmail

At a glance

  • 1The retrospective study analysed 3,451 MIMIC-IV records and 10,182 University of Minnesota records, with 10,662 epilepsy cases and 2,971 PNES cases in total.
  • 2Cross-institution AUC ranged from 0.829 to 0.875 when models trained at Minnesota were tested on MIMIC-IV, and from 0.966 to 0.980 in the reverse direction.
  • 3Eight clinicians reviewing 300 records improved from 0.763–0.817 accuracy without assistance to 0.877–0.893 with model predictions and sentence attributions, but the retrospective simulation did not test prospective decisions or patient outcomes.
Key themesEpilepsyPNESElectronic health recordsLarge language modelsDecision supportExternal validation

Research topic

Cross-institutional large-language-model screening of epilepsy versus psychogenic non-epileptic seizures from clinical notes

The Impact of AI research cover asking whether EHR-note AI can help distinguish epilepsy from PNES, with conceptual document layers and decision-support paths; it notes 13,633 records, two US datasets, eight clinicians and that the tool is not a diagnosis.
AI-generated editorial illustration. The document layers, AI processor, signal paths and neurological symbols are conceptual; they do not depict a patient, medical record, provider interface, brain scan, clinical diagnosis or approved product.

The direct answer: the notes contained a useful signal, but the system is not a diagnosis

A language model could distinguish epilepsy from psychogenic non-epileptic seizures, or PNES, in two retrospective US record sets and appeared to help clinicians in a controlled note-review exercise. The EpiScreen study used 3,451 MIMIC-IV records—2,651 labelled epilepsy and 800 PNES—and 10,182 University of Minnesota records—8,011 epilepsy and 2,171 PNES. Models were trained at one institution and evaluated at the other, an important test of whether performance survives different writing and care practices.

The strongest cross-institution result for models trained at Minnesota and tested on MIMIC-IV was an area under the curve of 0.875, with a 95% confidence interval from 0.848 to 0.899. In the reverse direction, the best AUC was 0.980, with a 95% interval from 0.974 to 0.985. These are discrimination results, not proof that EpiScreen identifies the cause of an individual event. The intended use is to support prioritisation before confirmatory video electroencephalography, not replace it.[1]

What was compared and how the records were labelled

Candidate cases were first found through diagnostic codes, then notes were manually reviewed so the retained label reflected the treating clinician’s primary discharge diagnosis. Patients documented with both epilepsy and PNES were excluded. The authors processed notes to represent time points before the final diagnosis, split each dataset 70% for training, 10% for validation and 20% for testing, and fine-tuned open-source models with quantised low-rank adaptation. Baselines included a feature system, BERT, ClinicalBERT and several models used only through prompting.

That pipeline is more demanding than testing a model on notes from the same institution, but it does not eliminate label leakage or circularity. Clinical narratives can contain preliminary suspicions, tests and treatment decisions that reflect the diagnostic pathway. The final label was usually a documented clinical conclusion, not video-EEG confirmation for every case. The study did run sensitivity analyses on 679 University of Minnesota test cases confirmed by long-term video-EEG and on a separate 300-case sample retaining potentially misleading preliminary information; the authors report consistent performance, but full prospective uncertainty remains.[1]

Cross-site performance was stronger than simpler comparators

When trained on Minnesota data and tested on MIMIC-IV, four fine-tuned EpiScreen variants produced AUCs from 0.829 to 0.875. The strongest simpler comparator, ClinicalBERT, reached 0.749. When trained on MIMIC-IV and tested at Minnesota, the EpiScreen variants ranged from 0.966 to 0.980, while ClinicalBERT reached 0.788. Five-fold analyses were also reported, and performance declined as model size fell or class imbalance became more severe. Those checks reduce the chance that one fortunate split explains the headline result.

The asymmetry between sites deserves attention. Minnesota was the larger, less imbalanced cohort and produced much higher AUCs when used as the external test set than MIMIC-IV did. Record composition, case difficulty, note structure and labels may all contribute. AUC also summarises ranking across thresholds; a hospital still needs a calibrated operating point tied to a specific action, such as expedited specialist review. The study does not report a live alert rate, staffing burden or the consequences of acting on false positives and false negatives.[1]

Eight clinicians improved in a 300-record review

Eight clinicians took part in the human–AI experiment: six neurology residents divided into three pairs and two senior neurologists. Each pair reached a consensus on a stratified sample of 300 MIMIC-IV test cases. Without AI support, the three resident groups recorded accuracies of 0.793, 0.763 and 0.767; the senior group reached 0.817. With EpiScreen’s predicted label and sentence-level Integrated Gradients attributions, the corresponding accuracies rose to 0.880, 0.890, 0.877 and 0.893. The largest absolute gain was 14.3 percentage points.

Assistance reduced false negatives from 36 to 14 in the illustrated resident comparison and false positives from 37 to 19 for the senior neurologists. This is promising evidence that a prediction plus highlighted evidence may work better than a score alone. It remains a small retrospective simulation in which clinicians reviewed prepared notes and could use internet resources. It did not measure whether explanations caused appropriate trust, whether clinicians could detect misleading model rationales, or whether real decisions about antiseizure treatment and video-EEG scheduling improved.[1]

The clinical boundary is especially important for PNES

Epilepsy and PNES can share descriptions such as shaking, loss of consciousness and unresponsiveness, while no single history feature settles the diagnosis. Misclassification has consequences in both directions: a person with epilepsy may wait for treatment, while a person with PNES may receive unnecessary antiseizure medication and delayed appropriate care. A screening model trained on narrative patterns could prioritise review, but it could also learn clinicians’ prior assumptions or documentation styles and reproduce stigma around functional neurological disorders.

The comparison group mainly contained PNES, not the wider range of seizure mimics such as syncope and sleep disorders. Real patients can also have both epilepsy and PNES, yet the study excluded those cases to create a clean binary label. That makes the reported task narrower than routine differential diagnosis. Any interface should show uncertainty, preserve the underlying evidence, support an abstain option and make clear that PNES is a genuine condition requiring respectful communication and appropriate care—not simply the absence of epilepsy.[1]

What would change the assessment

The decisive next test is prospective, silent evaluation in multiple health systems using a locked model and notes as they actually appear before confirmation. Investigators should prespecify thresholds, report calibration, missingness, subgroup performance, data drift and how often the model declines to answer. A later randomised or stepped-wedge trial could assess time to video-EEG, diagnostic delay, medication exposure, clinician workload, patient experience and errors across epilepsy, PNES, mixed diagnoses and other mimics. Secure local processing and governance are essential because the input is sensitive longitudinal text.

The peer-reviewed paper appeared on 8 October 2026, following an unreviewed March preprint; this article treats the journal version as the current evidence rather than presenting the underlying project as newly begun. The study was supported by three US National Institutes of Health grants. Rui Zhang is an associate editor of npj Digital Medicine but was not involved in the journal’s review or decision; the other authors declare no competing interests. EpiScreen is a substantial cross-site result with a useful human-assistance signal. It is not yet evidence of safer or faster diagnosis in live care.[1][2]

What this means for people

  • Earlier prioritisation could shorten uncertainty for people waiting for specialist assessment and video-EEG.
  • False labels could expose patients to unnecessary medication, missed epilepsy care or stigma around functional neurological symptoms.
  • Clinicians need access to the source evidence and authority to reject or defer a model output; patients need transparent data-use safeguards.

Global context

Both datasets came from US institutions, one from critical-care records and one from a Minnesota academic health system. Documentation, referral timing, language, video-EEG capacity and diagnostic coding differ internationally. Cross-institution transfer within the United States is stronger evidence than a single-site split, but it does not establish performance in community neurology, paediatrics, emergency care or lower-resource systems. Local prospective validation must precede deployment.

What the evidence does not yet show

  • The study was retrospective and did not test real-time workflow, treatment decisions or patient outcomes.
  • Most labels were manually confirmed discharge diagnoses rather than video-EEG gold standards, although the authors ran a video-EEG-confirmed sensitivity analysis.
  • The binary task excluded people with both epilepsy and PNES and did not include major alternative seizure mimics such as syncope or sleep disorders.
  • Clinical notes may encode prior diagnostic assumptions, documentation bias, institutional practice and sensitive information.
  • The clinician experiment involved eight clinicians and 300 prepared records; it was not a prospective clinical trial.

What to watch next

  • Prospective external validation using live, minimally curated notes before definitive testing.
  • Calibration and operational thresholds tied to video-EEG access, staffing and acceptable error costs.
  • Performance for mixed epilepsy/PNES diagnoses, other seizure mimics and demographic and language groups.
  • Patient-centred outcomes, privacy controls and evidence that explanations prevent both automation bias and underuse.

Living evidence record

Impact record IAI-190RTSL

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent or research support

Present

Record status

Monitoring

Last checked

8 October 2026

Source trail

2 direct sources across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Related-source reporting disclosure

This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.

Evidence trail

Sources used for this report

Links checked 8 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Health & Life Sciences

Can brain MRI reveal more than BMI?

A peer-reviewed study trained a deep-learning model on 45,702 MRI scans from six cohorts. Brain-derived features tracked BMI and separated several disease groups better than BMI alone—but the disease analysis stayed inside UK Biobank and cannot establish cause, diagnosis or clinical benefit.

9 min · 2 sources

Health & Life Sciences

Did clinicians prefer AI discharge summaries after long hospital stays?

In a retrospective 60-case comparison, 12 physicians usually preferred GPT-5.2 summaries and annotated fewer omissions. Reviewers knew which summary was AI-written, one hospital supplied the records, and no patient outcome or time saving was tested.

7 min · 3 sources

Health & Life Sciences

Can AI reliably predict facial growth?

A registered systematic review found 12 studies of AI-based craniofacial or mandibular growth prediction, with samples from 33 to 639 people. Seven studies were at high risk of bias and five at unclear risk; none had independent external validation.

9 min · 1 source

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Reader commentary

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Explore commentary across the portal →

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.