Can AI grade medical students without hiding mistakes?
A prospective US deployment processed 72,907 OSCE rubric scores for 222 students and sharply reduced human scoring passes, but validation focused on the low-scoring tail and physician adjudicators could see which score came from AI.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1MAPLES produced 72,907 retained item-level scores for 222 students across written notes, audio and video during a Fall 2025 pre-clerkship OSCE deployment.
- 2Within the routed low-scoring review set, tolerant agreement between AI and standardized-patient evaluators ranged from 85.7% to 92.4% by modality.
- 3Physicians matched AI on 76.0% of 616 selected disagreements, but adjudicators saw score provenance; the study was single-centre and did not validate every AI score independently.
Research topic
Whether a routed multimodal AI workflow can perform first-pass scoring of medical-student OSCE evidence while preserving human review of consequential cases

The direct answer: it reduced workload, but did not remove the need for human judgement
A multimodal AI system handled the first scoring pass for a real medical-school examination and routed lower-scoring cases to people, showing that this workflow can operate at meaningful scale. In the Fall 2025 deployment at UT Southwestern Medical Center, MAPLES processed written notes, audio and video from 222 students and retained 72,907 rubric-item scores. The reported number of human scoring passes was 92.3% lower than a modelled workflow in which a person scored every retained item once.
That is an operational result, not proof that AI can safely become an autonomous examiner. The study concentrated human validation in the low-scoring tail, where errors could affect remediation or progression. It did not obtain an independent human reference score for every one of the 72,907 items. The defensible conclusion is therefore that AI can triage and score within a supervised local workflow, while the fairness and accuracy of the much larger unreviewed set remain less directly observed.[1]
How the routed assessment worked
MAPLES is a rubric-driven, zero-shot multimodal large-language-model pipeline. It scored post-encounter notes, audio-derived communication evidence and video evidence of physical examination against item-specific criteria. Students whose aggregate scores fell below prespecified thresholds were routed for an independent rescore by standardized-patient evaluators. Those evaluators were blinded to AI scores, model rationales, aggregate scores and the reason for routing, and did not rescore encounters in which they had acted as the simulated patient.
The routed review contained 4,973 independently rescored items. Tolerant AI–human agreement ranged from 85.7% to 92.4% across modalities. Tolerant agreement allows a defined amount of score difference, so it should not be confused with exact identity. It is relevant to a rubric where adjacent scores may have similar educational meaning, but institutions need to decide in advance which disagreements can safely be tolerated and which can alter a student's outcome.[1]
The physician comparison favoured AI, with an important source of bias
Items outside the tolerance rule were escalated. After excluding disagreements attributed to rubric or data-input problems, the authors retained 616 genuine disagreements for the main physician-adjudication analysis: 376 note items, 207 audio items and 33 video items. Physicians matched the AI score on 76.0% of those items and the standardized-patient evaluator on 19.0%; the remainder matched neither. This suggests that AI often identified scoring evidence missed or interpreted differently by the first human reviewer.
However, physicians could see both scores and their sources. The adjudication was therefore selected and non-blinded. Knowing which value came from AI can influence judgement in either direction, especially in a team that designed and deployed the system. The 616 items also represent disagreements from a routed subset, not a random sample of all outputs. A stronger test would hide score provenance, sample cases across the full score distribution and use several independent adjudicators with prespecified tie-breaking rules.[1]
A 92.3% workload reduction is a modelled comparison
The AI-assisted workflow required 5,616 human scoring passes: 4,973 standardized-patient rescoring passes and 643 physician adjudication passes. The counterfactual comparator assumed one human pass for each of the 72,907 retained items. Against that benchmark, human passes fell by 92.3%, with reported reductions of 94.0% for notes, 88.3% for audio and 88.9% for video. That could release educator time for feedback, remediation and quality assurance.
It is not a measured time-and-cost trial against a parallel manual examination. A human pass on a simple checklist item and a physician adjudication are not interchangeable units of labour, and model operation, data preparation, oversight, appeals and system maintenance also consume resources. The study supports a plausible efficiency gain, but procurement decisions need elapsed time, staff grades, computing costs, failure recovery and the cost of correcting consequential errors.[1]
What this means for students and assessment leaders
For students, the central issue is not whether AI participates but whether the assessment remains contestable and educationally valid. Learners should know which evidence was processed, when a human reviews the result, how to request reconsideration and whether recordings are retained for model development. A low aggregate score should trigger careful review rather than automated labelling. Institutions also need subgroup audits to detect whether accent, speech pattern, disability, skin tone, camera framing or documentation style changes error rates.
Assessment leaders can treat the paper as evidence for a bounded pilot: AI first pass, prespecified routing, blinded secondary review where possible, audit samples from every score band and final human authority over high-stakes decisions. The system was evaluated at one US medical centre in one pre-clerkship OSCE context. Different stations, rubrics, languages, licensing standards and recording conditions could change performance. Local validation is not optional simply because the architecture is described as zero-shot.[1]
Commercial interests and the evidence that should come next
The work was supported by UT Southwestern institutional funds. The authors disclose that UT Southwestern filed patent applications related to AI-based assessment technologies described in the work and that several authors are named inventors. This does not negate the results, but it raises the value of independent replication and blinded adjudication. The authors also describe plans to make tools available for non-commercial academic use, which could help other teams reproduce the workflow.
The assessment would change most with a preregistered multicentre study that independently scores a random sample from the entire distribution, masks adjudicators to provenance and reports exact as well as tolerant agreement, pass/fail changes, appeals, subgroup errors and downstream remediation. A parallel resource evaluation should measure staff time and cost rather than infer them from scoring-pass counts. Until then, MAPLES is credible evidence that supervised AI scoring can be embedded in one medical school—not proof that an AI examiner is ready to make final decisions about students.[1]
What this means for people
- Students could receive faster scoring, but a hidden error could affect remediation or progression if review is too narrowly routed.
- Educators may spend less time on first-pass marking and more on difficult cases, feedback and quality assurance.
- Medical schools need an appealable, auditable process before using multimodal AI in consequential assessment.
Global context
OSCEs are used internationally, but rubrics, examiner practice, licensing consequences, languages and consent rules differ. The US single-centre deployment shows that a supervised multimodal workflow is technically and operationally possible. It does not establish portability to national licensing examinations or institutions without equivalent recording infrastructure, technical support and human adjudication capacity.
What the evidence does not yet show
- The deployment involved one US medical centre and one Fall 2025 pre-clerkship OSCE context.
- Independent human rescoring focused on the routed low-scoring tail rather than a random sample of all 72,907 retained item scores.
- The 616-item physician adjudication set was selected, and adjudicators could see which score came from AI and which from the standardized-patient evaluator.
- The 92.3% reduction compares scoring-pass counts with a modelled single-pass manual workflow, not measured staff time or total cost in a parallel trial.
- Only 33 video disagreements were included in the retained adjudication set, limiting precision for that modality.
- UT Southwestern filed patent applications related to the technology, with several authors named as inventors.
What to watch next
- Multicentre replication across different OSCE stations, institutions, languages and student populations.
- Blinded physician adjudication and audit samples drawn from the full score distribution.
- Exact agreement, pass/fail changes, appeals and subgroup error rates—not only tolerant agreement.
- Measured educator time, computing cost, governance burden and remediation outcomes.
- Clear student rights concerning notice, recording retention, human review and contesting an AI-assisted score.
Living evidence record
Impact record IAI-113TFZ4
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
8 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what npj Digital Medicine published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
Evidence trail
Sources used for this report
Links checked 8 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Education
Can AI make clinical skills exams fairer?
An Oxford study found that an examiner-aware AI second marker improved agreement with a panel-derived reference across 442 ratings from 120 students, especially for borderline fails. That is an audit result in virtual-reality exams—not proof of fairer real-world assessment or a licence for automated grading.
9 min · 1 source
Education
Are Kurdistan universities ready to govern generative AI?
Not visibly, according to a peer-reviewed audit of 30 university websites. Twenty-one institutions were in the study's lowest readiness stage and only two showed direct public GenAI guidance—but the audit measured published evidence in June 2026, not confidential policy or actual classroom practice.
9 min · 2 sources
Education
Can expert checks make AI study materials useful?
A two-cohort economics-course study associated tutor-verified AI materials with fewer marks below a UK degree boundary, but the design cannot isolate verification, rule out cohort differences or prove that use caused the result. It is an unreviewed preprint.
8 min · 1 source
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.