Can AI make clinical skills exams fairer?
An Oxford study found that an examiner-aware AI second marker improved agreement with a panel-derived reference across 442 ratings from 120 students, especially for borderline fails. That is an audit result in virtual-reality exams—not proof of fairer real-world assessment or a licence for automated grading.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1RAVEN combined first-person video, virtual-reality action logs, examiner marks and examiner verbal reasoning in a retrospective analysis of 120 students, 442 ratings and eight assessment domains.
- 2A confidence-gated human–AI hybrid increased pass/fail agreement with leave-one-examiner-out consensus by 6.2 percentage points and borderline-fail concordance by 16.3 points; five of eight domain improvements were statistically significant.
- 3The reference was agreement derived from examiners, not an external truth about competence. The authors position the system for audit and flagging followed by human adjudication, not autonomous grading.
Research topic
Whether a multimodal, examiner-conditioned AI system can improve the consistency and auditability of virtual-reality clinical skills assessment

The result supports a second look, not an automated final mark
The direct answer is that AI may help identify clinical-exam ratings that deserve a second human look, but this study does not show that it makes assessment fairer in the real world. University of Oxford researchers evaluated RAVEN, a multimodal system designed as a second marker for paediatric virtual-reality Objective Structured Clinical Examinations, or OSCEs. The retrospective dataset covered 120 students, 442 ratings and eight scoring domains.
When the system was allowed to contribute only when sufficiently confident, agreement with a leave-one-examiner-out panel reference improved. Pass-or-fail concordance rose by 6.2 percentage points, with a reported p value of 0.0004. Average domain-score agreement increased by 3.0 points, with statistically significant gains in five of eight domains. Borderline-fail cases showed the largest improvement: 16.3 points of added concordance with consensus.
Those are agreement results, not direct measurements of knowledge, patient safety or justice. The reference was derived from other examiners rather than an independent ground truth about each student's competence. RAVEN can therefore be judged on whether it better reproduces a panel-derived decision, while the panel itself can share conventions, blind spots or bias. The paper's defensible use is audit and flagging for human adjudication, not autonomous pass–fail decisions.[1]
Why a second marker could matter
OSCEs test clinical performance through structured stations, but a station is often scored by a single examiner. Training and rubrics reduce variation; they do not eliminate differences in thresholds, attention or interpretation. A student near the pass line can therefore be especially affected by which examiner observed the station. Adding a human second marker to every assessment would improve auditability but is expensive and difficult to scale.
Virtual-reality stations create another source of evidence. The system can record what the candidate did and when, while first-person video captures visual behaviour and the examiner's spoken reasoning supplies context about the mark. RAVEN combines those channels with the original scores and conditions its estimate on the examiner. That is intended to model rater-specific patterns rather than pretending every examiner uses a rubric identically.
A confidence gate is central to the proposed workflow. The AI does not need to replace every judgment; it can remain silent when uncertain and flag cases where the combined evidence supports review. That division of labour is more plausible than universal automated grading, although the paper still needs prospective evidence that the alerts are accurate, understandable and worth the additional adjudication work.[1]
What the researchers compared
The evaluation was retrospective, using data already collected in paediatric VR OSCE marking studies. Its independent human denominator was 120 students, not the larger number of ratings. The 442 ratings arose because students contributed multiple observations across domains or marking encounters. The outcome covered eight domains rather than a single global score, allowing the researchers to ask where agreement changed.
The authors used a leave-one-examiner-out consensus as the reference: the examiner whose result was being assessed did not define the comparison target. They reported chance-adjusted agreement—Gwet's AC1 for pass or fail and AC2 for domain scores—rather than simple percentage agreement alone. That is useful because raw agreement can be misleading when one outcome is much more common than another.
The hybrid joined human and model decisions through confidence gating. The reported 6.2-point pass/fail gain and 3.0-point average domain gain are changes in agreement coefficients, not a percentage reduction in unfairness or exam error. Five domains met the study's statistical-significance threshold; three did not. The abstract does not support treating the average as proof of uniform improvement across every clinical skill.[1]
Borderline cases are promising—and the most sensitive
The 16.3-point gain for borderline-fail ratings is the most practically interesting result. Borderline decisions are where an audit tool could prevent a consequential mark from resting on one ambiguous observation. They are also where changes in threshold, examiner reasoning or cohort composition can produce the largest swings, so the subgroup requires careful replication rather than promotional emphasis alone.
A review flag could help a school route limited second-marker capacity toward contested cases. But it could also create automation bias: an adjudicator might treat the flag as an objective verdict even though it was trained on human marks. Safe deployment would show the underlying evidence, preserve the examiner's opportunity to disagree, document the final reason and monitor whether certain student groups are flagged disproportionately.
The study also compared criteria inferred by the model with the written rubric and criteria articulated by examiners. It found implicit practices that were absent from marking guidance. That can expose useful tacit expertise, but it can also surface undocumented habits that should not become policy. An assessment team must decide whether an inferred criterion is educationally valid and disclosed to students before encoding it in a scoring system.[1]
Fairness was not directly measured
The headline question cannot be answered by overall agreement alone. A fairer examination would need evidence about consistency across examiners, equal opportunity to demonstrate competence, accessibility, subgroup performance, transparency, appeals and alignment with the intended curriculum. The paper reports improved concordance with consensus but does not establish those broader outcomes.
Conditioning on examiner identity may correct some rater variability, yet it can also reproduce an examiner's systematic tendency. If a rater is unusually strict or attends to a feature unrelated to competence, a model that learns that pattern could make the inconsistency more stable. Leave-one-examiner-out evaluation is a useful safeguard against directly copying one mark, but external validation with new examiners, institutions and station designs is still necessary.
The setting also matters. These were paediatric assessments conducted in virtual reality, where actions can be logged precisely. A bedside, simulated-patient or communication-heavy station produces different signals, privacy concerns and failure modes. Results should not be transferred to licensing examinations, other specialties or face-to-face clinical encounters without fresh validation.[1]
What this could mean for students and educators
Students could benefit if the tool gives borderline or internally inconsistent results a structured second review. A transparent audit trail may make appeals more evidence-based and help exam boards discover ambiguous rubric language. The same system could harm students if a model-generated judgment becomes difficult to challenge, if recorded video and voice are reused without clear governance or if performance differs across accents, movement patterns or disabilities.
For examiners, AI could act as quality assurance rather than surveillance, but implementation needs agreement about the purpose. Staff should know which data are recorded, how long they are retained, when a flag is generated and who makes the final decision. Monitoring should distinguish a legitimate difference in expert interpretation from an error and should avoid using disagreement itself as proof that the human is wrong.
For institutions, the economic comparison is not simply AI versus no second marker. It should include system development, secure storage, validation, staff training, adjudication time, appeals and the cost of wrong decisions. A prospective trial could compare normal marking, targeted human second marking and AI-assisted targeted review, measuring consistency, workload, student trust and decision reversals.[1]
Funding, interests and the evidence needed next
The paper is a peer-reviewed accepted article published early on 7 October 2026 and carrying a permanent DOI; the journal notes that editorial changes may occur before the final Version of Record. The work acknowledges Oxford Simulation, Teaching and Research and the Oxford Medical Simulation platform. Funding included an EPSRC Turing AI Fellowship and University of Oxford support. The authors state that funders had no role in design, collection, analysis, publication or manuscript preparation and declare no competing interests.
The next test should be prospective and pre-registered, with the model locked before marking begins. It should recruit new cohorts, examiners and institutions, report the number of students and stations as the primary denominators, compare results with independent expert adjudication and publish confidence intervals. Subgroup analyses should cover characteristics relevant to access and performance, while privacy and appeal processes should be evaluated with students.
Confidence would rise if targeted AI-assisted review reduces unexplained examiner variation without increasing false flags, workload or inequity—and if decisions remain auditable and reversible. Until then, RAVEN is best understood as a promising research prototype that improved agreement with a panel-derived reference in one retrospective VR setting. It does not yet prove fairer exams, better clinicians or safe automated grading.[1]
What this means for people
- Students may receive more consistent review of borderline marks, but an opaque AI flag could create a new barrier to appeal.
- Examiners could gain an audit aid while retaining final judgment; they also need protection from automation bias and inappropriate surveillance.
- Medical schools need evidence that targeted review improves assessment quality and equity before using it in high-stakes decisions.
Global context
The research comes from Oxford and uses paediatric virtual-reality clinical assessments. Medical schools worldwide face examiner variation, but their curricula, licensing rules, languages, technology and accessibility duties differ. Multisite validation should test those differences explicitly; a result in one VR environment should not be treated as a universal assessment standard.
What the evidence does not yet show
- This was a retrospective evaluation in paediatric virtual-reality OSCEs, not a prospective trial in routine high-stakes assessment.
- The reference was leave-one-examiner-out consensus rather than an independent ground truth about clinical competence.
- The independent student denominator was 120; the 442 ratings were repeated observations rather than 442 separate students.
- Agreement improved on average, but only five of eight domain changes were statistically significant and fairness outcomes were not directly measured.
- Transfer to other institutions, examiners, specialties, station types and in-person examinations remains unproven.
What to watch next
- Prospective, pre-registered validation using new students, examiners, institutions and a locked model.
- Comparison with independent expert adjudication and targeted human second marking, not consensus alone.
- Subgroup performance, accessibility, privacy, appeal outcomes and whether any students are disproportionately flagged.
- Effects on examiner workload, decision reversals, student trust and the cost of adjudication.
- Whether inferred examiner criteria are educationally valid, disclosed in rubrics and governed before use.
Living evidence record
Impact record IAI-1A7Z09E
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
7 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what npj Digital Medicine published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
Evidence trail
Sources used for this report
Links checked 7 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Education
Can expert checks make AI study materials useful?
A two-cohort economics-course study associated tutor-verified AI materials with fewer marks below a UK degree boundary, but the design cannot isolate verification, rule out cohort differences or prove that use caused the result. It is an unreviewed preprint.
8 min · 1 source
Education
What should design students know about AI?
Researchers in Guangzhou developed a 23-item scale covering technical skills, tool use, ethics and originality, perceived value, and independent evaluation of AI output. Two Chinese student samples supported the five-factor structure, but the instrument still needs cross-cultural and outcome validation.
8 min · 1 source
Education
Did ChatGPT improve medical training—or the whole teaching package?
A peer-reviewed randomised study at a Chinese hospital found higher examination scores after combining problem-based learning with ChatGPT. Because there was no PBL-only arm, the trial cannot isolate the AI's contribution.
8 min · 2 sources
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.