Is operating-room AI ready for clinical use?
Not on the published evidence yet. A peer-reviewed scoping review screened 3,020 records but found only one completed feasibility study with five analysed patients; four larger prospective studies had no results posted.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1Researchers searched five bibliographic databases and screened 3,020 records published from 2010 to July 2026; 2,535 unique records remained after 485 duplicates were removed, and five studies met the inclusion criteria.
- 2Only one study had completed clinical results: seven patients were enrolled and five analysed after two motion-artifact exclusions. The other four records were protocols or registries with 466 participants planned but no posted results.
- 3None of the five records reported formal bias audits or subgroup performance. The review's seven-domain Ethics-Implementation Scorecard is illustrative and unvalidated, not a regulatory threshold or proof of deployment readiness.
Research topic
The maturity, validation and ethical implementation reporting of AI-based clinical decision support used during surgery and interventional procedures

The direct answer: clinical interest is real, but the evidence is not deployment-ready
The most important result is the denominator. Luka Parikh and colleagues searched five electronic databases for human studies of AI-based clinical decision support used during surgery or an interventional procedure, starting with 3,020 records. After removing 485 duplicates, two reviewers screened 2,535 titles and abstracts. Eleven full texts or trial-registry records reached eligibility review; six were excluded and only five met the pre-specified criteria. Of those five, just one had completed clinical results, and that study analysed five patients. The remaining four were prospective registries or randomised-trial protocols without posted outcomes when the review was conducted.
That finding does not mean every operating-room algorithm is ineffective. It means there is too little completed, prospective clinical evidence in the review's defined category to support broad claims about patient benefit, safety or reliable transfer between hospitals. The systems covered different tasks—perfusion assessment, surgical-video analysis, anaesthetic dosing, nociception-guided opioid dosing and fluid management—and could not be pooled into a single performance estimate. For a technology intended to influence time-sensitive decisions while a procedure is under way, technical promise and a planned trial are not substitutes for measured clinical outcomes.[1]
What the review actually searched and selected
An information specialist searched MEDLINE, MEDLINE ePubs and In-Process Citations, Embase, the Cochrane Central Register of Controlled Trials, the Cochrane Database of Systematic Reviews and LILACS through the Ovid platform. The searches were first run on 25 November 2025 and updated on 6 July 2026. Records had to concern humans from 2010 onward, place an AI-based decision-support system inside a real surgical or interventional workflow, and describe empirical or planned clinical evaluation of actionable outputs such as alerts, risk predictions or treatment guidance. Retrospective studies, non-human work, endoscopy-suite studies, automation without decision support and non-original publications were excluded.
Two reviewers independently screened records and full texts, with disagreement resolved by consensus and a third reviewer available for unresolved cases. Two reviewers also independently extracted study design, AI modality, clinical context, autonomy, validation, diversity, governance and reported limitations. This is a more disciplined process than a narrative overview, but it remains a scoping review: the aim was to map the field, not to calculate comparative effectiveness. The authors explicitly did not perform a formal risk-of-bias assessment, so the review cannot turn the five heterogeneous records into a ranked list of clinically superior systems.[1]
The only completed study was a five-patient analysis
The sole completed feasibility study used an open-source semi-automated two-dimensional perfusion tool during infrapopliteal angioplasty for critical limb ischaemia in Greece. Seven patients were enrolled, but two were excluded because motion artefacts prevented analysis, leaving five. The software quantified changes in digital-subtraction angiography perfusion measures before and after intervention. The review reports directional technical changes and notes that four of the five analysed patients had limb salvage or wound healing at six months. With no comparator group, a tiny analysed sample and two technical exclusions, those outcomes cannot establish that the software caused the clinical result or would perform reliably across devices, teams and populations.
The other four studies were more ambitious but prospective. An Italian registry planned to recruit 100 people for surgical-video and indocyanine-green analysis in gynaecologic oncology. A Chinese proof-of-concept randomised trial planned 40 participants for reinforcement-learning anaesthetic recommendations. A Chilean registry planned 150 participants for a multimodal physiological 'copilot', while an Italian multicentre randomised trial planned 176 participants for AI-assisted fluid management in major cancer surgery. Together they planned 466 participants, but none had posted results at extraction. A trial protocol is evidence that a question is being tested, not evidence that the intervention works.[1]
Human oversight was common; auditable governance was not
All five systems were advisory rather than autonomous. Four explicitly kept clinicians in the loop, and the fifth described a decision-support platform without claiming independent control. No system executed a closed-loop clinical action without clinician confirmation. That is an important safeguard, but human presence alone does not show appropriate reliance. None of the records measured clinician trust or adherence, logged override behaviour as an outcome, or tested whether staff could recognise a wrong or uncertain recommendation under operating-room pressure. A nominal right to override is only useful if the interface, workflow and institutional policy make an override practical and visible.
Reporting was particularly thin on data access, retention, auditing, model updates and responsibility for an AI-mediated recommendation. Four records described some form of consent, but the review found no formal subgroup-performance analysis, fairness assessment or bias audit in any study. No record reported calibrated uncertainty displays or comprehensive operational safety-monitoring frameworks. Those omissions matter because surgical video, physiological streams and invasive blood-pressure data can be sensitive, and because an average performance score can hide failures concentrated in a patient group, device type or clinical setting.[1]
What the proposed scorecard can—and cannot—do
The authors translate existing medical-AI reporting and governance guidance into an Ethics-Implementation Scorecard covering seven domains: safety and risk management; autonomy and human oversight; transparency and explainability; regulatory compliance and auditability; equity and representativeness; governance readiness; and clinical translation readiness. Each domain contains three operational elements and is counted as present when at least two are reported. Applied to the five records, the exercise highlighted relatively stronger descriptions of intended function and clinical translation, and weaker reporting of regulatory compliance, auditability and governance.
The scorecard is a reporting framework proposed by the same team that conducted the review. It has not been externally validated, is not a regulatory standard and does not define an acceptable score for deployment. The authors describe it as hypothesis-generating. Evidence that would change the assessment includes completed prospective multicentre trials reporting patient outcomes and failures, external validation across devices and hospitals, subgroup results, documented override behaviour, pre-specified safety monitoring and operational data-governance procedures. Until then, the review supports careful trials and transparent reporting—not routine adoption on the assumption that a human clinician will absorb every residual risk.[1]
What this means for people
- Patients may eventually benefit from additional real-time information during surgery, but current published evidence is too sparse to quantify that benefit or its risks.
- Clinicians remain responsible for decisions in the systems reviewed, yet the studies did not measure whether staff used, challenged or ignored the recommendations appropriately.
- Hospitals considering trials need governance for sensitive video and physiological data, traceable updates and a clear account of responsibility when system advice conflicts with clinical judgment.
Global context
The five eligible studies were located in Greece, Italy, China and Chile, with none from North America, Africa or Oceania despite the review team being based in Canada. That distribution is descriptive rather than a census of every surgical-AI project, because the eligibility criteria required real-time clinical decision support and excluded retrospective development work. It nevertheless shows how little published prospective evidence spans regions. Local devices, clinical teams, patient populations, consent rules and regulatory systems may all alter performance and acceptability, so results from one hospital or country should not be treated as global validation.
What the evidence does not yet show
- A scoping review maps evidence but does not estimate a pooled effect or determine comparative effectiveness; no formal risk-of-bias assessment was performed.
- Only five heterogeneous records met the inclusion criteria, and four were protocols or registries without results.
- The single completed study analysed five patients after two motion-artifact exclusions and had no comparator group.
- The review's search ended on 6 July 2026, so subsequently posted trial results would not be included.
- The proposed scorecard was developed and applied by the review authors and requires external validation before operational or regulatory use.
- One author disclosed consultancy work with Johnson & Johnson; the other authors declared no competing interests. The article page does not list a separate funding statement.
What to watch next
- Results from the 100-participant SLN registry, 40-participant RL-PRAIS trial, 150-participant SEASCAPE registry and 176-participant FOCUS-AFM trial.
- Prospective patient-safety and clinical-outcome results rather than only technical accuracy or feasibility measures.
- External validation across hospitals, devices, specialties and patient subgroups, with formal bias and fairness analysis.
- Measured clinician reliance, overrides, usability and recognition of incorrect or uncertain recommendations.
- Independent validation of the proposed Ethics-Implementation Scorecard and clearer regulatory and audit expectations.
Living evidence record
Impact record IAI-1OK5UFB
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
8 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what npj Digital Surgery published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
Evidence trail
Sources used for this report
Links checked 8 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Health & Life Sciences
Can clinical AI recognise when the patient record does not support an answer?
A new clinical-agent benchmark raises a practical question for health systems: can an assistant explain what the record cannot establish? Our analysis examines evidence, local testing and the burden on staff.
6 min · 2 sources
Health & Life Sciences
Can a medical AI support a clinical board without becoming the decision-maker?
A peer-reviewed German study put OpenEvidence into discussions of 100 real inflammatory-disease cases. Specialists found the answers useful, but complete agreement with the board was limited, prompt tuning was inconclusive and outputs changed over time. The study tested workflow feasibility—not patient benefit or safety.
9 min · 1 source
Health & Life Sciences
Did clinicians prefer AI discharge summaries after long hospital stays?
In a retrospective 60-case comparison, 12 physicians usually preferred GPT-5.2 summaries and annotated fewer omissions. Reviewers knew which summary was AI-written, one hospital supplied the records, and no patient outcome or time saving was tested.
7 min · 3 sources
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.