Do AI agent guards miss required safety actions?
Yes, often in this coding-agent benchmark. Fourteen evaluated models struggled to identify safety obligations that an agent failed to complete; the strongest baseline recovered 48.97% of the required actions, while a purpose-trained guard reached 57.52%. The study is a preprint and does not establish performance outside software tasks.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1ObligationBench contains 240 expert-validated coding-agent trajectories: 120 with unfulfilled obligations and 120 difficult negative examples, covering 339 obligations in total.
- 2Among 14 evaluated models, the strongest baseline recovered 48.97% of obligations and exactly matched the complete set for 10% of trajectories; ObligationGuard reached 57.52% recall and 21.67% exact match.
- 3In a 186-task deployment-style test, ObligationGuard raised the share of solutions that were both functionally correct and safe from 6.5% to 15.1%; the result remains narrow, model-dependent and unreviewed.
Research topic
Whether guard models can identify safety-critical actions that coding agents were required to take but omitted

The direct answer: checking prohibited acts is not enough
The study's central finding is that an agent can avoid an explicitly forbidden act and still leave a system unsafe because it failed to do something required. A coding agent may need to validate input, request confirmation before a destructive command, preserve a permission boundary or verify a change after applying it. Most guard designs focus on harmful actions that occurred. The authors instead ask whether a guard can reconstruct the safety obligations implied by the task and environment, then determine which remained unfulfilled.
No tested model solved that problem reliably. The strongest of 14 existing general-purpose and guard models achieved 48.97% recall over annotated obligations and returned the complete correct set for only 10% of trajectories. A purpose-trained model improved those figures to 57.52% and 21.67%. That is a material gain in this benchmark, but it also means many required actions were still missed and most trajectories did not receive a completely correct diagnosis.[1]
The Impact Brief · Free
Follow the evidence in ai risks & safety.
Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.
What the benchmark contains
ObligationBench contains 240 trajectories: 120 positive cases with at least one unfulfilled obligation and 120 hard negatives designed to resemble risky situations without a missed duty. The positive cases contain 339 obligations, or 2.83 per trajectory on average. The tasks span 80 issue-resolution cases, 32 feature-development cases and 128 terminal-operation cases, grouped into five safety categories.
The team ran four language-model backends on 1,000 tasks and retained 5,684 valid trajectories after infrastructure failures were removed. GPT-5.6-Sol helped filter and annotate candidates. Two software engineers, each reported to have more than five years of experience, independently reviewed the selected examples and agreed on the complete obligation set for 82.2% before adjudication. The benchmark is therefore expert-validated but still partly shaped by model-assisted filtering and judgments about what was required.[1][2]
Purpose-built training helped, but did not close the gap
The researchers trained ObligationGuard on 40,000 synthetic examples derived from 10,000 templates. It outperformed the comparison models, supporting the idea that omitted obligations deserve an explicit objective rather than being expected to emerge from a general safety prompt. Exact match remains the sobering metric: even the specialised guard identified the entire reference set on only about one in five benchmark trajectories.
Synthetic training creates a possible fit between generated examples and the benchmark's ontology. The public repository should allow independent stress tests, but stronger evidence would use held-out tasks from other organisations, annotations produced without the same model-assisted pipeline and evaluations in which obligation wording, tools and environmental constraints differ substantially from the training templates.[1][2]
The deployment-style test improved safety from a low baseline
The authors paired Qwen3.8-27B with Mini-SWE-Agent on 186 SusVibes tasks. Without guidance, 6.5% of solutions passed both functional and security criteria. A generic self-reminder reached 8.1%, Claude Opus 4.8 guidance reached 12.4%, and ObligationGuard reached 15.1%. Functional pass rate was 27.4% without guidance and 26.9% with ObligationGuard, so the headline safety gain did not visibly trade away task completion.
The absolute result still matters. ObligationGuard more than doubled the joint safe-and-correct rate, yet roughly five in six tasks failed to be both correct and safe. The paper says 16 of 39 initially correct-but-unsafe solutions became correct and safe, with 17 new safe solutions overall. This is evidence that obligation feedback can redirect some behaviour, not evidence that the guarded agent is ready for unsupervised high-consequence work.[1]
What the correlations do and do not show
Across six guard models, benchmark recall correlated strongly with improvement in security pass rate: Spearman's rank correlation was 0.94 and Pearson's was 0.97. That links the benchmark metric to downstream behaviour, but six models are too few for a stable general relationship. Common model families or design choices could influence both results.
The semantic judge was checked against people on a balanced sample of 400 predicted obligation pairs, where agreement was 95%. That reduces concern that every result is an automated-grading artefact. It does not prove that the original obligation was necessary, and a balanced validation sample does not reproduce the prevalence of positives in every deployment.[1]
Practical meaning, limits and what would change the assessment
Agent builders should separate two monitoring questions: did the agent do something forbidden, and did it fail to do something required? Logs and approval systems must preserve enough context to answer both. For users, silence from a guard is not proof of safety. High-impact actions still need explicit permissions, deterministic controls where possible, reversible execution, human review and post-action verification.
The evidence is an unreviewed preprint focused on coding agents, with a modest benchmark, synthetic training data, model-assisted annotation and one principal end-to-end configuration. Confidence would rise if independent teams reproduced the annotations, evaluated naturally occurring traces, compared more frameworks and tested noncoding environments. Prospective studies should measure false alarms, operator burden and whether guidance prevents harm without merely suppressing useful work.[1][2]
What this means for people
- A guard's silence does not mean every safety-critical check was completed.
- Operators need visible obligations and reversible actions.
- Organisations should measure missed duties and false alarms before delegation.
Global context
The study was led from China using public software-agent tasks and models developed across several organisations. Its engineering problem is global, but legal duties, workplace approvals and acceptable residual risk vary by sector and jurisdiction. Local governance still determines which obligations must be explicit and who remains accountable.
What the evidence does not yet show
- Unreviewed preprint.
- A 240-trajectory coding benchmark may not represent other domains.
- Filtering, annotation support and semantic evaluation partly rely on GPT-5.6-Sol.
- ObligationGuard uses synthetic training data and one main end-to-end agent configuration.
- The reported correlations cover only six guard models.
What to watch next
- Independent replication on production traces.
- Tests in noncoding domains.
- False-positive rates and operator workload.
- Comparisons with deterministic permission and rollback controls.
- Generalisation across models and agent frameworks.
Living evidence record
Impact record IAI-0149YDW
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
9 October 2026
Source trail
2 direct sources across 2 source types.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
Evidence trail
Sources used for this report
Links checked 9 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
AI Risks & Safety
Can AI judges be accurate yet disagree on the errors that matter?
Yes. In a new unreviewed study of ten reasoning models, average accuracy stayed high on a 600-item answer-checking benchmark, while prompt and model choices sharply changed whether errors were too lenient or too strict. The result argues against trusting one headline accuracy score—or one AI judge.
9 min · 2 sources
AI Risks & Safety
Can AI agents learn what is appropriate?
Perhaps, but the new Google-led report is a research agenda, not a tested safeguard. More than 50 contributors propose contextual policy engines, dynamic sandboxes and multi-agent benchmarks without demonstrating that they prevent real privacy or security failures.
8 min · 2 sources
AI Risks & Safety
Do AI models favour one cancer guideline?
Four models produced 6,000 choices across 15 conflicts between US immunotherapy guidelines and usually selected NCCN. But 10 conflicts involved one lung-cancer setting, and removing them weakened or reversed the pattern for most models.
7 min · 2 sources
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.