Back to the news portal
Work & SkillsNew analysis today · source 6 October 2026Primary sourceResearchSource analysisUnited StatesNorth AmericaGlobal
Source record 1. OpenAI

Can AI agents configure complex contract workflows?

OpenAI reports that GPT-6 Astra scored 55.0% across 11 Ironclad tasks versus 41.6% for GPT-5.6 Sol, while simulated time fell from 37.0 to 19.2 minutes. The company-run benchmark shows progress—not autonomous legal work or measured customer productivity.

By The Impact of AI Editorial DeskReleased 7 October 2026 at 08:04 BST7 min read1 source

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

Share
Social links
LinkedInXBlueskyRedditEmail

At a glance

  • 1OpenAI and Ironclad defined 11 legal, commercial and procurement tasks, each assessed against 8 to 50 task-specific criteria.
  • 2GPT-6 Astra averaged 55.0% versus 41.6% for GPT-5.6 Sol; estimated time per attempt fell from 37.0 to 19.2 minutes.
  • 3The tasks, rubrics, training environment and evaluation were company-run, and the time figures are simulations rather than observed employee or customer savings.
Key themesAI agentsLegal operationsContract workflowsComputer useWorkplace automationBenchmarking

Research topic

Whether a frontier computer-use model can configure multi-step contracting workflows more completely and with lower simulated task time than an earlier model in a company-designed evaluation

The Impact of AI research cover asking whether AI agents can configure contract workflows, with conceptual contract pages, approval nodes and a rule-checking path.
AI-generated editorial illustration. The contract, approval workflow and document-check symbols are conceptual and do not reproduce an Ironclad interface, customer contract, legal record or live deployment.

The answer is partial capability, not autonomous legal operations

GPT-6 Astra completed more of the specified requirements than GPT-5.6 Sol on an 11-task contracting benchmark, but its mean rubric score was 55.0%. That means the stronger model still missed a substantial share of the criteria in the evaluation designed by OpenAI with Ironclad. The result supports measurable progress on long, rule-bound computer tasks; it does not support handing an agent unsupervised authority over contracts, approvals or procurement.

The comparison is useful because the tasks went beyond drafting text. They required a model to configure forms, templates, approval rules and reusable clauses while preserving conditions across a workflow. Those are closer to operational work than a one-turn question. They are still research tasks inside a hosted software environment, not observations of live customer teams, signed agreements or financial and legal outcomes.[1]

Eleven tasks were scored against detailed requirements

Ironclad employees and people who use Ironclad at OpenAI helped identify 11 tasks spanning legal, commercial and procurement work. Examples included setting up a nondisclosure-agreement workflow, creating a procurement approval process and updating a reusable clause so it changed with a requester’s jurisdiction. OpenAI estimated that an experienced user would need about 30 to 40 minutes for each task.

Each task had between 8 and 50 criteria. This is more informative than declaring a task simply passed or failed: an agent may create the right form but miss a security review, spending threshold or jurisdiction-specific branch. In contract operations, one missed condition can matter more than several correctly completed cosmetic steps, so future reporting should distinguish critical controls from lower-consequence rubric items rather than relying only on an unweighted average.[1]

Astra improved the score, but both models were incomplete

OpenAI compared GPT-6 Astra at Max reasoning with GPT-5.6 Sol at High reasoning, the setting where each model performed best. Astra’s mean rubric score was 55.0%, compared with 41.6% for Sol—a relative increase of about 32%. An internal model used during Astra’s development scored 63.7%, but that model was not the deployed comparison and should not be treated as a product result.

A single illustrated task looked much stronger: Astra met about 94% of its criteria and Sol about 85%. That example cannot replace the 11-task mean. Highlighting the best-looking workflow may obscure harder failures elsewhere. Buyers and workers need the distribution of scores, repeated-run variability and the specific types of requirement missed, especially where an omission could bypass an approval or apply the wrong contractual term.[1]

The time result is simulated, not saved staff time

Estimated average time per attempt fell from 37.0 minutes for Sol to 19.2 minutes for Astra, a reported reduction of 48%. OpenAI explicitly says those figures are simulations based on assumed model processing and generation speeds. They are not stopwatch measurements of employees using the systems, and they exclude time required to review errors, correct configuration and investigate uncertain results.

That qualification changes the practical interpretation. A faster attempt that misses a material approval rule may create more downstream work than a slower human setup. A credible productivity study would measure the complete human-agent workflow: initial instruction, agent execution, review, correction, validation, deployment and later rework. It would also record error severity and whether reviewers detect every defect before a workflow reaches production.[1]

Training and evaluation were closely linked to one platform

Ironclad provided hosted environments in which the models could practise. OpenAI created synthetic training tasks around representative workflows and used reinforcement learning. It says the simulated contracts came from public SEC EDGAR filings after filters intended to remove personal information, and that it did not use OpenAI customer data, internal contracts or non-public Ironclad customer contracts for training or evaluation.

The partnership gives the benchmark realism, but also limits independence. The platform provider helped define the tasks and success criteria, while the model developer trained and evaluated the systems. There is no peer review, external replication or comparison with other agents and platforms. Performance may depend on interface structure, environment stability and familiarity created during training. Generalisation to another contract-management system remains untested.[1]

The practical impact is job redesign before job replacement

If performance improves, legal-operations and procurement staff may spend less time configuring repetitive forms and routing rules. Their work would shift toward specifying requirements, reviewing exceptions, testing workflows and deciding when automated action is inappropriate. The 55% average, however, is a reminder that domain expertise is still needed to recognise what the agent omitted and to understand the legal or commercial consequence.

Organisations should begin with bounded, reversible tasks and test the agent against historical scenarios before enabling production changes. Approval authority, audit logs, separation of duties and rollback should remain human-controlled. Staff need time and training for verification; presenting review as effortless would hide the most important new workload created by partial automation.[1]

What would change the assessment

Confidence would rise if OpenAI or an independent evaluator released the full task specifications, criterion weights, repeated-run results and error taxonomy. A held-out benchmark created after training would help separate genuine workflow reasoning from adaptation to a familiar task family. Cross-platform tests and comparisons with experienced users, other agents and non-agent automation would show where the capability actually adds value.

The decisive evidence is a prospective workplace study measuring total human time, critical error rates, caught and uncaught defects, rework, user trust and downstream contract outcomes. Results should be segmented by task complexity and reviewer expertise. Until then, the benchmark is evidence that a newer model handled more requirements with lower simulated time—not that AI has automated contracting or delivered customer productivity gains.[1]

What this means for people

  • Legal-operations and procurement staff may spend less time on repetitive configuration but more time specifying and validating rules.
  • Workers cannot safely review an agent’s output without enough domain knowledge and protected time to test exceptions.
  • Customers and counterparties need accountable humans when an automated workflow routes, approves or records a consequential contract decision.

Global context

The research was conducted by US-based companies using public SEC EDGAR material and one commercial contracting platform. Contract law, procurement controls, privacy requirements and approval practices vary across jurisdictions. A workflow that is correct under one organisation’s rules may be wrong elsewhere, making local legal review and jurisdiction-specific testing essential before deployment.

What the evidence does not yet show

  • The evaluation covered 11 tasks on one contract-management platform and was designed with the platform provider.
  • The model developer trained and evaluated the systems; there is no reported peer review or independent replication.
  • Astra’s mean score was 55.0%, so many task criteria remained unmet even in the research environment.
  • Time per attempt was simulated and does not include human review, correction, deployment or rework.
  • The publication does not report repeated-run variance, a severity-weighted error analysis or live customer outcomes.

What to watch next

  • Release of complete tasks, scoring rules, failure categories and repeated-run results.
  • Independent replication on held-out workflows and other contract-management platforms.
  • Prospective measurement of total human-agent time and critical errors rather than simulated execution time.
  • Evidence that staff detect omissions before configurations affect approvals, obligations or signed agreements.

Living evidence record

Impact record IAI-059W0A1

Explore the full tracker

Evidence stage

Announced

Confidence

Developing

Reporting basis

Source analysis

Independent or research support

Not yet

Record status

Monitoring

Last checked

7 October 2026

Source trail

1 direct source across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Single-source reporting disclosure

This record analyses one direct source. It can establish what OpenAI published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.

Evidence trail

Sources used for this report

Links checked 7 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Work & Skills

Who is gaining from AI at work?

A Gallup survey of 15,482 US employees finds that regular AI users report more speed, creativity and work quality, but frequent use is concentrated among graduates, managers and workers who already have better jobs. The results are self-reported associations, not a productivity trial.

8 min · 2 sources

Work & Skills

Should AI wait until people think first?

In an unreviewed study of 398 adults across five countries, a writing assistant withheld direct generation until users stated a position and supporting argument. The design shifted effort earlier and improved one conditional error-classification measure, but did not significantly improve overall judgment accuracy.

8 min · 2 sources

Work & Skills

Does an AI label bias code review—or does job rank matter more?

In a randomized vignette experiment, 447 Microsoft engineers did not penalize an AI-use disclosure on identical code, but rated the same code and author more favourably when the fictional author was labelled Principal rather than Level 1. One AI-normalized US workplace cannot settle how other teams respond.

9 min · 2 sources

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Reader commentary

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Explore commentary across the portal →

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.