THE SIGNAL IN ONE SENTENCE
Contracts are full of tiny switches that can become expensive doors. A purchase above a certain amount needs Finance approval. A request involving sensitive data needs Security review. An unusual clause needs Legal. A reusable template may need to change when the requester selects another jurisdiction. The person configuring that system is not merely clicking through screens. They are translating business rules into a machine that must behave correctly when normal cases, edge cases and awkward exceptions arrive. OpenAI and contract-software company Ironclad have turned 11 examples of that work into research tasks for computer-using AI agents. GPT-6 Astra averaged 55.0 percent on the resulting rubrics. GPT-5.6 Sol averaged 41.6 percent. OpenAI also estimates that Astra used 19.2 minutes per attempt, compared with 37.0 minutes for Sol. That is meaningful progress. It is also a score of 55 percent. The headline version of agent progress tends to celebrate the clicks. The model opened the application, found the right menu, created an approval step and saved the workflow. Lovely. The contract version of success starts after the clicking stops. Does a request one dollar below the threshold avoid Finance while a request one dollar above it triggers the correct approver? Does a nonstandard clause route to Legal? Does the reusable language change for the selected jurisdiction without damaging the rest of the template? Can an unauthorized user alter the rule? Does the final agreement record preserve what happened? OpenAI's own research description makes this distinction clearly. Getting individual steps right is not enough. The finished process must work across the situations it was designed to handle. That is the signal. Computer-using agents are getting better at professional software, but businesses should evaluate the state they leave behind, not how convincingly they move through the interface. Ironclad employees and people who use Ironclad at OpenAI selected the 11 tasks. They span legal, commercial and procurement work, including setting up nondisclosure agreements, building procurement approval processes and updating reusable clauses according to jurisdiction. Each task has between 8 and 50 criteria. That is an important design choice. A vague judgment such as "the workflow looks right" would hide too much. A long rubric can award credit for the parts an agent completed and reveal the requirement it quietly forgot. OpenAI says the work would take an experienced user about 30 to 40 minutes per task on average. Ironclad supplied hosted versions of its software where models could practice. OpenAI researchers created synthetic training tasks based on contracts from the Securities and Exchange Commission's public EDGAR database, after applying filters intended to remove personal information. The companies say they did not use OpenAI customer data, OpenAI's internal contracts or nonpublic Ironclad customer contracts for training or evaluation. Researchers then used reinforcement learning so models could improve through practice and feedback. The resulting comparison used each model's strongest reported reasoning setting: Max for Astra and High for Sol. Astra's 55.0 percent score is 13.4 percentage points above Sol's 41.6 percent. OpenAI describes that as a 32 percent relative improvement. Those are two different ways of stating the same comparison. Percentage points describe the direct gap. Relative improvement divides that gap by the earlier score and produces the larger-sounding number. Both are mathematically legitimate. The direct scores are more useful for understanding the unfinished work. An internal model used while developing Astra reached 63.7 percent. That result suggests more progress may be possible. It is not a public model that customers can test today. The time numbers need even more care. OpenAI calls them simulated estimates based on assumed processing and generation speeds. They are not measured customer time savings. The evaluation does not report an employee completing the same work beside the models under identical production conditions, nor does it include the time required to inspect and repair the agents' mistakes. If an agent saves 18 minutes and creates a broken approval path that takes an hour to diagnose, the stopwatch has played a prank on the business. The evaluation is small by design. It covers 11 research tasks, not all Ironclad workflows and certainly not all contract operations. The partners chose the tasks, built the environment, trained the model and reported the results. There is no independent replication in the announcement. That does not make the work useless. It tells us exactly what it is: a promising partner-led research evaluation, not proof that autonomous contract administration is ready for unsupervised production. One demo task illustrates another trap. Astra met about 94 percent of its criteria in an estimated 20 minutes, while Sol met about 85 percent in an estimated 32 minutes. Ninety-four percent sounds close to done. In a contract workflow, the missing six percent could be a label color. It could also be the approval rule standing between a purchase request and unauthorized spending. A single average cannot express the consequence of the missing item. That is why rubrics need severity. Every criterion should carry a failure class. Cosmetic errors can enter a repair queue. Missing metadata may block publication. A broken permission, approval threshold, jurisdiction rule or signature condition should fail the whole task until a person reviews it. The agent also needs tests that attack the workflow it created. For a spending threshold, run cases immediately below, at and above the boundary. For a jurisdiction rule, try every supported region plus an unknown value. For roles, test a standard user, an approver and an unauthorized account. For an exception route, confirm both the trigger and the ordinary path. This is end-state verification: inspecting whether the completed system satisfies its requirements after the agent stops acting. It should happen before the workflow reaches employees. A safe agent can generate its own test plan from the original request, but the acceptance rules should come from people who own the process. The agent should not be allowed to quietly redefine success around the result it happened to produce. The strongest production pattern is not "agent does the job." It is "agent proposes, system tests, person approves." The agent works in a sandbox or draft environment. The platform records every action and captures a structured diff of the final configuration. Automated checks run ordinary and adversarial cases. A knowledgeable reviewer sees the original requirements, changed rules, failed tests and unresolved ambiguity in one place. Only then can the workflow move into production. Rollback matters too. Contract systems evolve. A team may ask an agent to update an existing template or approval process rather than build a clean one. The platform should snapshot the previous state, show dependencies and restore the last approved configuration if the new rules behave badly. Auditability cannot be a video replay of mouse movements. The useful record says which requirement produced each change, which tests passed, which policy granted access, which model and settings were used, who approved the result and what changed after approval. That record should survive even if the interface changes. Human oversight is not a ceremonial final click. The reviewer needs enough context to find the dangerous omission without repeating the entire task manually. If checking an agent takes longer than doing the work, the product has moved labor rather than removed it. This is where Ironclad's participation adds value. Software companies understand the hidden rules that make a professional workflow valid. They can turn customer pain into specific tasks, provide safe practice environments and define what a correct end state looks like. There is also an incentive problem. The software partner may eventually sell more agent features. The model company wants evidence that its frontier model can perform professional work. Their expertise improves the evaluation, while their shared commercial interest makes outside testing more important. Future reports should publish more than an average. Show pass rates by task, criterion severity, model variability across repeated attempts, repair time, human review time, permission failures, rollback frequency and performance on held-out workflows created after training. Report whether an agent recognizes ambiguity and stops instead of inventing a business rule. Test the boring disasters. Use two employees with the same name. Remove an expected approver. Change a threshold after the draft is created. Give the agent a request containing contradictory instructions. Interrupt a session. Revoke access halfway through. Ask it to modify a template with an active dependent workflow. Professional work is where normal cases go to meet organizational history. OpenAI says it is inviting a small number of other software companies to bring concrete agent failures, knowledgeable practitioners, secure test environments and safely usable data into similar collaborations. That could produce better agents and better evaluations. The danger is turning each partner's application into another tailored exam that the model practices until it passes, while general reliability remains uncertain. The next step should include held-out tasks, outside evaluators and customer-controlled acceptance suites that the model developer never sees. There is no shame in a 55 percent research score. It is evidence of a difficult system becoming more capable. It is also a bright red reminder that successful-looking computer use can leave half the job hiding in the walls. Let the agent move quickly. Make the workflow prove it works.
01
WHAT ACTUALLY CHANGED
OpenAI and Ironclad created 11 hosted research tasks covering legal, commercial and procurement workflows
Each task was evaluated against 8 to 50 criteria rather than a single pass or fail judgment
GPT-6 Astra averaged 55.0 percent compared with 41.6 percent for GPT-5.6 Sol using each model’s strongest reported reasoning setting
Estimated time per attempt fell from 37.0 minutes to 19.2 minutes, but OpenAI says those figures are simulated rather than measured customer savings
An internal development model reached 63.7 percent, while OpenAI is seeking additional software partners for similar research
02
WHY THIS MATTERS
A computer-using agent can complete visible steps while leaving a critical business rule or exception broken
Average rubric scores hide whether the missing criterion is cosmetic or a consequential approval, permission or jurisdiction rule
Simulated execution time excludes the human effort needed to inspect, repair and approve an agent’s work
Partner-designed tasks improve realism but also require independent and customer-controlled evaluation
Enterprise agents need end-state tests, audit records, review queues and rollback before they can safely change production workflows
03
WHERE IT COULD HELP
- Build acceptance suites that test workflow thresholds immediately below, at and above each boundary
- Require structured diffs linking every agent-made configuration change to an original business requirement
- Classify rubric criteria by consequence so critical permission or approval failures block deployment
- Run agents in draft environments with snapshots, dependency checks and one-click rollback
- Measure review and repair time alongside agent execution time when calculating business value
- Maintain held-out customer test cases that model vendors and training teams cannot optimize against directly
KEEP A HAND ON THE WHEEL
Watch for independent reproduction of the 11-task evaluation, per-task and severity-weighted scores, repeated-run reliability, measured human review and repair time, customer-controlled acceptance suites, held-out workflows, permission and rollback tests, production outcome evidence, details about additional software partners, and a clear distinction between research environments and released Ironclad product capabilities.
04
TERMS WORTH KEEPING
OPEN GLOSSARY CARD
End-state verification
Checking whether the system left behind by an agent satisfies the original requirements after the agent stops acting.
OPEN GLOSSARY CARD
Rubric score
A result calculated from a defined set of criteria used to judge how much of a task was completed correctly.
OPEN GLOSSARY CARD
Agentic AI
AI software designed to pursue a goal through several steps, including choosing actions and using approved tools with limited supervision.
SOURCES AND VERIFICATION STATUS
This article was written from the materials below. Product claims and dates were checked against those sources on October 7, 2026.
THE PUBLICATION ENGINE
WANT A SIGNAL OF YOUR OWN?
We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.
WORK WITH US