🔍 Read the full analysis: OpenAI Is Training Agents Inside Software. Here’s Why Ironclad Matters on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
OpenAI’s Oct. 6 post describes training GPT-6 Astra in hosted copies of contract-management software Ironclad, using 11 legal, commercial and procurement tasks. Astra met an average 55% of task criteria; OpenAI’s time estimates were simulated, and the company says human oversight remains important.
OpenAI said on Oct. 6 that it trained and evaluated its GPT-6 Astra model on selected legal, commercial and procurement workflows inside hosted copies of Ironclad’s contract-management software. The reported results show progress on specialized software tasks, but Astra met an average of 55% of the criteria in the test, and OpenAI says its estimates of task time were simulated rather than measured customer savings.
OpenAI said employees from both companies selected 11 tasks, including creating nondisclosure agreements, configuring procurement approval processes and updating a reusable contract clause based on a selected jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task. The tasks were scored against rubrics containing between 8 and 50 criteria, depending on complexity.
Ironclad supplied hosted copies of its product for model practice. OpenAI said it generated synthetic training tasks using publicly filed contracts from the SEC’s EDGAR database, filtered to remove personal information. It said the work did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.
In OpenAI’s reported comparison, GPT-6 Astra met 55.0% of rubric criteria on average, compared with 41.6% for GPT-5.6 Sol at the high setting. OpenAI also reported an average estimated attempt time of 19.2 minutes for Astra and 37.0 minutes for Sol. An internal model used in Astra’s development scored 63.7%. On one highlighted task, Astra met about 94% of the criteria. These are criteria scores across the test, not percentages of tasks completed successfully.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Contract Workflow Accuracy Matters
The test addresses a practical challenge for businesses: agents must follow company-specific rules while working through several steps in specialized software. A procurement workflow, for example, may require Finance approval above a spending limit, a Security review for some requests and Legal review for nonstandard terms. Missing one condition could undermine the process even if the agent completes other steps correctly.
That makes the average 55% criteria score a measure of partial performance, not evidence that the agent can independently handle contracts. OpenAI’s post acknowledges that losing track of a rule limits what a software company can confidently delegate and says human oversight still matters. The reported results suggest a way to test agents against structured work, but do not establish that the system is ready for unsupervised business use.
For software vendors, the collaboration also points to a possible role as training and evaluation environments for AI models. Better agents could make a vendor’s product more useful. At the same time, if customers increasingly give instructions through an agent rather than use the software’s screens, a vendor’s durable value may depend more on its underlying rules, records and controls than on its interface. That is an implication of the approach, not a demonstrated outcome of this test.
As an affiliate, we earn on qualifying purchases.
How OpenAI Tested Ironclad Workflows
The Oct. 6 post describes a research collaboration, not the launch of a generally available Ironclad agent. The work focused on a small set of 11 selected tasks and used Ironclad’s hosted software copies as the environment in which models practised. Rubrics provided a way to assess whether particular requirements were met, with the number of criteria varying by task.
OpenAI framed the project as an effort to train models to understand business rules, complete multi-step work in specialized software and check the result against the original requirements. It also invited a small number of software companies to propose difficult tasks agents cannot reliably complete. The post asks prospective partners to provide concrete failure examples, subject-matter expertise, a secure testing environment and data suitable for research.
The reported time figures need separate treatment from the criteria scores. OpenAI’s footnote says they are simulated estimates based on assumed processing and generation speeds, not observed time savings for customers. They apply to the 11 research tasks and should not be read as a measurement of productivity across Ironclad’s full product or customer workflows.
As an affiliate, we earn on qualifying purchases.
What the Ironclad Test Cannot Show
The post does not establish how Astra would perform across Ironclad’s broader range of workflows, with live customer data, or under routine production conditions. The 11 tasks were selected for the research, and the average criteria score does not show which individual requirements were missed on every task. OpenAI’s highlighted result of about 94% on one task does not resolve that broader limitation.
It is also unclear whether the reported results have been independently evaluated or how the scoring rubrics compare with the standards businesses use before approving real contracts or purchases. The source material describes no measured customer outcomes, deployment timeline or evidence that Astra can safely complete the tested work without a person checking it. OpenAI’s reported time estimates are simulations, so they cannot establish actual time saved.
The post does not name additional software-company partners or say when further evaluations might happen. It also does not report whether the collaboration will lead to a product feature, a customer-facing agent or changes to Ironclad’s commercial offering.
As an affiliate, we earn on qualifying purchases.
Further Software Partnerships Ahead
OpenAI says it is seeking a small number of software-company partners to identify tasks that current agents fail to complete reliably. Companies interested in the work are asked to bring a specific example, people with detailed knowledge of the work, a secure test environment and research-appropriate data. The post does not give a timetable or identify the next partner.
The next useful evidence would include task-by-task results, clearer reporting of which controls were missed, and tests showing how agents perform when their work is checked in realistic operating conditions. Until those details are available, the Ironclad project is best read as a research test of agent performance in specialized software—not proof of dependable autonomous contract work or confirmed customer productivity gains.
procurement approval workflow software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Ironclad in this story?
Ironclad is a contract-management software company, not an AI agent framework. OpenAI’s post describes training and testing a model in hosted copies of its software.
What did GPT-6 Astra achieve?
OpenAI reported that Astra met an average 55.0% of rubric criteria across 11 selected tasks. That figure is the share of criteria met, not the share of tasks completed successfully.
Did the test prove that customers will save time?
No. OpenAI said the reported attempt times—19.2 minutes for Astra on average—were simulated estimates, not measured customer time savings. The estimates covered the research tasks, not Ironclad workflows generally.
What data did OpenAI say it used?
OpenAI said it created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, with personal information filtered out. It said it did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data.
Can businesses use Astra to manage contracts without human review?
The source does not establish that. OpenAI’s post says human oversight remains important, and the reported average criteria score leaves unanswered which requirements may be missed on particular tasks.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
