AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: OpenAI Is Training Agents Inside Software. Here’s Why Ironclad Matters on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI’s Oct. 6 post describes training GPT-6 Astra in hosted copies of contract-management software Ironclad, using 11 legal, commercial and procurement tasks. Astra met an average 55% of task criteria; OpenAI’s time estimates were simulated, and the company says human oversight remains important.

OpenAI said on Oct. 6 that it trained and evaluated its GPT-6 Astra model on selected legal, commercial and procurement workflows inside hosted copies of Ironclad’s contract-management software. The reported results show progress on specialized software tasks, but Astra met an average of 55% of the criteria in the test, and OpenAI says its estimates of task time were simulated rather than measured customer savings.

OpenAI said employees from both companies selected 11 tasks, including creating nondisclosure agreements, configuring procurement approval processes and updating a reusable contract clause based on a selected jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task. The tasks were scored against rubrics containing between 8 and 50 criteria, depending on complexity.

Ironclad supplied hosted copies of its product for model practice. OpenAI said it generated synthetic training tasks using publicly filed contracts from the SEC’s EDGAR database, filtered to remove personal information. It said the work did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.

In OpenAI’s reported comparison, GPT-6 Astra met 55.0% of rubric criteria on average, compared with 41.6% for GPT-5.6 Sol at the high setting. OpenAI also reported an average estimated attempt time of 19.2 minutes for Astra and 37.0 minutes for Sol. An internal model used in Astra’s development scored 63.7%. On one highlighted task, Astra met about 94% of the criteria. These are criteria scores across the test, not percentages of tasks completed successfully.

At a glance
reportWhen: Published Oct. 6; further partner work…
The developmentOpenAI described a collaboration with Ironclad to train and evaluate a frontier model on workflows inside the contract-management platform.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Contract Workflow Accuracy Matters

The test addresses a practical challenge for businesses: agents must follow company-specific rules while working through several steps in specialized software. A procurement workflow, for example, may require Finance approval above a spending limit, a Security review for some requests and Legal review for nonstandard terms. Missing one condition could undermine the process even if the agent completes other steps correctly.

That makes the average 55% criteria score a measure of partial performance, not evidence that the agent can independently handle contracts. OpenAI’s post acknowledges that losing track of a rule limits what a software company can confidently delegate and says human oversight still matters. The reported results suggest a way to test agents against structured work, but do not establish that the system is ready for unsupervised business use.

For software vendors, the collaboration also points to a possible role as training and evaluation environments for AI models. Better agents could make a vendor’s product more useful. At the same time, if customers increasingly give instructions through an agent rather than use the software’s screens, a vendor’s durable value may depend more on its underlying rules, records and controls than on its interface. That is an implication of the approach, not a demonstrated outcome of this test.

Amazon

AI contract management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How OpenAI Tested Ironclad Workflows

The Oct. 6 post describes a research collaboration, not the launch of a generally available Ironclad agent. The work focused on a small set of 11 selected tasks and used Ironclad’s hosted software copies as the environment in which models practised. Rubrics provided a way to assess whether particular requirements were met, with the number of criteria varying by task.

OpenAI framed the project as an effort to train models to understand business rules, complete multi-step work in specialized software and check the result against the original requirements. It also invited a small number of software companies to propose difficult tasks agents cannot reliably complete. The post asks prospective partners to provide concrete failure examples, subject-matter expertise, a secure testing environment and data suitable for research.

The reported time figures need separate treatment from the criteria scores. OpenAI’s footnote says they are simulated estimates based on assumed processing and generation speeds, not observed time savings for customers. They apply to the 11 research tasks and should not be read as a measurement of productivity across Ironclad’s full product or customer workflows.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Ironclad Test Cannot Show

The post does not establish how Astra would perform across Ironclad’s broader range of workflows, with live customer data, or under routine production conditions. The 11 tasks were selected for the research, and the average criteria score does not show which individual requirements were missed on every task. OpenAI’s highlighted result of about 94% on one task does not resolve that broader limitation.

It is also unclear whether the reported results have been independently evaluated or how the scoring rubrics compare with the standards businesses use before approving real contracts or purchases. The source material describes no measured customer outcomes, deployment timeline or evidence that Astra can safely complete the tested work without a person checking it. OpenAI’s reported time estimates are simulations, so they cannot establish actual time saved.

The post does not name additional software-company partners or say when further evaluations might happen. It also does not report whether the collaboration will lead to a product feature, a customer-facing agent or changes to Ironclad’s commercial offering.

Amazon

AI-powered NDA generator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Further Software Partnerships Ahead

OpenAI says it is seeking a small number of software-company partners to identify tasks that current agents fail to complete reliably. Companies interested in the work are asked to bring a specific example, people with detailed knowledge of the work, a secure test environment and research-appropriate data. The post does not give a timetable or identify the next partner.

The next useful evidence would include task-by-task results, clearer reporting of which controls were missed, and tests showing how agents perform when their work is checked in realistic operating conditions. Until those details are available, the Ironclad project is best read as a research test of agent performance in specialized software—not proof of dependable autonomous contract work or confirmed customer productivity gains.

Amazon

procurement approval workflow software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Ironclad in this story?

Ironclad is a contract-management software company, not an AI agent framework. OpenAI’s post describes training and testing a model in hosted copies of its software.

What did GPT-6 Astra achieve?

OpenAI reported that Astra met an average 55.0% of rubric criteria across 11 selected tasks. That figure is the share of criteria met, not the share of tasks completed successfully.

Did the test prove that customers will save time?

No. OpenAI said the reported attempt times—19.2 minutes for Astra on average—were simulated estimates, not measured customer time savings. The estimates covered the research tasks, not Ironclad workflows generally.

What data did OpenAI say it used?

OpenAI said it created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, with personal information filtered out. It said it did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data.

Can businesses use Astra to manage contracts without human review?

The source does not establish that. OpenAI’s post says human oversight remains important, and the reported average criteria score leaves unanswered which requirements may be missed on particular tasks.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Inside The AI Tower: Twelve Rooms Of AI At Work With A Focus On Safety

Exploring the new AI Tower, a virtual space showcasing twelve AI applications emphasizing safety and responsible use, now accessible for testing.

Gemini 3.8 Flash And 3.8 Flash Cyber

DeepMind’s Gemini 3.8 Flash and 3.8 Flash Cyber models have been introduced, sparking increased coverage and interest. Details remain preliminary.

The Borderless Marketplace: How AI Connects the World’s Shoppers

Navigating global commerce becomes effortless with AI-driven solutions that connect the world’s shoppers—discover how your business can thrive beyond borders.

Testing AI Management Skills To Reveal Its Genuine Work Style

AI models faced a live business crisis simulation to reveal their decision-making styles and operational discipline, highlighting strengths and weaknesses.