OpenAI Training Agents In Software: The Terms To Examine At Ironclad
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: OpenAI Training Agents In Software: The Terms To Examine At Ironclad on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on office and shipping supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI described training GPT-6 Astra in hosted copies of Ironclad’s contract-management software, using 11 selected legal, commercial and procurement tasks. Astra met an average 55% of task rubric criteria, while OpenAI’s estimated completion times were simulations rather than measured customer results.

OpenAI said on October 6 that it trained and evaluated GPT-6 Astra on selected workflows in hosted copies of Ironclad’s contract-management software, offering a concrete example of frontier-model training inside a specialized business product. Astra met an average 55% of rubric criteria across 11 tasks, according to OpenAI; the company’s estimated task times are simulations, not measured customer savings.

The tasks were selected by Ironclad staff and OpenAI employees familiar with the product. They covered legal, commercial and procurement work, including setting up nondisclosure agreements, creating procurement approval processes and changing a reusable contract clause based on a requester’s selected jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task.

OpenAI scored each task against a rubric of 8 to 50 criteria, depending on its complexity. It reported that GPT-6 Astra met an average 55% of criteria, compared with 41.6% for GPT-5.6 Sol in the high setting. An internal model used during Astra’s development reached 63.7%. On one showcase task, Astra met about 94% of criteria. These are rubric scores, not percentages of tasks completed successfully.

For training, Ironclad supplied hosted product copies where models could practise. OpenAI said it created synthetic tasks from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information. It said it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data. OpenAI also reported estimated times of 19.2 minutes per Astra attempt and 37 minutes for GPT-5.6 Sol, while explicitly describing those figures as simulated estimates based on assumed processing and generation speeds.

At a glance
reportWhen: Published October 6; further partner pa…
The developmentOpenAI published details of a collaboration with Ironclad to train and evaluate a frontier model on workflows inside contract-management software.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Contract Workflow Accuracy Matters

The reported results show both the promise and the limits of asking AI agents to perform work inside business software. A model may complete multiple steps, but a partial score can hide a consequential omission. In procurement, for example, missing a required Finance approval above a spending threshold or skipping a Security review could leave a purchase outside the company’s rules. In contract work, getting most rubric items right does not establish that the final workflow is safe to use.

OpenAI’s account points to a developing role for software companies as training and evaluation partners: they can supply difficult workflows, domain experts and secure test environments. For buyers, the practical question is not only whether an agent can operate an interface, but whether it follows business rules, records its actions and flags work requiring human review. The results do not establish that agents are ready to run these processes without oversight.

There is also a strategic question for vendors. If agents increasingly mediate how customers use a product, a vendor’s value may depend less on its screens and more on its business rules, data structures, audit records and controls. OpenAI’s post frames the work as evidence for the continuing need for a full contracting platform. That is a rationale in the company’s account, not an independently demonstrated market outcome.

Amazon

contract management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Trial Was Set Up

The announcement concerns a specific research effort, not a general claim that OpenAI agents now operate contract systems autonomously. The evaluation involved 11 selected tasks in one specialized product. OpenAI estimated that an experienced user could complete each in 30 to 40 minutes; the model-time figures were calculated simulations and do not measure how long customers actually took or how much time a deployed system would save.

The scoring method also matters. Each task had multiple criteria, with the number varying from 8 to 50. OpenAI’s reported 55% is the average share of those criteria met, not a 55% task success rate and not a guarantee that a workflow met every required control. The company itself noted that losing track of a business rule can constrain what a software company should ask an agent to do, and that human oversight remains important.

OpenAI said it is inviting a small number of software companies to partner on workflows that current agents cannot reliably complete. It described the requested contributions as a concrete failing example, people with deep knowledge of the work, a secure test environment and data that can safely be used for research. The source account says OpenAI published another item on October 6 about mathematics manuscripts, but this article focuses on the Ironclad development and its reported evaluation.

What the Evaluation Cannot Establish

The reported findings do not establish how Astra would perform across the full range of Ironclad customers, contracts or unusual cases. The work covered 11 selected tasks, and the source material does not provide a detailed breakdown of which criteria the model missed on each task. Without that information, readers cannot tell whether the average shortfall involved minor formatting issues or errors in required approvals and other controls.

It is also unclear how performance would change with live customer data, varied company policies, system changes or repeated use in production. OpenAI said it did not use non-public Ironclad customer data, and its time figures were simulated. The source does not give a customer deployment, independently verified results, or a timeline for broader availability. The reported scores should not be read as proof of safe use without review.

What Partner Testing Could Show

OpenAI’s stated next step is to work with a small number of software companies that can bring workflows current agents struggle to complete, along with domain experts, secure test environments and research-safe data. Further partner evaluations could show whether models improve on harder tasks and which kinds of errors persist. The company has not supplied a schedule or named additional partners in the source material.

For software vendors and business buyers, useful follow-up evidence would include task-by-task results, the criteria missed, how often agents need human correction, and performance on live or representative workflows. Any claimed time savings would need to be measured in actual use rather than inferred from simulated processing speeds. Until such evidence is available, the Ironclad results describe an early evaluation, not a demonstrated replacement for human review.

Key Questions

What did OpenAI announce about Ironclad?

OpenAI published details of training and evaluating GPT-6 Astra in hosted copies of Ironclad’s contract-management product, using 11 legal, commercial and procurement tasks.

What does Astra’s 55% score mean?

It is the average share of rubric criteria met across the evaluated tasks. It is not the share of tasks completed and does not show that each workflow met all required business controls.

Did OpenAI demonstrate customer time savings?

No. OpenAI described its per-task time figures as simulated estimates based on assumed processing and generation speeds, not measured customer results.

What data did OpenAI say it used?

OpenAI said it generated synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information. It said it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.

Is Astra ready to run contract workflows without people checking its work?

The reported evaluation does not show that. Astra met an average 55% of criteria, and the source account says human oversight remains important. OpenAI has not established from these results that the model can safely handle these workflows without review.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Trade and supply-chain operations signal monitor: U.S. strikes Iranian military sites after ship was hit in Strait of Hormuz

The US has launched strikes on Iranian military targets following an attack on a ship in the Strait of Hormuz, escalating regional tensions and affecting trade routes.

New York City To To Ban Deceptive Subscription Practices

NYC announces a new ban on deceptive subscription practices to protect consumers from misleading tactics. Implementation details are forthcoming.

ClearSign Technologies Corp Files 8-K: Material Agreement

ClearSign Technologies announced a material agreement via an 8-K filing, impacting its strategic operations. Details are confirmed, but some specifics remain undisclosed.

The European Union: Rules First, Cushion Always

The EU’s AI Act and social policies exemplify its strategy to regulate and cushion labor market shifts, emphasizing rules over ownership.