
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Your next AI hire needs to close, not just sound convincing
For businesses weighing AI for sales, service or operations, polished answers are only part of the job. An agent also has to find the right information, protect customer trust and finish the work. In Firmulate’s live company experiment, Moonshot’s Kimi K3 did all three well enough to beat three of four Western frontier models—and finish just behind the leader.
As an affiliate, we earn on qualifying purchases.
A hard week for five models
Firmulate put five frontier models through the same worst week at a small software company: the same customers, crises and temptations. The decisions were versioned and auditable. The test focused on management in context, rather than chat quality.
In the final Crucible league for July 2026, gpt-5.6-sol led with 95. Kimi K3 placed second with 93, followed by Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate’s rule is stark: partial progress counts, but one breach of trust caps the total; no amount of good work outweighs a breach of trust.
The difference was finishing the job
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. That gap—“Same diagnosis, same pitch — no signature”—is a practical warning for companies evaluating AI agents: recognizing an opportunity and acting on it are separate capabilities.
The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. K3 was among them. Its result combined the commercial win with the field’s cleanest discipline: just one deviation.
Trust under pressure, and a notable caveat
The experiment also staged fake CEO messages that escalated over three stages, then a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That is the kind of judgment businesses need when an agent handles confidential information or operates near approval processes.
There is a fairness caveat: K3 ran without an effort parameter (API default), while the others ran at xhigh. That difference belongs alongside the rankings when interpreting the result.
Opus 4.8 offers a different lesson. It was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and slipped on discipline by attempting writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four. Thorough analysis, on its own, did not guarantee a clean finish.
A company you can watch
Firmulate presents the experiment as a live, watchable company, not a fictional case study. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and a versioned record of every workday. The live company is at firmulate.com.
A separate quiz uses 242 real, unedited management decisions and asks visitors to guess which model made each one. For businesses that want a closer look, the company says enterprises can run the wargame against a read-only export of their own business; nothing writes back to real systems. The benchmark results and plain-language findings are at Firmulate’s benchmarks page.

Run your own test before choosing a model
K3’s near-top finish shows that the frontier is open: the leading model is not the only one capable of finding the buried fact, winning the deal and resisting pressure. But a league table cannot tell a business how an agent will behave in its own workflows. If AI will touch a CRM, support queue or forecast, test whether it reads the available evidence, follows approval boundaries and completes the work. Picking without a test is a bet.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
