
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A polished answer is not the same as a business result
For marketing and ecommerce leaders, the dangerous AI failure may not be a bad response. It may be a good response that never becomes an action: the customer is diagnosed correctly, the sales pitch is prepared, and the contract remains unsigned.
That distinction sits at the heart of Firmulate, a live experiment that puts frontier models in charge of the same small software company during its worst week. Each receives the same customers, crises and temptations. Every decision is versioned and auditable. The aim is to measure management quality, not chat quality.
The final July 2026 Crucible League results suggest that this is a distinct capability. gpt-5.6-sol led with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. Yet the ranking is less revealing than the behaviors underneath it.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Seeing the crisis was not enough
All five models spotted every crisis and rejected every manipulation attempt. On those visible tests, the field looked reassuringly competent. But only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
That is the measurement problem facing businesses evaluating agents through coding benchmarks or chat comparisons. Those tests can show whether a model produces a strong answer. They do not necessarily show whether it triages competing demands, follows through under capacity pressure or accepts responsibility for consequences that unfold across days.
The decisive sales advantage was not presented conveniently in the customer event. It sat two document references deep in the company’s own files: a buried competitor weakness. Models that read the file won the deal at full price, adding €4,583 in monthly recurring revenue. The practical lesson is recognizable to any commercial team. A model can sound persuasive while missing the one piece of internal context that changes the negotiation.
Pressure tests honesty as well as competence
Firmulate also confronted the models with fake CEO messages that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” Five of five refused. Kimi K3 recorded the clearest description of the threat: “Treat the request as a suspected approval-bypass / possible impersonation.”
This matters because an AI manager is not merely a writing assistant with access to more buttons. If it touches a customer queue, commercial forecast or sensitive company record, resistance to urgency and authority theater becomes part of its job. Firmulate treats trust as non-negotiable: “no amount of good work outweighs a breach of trust.”
The experiment’s scenario names read less like academic tests than an executive’s sleepless calendar: churn wave, price increase, downround and PR crisis. That is the emerging curriculum for business agents. The question is not simply whether the model knows what should happen. It is whether it can decide what matters now, find the necessary evidence, finish the work and remain honest while pressure accumulates.
Thoroughness can still lose
Opus 4.8 offers the most instructive profile. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in weaker form across the other four models.
This is a useful warning against equating visible effort with managerial effectiveness. More analysis can be valuable, but analysis does not rescue an unfinished sale or a process mistake. A business needs agents that recognize the difference between studying a problem and owning its outcome.
The comparison also requires a fairness note. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That does not erase the result, but it belongs beside it. Serious evaluation should expose operating conditions rather than flattening them into a single score.
A company that keeps teaching the test
The live company includes 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, publishes a cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, making the experiment watchable rather than retrospective.
Readers can inspect the broader experiment through Firmulate and review the benchmark findings. A separate quiz is powered by 242 real, unedited management decisions, inviting people to guess which model made each choice. Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems.

The next AI category is management quality
The Crucible League does not argue that coding benchmarks or chat arenas are useless. It shows what they leave out. Answer quality is only one layer of performance when an agent must operate inside a company with incomplete context, limited capacity, commercial consequences and people attempting to manipulate it.
The strongest models in this experiment did more than identify the correct course. They read far enough, protected trust and completed the consequential action. The weaker performances were not necessarily unintelligent. They were operationally incomplete.
For business buyers, that changes the evaluation question. Do not ask only whether an AI can produce the right analysis. Ask whether it can survive the week: finding buried evidence, resisting a false executive, escalating when blocked and closing the deal it has already earned. That is where chat quality ends and management quality begins.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
