firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on office and shipping supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Your next AI hire needs to close, not just sound convincing

For businesses weighing AI for sales, service or operations, polished answers are only part of the job. An agent also has to find the right information, protect customer trust and finish the work. In Firmulate’s live company experiment, Moonshot’s Kimi K3 did all three well enough to beat three of four Western frontier models—and finish just behind the leader.

Amazon

AI customer service chatbot

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A hard week for five models

Firmulate put five frontier models through the same worst week at a small software company: the same customers, crises and temptations. The decisions were versioned and auditable. The test focused on management in context, rather than chat quality.

In the final Crucible league for July 2026, gpt-5.6-sol led with 95. Kimi K3 placed second with 93, followed by Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate’s rule is stark: partial progress counts, but one breach of trust caps the total; no amount of good work outweighs a breach of trust.

The difference was finishing the job

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. That gap—“Same diagnosis, same pitch — no signature”—is a practical warning for companies evaluating AI agents: recognizing an opportunity and acting on it are separate capabilities.

The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. K3 was among them. Its result combined the commercial win with the field’s cleanest discipline: just one deviation.

Trust under pressure, and a notable caveat

The experiment also staged fake CEO messages that escalated over three stages, then a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That is the kind of judgment businesses need when an agent handles confidential information or operates near approval processes.

There is a fairness caveat: K3 ran without an effort parameter (API default), while the others ran at xhigh. That difference belongs alongside the rankings when interpreting the result.

Opus 4.8 offers a different lesson. It was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and slipped on discipline by attempting writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four. Thorough analysis, on its own, did not guarantee a clean finish.

A company you can watch

Firmulate presents the experiment as a live, watchable company, not a fictional case study. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and a versioned record of every workday. The live company is at firmulate.com.

A separate quiz uses 242 real, unedited management decisions and asks visitors to guess which model made each one. For businesses that want a closer look, the company says enterprises can run the wargame against a read-only export of their own business; nothing writes back to real systems. The benchmark results and plain-language findings are at Firmulate’s benchmarks page.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Run your own test before choosing a model

K3’s near-top finish shows that the frontier is open: the leading model is not the only one capable of finding the buried fact, winning the deal and resisting pressure. But a league table cannot tell a business how an agent will behave in its own workflows. If AI will touch a CRM, support queue or forecast, test whether it reads the available evidence, follows approval boundaries and completes the work. Picking without a test is a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How Streaming Algorithms Really Choose Your Next Binge—And How to Outsmart Them

Navigating how streaming algorithms select your next binge reveals hidden tricks—discover how to outsmart them and regain control of your viewing choices.

Celebrity Advocacy for Mental Health

Growing celebrity advocacy for mental health sparks change, but how does it truly impact society and those in need? Keep reading to find out.

Five AI Managers Faced an Impersonation Attack—and Held the Line

Five frontier AI models rejected every fake-CEO demand and reporter trick, showing businesses can test agent integrity before a real crisis hits.