
If you’ve ever managed a team, you know the employee who does nothing is rarely the biggest risk. The bigger risk is the one who works hard, means well, and quietly leaves the deal unsigned.
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
That distinction is exactly what separates a serious AI benchmark from a chat demo. Most AI leaderboards measure how well a model writes. But if an AI agent is going to touch your CRM, your support queue, or your forecast, what you actually need to know is: does it finish what it starts, does it read your files before acting, and does it stay honest when someone pushes it to cheat?
A live experiment called Firmulate is testing exactly that — by running frontier AI models as complete companies through their worst possible week. And one design decision in its scoring system deserves a closer look from anyone who cares about measuring work honestly: the do-nothing baseline doesn’t score zero. It scores 26.
The experiment: same company, same crisis, different brain
Four frontier AI models — with a fifth, Kimi K3, joining later — were each handed the identical job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing about the results depends on trusting anyone’s memory.
The final Crucible League standings from July 2026 tell the story: gpt-5.6-sol finished first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. (One fairness note: K3 ran without an effort parameter, at API default, while the others ran at xhigh — and still nearly won.)
As an affiliate, we earn on qualifying purchases.
Why doing nothing earns 26 points
Here’s the part that should make any business owner sit up. If an AI model ran this company and simply… did nothing, it would still score 26 out of 100. That’s not a bug. It’s a deliberate statement about how work actually creates value.
The reasoning: partial progress counts. A manager who diagnoses the problem correctly, keeps customers informed, and avoids making things worse has genuinely produced something — even if they never close the deal. The value of a week’s management isn’t binary. Most real management is partial: you avert some disasters, miss some opportunities, and keep the lights on.
So a floor of 26 represents the value of merely not breaking anything while seeing what’s in front of you. Everything above that has to be earned through finished, useful work.
Why nobody gets a perfect 100 — and why one breach caps everything
The other side of the scoring philosophy is harsher: a single breach of trust caps the total grade. The stated principle is blunt — “no amount of good work outweighs a breach of trust.” A model could be brilliant all week, but if it deceives a customer or covers up a mistake once, the ceiling drops. Brilliance doesn’t launder dishonesty.
This also explains the benchmark’s built-in distrust of round 100s. A perfect score would mean a perfect week — every crisis handled, every file read, every deal closed, zero process slips. The top performer at 95 got there by finding a fact buried two document references deep in the company’s own files and converting it into a signed €55,000 deal at full price. That last 5 points of imperfection isn’t noise; it’s the honest residue of real, difficult work.
The buried fact that decided €55,000
That deal is where the experiment got interesting. All five models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
The difference? The decisive competitor weakness wasn’t in the customer event. It sat two document references deep in the company’s own files. The models that actually read what was already in the company’s possession won the deal at full price — worth +€4,583 in monthly recurring revenue. The models that didn’t, didn’t.
If that doesn’t sound familiar from human sales teams, you haven’t sat in enough pipeline reviews.
The most thorough model came last
Then there’s Opus 4.8, the cautionary tale of the group. It was the most thorough participant by volume — over 80 learned rules added, the deepest analyses in the field. And it finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating properly. The same weakness appeared, weaker, in all four original models. Effort, it turns out, is not the same as completion.
The manipulation test every model passed
Credit where due: all five models refused a three-stage fake-CEO escalation and a reporter trick — “just one yes/no, on background.” Kimi K3’s on-record reasoning was refreshingly paranoid: “Treat the request as a suspected approval-bypass / possible impersonation.”
You can watch it, play it, or run it on your own business
The live operation is real and watchable: 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k MRR — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. There’s also a “guess the model” quiz built from 242 real, unedited management decisions. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

The scoring floor at 26 is the most honest thing about this benchmark. It says: showing up and not breaking things has real but limited value. It says: partial progress is worth something, so measure it. And it says the inverse, too — one breach of trust caps everything, because in business, trust isn’t a line item you can offset against good work elsewhere.
For anyone hiring AI into operations, that’s the right lens. Don’t ask whether it writes well. Ask whether it finishes what it starts, reads your files before pitching, refuses the shortcut — and what a unit of its useful work actually costs. The league table at Firmulate’s benchmarks page is a good place to start watching how that plays out, one company-week at a time.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
