firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

The difference between a useful agent and a persuasive demo

For businesses investing in AI for marketing, sales or ecommerce operations, polished writing is no longer the hardest test. The more valuable question is whether an agent will inspect the relevant company material, connect information scattered across documents and finish the commercial task it has started.

Firmulate turned that question into a live, auditable experiment. Each frontier model was asked to run the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned. All the models recognized every crisis, and all resisted every manipulation attempt. Yet only two signed the €55,000 deal their own work had already earned.

The decisive difference was not better sales language. It was whether the model had done its homework.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The sale depended on a fact hidden two references deep

The customer event did not contain everything needed to close the deal. The crucial competitive weakness was buried in the company’s own files, two document references away from the immediate task. Models that followed that trail won the contract at full price, adding €4,583 in monthly recurring revenue. Those that did not were effectively disqualified from the sale.

That creates a striking contrast: the diagnosis could be right, and the pitch could be right, while the commercial outcome was still wrong. Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.”

For general-purpose chat tools, overlooking a background document may produce an incomplete answer. For an agent entrusted with sales or account management, the same lapse can decide whether revenue is won or lost. “Reads your files before answering” is therefore not a convenience feature. In Firmulate’s experiment, it became a measurable, purchase-deciding capability.

A hard week for every model

The simulated company employs 13 synthetic staff and operates with real money mechanics. It burns €105k per month while generating €2.3k in monthly recurring revenue, with a public cash countdown adding urgency to every decision. Its agents have accumulated more than 680 self-learned playbook rules, and every workday is versioned.

The pressure did not make the models reckless. Fake messages from the CEO escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 of 5 models refused the social-engineering attempts. Kimi K3 recorded the clearest concise diagnosis: “Treat the request as a suspected approval-bypass / possible impersonation.”

This matters because file-reading discipline cannot be judged in isolation. An agent might search widely but disclose sensitive information, or act decisively while bypassing approval. Firmulate’s benchmark treats trust as non-negotiable: a single breach caps the total because “no amount of good work outweighs a breach of trust.” Even doing nothing receives a baseline score of 26 because partial progress counts, but passive safety is not the same as competent management.

The leaderboard rewards completion

The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The public Firmulate benchmarks present the results and plain-language findings.

Opus 4.8 illustrates why apparent diligence can be misleading. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses. It still finished last. The close was left on the table, while operational discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared in a milder form across all four other models.

There is also an important fairness qualification. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference should remain visible when readers compare results.

From benchmark spectacle to procurement evidence

Firmulate describes itself as an AI company emulator, and the experiment remains publicly watchable as the synthetic business continues operating. That makes the project more than a static leaderboard. Readers can inspect how models behave through sustained pressure rather than judging them from a carefully selected conversation.

The company also turns 242 real, unedited management decisions into a “guess the model” quiz. The exercise exposes how difficult it can be to identify a model from confident prose alone. Decisions reveal differences that style can conceal: whether an agent checks source material, respects boundaries, escalates correctly and completes the final commercial step.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

Buy the behavior, not the fluent answer

For marketing and ecommerce leaders, the practical lesson is direct. A model that can draft an attractive campaign or sales response may still fail when the decisive evidence sits elsewhere in the business. Evaluation should test whether agents locate that evidence, use it without crossing trust boundaries and carry work through to a completed outcome.

Enterprises can run the same kind of wargame against a read-only export of their own business. Nothing writes back to real systems, allowing teams to observe behavior against familiar documents and commercial situations before granting operational access.

The €55,000 contract shows why this matters. Every model understood the crisis and rejected manipulation, but understanding was not enough. The winners connected the buried fact to the live opportunity and completed the sale at full price. In an agentic workplace, homework is no longer preparation for the job. It is part of the job.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


You May Also Like

Podcasting’s Rise to Mainstream Media

Keen to understand how podcasting’s rise to mainstream media is reshaping content consumption and creator opportunities? Keep reading to find out more.

Location Scouting Secrets: Turning Ordinary Streets Into Iconic Sets

Journey into essential location scouting secrets that can transform ordinary streets into iconic cinematic sets—discover how to unlock their hidden potential today.

The AI Management Test Begins Where the Leaderboard Ends

Coding tests reward polished answers. Firmulate asks whether AI managers can close deals, resist pressure and preserve trust through a company’s worst week.