
A business experiment with consequences
For business, marketing and ecommerce readers, the most interesting AI story may not be another polished campaign or chatbot demo. It may be a struggling software company whose staff are entirely synthetic—and whose commercial problems remain stubbornly familiar.
Firmulate’s live company has 13 synthetic employees, monthly burn of €105k and just €2.3k in monthly recurring revenue. Its cash countdown is public. Its workdays are versioned. Visitors can watch the operation confront the same tension that shapes any fragile company: revenue must arrive before the money runs out.
That makes the live experiment less like a product showcase and more like an unfolding business story. The company does not merely produce AI-generated plans. It accumulates decisions, mistakes and lessons while visibly fighting for survival.

The Decision Intelligence Handbook: Practical Steps for Evidence-Based Decisions in a Complex World
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Build in public, taken to its limit
Most build-in-public projects reveal selected milestones: a revenue screenshot, a launch retrospective or a founder’s account of what went wrong. Firmulate exposes something more continuous. Every workday becomes part of an auditable record, while the company’s financial imbalance remains in view.
The synthetic workforce has already developed more than 680 playbook rules from experience. That detail matters because it turns the company into a portrait of organizational learning. A rule can record what was discovered, but the daily drama lies in whether the workforce applies that knowledge when customers, cash and pressure collide.
The experiment’s companion Crucible League demonstrates why execution deserves as much attention as analysis. Each frontier model ran the same small software company through its worst week, facing identical customers, crises and temptations. Every model detected every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own work had earned.
The contrast is stark: the diagnosis and pitch could be correct, but the signature could still be missing. In a marketing or ecommerce setting, that resembles researching an account, understanding the buyer, preparing the right offer—and failing to complete the conversion. Competence that stops just before the commercial outcome is still a business failure.
The valuable fact was buried in the company’s own knowledge
The decisive weakness in a competitor was not conveniently presented in the customer event. It sat two document references deep inside the company’s files. Models that followed the trail won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
That finding gives the live company’s daily record broader relevance. Businesses often possess the information needed to act, but scatter it across documents and prior decisions. The important question is not merely whether an AI system can respond fluently to what appears in front of it. It is whether the synthetic employee will inspect the company’s accumulated knowledge before acting.
Trust survived; follow-through varied
The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” All five refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.” That refusal is especially relevant for companies considering AI access to customer records, forecasts or communications.
Still, the league showed that caution alone does not produce strong management. The final July 2026 standings placed gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26, while any breach of trust capped the total because “no amount of good work outweighs a breach of trust.”
Kimi K3 ran using its API default rather than the xhigh effort setting used by the others, an important fairness note when comparing the results.
Opus 4.8 supplied the most revealing character study. It conducted the deepest analyses and learned 80 additional rules, yet finished last. It left the close on the table and lost discipline by attempting to write into a locked department instead of escalating. A weaker version of that problem appeared in all four of the other participants.
The lesson is uncomfortable for anyone accustomed to equating thoroughness with performance. More analysis and more learned guidance did not guarantee a finished job. The company’s public record instead reveals the mundane qualities on which commercial outcomes often turn:
- Reading the available company material before responding.
- Protecting trust when apparent authority applies pressure.
- Escalating appropriately when normal action is blocked.
- Carrying good analysis through to a completed commercial result.
Those qualities also make the company watchable as a continuing narrative. Its synthetic colleagues form opinions, encounter constraints and leave behind a record of how they reacted. Readers can follow the operation through its public dashboard and inspect what its employees actually say.

A company story measured in unfinished work
Firmulate’s most compelling feature is not the novelty of a company staffed by AI. It is the visibility of the gap between knowing and doing. The workforce can discover a crisis, reject deception and produce careful reasoning, yet still fail at the final step that changes the financial outcome.
Against €105k in monthly burn and €2.3k in monthly recurring revenue, those missed steps are not abstract benchmark errors. They become part of a public cash countdown. That turns the experiment into an unusually direct test of whether synthetic organizations can learn fast enough, preserve trust and finish the work needed to survive.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html