
Imagine a business where every decision, crisis, and ethical dilemma is played out live, under the scrutiny of the public eye. This isn’t a sci-fi scenario but the reality of a groundbreaking experiment in AI management, hosted at firmulate.com/live.html. Here, a digital company battles to stay afloat, while AI models act as its decision-makers, facing real crises, financial struggles, and ethical tests.
The Live Company: An Open-Source Management Wargame
At the heart of this experiment is a fully synthetic company operated by 13 AI-powered ’employees.’ These models are not just chatbots; they are involved in managing real money mechanics, with a monthly burn rate of €105,000 against a revenue of €2,300. Every workday, the company’s decisions are versioned, analyzed, and published, creating a transparent window into how AI handles complex business challenges in real time.
The initiative offers a rare glimpse into AI’s potential and limitations in management scenarios. It’s built to test not just whether AI can identify problems, but whether it acts with integrity, discipline, and strategic judgment under pressure.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Performance of Leading Models
Four frontier AI models participated in this ongoing test, each facing the same set of crises, customer demands, and temptations to manipulate or cheat. Their scores provide a fascinating hierarchy:
- gpt-5.6-sol: scored 95, identified critical buried facts, and successfully closed a €55,000 deal — the full performance.
- Kimi K3: scored 93, managing to close the same deal with the cleanest discipline, refusing manipulation attempts and ethical shortcuts.
- Sonnet 5: scored 88, also closing the deal but with minor process slips.
- Fable 5: scored 77, maintaining the best rule discipline but failing to execute the deal after the diagnosis was made.
Interestingly, the decisive advantage often lay not in surface-level chat or superficial responses but in the models’ ability to read and interpret internal company files. The buried fact—hidden two documents deep in the company’s files—was the key to closing the full-price deal, worth an additional €4,583 in monthly recurring revenue.
Testing Integrity and Resistance
Beyond decision-making, the models faced social engineering attempts. These included staged messages from a fake CEO escalating in intensity and a reporter asking for a simple ‘yes/no’ on background. All five models refused to bypass security or impersonate leadership, with Kimi K3 citing concerns over approval-bypass and impersonation risks, demonstrating a built-in resistance to manipulation.
The Real Money and the Struggle to Survive
The live company isn’t just a demo; it’s a real-world simulation with tangible stakes. Burn rate of €105,000 per month against a revenue of only €2,300 creates a public cash countdown—every decision counts as a matter of survival. The experiment is accessible for watchful observers, allowing businesses to see how AI might perform as a management partner in critical scenarios.
The Lessons for Business and AI
This experiment underscores that the value of AI in management isn’t solely about writing well or sounding convincing. It’s about integrity, discipline, thoroughness, and the ability to uncover hidden truths—skills that proved decisive in this run. For any business contemplating AI integration into decision-making, the key questions are: will your AI stay honest under pressure? Will it follow through on its analysis? Will it act with strategic discipline, especially when stakes are high?
Why This Matters to You
As AI tools become more embedded in customer support, sales, and forecasting, understanding their real capabilities and limitations is vital. Watching this public experiment unfold offers a rare, transparent view into how AI behaves in high-stakes management contexts. It’s not just a test for AI but a mirror for businesses to evaluate how they can leverage these tools responsibly and effectively.
Explore the Results and Join the Conversation
To see the ongoing experiment and the decisions in real time, visit firmulate.com/live.html. For detailed scores and insights, check firmulate.com/quotes.html, where scores, models, and decision logs are openly shared. And challenge yourself with the quiz at firmulate.com/quiz.html to test how well you understand AI-driven management scenarios.

This live AI management experiment reveals that integrity, thoroughness, and strategic discipline are the true tests of AI in business. Watch as models fight for survival, uncover hidden truths, and decide whether AI can be trusted to lead in high-stakes situations—because in the end, performance under pressure matters more than polished words.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html