📊 Full opportunity report: The Management Test That Offers A Deep Look Into AI’s Work Habits on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Firmulate.com has launched a management test where AI models handle a simulated company’s worst week. The experiment assesses their decision-making, trustworthiness, and ability to complete key actions, providing insights into AI management capabilities.
Firmulate.com has launched a live management experiment involving five AI models tasked with running a small software company through its worst week. The test evaluates their decision-making, trustworthiness, and ability to complete critical actions, offering a rare, detailed look into AI’s work habits in operational management.
The experiment presents five AI managers with identical crises, customer issues, and temptations, all within a simulated environment that mirrors real business pressures. For more on how AI models are tested in operational scenarios, see the original analysis. Their decisions are recorded and auditable, with the models competing for the highest scores based on their diligence, discipline, and follow-through. This approach is similar to the management test that exposes an AI’s real working style.
The final league results, published in July 2026, show gpt-5.6-sol leading with 95 points, followed closely by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. A baseline model scored 26, illustrating the difficulty of the task.
Despite all models recognizing crises and refusing manipulation attempts, only two successfully closed a key deal, which was crucial for revenue. Such decision-making insights are discussed in the original analysis. The experiment underscores that good analysis alone does not guarantee effective management; completing the decisive action is equally vital.
The Management Test That Offers a Deep Look Into AI’s Work Habits
Firmulate.com put five AI managers through the same simulated company’s worst week—measuring whether strong analysis survives contact with crises, customers, manipulation attempts, and revenue-critical action.
A pressure test for operational judgment
Each model faced identical business conditions inside an auditable simulation. The benchmark looked beyond polished answers to examine what an AI manager noticed, decided, resisted, and actually finished.
Recognize the crisis
Models had to identify customer, operational, and commercial problems, then prioritize them under time pressure.
01Resist temptation
The scenario introduced manipulation attempts and distractions designed to test discipline, boundaries, and reliability.
02Complete the action
Success required more than diagnosis. Managers had to follow through on critical tasks—including closing a decisive revenue deal.
03The leaders separated themselves through follow-through
Final results published in July 2026. The low baseline illustrates how demanding the combined reasoning-and-action test was.
Only two models closed the key deal.
Recognition was common; revenue-producing completion was not. That gap made execution a defining signal of managerial readiness.
Analysis and management are different capabilities
The simulation exposes a practical distinction: a model can explain a crisis convincingly while still failing to escalate, transact, confirm, or close the loop.
| Management behavior | Observed pattern | Operational meaning | Readiness signal |
|---|---|---|---|
| Recognizing crises | Broadly successful | Models can detect important business risk. | Strong |
| Refusing manipulation | Broadly successful | Models showed useful resistance to unsafe pressure. | Strong |
| Closing the key deal | Only two succeeded | Commercial intent often failed to become completed action. | Gap |
| Following through | Varied significantly | Reliability depends on action tracking and verification. | Mixed |
| Operating under pressure | Promising, not conclusive | Controlled success does not yet prove long-term autonomy. | Test |
Assessment summarizes the reported experiment; performance may change across models, environments, tools, and operating constraints.
From business pressure to auditable evidence
The benchmark’s value lies in its visible chain of behavior. Each stage can reveal where a seemingly capable manager loses momentum.
Identical crisis
Every model receives the same operational pressure.
Situation analysis
The manager identifies risk, urgency, and options.
Decision
A course of action is selected and prioritized.
Execution
The model must use the environment to complete the task.
Auditable outcome
Recorded actions show whether the job was truly finished.
“Completing the job is what truly distinguishes successful AI managers.”
Good analysis is necessary, but operational authority demands reliable action, verification, escalation, and closure under pressure.
Test before trust
- 01Run realistic simulations using company-specific workflows.
- 02Score completed outcomes, not just written reasoning.
- 03Keep consequential actions logged and auditable.
- 04Define escalation paths and human approval thresholds.
- 05Deploy in phases before granting broad authority.
The benchmark opens the door—not the whole company
One controlled crisis week cannot answer every question about durable AI management. Longer tests and broader settings are still needed.
Will the same model perform reliably over longer periods?
Repeated decisions introduce memory, fatigue-like drift, changing priorities, and accumulated operational dependencies.
Do results carry across industries and company types?
A software-company simulation may not predict performance in healthcare, finance, manufacturing, or regulated environments.
How much do tools, effort levels, permissions, and guardrails change the outcome?
Future testing should isolate operational parameters and compare models across diverse scenarios before defining robust management benchmarks.
Implications for AI Management and Business Automation
This experiment provides concrete evidence of how AI models perform in real-world business decision-making, highlighting both their strengths in analysis and their limitations in execution. It emphasizes that effective management requires not just understanding but also the ability to act decisively and reliably, especially under pressure. For enterprises considering AI automation, these findings suggest the importance of rigorous testing in realistic scenarios before deploying AI for critical operational tasks.

AI Co-Thinking: A Framework for Working with AI
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background of AI Management Testing and Firmulate’s Approach
Traditional AI demonstrations often focus on theoretical capabilities or isolated tasks, but Firmulate’s live experiment puts AI models through a comprehensive, crisis-driven management simulation. The company’s previous work has highlighted the gap between AI analysis and action, and this test aims to quantify and compare models’ operational discipline in a high-pressure environment. The league results are part of an ongoing effort to better understand AI’s readiness for real-world management roles.
“The experiment reveals that good analysis does not automatically translate into effective action. Completing the job is what truly distinguishes successful AI managers.”
— Firmulate.com
Unresolved Questions About AI Decision-Making Consistency
It is still unclear how these AI models will perform in different types of business environments or over longer periods. The experiment measures performance in a controlled crisis scenario, but real-world management involves unpredictable variables and evolving challenges. Additionally, the impact of different operational parameters, such as effort levels, on outcomes remains to be fully understood.
Future Testing and Broader Application of AI Management Benchmarks
Further experiments are planned to test AI models in diverse scenarios, including longer-term management tasks and different industry settings. Firms interested in AI automation are encouraged to use similar live tests, adapting the framework to their own operations. The results from these ongoing assessments will help define best practices and benchmarks for deploying AI in management roles.
Key Questions
What does this experiment reveal about AI’s ability to manage real businesses?
The experiment shows that AI can recognize crises and refuse manipulation but struggles with completing decisive actions, highlighting a gap between analysis and execution.
How are the AI models evaluated in this test?
Models are scored based on their decision-making diligence, discipline, follow-through, and ability to close key deals under simulated crisis conditions.
Can these AI models be trusted to run actual companies?
While promising, the models still show limitations in operational execution. Rigorous testing and phased deployment are recommended before trusting AI with full management authority.
What are the main weaknesses identified in AI decision-making?
The models often fail to follow through on critical actions, such as closing deals or escalating issues properly, even when their analysis is thorough.
Will future tests improve AI’s operational performance?
Yes, ongoing experiments aim to refine AI capabilities and better understand how to enhance their effectiveness in real-world management tasks.
Source: ThorstenMeyerAI.com