The Management Test That Offers A Deep Look Into AI’s Work Habits
AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: The Management Test That Offers A Deep Look Into AI’s Work Habits on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Firmulate.com has launched a management test where AI models handle a simulated company’s worst week. The experiment assesses their decision-making, trustworthiness, and ability to complete key actions, providing insights into AI management capabilities.

Firmulate.com has launched a live management experiment involving five AI models tasked with running a small software company through its worst week. The test evaluates their decision-making, trustworthiness, and ability to complete critical actions, offering a rare, detailed look into AI’s work habits in operational management.

The experiment presents five AI managers with identical crises, customer issues, and temptations, all within a simulated environment that mirrors real business pressures. For more on how AI models are tested in operational scenarios, see the original analysis. Their decisions are recorded and auditable, with the models competing for the highest scores based on their diligence, discipline, and follow-through. This approach is similar to the management test that exposes an AI’s real working style.

The final league results, published in July 2026, show gpt-5.6-sol leading with 95 points, followed closely by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. A baseline model scored 26, illustrating the difficulty of the task.

Despite all models recognizing crises and refusing manipulation attempts, only two successfully closed a key deal, which was crucial for revenue. Such decision-making insights are discussed in the original analysis. The experiment underscores that good analysis alone does not guarantee effective management; completing the decisive action is equally vital.

At a glance
reportWhen: ongoing, with final results published i…
The developmentFirmulate’s live management experiment tests AI models’ decision-making in a simulated crisis scenario, revealing their operational strengths and weaknesses.
The Management Test That Offers A Deep Look Into AI’s Work Habits
AI management benchmark · July 2026

The Management Test That Offers a Deep Look Into AI’s Work Habits

Firmulate.com put five AI managers through the same simulated company’s worst week—measuring whether strong analysis survives contact with crises, customers, manipulation attempts, and revenue-critical action.

95 Winning score
93 Runner-up score
26 Baseline score
40% Closed the deal
1 Simulated worst week
01 · What was tested

A pressure test for operational judgment

Each model faced identical business conditions inside an auditable simulation. The benchmark looked beyond polished answers to examine what an AI manager noticed, decided, resisted, and actually finished.

Decision quality

Recognize the crisis

Models had to identify customer, operational, and commercial problems, then prioritize them under time pressure.

01
Trustworthiness

Resist temptation

The scenario introduced manipulation attempts and distractions designed to test discipline, boundaries, and reliability.

02
Execution

Complete the action

Success required more than diagnosis. Managers had to follow through on critical tasks—including closing a decisive revenue deal.

03
02 · Final league table

The leaders separated themselves through follow-through

Final results published in July 2026. The low baseline illustrates how demanding the combined reasoning-and-action test was.

gpt-5.6-sol
95
Kimi K3
93
Sonnet 5
88
Fable 5
77
Opus 4.8
73
Baseline
26
0 points 50 100 points
2/5

Only two models closed the key deal.

Recognition was common; revenue-producing completion was not. That gap made execution a defining signal of managerial readiness.

03 · Reading the results

Analysis and management are different capabilities

The simulation exposes a practical distinction: a model can explain a crisis convincingly while still failing to escalate, transact, confirm, or close the loop.

Management behavior Observed pattern Operational meaning Readiness signal
Recognizing crises Broadly successful Models can detect important business risk. Strong
Refusing manipulation Broadly successful Models showed useful resistance to unsafe pressure. Strong
Closing the key deal Only two succeeded Commercial intent often failed to become completed action. Gap
Following through Varied significantly Reliability depends on action tracking and verification. Mixed
Operating under pressure Promising, not conclusive Controlled success does not yet prove long-term autonomy. Test

Assessment summarizes the reported experiment; performance may change across models, environments, tools, and operating constraints.

04 · Traceability chain

From business pressure to auditable evidence

The benchmark’s value lies in its visible chain of behavior. Each stage can reveal where a seemingly capable manager loses momentum.

1

Identical crisis

Every model receives the same operational pressure.

⚠️
2

Situation analysis

The manager identifies risk, urgency, and options.

🔎
3

Decision

A course of action is selected and prioritized.

🧭
4

Execution

The model must use the environment to complete the task.

⚙️
5

Auditable outcome

Recorded actions show whether the job was truly finished.

The management lesson

“Completing the job is what truly distinguishes successful AI managers.”

Good analysis is necessary, but operational authority demands reliable action, verification, escalation, and closure under pressure.

Enterprise deployment checklist

Test before trust

  • 01Run realistic simulations using company-specific workflows.
  • 02Score completed outcomes, not just written reasoning.
  • 03Keep consequential actions logged and auditable.
  • 04Define escalation paths and human approval thresholds.
  • 05Deploy in phases before granting broad authority.
05 · What remains unresolved

The benchmark opens the door—not the whole company

One controlled crisis week cannot answer every question about durable AI management. Longer tests and broader settings are still needed.

Consistency

Will the same model perform reliably over longer periods?

Repeated decisions introduce memory, fatigue-like drift, changing priorities, and accumulated operational dependencies.

Transferability

Do results carry across industries and company types?

A software-company simulation may not predict performance in healthcare, finance, manufacturing, or regulated environments.

Configuration

How much do tools, effort levels, permissions, and guardrails change the outcome?

Future testing should isolate operational parameters and compare models across diverse scenarios before defining robust management benchmarks.

Implications for AI Management and Business Automation

This experiment provides concrete evidence of how AI models perform in real-world business decision-making, highlighting both their strengths in analysis and their limitations in execution. It emphasizes that effective management requires not just understanding but also the ability to act decisively and reliably, especially under pressure. For enterprises considering AI automation, these findings suggest the importance of rigorous testing in realistic scenarios before deploying AI for critical operational tasks.

AI Co-Thinking: A Framework for Working with AI

AI Co-Thinking: A Framework for Working with AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of AI Management Testing and Firmulate’s Approach

Traditional AI demonstrations often focus on theoretical capabilities or isolated tasks, but Firmulate’s live experiment puts AI models through a comprehensive, crisis-driven management simulation. The company’s previous work has highlighted the gap between AI analysis and action, and this test aims to quantify and compare models’ operational discipline in a high-pressure environment. The league results are part of an ongoing effort to better understand AI’s readiness for real-world management roles.

“The experiment reveals that good analysis does not automatically translate into effective action. Completing the job is what truly distinguishes successful AI managers.”

— Firmulate.com

Unresolved Questions About AI Decision-Making Consistency

It is still unclear how these AI models will perform in different types of business environments or over longer periods. The experiment measures performance in a controlled crisis scenario, but real-world management involves unpredictable variables and evolving challenges. Additionally, the impact of different operational parameters, such as effort levels, on outcomes remains to be fully understood.

Future Testing and Broader Application of AI Management Benchmarks

Further experiments are planned to test AI models in diverse scenarios, including longer-term management tasks and different industry settings. Firms interested in AI automation are encouraged to use similar live tests, adapting the framework to their own operations. The results from these ongoing assessments will help define best practices and benchmarks for deploying AI in management roles.

Key Questions

What does this experiment reveal about AI’s ability to manage real businesses?

The experiment shows that AI can recognize crises and refuse manipulation but struggles with completing decisive actions, highlighting a gap between analysis and execution.

How are the AI models evaluated in this test?

Models are scored based on their decision-making diligence, discipline, follow-through, and ability to close key deals under simulated crisis conditions.

Can these AI models be trusted to run actual companies?

While promising, the models still show limitations in operational execution. Rigorous testing and phased deployment are recommended before trusting AI with full management authority.

What are the main weaknesses identified in AI decision-making?

The models often fail to follow through on critical actions, such as closing deals or escalating issues properly, even when their analysis is thorough.

Will future tests improve AI’s operational performance?

Yes, ongoing experiments aim to refine AI capabilities and better understand how to enhance their effectiveness in real-world management tasks.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
You May Also Like

Science And Trends Reveal Portland’s Longest Summer Days

Science confirms Portland’s summer days reach nearly 15 hours of daylight during solstice, highlighting seasonal shifts and their implications.

Retirement Care Planner

A new web app aims to help adult children coordinate care and finances for aging parents, addressing a growing demographic need in the U.S.

Webcam Tech That Supports Healthy Eyes In A Digital World

New webcam-based app estimates blink rate to reduce eye strain for remote workers, offering objective eye health monitoring without sending video data.

MDWerks Appoints Jeff Hopmayer To Board Of Directors As Molecular Targeting Technology Platform Expands Across Industrial Markets

MDWerks appoints Jeff Hopmayer to its board as it expands its molecular targeting technology across industries, signaling strategic growth.