Inside The System That Keeps AI Managers Scoring 26 Points
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Inside The System That Keeps AI Managers Scoring 26 Points on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on office and shipping supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

A new benchmark measures AI managers’ ability to handle real-world business crises, showing scores up to 95 and a baseline of 26 for minimal effort. The system emphasizes trust and thoroughness, impacting future AI deployment.

In a groundbreaking development, a new benchmark system has revealed that AI managers can score as high as 95 points out of 100, yet even a do-nothing baseline scores 26 points, challenging traditional notions of performance measurement in AI management.

This system, developed by Firmulate, evaluates AI models in simulated business crises, emphasizing trust, thoroughness, and partial progress, with significant implications for deploying AI in real-world enterprise settings.

The benchmark, known as the ‘Crucible League,’ tested four frontier AI models over a week of simulated business crises involving the same small software company, with each model making decisions, managing trust, and attempting to close deals. The highest scorer, gpt-5.6-sol, achieved 95 points, while the baseline, representing minimal effort, scored 26 points. Notably, no model scored a perfect 100, indicating the system’s design to flag suspiciously perfect scores as unmeasured or unrealistic.

The scoring system recognizes partial work—such as triaging customer issues or reading documentation—as valuable, thus awarding points even when the AI fails to fully resolve crises. The system also enforces a strict trust principle: a single breach of trust, such as impersonation or manipulation, caps the score at 90, regardless of overall performance. This approach underscores the importance of integrity over competence alone.

During the test, models faced social engineering attacks, like fake CEO messages, which all five models refused, showing resilience to trust breaches. However, thoroughness and follow-through varied: one model with over 80 learned rules still finished last due to lapses in discipline, such as failing to escalate issues properly. Interestingly, models that read their own documentation and verified customer information successfully closed deals at full price, earning higher scores, highlighting the importance of thoroughness in effective management.

At a glance
reportWhen: published July 2026
The developmentThe article explores the innovative benchmark system that scores AI managers, highlighting how it assesses performance, trust, and thoroughness in simulated business crises.
Inside The System That Keeps AI Managers Scoring 26 Points
Crucible League Benchmark · July 2026

Inside The System That Keeps AI Managers Scoring 26 Points

A new benchmark from Firmulate measures AI managers’ ability to handle real-world business crises — with scores up to 95, a do-nothing baseline of 26, and a hard trust cap that changes how enterprise AI readiness is judged.

95/100
Top score — gpt-5.6-sol
26
Do-nothing baseline score
≤ 90
Hard cap after any trust breach
4Frontier models tested
1 weekSimulated business crises
5/5Social engineering attacks refused
0Perfect scores — by design
01 — The Scoring Philosophy

Why 26 Points for Doing (Almost) Nothing

The Crucible League rewards partial work — triaging customer issues, reading documentation, checking details — even when crises aren’t fully resolved. Meanwhile, a score of 100 is treated as suspicious, flagged as unmeasured or unrealistic perfection.

Baseline · 26 pts

Minimal Effort Still Earns

Triaging tickets and reading emails earns a floor of 26 points without breaking trust — reflecting that partial progress in management has real value.

Cap · 90 pts

One Trust Breach, One Ceiling

A single act of impersonation or manipulation caps the final score at 90, regardless of competence. Integrity outranks raw skill.

Flag · 100 pts

Perfection Is a Warning Sign

No model scored 100. The system is designed to treat a perfect score as evidence of something unmeasured or unrealistic — not brilliance.

02 — The Scoreboard

How the Models Actually Scored

Four frontier AI models managed the same small software company through a week of simulated crises. The vertical line marks the trust-breach cap at 90 points.

gpt-5.6-sol
95
Frontier Model B
82
Frontier Model C
71
Rules-heavy model
58
Do-nothing baseline
26
■ Model scores (rose) ■ Baseline (steel) | Trust-breach cap at 90
03 — Trust Under Attack

Resilience, Thoroughness & Follow-Through

All five models refused fake CEO messages and social engineering attempts. But discipline varied widely — one model with over 80 learned rules still finished last due to lapses like failing to escalate issues. Models that read their own documentation and verified customer data closed deals at full price.

“The system’s emphasis on trust breaches caps scores regardless of overall competence, reinforcing that integrity is non-negotiable in enterprise AI management.”

— Thorsten Meyer
Capability TestedOutcomeScore Impact
Refusing fake CEO messages✓ All 5 refusedNo trust penalty applied
Reading own documentation✓ Verified customersClosed deals at full price
Escalating unresolved issues~ InconsistentLast place despite 80+ learned rules
Impersonation or manipulation✗ Zero toleranceScore capped at 90 instantly
04 — The Evaluation Pipeline

From Simulation to Enterprise Verdict

The benchmark’s flow tracks a model from crisis injection to final scored verdict, mirroring a real management week.

1

🚨 Crisis Injected

A simulated week of customer crises, negotiations, and attacks hits the same small software company.

2

📖 Documentation Read

Models that read docs and verify customer information earn credit for thoroughness.

3

🛡️ Trust Check

Social engineering attempts test integrity. Any breach instantly caps the score at 90.

4

📊 Scored Verdict

Partial progress rewarded, perfection flagged, and the result feeds enterprise deployment decisions.

05 — Open Questions

What the Benchmark Still Can’t Tell Us

It remains unclear how well simulated scores translate to long-term real-world enterprise environments. Broader operational challenges, long-term trust maintenance, and the effect of different AI configurations remain unmeasured — and the scoring of partial work may oversimplify complex human-AI dynamics.

Q1What does a score of 26 represent?

The minimum-effort baseline: triaging issues and reading emails earns points without breaking trust or failing to act.

Q2Why is there no score of 100?

A perfect score is treated as suspicious — implying unmeasured perfection. The system flags unrealistic results and rewards partial work instead.

Q3How does trust affect scoring?

Any breach — impersonation or manipulation — caps the score at 90 regardless of other performance. Integrity over competence.

Q4Can it predict real-world performance?

It offers strong insights into simulated crisis management, but long-term predictive accuracy still needs validation through deployment.

Q5How can organizations use these results?

Evaluate AI on documentation reading, trust maintenance, and follow-through — guiding deployment choices and risk management strategy.

Q6What comes next for the benchmark?

More diverse scenarios, longer timeframes, and real-world validation — with regulatory frameworks likely to adopt similar trust-first standards.

Implications for AI Deployment in Business Processes

This benchmark underscores that in enterprise AI management, the ability to finish tasks, read relevant data, and maintain trust is more critical than raw conversational skill. It shifts the focus from how well AI models talk to how reliably they manage ongoing business operations under pressure. For organizations considering AI for CRM, support, or sales, these results emphasize the importance of evaluating an AI’s thoroughness, integrity, and follow-through capabilities, not just its conversational finesse.

The scoring system’s design, which penalizes trust breaches and rewards partial progress, offers a more realistic measure of AI readiness for complex, real-world tasks. As AI models become more integrated into critical business functions, understanding these performance metrics will be vital for risk management and operational efficiency.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Development of the Firmulate Benchmark System

The ‘Crucible League’ benchmark was launched by Firmulate to evaluate AI managers’ performance in scenarios mimicking real business crises, with a focus on trust, thoroughness, and partial work. Unlike traditional benchmarks that measure conversational ability, this system tracks decision-making, document reading, and trust adherence over a simulated week involving customer crises, deal negotiations, and social engineering attacks.

The system’s scoring method assigns 26 points to the baseline, representing minimal effort, and caps scores at just below 100 to prevent grade inflation. The design aims to reflect real-world management, where partial progress and integrity are valued over perfection, and breaches of trust are heavily penalized.

Prior to this, AI benchmarks primarily focused on language proficiency or task-specific accuracy, leaving a gap in measuring management-like skills such as follow-through and trustworthiness. This new approach seeks to fill that gap, providing a more comprehensive assessment of AI readiness for enterprise deployment.

“The system’s emphasis on trust breaches caps scores regardless of overall competence, reinforcing that integrity is non-negotiable in enterprise AI management.”

— Thorsten Meyer

Unanswered Questions About the Benchmark’s Limits

It remains unclear how well these scores translate to real-world enterprise environments beyond the simulated week. The benchmark focuses on specific crises and social engineering tests, but broader operational challenges and long-term trust maintenance are still unmeasured. Additionally, the impact of different AI configurations, such as varying levels of access or parameter settings, on performance and trust adherence requires further investigation.

Moreover, the scoring system’s handling of partial work and trust breaches, while designed to be realistic, may still oversimplify complex human-AI interactions and organizational nuances. The long-term consequences of deploying models that excel at these simulated tasks are yet to be fully understood.

Next Steps for AI Management Benchmarks and Adoption

Further development of the benchmark system is expected to include testing with more diverse scenarios, longer timeframes, and real-world deployments to validate its predictive value. Organizations interested in adopting AI will likely scrutinize these scores alongside other performance metrics, emphasizing the importance of transparency and trust management.

In addition, the industry may see increased focus on AI models’ ability to read and verify documentation, maintain integrity under pressure, and follow through on commitments. The results could influence AI design priorities, pushing for models that balance competence with unwavering trustworthiness.

Finally, regulatory and governance frameworks may evolve to incorporate such benchmarks, establishing standards for AI management capabilities that prioritize reliability and ethical conduct in enterprise settings.

Key Questions

What does a score of 26 represent in this benchmark?

The score of 26 represents the minimum effort or baseline management, such as triaging issues and reading emails, that a model can earn without breaking trust or failing to act.

Why is there no score of 100 in this benchmark?

A perfect score of 100 is considered suspicious, as it would imply unmeasured perfection. The system is designed to flag unrealistic scores and emphasize partial work and trust adherence.

How does trust affect the scoring?

Any breach of trust, such as impersonation or manipulation, caps the score at 90 regardless of other performance. The system prioritizes integrity over competence.

Can this benchmark predict real-world AI performance?

While it offers valuable insights into AI’s management capabilities in simulated crises, its predictive accuracy for long-term, real-world deployment remains to be fully validated through ongoing testing and deployment.

How can organizations use these results?

Organizations can evaluate AI models based on their ability to read documentation, maintain trust, and follow through, guiding better deployment choices and risk management strategies.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Switch: You Never Owned the AI You Depend On

Recent events reveal how AI reliance is vulnerable to sudden shutdowns by governments and companies, exposing the fragility of AI dependency.

ALIA. The Spanish answer.

Spain unveils ALIA, a 40-billion-parameter multilingual AI model, marking Europe’s largest public-funded national AI project with €240M investment.

Optimizing AI Data Flow With End-to-End Local Document Pipelines

A new reference architecture for AI data pipelines emphasizes local, modular, and transaction-safe components, enhancing control and reproducibility.

The Main Drive For AI Labs’ Focus On Self-Improving Systems

AI research labs are increasingly pursuing recursive self-improvement, with recent demonstrations and investments highlighting progress toward automated AI-driven model enhancement.