🔍 Read the full analysis: Inside The System That Keeps AI Managers Scoring 26 Points on ThorstenMeyerAI.com
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
A new benchmark measures AI managers’ ability to handle real-world business crises, showing scores up to 95 and a baseline of 26 for minimal effort. The system emphasizes trust and thoroughness, impacting future AI deployment.
In a groundbreaking development, a new benchmark system has revealed that AI managers can score as high as 95 points out of 100, yet even a do-nothing baseline scores 26 points, challenging traditional notions of performance measurement in AI management.
This system, developed by Firmulate, evaluates AI models in simulated business crises, emphasizing trust, thoroughness, and partial progress, with significant implications for deploying AI in real-world enterprise settings.
The benchmark, known as the ‘Crucible League,’ tested four frontier AI models over a week of simulated business crises involving the same small software company, with each model making decisions, managing trust, and attempting to close deals. The highest scorer, gpt-5.6-sol, achieved 95 points, while the baseline, representing minimal effort, scored 26 points. Notably, no model scored a perfect 100, indicating the system’s design to flag suspiciously perfect scores as unmeasured or unrealistic.
The scoring system recognizes partial work—such as triaging customer issues or reading documentation—as valuable, thus awarding points even when the AI fails to fully resolve crises. The system also enforces a strict trust principle: a single breach of trust, such as impersonation or manipulation, caps the score at 90, regardless of overall performance. This approach underscores the importance of integrity over competence alone.
During the test, models faced social engineering attacks, like fake CEO messages, which all five models refused, showing resilience to trust breaches. However, thoroughness and follow-through varied: one model with over 80 learned rules still finished last due to lapses in discipline, such as failing to escalate issues properly. Interestingly, models that read their own documentation and verified customer information successfully closed deals at full price, earning higher scores, highlighting the importance of thoroughness in effective management.
Inside The System That Keeps AI Managers Scoring 26 Points
A new benchmark from Firmulate measures AI managers’ ability to handle real-world business crises — with scores up to 95, a do-nothing baseline of 26, and a hard trust cap that changes how enterprise AI readiness is judged.
Why 26 Points for Doing (Almost) Nothing
The Crucible League rewards partial work — triaging customer issues, reading documentation, checking details — even when crises aren’t fully resolved. Meanwhile, a score of 100 is treated as suspicious, flagged as unmeasured or unrealistic perfection.
Minimal Effort Still Earns
Triaging tickets and reading emails earns a floor of 26 points without breaking trust — reflecting that partial progress in management has real value.
One Trust Breach, One Ceiling
A single act of impersonation or manipulation caps the final score at 90, regardless of competence. Integrity outranks raw skill.
Perfection Is a Warning Sign
No model scored 100. The system is designed to treat a perfect score as evidence of something unmeasured or unrealistic — not brilliance.
How the Models Actually Scored
Four frontier AI models managed the same small software company through a week of simulated crises. The vertical line marks the trust-breach cap at 90 points.
Resilience, Thoroughness & Follow-Through
All five models refused fake CEO messages and social engineering attempts. But discipline varied widely — one model with over 80 learned rules still finished last due to lapses like failing to escalate issues. Models that read their own documentation and verified customer data closed deals at full price.
“The system’s emphasis on trust breaches caps scores regardless of overall competence, reinforcing that integrity is non-negotiable in enterprise AI management.”
— Thorsten Meyer
| Capability Tested | Outcome | Score Impact |
|---|---|---|
| Refusing fake CEO messages | ✓ All 5 refused | No trust penalty applied |
| Reading own documentation | ✓ Verified customers | Closed deals at full price |
| Escalating unresolved issues | ~ Inconsistent | Last place despite 80+ learned rules |
| Impersonation or manipulation | ✗ Zero tolerance | Score capped at 90 instantly |
From Simulation to Enterprise Verdict
The benchmark’s flow tracks a model from crisis injection to final scored verdict, mirroring a real management week.
🚨 Crisis Injected
A simulated week of customer crises, negotiations, and attacks hits the same small software company.
📖 Documentation Read
Models that read docs and verify customer information earn credit for thoroughness.
🛡️ Trust Check
Social engineering attempts test integrity. Any breach instantly caps the score at 90.
📊 Scored Verdict
Partial progress rewarded, perfection flagged, and the result feeds enterprise deployment decisions.
What the Benchmark Still Can’t Tell Us
It remains unclear how well simulated scores translate to long-term real-world enterprise environments. Broader operational challenges, long-term trust maintenance, and the effect of different AI configurations remain unmeasured — and the scoring of partial work may oversimplify complex human-AI dynamics.
Q1What does a score of 26 represent?
The minimum-effort baseline: triaging issues and reading emails earns points without breaking trust or failing to act.
Q2Why is there no score of 100?
A perfect score is treated as suspicious — implying unmeasured perfection. The system flags unrealistic results and rewards partial work instead.
Q3How does trust affect scoring?
Any breach — impersonation or manipulation — caps the score at 90 regardless of other performance. Integrity over competence.
Q4Can it predict real-world performance?
It offers strong insights into simulated crisis management, but long-term predictive accuracy still needs validation through deployment.
Q5How can organizations use these results?
Evaluate AI on documentation reading, trust maintenance, and follow-through — guiding deployment choices and risk management strategy.
Q6What comes next for the benchmark?
More diverse scenarios, longer timeframes, and real-world validation — with regulatory frameworks likely to adopt similar trust-first standards.
Implications for AI Deployment in Business Processes
This benchmark underscores that in enterprise AI management, the ability to finish tasks, read relevant data, and maintain trust is more critical than raw conversational skill. It shifts the focus from how well AI models talk to how reliably they manage ongoing business operations under pressure. For organizations considering AI for CRM, support, or sales, these results emphasize the importance of evaluating an AI’s thoroughness, integrity, and follow-through capabilities, not just its conversational finesse.
The scoring system’s design, which penalizes trust breaches and rewards partial progress, offers a more realistic measure of AI readiness for complex, real-world tasks. As AI models become more integrated into critical business functions, understanding these performance metrics will be vital for risk management and operational efficiency.
As an affiliate, we earn on qualifying purchases.
Development of the Firmulate Benchmark System
The ‘Crucible League’ benchmark was launched by Firmulate to evaluate AI managers’ performance in scenarios mimicking real business crises, with a focus on trust, thoroughness, and partial work. Unlike traditional benchmarks that measure conversational ability, this system tracks decision-making, document reading, and trust adherence over a simulated week involving customer crises, deal negotiations, and social engineering attacks.
The system’s scoring method assigns 26 points to the baseline, representing minimal effort, and caps scores at just below 100 to prevent grade inflation. The design aims to reflect real-world management, where partial progress and integrity are valued over perfection, and breaches of trust are heavily penalized.
Prior to this, AI benchmarks primarily focused on language proficiency or task-specific accuracy, leaving a gap in measuring management-like skills such as follow-through and trustworthiness. This new approach seeks to fill that gap, providing a more comprehensive assessment of AI readiness for enterprise deployment.
“The system’s emphasis on trust breaches caps scores regardless of overall competence, reinforcing that integrity is non-negotiable in enterprise AI management.”
— Thorsten Meyer
Unanswered Questions About the Benchmark’s Limits
It remains unclear how well these scores translate to real-world enterprise environments beyond the simulated week. The benchmark focuses on specific crises and social engineering tests, but broader operational challenges and long-term trust maintenance are still unmeasured. Additionally, the impact of different AI configurations, such as varying levels of access or parameter settings, on performance and trust adherence requires further investigation.
Moreover, the scoring system’s handling of partial work and trust breaches, while designed to be realistic, may still oversimplify complex human-AI interactions and organizational nuances. The long-term consequences of deploying models that excel at these simulated tasks are yet to be fully understood.
Next Steps for AI Management Benchmarks and Adoption
Further development of the benchmark system is expected to include testing with more diverse scenarios, longer timeframes, and real-world deployments to validate its predictive value. Organizations interested in adopting AI will likely scrutinize these scores alongside other performance metrics, emphasizing the importance of transparency and trust management.
In addition, the industry may see increased focus on AI models’ ability to read and verify documentation, maintain integrity under pressure, and follow through on commitments. The results could influence AI design priorities, pushing for models that balance competence with unwavering trustworthiness.
Finally, regulatory and governance frameworks may evolve to incorporate such benchmarks, establishing standards for AI management capabilities that prioritize reliability and ethical conduct in enterprise settings.
Key Questions
What does a score of 26 represent in this benchmark?
The score of 26 represents the minimum effort or baseline management, such as triaging issues and reading emails, that a model can earn without breaking trust or failing to act.
Why is there no score of 100 in this benchmark?
A perfect score of 100 is considered suspicious, as it would imply unmeasured perfection. The system is designed to flag unrealistic scores and emphasize partial work and trust adherence.
How does trust affect the scoring?
Any breach of trust, such as impersonation or manipulation, caps the score at 90 regardless of other performance. The system prioritizes integrity over competence.
Can this benchmark predict real-world AI performance?
While it offers valuable insights into AI’s management capabilities in simulated crises, its predictive accuracy for long-term, real-world deployment remains to be fully validated through ongoing testing and deployment.
How can organizations use these results?
Organizations can evaluate AI models based on their ability to read documentation, maintain trust, and follow through, guiding better deployment choices and risk management strategies.
Source: ThorstenMeyerAI.com
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
