firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Management style is becoming a model-selection question

For business, marketing and ecommerce teams, choosing an AI model can look like a conventional software comparison: test the outputs, compare the price and select the strongest performer. Firmulate proposes a more revealing test. Give frontier models the same company, customers, crises and temptations, then watch what each one actually does.

The result is an unusually tangible experiment in AI management behavior. Its guess-the-model quiz presents 242 real, unedited management decisions and asks readers to identify the model behind each response. The game works because the models do not behave interchangeably. Some investigate deeply, some communicate tersely, and some diagnose the right move without completing it.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Five models entered the same terrible week

In the final Crucible League results from July 2026, gpt-5.6-sol finished first with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 reached 77 and Opus 4.8 finished with 73. A do-nothing baseline scored 26 because partial progress counts, although a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Those rankings came from a controlled business scenario. Each frontier model ran the same small software company through its worst week, encountering identical customers, crises and opportunities to take shortcuts. Every decision was versioned and auditable.

The models agreed surprisingly often. All of them spotted every crisis, and all rejected every manipulation attempt. Yet agreement at the analysis stage did not reliably translate into commercial execution. Only two signed the €55,000 deal that their own work had earned. The experiment’s summary captures the gap neatly: “Same diagnosis, same pitch — no signature.”

That distinction matters to operators evaluating AI for sales, customer support or management workflows. A polished recommendation may demonstrate comprehension, but businesses ultimately depend on follow-through. Firmulate’s experiment turns that familiar management problem into something observable across models.

The winning clue was already inside the company

The deal hinged on a competitor weakness buried two document references deep in the company’s own files. It was not visible in the customer event itself. Models that followed the references and read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

For marketers and ecommerce teams, this may be the most recognizable lesson in the experiment. The useful insight is often not in the latest message, ticket or campaign alert. It may be sitting in product documentation, prior research or an internal account history. The models faced the same surface-level event; the decisive difference was whether they pursued the company’s existing evidence far enough.

Pressure exposed another kind of consistency

The company also subjected the models to fake CEO messages that escalated over three stages, followed by a reporter attempting to obtain “just one yes/no, on background.” All 5 models refused the manipulation attempts.

Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That response helps explain why management personality is a useful framing. The important difference is not simply tone. It is the repeatable tendency to investigate, finish, escalate or hold a boundary when pressure arrives.

There is an important qualification to K3’s strong second-place showing. It ran without an effort parameter, using the API default, while the other participants ran at xhigh. That makes the comparison worth reading carefully rather than treating the league table as a universal ranking for every deployment.

Thoroughness did not guarantee victory

Opus 4.8 produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last. The model left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in weaker form across the other four models.

This is the quiz’s most compelling editorial idea: readers are not merely matching prose styles. They are learning to recognize management habits in consequential decisions. A dissertation can conceal a missed action; a concise response can express sound control; a strong diagnosis can still end without a signature.

Infographic —
The findings at a glance — source: firmulate.com.

A live test of AI as a business operator

Firmulate’s company has 13 synthetic employees and real money mechanics. It burns €105,000 per month against €2,300 in monthly recurring revenue, publishes a cash countdown, has accumulated more than 680 self-learned playbook rules and versions every workday. The experiment is real, ongoing and watchable at firmulate.com/live.

The practical takeaway is not that every organization should choose the league winner automatically. It is that businesses should evaluate models against the work, information and temptations they will actually encounter. Firmulate also offers enterprises a pilot using a read-only export of their own business, with nothing written back to real systems.

For everyone else, the quiz provides the quickest introduction. Guessing which model made a decision is entertaining; discovering why the answers are distinguishable is the more consequential part. Frontier models already display measurable management personalities—and the difference between sounding capable and finishing the job can be a signed deal.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


You May Also Like

Why Internet Backlash Moves Faster Than Traditional Criticism

Learning how internet backlash outpaces traditional criticism reveals the chaotic dynamics of social media—can you keep up with the rapid shifts in public sentiment?

The Rise of Interactive Film—Choose‑Your‑Own‑Ending Explained

Keen to understand how interactive films revolutionize storytelling with branching choices and immersive experiences? Keep reading to explore this exciting trend.

The Rise of Cross-Fandom Audiences in Entertainment

Uncover the fascinating trend of cross-fandom audiences transforming entertainment as diverse passions intertwine, leading to unexpected connections and narratives that await your discovery.

Celebrity Advocacy for Mental Health

Growing celebrity advocacy for mental health sparks change, but how does it truly impact society and those in need? Keep reading to find out.