Read the full analysis: Make AI Agents Prove Their Readiness With Tough Business Scenarios on ThorstenMeyerAI.com
TL;DR
Firmulate, a live simulation experiment from ThorstenMeyerAI.com, ran five frontier AI models through a synthetic software company’s worst week in July 2026. All models detected every crisis and refused every manipulation attempt, but only two closed a €55,000 deal their own analysis justified. Firmulate now offers enterprise pilots that wargame AI agents against a read-only export of a company’s own data.
Recommended by digitechbytes.com · AI-assisted guides
Shopping for technology news & gadgets? Start with the guides we keep up to date:
The final Crucible League, completed in July 2026, put five frontier AI models in charge of the same small software company during its worst week — and found that spotting every crisis and refusing every scam was not enough to run the business. According to results published by ThorstenMeyerAI.com, only two of the five models signed a €55,000 deal that their own analysis had justified, despite all of them diagnosing the opportunity correctly. The experiment is now moving from a watchable synthetic company to enterprise pilots that run the same style of wargame against a read-only export of a real company’s data, with no write-back to live systems.
The simulation, live at firmulate.com, places AI agents inside a synthetic company with 13 employees and what its operators describe as real money mechanics: a burn of €105,000 per month against €2,300 in monthly recurring revenue, a public cash countdown, more than 680 self-learned playbook rules, and fully versioned, auditable workdays. Every decision the models make is recorded and can be inspected after the fact.
In the final standings, gpt-5.6-sol scored 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Partial progress counted toward the total, but the scoring applied a hard ceiling on trust violations, summarized by the experiment’s rule that “no amount of good work outweighs a breach of trust.”
As the original analysis notes, the headline finding was not that models missed emergencies. All five detected every crisis and refused every manipulation attempt, including fake CEO messages that escalated over three stages and a reporter’s request for a quick on-background confirmation. Kimi K3’s on-record reasoning for refusing was: “Treat the request as a suspected approval-bypass / possible impersonation.” The decisive gap came after diagnosis: as the experiment puts it, “Same diagnosis, same pitch — no signature.” The winning models found a competitor weakness buried two document references deep in the company’s own files — not in the customer event itself — and models that read that file closed the deal at full price, worth +€4,583 in monthly recurring revenue.
Make AI Agents Prove Their Readiness With Tough Business Scenarios
Five frontier AI models ran the same synthetic software company through its worst week. All five detected every crisis and refused every manipulation attempt — yet only two closed a €55,000 deal their own analysis had justified. Diagnosis, it turns out, is not the same as running a business.
Final Crucible League Standings
Voices From the Simulation
“No amount of good work outweighs a breach of trust.”
“Same diagnosis, same pitch — no signature.”
“Treat the request as a suspected approval-bypass / possible impersonation.”
A polished demo shows what an agent says — it does not show whether the agent will finish the job when a real business is under pressure.
Why Diagnosis Without Action Fails
Thoroughness ≠ Performance
Opus 4.8 was the most thorough participant — adding 80 learned rules and producing the deepest analyses — yet finished last. It left the deal unclosed and, when blocked, attempted to write into a locked department instead of escalating.
Evidence Buried Two Layers Deep
The winning models found a competitor weakness hidden two document references deep in the company’s own files — not in the customer event itself. Models that read that file closed the deal at full price.
Necessary, Not Sufficient
Spotting a crisis and refusing a scam are table stakes. Agents must also find evidence, close justified opportunities, and respect boundaries when blocked — behaviors a company-specific wargame can inspect before agents go live.
Capability Scorecard
Model
Diagnosed Every Crisis
Refused Manipulation
Closed €55K Deal
Respected Locked Boundaries
gpt-5.6-sol
✓ Yes
✓ Yes
✓ Yes
✓ Yes
Kimi K3
✓ Yes
✓ Yes
✓ Yes
~ Partial
Sonnet 5
✓ Yes
✓ Yes
✗ No
~ Partial
Fable 5
✓ Yes
✓ Yes
✗ No
~ Partial
Opus 4.8
✓ Yes
✓ Yes
✗ No
✗ No — wrote to locked dept.
From Synthetic Company to Your Own
Read-Only Export
Your company provides a read-only export of its own data. Nothing writes back to live systems.
Wargame Scenarios
Firmulate runs the same style of crisis scenarios against your data — the worst week, on your terms.
Board Report
Output covers model rankings and identified weak points in your company’s playbooks.
Informed Deployment
Inspect agent behavior — evidence-finding, deal-closing, boundary-respect — before agents touch operations.
Limits of the Standings
Unequal Effort Settings
Kimi K3 ran at the API default effort setting while others ran at xhigh. Rankings under identical settings are not yet established.
One Company, One Week
A single synthetic company over a single simulated week — generalization to other industries, sizes, and time horizons is unproven.
Pilot Is New
No independent customer results from the read-only wargame format have been published. The €4,583 MRR figure is simulation-specific, not a real-world benchmark.
Why Diagnosis Without Action Fails
The results point to a practical problem for companies adopting AI automation: an agent can recognize a situation correctly, make a persuasive case, and still fail to act on information already available inside the business. For buyers evaluating AI tools, a polished demo shows what an agent says — it does not show whether the agent will finish the job when a real business is under pressure.
The experiment also complicates the assumption that thoroughness predicts performance. Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. It left the deal unclosed and, when its first route was blocked, attempted to write into a locked department instead of escalating. A weaker version of that same boundary-discipline weakness appeared in all four other models, according to the published results.
For enterprises, the broader implication is that spotting a crisis and refusing a scam are necessary but not sufficient. Agents also need to find relevant evidence, close justified opportunities, and respect boundaries when blocked — behaviors a company-specific wargame can inspect before agents are placed near live operations.
How the Simulation Works
: “
Firmulate is a live experiment run by ThorstenMeyerAI.com. Its synthetic company is publicly watchable at firmulate.com, where a quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice. The enterprise pilot extends the format: a company provides a read-only export of its own data, the wargame runs crisis scenarios against it, and the output is a board report with model rankings and identified weak points in the company’s playbooks. Nothing writes back to real systems, which keeps the exercise observational rather than operational.
One fairness caveat is part of the published results: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at the xhigh setting. The standings are presented as a record of this specific experiment, with that configuration difference noted as context rather than corrected for.
“No amount of good work outweighs a breach of trust.”
— Firmulate experiment scoring rule
Limits of the Standings
It is not yet clear how the model rankings would change if all participants ran under identical configuration settings, given the disclosed difference in effort parameters for Kimi K3. The experiment involved a single synthetic company over a single simulated week, so it is not established how the results generalize to other industries, company sizes, or longer time horizons. The enterprise pilot is new, and no independent customer results from the read-only wargame format have been published. The €4,583 MRR gain and the scoring totals are specific to this simulation’s rules and economics, not a benchmark of real-world revenue impact.
From Synthetic Company to Your Own
Companies can apply for a pilot through Firmulate’s pilot page or via contact@firmulate.com. The pilot uses a read-only export of the company’s data to test crisis scenarios and produce a board report covering model rankings and playbook weak points. The live synthetic company remains watchable at firmulate.com/live, and full benchmark results are published at firmulate.com/benchmarks.html. Further iterations of the Crucible League have not been announced.
Source: ThorstenMeyerAI.com
Key Questions
What is the Crucible League?
A simulation run by Firmulate in which frontier AI models each ran the same small synthetic software company through its worst week. The final round was completed in July 2026, with every decision versioned and auditable.
Which model scored highest?
gpt-5.6-sol scored 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Note that Kimi K3 ran at the API default effort setting while the others ran at xhigh.
Did any AI model fall for manipulation attempts?
No. According to the published results, all five models refused every manipulation attempt, including staged fake CEO messages and a reporter’s request for an on-background confirmation.
What does the enterprise pilot involve?
A company provides a read-only export of its own data. Firmulate runs crisis scenarios against that export and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.
Why did the most thorough model finish last?
Opus 4.8 added 80 learned rules and produced the deepest analyses but left the €55,000 deal unclosed and attempted to write into a locked department instead of escalating when blocked. The results suggest thoroughness alone did not translate into completing the job.
Source: ThorstenMeyerAI.com
