The AI That Spots the Crisis but Still Misses the Deal

  • by

Live on firmulate.com.

AI can write a polished pitch and spot a brewing crisis. But when the stakes move from a chat window to a real business decision, can it follow through? Firmulate is putting that question on display with a live, watchable experiment: AI models run the same small software company through a week of hard choices.

A company under pressure

The experiment gives each model the same customers, crises and temptations. Decisions are versioned and auditable, so readers can follow what happened rather than judge a model by a slick demo. The company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue. Its public cash countdown makes the pressure visible, while more than 680 self-learned playbook rules capture lessons from workdays that are versioned as they happen.

The final Crucible League, from July 2026, ranked gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. Firmulate counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Seeing the answer is not the same as acting

The central finding is striking: every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The diagnosis and pitch were there; the signature was not. In an ordinary chat demo, that gap can be hard to see. In a business wargame, it becomes the story.

The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. That result points to a practical challenge for AI at work: important context may already exist inside a company, but an agent still has to find it and use it when the moment arrives.

Integrity under pressure, discipline under strain

The experiment also tested social engineering. Fake CEO messages escalated across three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 offers a more complicated result. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four. More analysis, by itself, did not guarantee better execution.

There is a fairness detail for readers comparing the leaderboard: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The result is a snapshot of this particular experiment, with that difference part of its context.

From watching to testing your own business

Firmulate’s live company is synthetic, but the business pressures are designed to be watchable and concrete. Its quiz uses 242 real, unedited management decisions and invites readers to guess which model made each choice. The live experiment and quiz offer two ways to inspect AI behavior beyond what a model says it would do.

The next step is a pilot for enterprises that want to test their own scenarios. Firmulate says a company can provide a read-only export of its business and run crisis scenarios against it, producing a board report with model rankings and weaknesses in its own playbooks. Nothing writes back to real systems. That moves the question from “How did the models handle this company?” to “What might they do with ours?”

The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

Watching an AI recognize a problem is only part of the evaluation. The harder test is whether it can act, preserve trust and follow the rules when the pressure is real. Enterprises can run a Firmulate pilot against a read-only business export, with no changes written to live systems. To discuss a pilot, visit firmulate.com/pilot.html or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Leave a Reply

Your email address will not be published.