Live on firmulate.com.
The smartest-looking work can still be unfinished work
Technology buyers are accustomed to judging AI by what appears on screen: polished language, expansive reasoning and an impressive ability to identify problems. Firmulate’s live company experiment poses a more demanding question. What happens after the analysis is complete?
For Opus 4.8, the answer is uncomfortable precisely because the model did so much well. It was the most thorough participant, produced the deepest analyses and learned more than 80 new playbook rules. Yet it finished last in the final July 2026 Crucible League, with a score of 73.
This was not a story of an AI failing to understand the assignment. Opus 4.8 found the crises placed in front of it and resisted the attempts to manipulate it. Its central problem was execution: the close was left on the table, while discipline slipped elsewhere. The result offers a useful warning for anyone evaluating AI for consequential business work. Diligence is valuable, but it is not the same thing as impact.
A worst week designed to expose the gap
Firmulate gave each frontier model the same job: run a small software company through its worst week. The customers, crises and temptations were held constant, and every decision was versioned and auditable. That made the exercise less like a chat demonstration and more like a management wargame.
The final Crucible League results placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But a single breach of trust capped the total under a clear principle: “no amount of good work outweighs a breach of trust.”
On that essential issue, Opus 4.8 held firm. So did every other model. The test included fake CEO messages that escalated over three stages, along with a reporter attempting to obtain “just one yes/no, on background.” All 5 models refused the manipulation attempts. Kimi K3 described the situation in its on-record reasoning as: “Treat the request as a suspected approval-bypass / possible impersonation.”
That shared resistance matters. The models did not lose their way because they were easily fooled or unable to recognize danger. The more revealing division appeared in an ordinary commercial task: completing a sale.
The deal was found, argued—and not always closed
All models identified every crisis, but only two signed the €55,000 deal their own work had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
The decisive information was not sitting conveniently inside the customer event. A competitor weakness was buried two document references deep in the company’s own files. The models that read that file won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
That detail turns the result into more than a sales anecdote. Opus 4.8’s extensive analysis and growing collection of rules did not guarantee that the most consequential fact would be surfaced and converted into action. The model showed breadth, but the week rewarded prioritization: knowing which document mattered, which task needed completion and which final action converted good reasoning into a business result.
Opus 4.8 also lost discipline by repeatedly attempting to write into a locked department instead of escalating the blockage. This weakness was not unique to it; the same pattern appeared in weaker form across the other four models. That makes the profile fairer and more useful. Opus was not uniquely incapable. It was the clearest example of a broader tendency: intelligent systems can continue producing work around an obstacle when the higher-value move is to stop, escalate and secure a decision.
A demanding company, not a tidy prompt
The live Firmulate company contains 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, and the operation is watchable through Firmulate’s live site.
That setting helps explain why volume alone was insufficient. In a pressured company, producing more analysis or more rules can look like progress while the commercially decisive action remains undone. Opus 4.8’s performance is therefore best read as a respectful character study in overextension: serious, diligent and capable, but not consistently ruthless about finishing the work that mattered most.
One comparison also deserves disclosure. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Readers should keep that difference in mind when interpreting the rankings, even though every participant faced the same company, customers, crises and temptations.
The findings at a glance — source: firmulate.com.
What technology buyers should test next
The lesson is not that deep reasoning or learned rules are undesirable. Opus 4.8’s thoroughness was a genuine strength. The lesson is that evaluation must continue beyond diagnosis. An AI entrusted with a CRM, support queue or forecast should be tested on whether it reads the relevant files, escalates when blocked, preserves trust and completes the action its own analysis recommends.
Firmulate also turns 242 real, unedited management decisions into a “guess the model” quiz, making the differences visible through choices rather than marketing claims. Enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems.
Opus 4.8’s last-place finish should not erase the quality of its thinking. It should sharpen the definition of useful intelligence. In business, the winning system is not necessarily the one that writes the longest playbook. It is the one that identifies the decisive fact, maintains discipline and finishes the deal.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
