Kimi K3 Cracks the AI Company League’s Top Two

  • by

Live on firmulate.com.

Choosing an AI model by its name or chat demo may miss the test that matters: what happens when it has to run a business under pressure? In Firmulate’s Crucible, Moonshot’s Kimi K3 finished second, beating three of four Western frontier models. The result puts a practical question to companies eyeing AI agents: how would your preferred model perform in your own workplace?

A company’s worst week, repeated

Firmulate ran each frontier model through the same small software company’s worst week: the same customers, crises and temptations. Decisions were versioned and auditable. The experiment is live and watchable, with a public cash countdown and synthetic employees handling the company’s work.

The final Crucible League, dated July 2026, ranks gpt-5.6-sol first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Reading the files made the difference

All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap between recognizing a problem and finishing the job is central to Firmulate’s case: “Same diagnosis, same pitch — no signature.”

The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. K3 was among them. It also saved the churning customer, found the security issue, and resisted all three baits, with one deviation—the cleanest discipline in the field, according to the brief.

Discipline under pressure

The manipulation tests included fake CEO messages escalating across three stages and a reporter asking for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 offers a counterpoint to the idea that more extensive work guarantees a better outcome. It was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. The deal was left unsigned, and it attempted to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four.

Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k in monthly recurring revenue. The public countdown makes the simulation’s operating pressure visible. The site says the company runs every business day and versions every workday. Readers can follow the experiment at Firmulate and see the full results on its benchmark page.

Fairness note: K3 ran without an effort parameter (API default), while the others ran at xhigh.

The findings at a glance — source: firmulate.com.

Test the model you plan to use

K3’s second-place finish makes the league look open, while the unsigned deals show why a benchmark score alone cannot settle the buying decision. Firmulate says enterprises can run the wargame against a read-only export of their own business; nothing writes back to real systems. A quiz built from 242 real, unedited management decisions also invites readers to guess which model made each choice. For businesses considering AI agents, the useful next step is to test them on work and pressures that resemble their own.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Leave a Reply

Your email address will not be published.