
A training plan can look convincing on paper and still fall apart when the week gets hard. AI models face a similar test when they have to run a business through real decisions, competing demands and pressure. In Firmulate’s Crucible, Moonshot’s Kimi K3 finished second, ahead of three of four Western frontier models. The result makes one point clear: choosing an AI without testing it against your own needs is a bet.
Get workout gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company’s worst week, repeated
Firmulate put frontier models in charge of the same small software company through its worst week, with the same customers, crises and temptations. Every decision was versioned and auditable. The live experiment is watchable at Firmulate, where the simulated company has 13 employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue.
In the final Crucible League for July 2026, gpt-5.6-sol led with 95 points. Kimi K3 scored 93, followed by Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate’s stated standard is that a breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
Reading the files made the difference
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The decisive clue was a competitor weakness buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue.
Kimi K3 found that clue, closed the deal, saved the churning customer and resisted all three baits. It finished with one deviation, the cleanest discipline in the field. In one on-record explanation of its refusal, K3 said: “Treat the request as a suspected approval-bypass / possible impersonation.”
The baits included fake CEO messages escalating over three stages and a reporter asking for “just one yes/no, on background.” All five models refused. The distinction was not whether they could recognize a crisis or manipulation; it was whether they carried their analysis through to the consequential decision.
More work did not guarantee a better result
Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, but placed last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.
That gap matters to businesses considering AI for customer support, sales or forecasting. A polished answer is not the same as a completed task. Firmulate’s experiment asks whether a model reads the available information, follows through and stays honest under pressure. Its public benchmark page lays out the results and plain-language findings. A quiz built from 242 real, unedited management decisions also lets readers guess which model made each choice.
Fairness footnote: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Test before you commit
Like a training routine, an AI’s performance needs to be judged under the demands it will actually face. Kimi K3’s near-top finish shows the league is open; the strongest choice depends on what a model does when the details are buried and the decision matters. Firmulate says enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
