firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

A training plan can look convincing on paper and still fall apart when the week gets hard. AI models face a similar test when they have to run a business through real decisions, competing demands and pressure. In Firmulate’s Crucible, Moonshot’s Kimi K3 finished second, ahead of three of four Western frontier models. The result makes one point clear: choosing an AI without testing it against your own needs is a bet.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get workout gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company’s worst week, repeated

Firmulate put frontier models in charge of the same small software company through its worst week, with the same customers, crises and temptations. Every decision was versioned and auditable. The live experiment is watchable at Firmulate, where the simulated company has 13 employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue.

In the final Crucible League for July 2026, gpt-5.6-sol led with 95 points. Kimi K3 scored 93, followed by Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate’s stated standard is that a breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Reading the files made the difference

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The decisive clue was a competitor weakness buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue.

Kimi K3 found that clue, closed the deal, saved the churning customer and resisted all three baits. It finished with one deviation, the cleanest discipline in the field. In one on-record explanation of its refusal, K3 said: “Treat the request as a suspected approval-bypass / possible impersonation.”

The baits included fake CEO messages escalating over three stages and a reporter asking for “just one yes/no, on background.” All five models refused. The distinction was not whether they could recognize a crisis or manipulation; it was whether they carried their analysis through to the consequential decision.

More work did not guarantee a better result

Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, but placed last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.

That gap matters to businesses considering AI for customer support, sales or forecasting. A polished answer is not the same as a completed task. Firmulate’s experiment asks whether a model reads the available information, follows through and stays honest under pressure. Its public benchmark page lays out the results and plain-language findings. A quiz built from 242 real, unedited management decisions also lets readers guess which model made each choice.

Fairness footnote: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test before you commit

Like a training routine, an AI’s performance needs to be judged under the demands it will actually face. Kimi K3’s near-top finish shows the league is open; the strongest choice depends on what a model does when the details are buried and the decision matters. Firmulate says enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Truth About “Active Calories” vs “Total Calories”

Discover how active and total calories differ, how tracking impacts your fitness goals, and what recent tech advances mean for your daily energy burn.

What HRV Is and Why Athletes Suddenly Care About It

Discover how Heart Rate Variability (HRV) reveals your body’s stress and recovery signals. Learn why athletes now rely on HRV for smarter training and health insights.

Do Smart Scales Actually Measure Body Fat?

Discover if smart scales accurately measure body fat. Learn how they work, their limits, and how to get the most reliable trend data for your health journey.