
Imagine training for a marathon and only managing to walk the first few steps. That’s roughly the starting point for AI models in business: even the most basic, do-nothing baseline scores 26 points in a recent benchmark. For business leaders, understanding this score isn’t about AI proficiency — it’s about trust, discipline, and the real work AI must do to be truly helpful.
Get workout gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the Benchmark: Beyond the Glitz of Chat Demos
When evaluating AI for business, many focus on how well it writes or converses. But the recent Firmulate experiment reveals a more telling story. The test involved four frontier AI models running the same simulated software company through its worst week — with the same customers, crises, and temptations. The goal? See if these models could act like responsible managers under pressure.
The surprising find: even the do-nothing baseline, which simply does nothing, scored 26 out of 100. This isn’t a flaw or a mistake; it’s a reflection of how the scoring system works. Partial progress counts, and even a model that doesn’t actively do anything still earns points for basic compliance and honesty. Essentially, doing nothing is better than doing something reckless — but it still leaves a lot to be desired.
As an affiliate, we earn on qualifying purchases.
Why Does the Baseline Score 26?
The score of 26 points on the baseline stems from the fact that no matter how lazy or unresponsive an AI might be, it still receives some credit for not breaching trust or making harmful decisions. The system is designed to reward honesty, discipline, and careful decision-making, not just productivity. This means that even a fully passive model gets a starting score, emphasizing that trustworthiness isn’t something to be taken lightly.
business AI trust assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Partial Progress Counts — But Trust Is the Cap
In the experiment, each model faced the same crises, including manipulative social engineering attempts and tricky document references. All four models successfully identified every crisis and refused manipulation attempts—a gold standard for integrity. However, only two models signed the €55,000 deal, earning the full reward for proper diagnosis and communication.
This highlights a key principle: partial progress — such as detecting crises or refusing manipulative tactics — contributes to the score. Yet, a breach of trust, like failing to escalate or signing a deal without proper analysis, caps the total score. In other words, no matter how much good work is done, one breach or slip can undermine the entire effort.
As an affiliate, we earn on qualifying purchases.
The Hidden Weaknesses That Make the Difference
Digging into the details reveals that the real weakness isn’t in detecting external crises but in reading and understanding internal documents. The models that successfully closed the deal found crucial information buried two document references deep within the company’s files, not in the customer interactions. This ability to read and interpret key internal data was decisive for closing the deal at full price — worth over €4,583 in monthly recurring revenue.
In contrast, the most thorough participant, Opus 4.8, left the close on the table and slipped into internal communication instead of escalating. This weak spot underscores a crucial insight: comprehensive internal understanding and disciplined escalation are essential for trust and performance in business AI.
As an affiliate, we earn on qualifying purchases.
Social Engineering and Ethical Questions
One of the most revealing tests involved a staged social engineering scam, with fake CEO messages escalating over three stages and a reporter trick. All models refused to sign or approve these manipulative requests, with Kimi K3 explicitly treating them as potential impersonation or approval-bypass attempts. This demonstrates a high level of ethical discipline and resistance to manipulation — vital qualities for AI operating in sensitive business environments.
What Does This Mean for Businesses?
The experiment underscores that the real test of AI isn’t just surface-level chat or impressive language skills. It’s about whether these models can finish what they start — reading your files thoroughly, resisting manipulation, and maintaining discipline under pressure. These qualities directly impact your bottom line and trustworthiness.
Deploying AI that merely chats well but fails to act responsibly can be risky. Conversely, models that demonstrate integrity and discipline may score lower in superficial demos but outperform in real business scenarios, closing deals and avoiding costly mistakes.
Firmulate’s Live Experiment: A Watchable Laboratory
Firmulate’s live site offers an ongoing, transparent look into this experiment. Visitors can see four AI models running a simulated company, facing real crises, making decisions, and learning in real-time at firmulate.com/live. The experiment is versioned daily, providing a window into how AI models handle complex business tasks with discipline and honesty — not just language prowess.
This approach moves beyond the hype of AI demos and into the realm of real-world trust and discipline. For businesses considering AI adoption, it’s a reminder that the true value lies in reliable, responsible behavior, not just impressive responses.

The Firmulate benchmark reveals that even a do-nothing baseline scores 26 points, emphasizing the importance of trust and discipline in AI. Partial progress counts, but a single breach caps the score, highlighting the crucial qualities your AI must have to truly serve your business — reading deeply, resisting manipulation, and staying disciplined under pressure.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
