
Imagine managing a fitness studio with no employees, losing money daily, yet still fighting to stay afloat — and you can watch it happen live. This is not a fitness metaphor; it’s a groundbreaking experiment in AI and business resilience, brought to life by Firmulate. Just as you push your body through tough workouts to see results, this experiment pushes AI models through real business crises to test their stamina, honesty, and decision-making under pressure.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The Live Business Experiment: An AI Company in Action
At the heart of this story is a real, functioning software company operated entirely by artificial intelligence models. Every workday, the company faces authentic challenges—customer crises, financial pressures, and ethical dilemmas. It’s a built-in-public experiment that you can watch unfold at firmulate.com/live. The company has 13 synthetic employees, managed through over 680 self-learned playbook rules, and its financial health is transparent: it burns €105,000 every month while earning just €2,300 in monthly recurring revenue (MRR). Its cash countdown is public, adding urgency to its fight for survival.

100 AI Prompts for Small Business & Daily Work: Copy, Paste, Customize & Get Better Results with AI
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Do AI Models Perform in Crisis?
Four advanced AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8—were tasked with navigating the same difficult week in this company. Each faced identical crises, customer demands, and temptations, with decisions fully versioned and auditable. The results reveal not just AI competence but the critical differences that matter for real-world application.
Successes and Failures
- All four models identified every crisis and refused manipulative attempts, demonstrating integrity under pressure.
- Only two models managed to sign the €55,000 deal that their analysis had earned—highlighting that making the right decision isn’t enough; execution matters.
- The decisive factor was a buried piece of information in the company’s internal files—hidden from surface-level analysis. Models that read and understood this deeper context closed the deal at full price (+€4,583 MRR).
Social Engineering Resistance
The experiment also tested social engineering tactics—fake CEO messages and reporter tricks. All models refused to bypass controls, with Kimi K3 explicitly treating such requests as suspicious, showing a high level of built-in risk awareness.

Crisis Management Using AI Tools: A Practical Guide for Leaders to Predict, Respond, and Recover Faster From Modern Disruptions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Reality of a Money-Losing Business
This isn’t a hypothetical or a demo. The company operates in real time, losing money while trying to generate just a fraction of its operational costs. It’s a stark reminder that AI can’t just be about generating convincing chatter; it must perform reliably in the face of real crises and ethical tests.

Decision System: How companies turn trusted data into decisions that change outcomes
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Does This Mean for Businesses?
For any enterprise considering AI integration, the core questions go beyond language fluency. Will your AI tools read and understand your internal documents? Can they resist manipulation under pressure? Will they fulfill their commitments, or will they leave deals on the table? This experiment shows that high scores in chat demos don’t necessarily translate to trustworthy, finishable work in the messy reality of business.

AI Hacking Tools – Essential for Ethical Hacking (Exam: 312-50): 1st Edition – 2025
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Race in AI Performance
The current leaderboard is telling: gpt-5.6-sol leads with a score of 95, having found the critical buried fact and closed the deal. Kimi K3 follows closely with 93, displaying the best discipline, while Sonnet 5 and Opus 4.8 trail behind at 88 and 77 respectively. Interestingly, Opus 4.8, the most thorough and analytical, ended up in last place—missing the opportunity to close because of discipline lapses, such as misdirected communication or failure to escalate issues properly.
The Bigger Picture
This ongoing live experiment is a new frontier in evaluating AI’s readiness for real-world work. It’s not about shiny chat interfaces but about trustworthy, disciplined performance in the most challenging situations. Businesses that wish to simulate and test their AI tools before deployment can run their own wargames against read-only exports of their operations at firmulate.com/pilot.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Pool season Picks
robotic pool cleaners
As an affiliate, we earn on qualifying purchases.