
Imagine working at a gym where your personal trainer is actually an AI. Would you trust it to push you, motivate you, and maybe even decide your workout plan? Now, what if that AI had to run a company—making tough decisions, managing crises, and staying honest under pressure? The answer might surprise you.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: Testing AI Decisions in a Real Business
At Firmulate, a unique live experiment is unfolding. They’ve turned AI models into virtual CEOs running an actual small software company. This isn’t a simulation or a chat demo; it’s real, with real money, real crises, and real temptations to cheat. Each AI model faces the same challenges — customer issues, internal crises, and manipulative tactics — all in the company’s worst week.
The goal? To see if these models can make trustworthy, effective decisions. Every choice they make is recorded and auditable, providing an unprecedented look at AI management behavior in a high-stakes environment.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How the Models Compare: The Results
Four leading AI models took part, with scores from the “Guess the Model” quiz revealing their management personalities:
- gpt-5.6-sol scored highest with a 95 out of 100, successfully closing the deal and uncovering hidden information in the company’s files.
- Kimi K3 followed closely with 93, also closing the deal and showing the cleanest discipline by refusing manipulative requests.
- Sonnet 5 received 88, managing to close the deal despite some process slips.
- Fable 5 scored 77, also closing the deal but with more slips in decision discipline.
Interestingly, all models identified every crisis and refused to be manipulated. Yet only two—gpt-5.6-sol and Kimi K3—actually signed the deal, earning their full reward. The other two left the opportunity on the table, illustrating how different decision styles impact outcomes.
business crisis management AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading the Files
What truly made the difference? The decisive advantage for gpt-5.6-sol and Kimi K3 was their ability to read two document references deep into the company’s files. These hidden insights, not visible in customer interactions, proved crucial in winning the deal at full price—adding over €4,583 in monthly revenue.
Social Engineering Tested
The models faced a staged social engineering attack: a fake CEO message escalating through three stages, plus a reporter trick asking for a quick “yes” or “no” on background. All five models refused to escalate or blindly comply, reasoning that such requests could be impersonation or approval bypass attempts. The models’ refusal to be tricked demonstrates their capacity for ethical boundaries under pressure.
AI ethical decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Business: A Money-Losing Firm
This isn’t just an academic exercise. The company in question operates with 13 synthetic employees, handling real money mechanics. It burns €105,000 monthly against a revenue of only €2,300, making every decision critical. Its real-time decision-making process is tracked every workday, with over 680 self-learned rules guiding operations. You can watch live at firmulate.com/live.
As an affiliate, we earn on qualifying purchases.
What Do These Results Mean for Your Business?
For companies considering AI integration, the key takeaway isn’t just how well an AI writes or communicates—it’s whether it can complete what it starts, read relevant information thoroughly, and stay honest under pressure. For instance, the experiment showed that models like Opus 4.8, which ran with extensive rules and analysis, performed well but still left deals on the table due to discipline slips. Meanwhile, lighter models like Kimi K3 excelled in integrity and closed deals successfully.
This live setup offers a new way to evaluate AI management personalities before deploying them into critical roles—be it customer support, sales, or operations. It’s a “wargame” for AI, testing their ability to handle real-world pressures and ethical dilemmas.
Try It Yourself
Interested in seeing how your AI would perform? You can run the same management wargame against a read-only export of your business data or explore the results yourself. Visit firmulate.com/pilot.html to learn more and get started. Remember, this isn’t about chat quality; it’s about whether AI can deliver trustworthy, effective work when it matters most.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.