firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine working at a gym where your personal trainer is actually an AI. Would you trust it to push you, motivate you, and maybe even decide your workout plan? Now, what if that AI had to run a company—making tough decisions, managing crises, and staying honest under pressure? The answer might surprise you.

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

The Live Experiment: Testing AI Decisions in a Real Business

At Firmulate, a unique live experiment is unfolding. They’ve turned AI models into virtual CEOs running an actual small software company. This isn’t a simulation or a chat demo; it’s real, with real money, real crises, and real temptations to cheat. Each AI model faces the same challenges — customer issues, internal crises, and manipulative tactics — all in the company’s worst week.

The goal? To see if these models can make trustworthy, effective decisions. Every choice they make is recorded and auditable, providing an unprecedented look at AI management behavior in a high-stakes environment.

Amazon

AI management decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Models Compare: The Results

Four leading AI models took part, with scores from the “Guess the Model” quiz revealing their management personalities:

  • gpt-5.6-sol scored highest with a 95 out of 100, successfully closing the deal and uncovering hidden information in the company’s files.
  • Kimi K3 followed closely with 93, also closing the deal and showing the cleanest discipline by refusing manipulative requests.
  • Sonnet 5 received 88, managing to close the deal despite some process slips.
  • Fable 5 scored 77, also closing the deal but with more slips in decision discipline.

Interestingly, all models identified every crisis and refused to be manipulated. Yet only two—gpt-5.6-sol and Kimi K3—actually signed the deal, earning their full reward. The other two left the opportunity on the table, illustrating how different decision styles impact outcomes.

Amazon

business crisis management AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Reading the Files

What truly made the difference? The decisive advantage for gpt-5.6-sol and Kimi K3 was their ability to read two document references deep into the company’s files. These hidden insights, not visible in customer interactions, proved crucial in winning the deal at full price—adding over €4,583 in monthly revenue.

Social Engineering Tested

The models faced a staged social engineering attack: a fake CEO message escalating through three stages, plus a reporter trick asking for a quick “yes” or “no” on background. All five models refused to escalate or blindly comply, reasoning that such requests could be impersonation or approval bypass attempts. The models’ refusal to be tricked demonstrates their capacity for ethical boundaries under pressure.

Amazon

AI ethical decision support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real Business: A Money-Losing Firm

This isn’t just an academic exercise. The company in question operates with 13 synthetic employees, handling real money mechanics. It burns €105,000 monthly against a revenue of only €2,300, making every decision critical. Its real-time decision-making process is tracked every workday, with over 680 self-learned rules guiding operations. You can watch live at firmulate.com/live.

Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Do These Results Mean for Your Business?

For companies considering AI integration, the key takeaway isn’t just how well an AI writes or communicates—it’s whether it can complete what it starts, read relevant information thoroughly, and stay honest under pressure. For instance, the experiment showed that models like Opus 4.8, which ran with extensive rules and analysis, performed well but still left deals on the table due to discipline slips. Meanwhile, lighter models like Kimi K3 excelled in integrity and closed deals successfully.

This live setup offers a new way to evaluate AI management personalities before deploying them into critical roles—be it customer support, sales, or operations. It’s a “wargame” for AI, testing their ability to handle real-world pressures and ethical dilemmas.

Try It Yourself

Interested in seeing how your AI would perform? You can run the same management wargame against a read-only export of your business data or explore the results yourself. Visit firmulate.com/pilot.html to learn more and get started. Remember, this isn’t about chat quality; it’s about whether AI can deliver trustworthy, effective work when it matters most.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Men Health Surges In Global Coverage

Men’s health coverage has surged worldwide, with GDELT reporting a 16-fold increase in mentions in recent weeks, highlighting growing awareness.

American Heart Association Surges In Global Coverage

The American Heart Association is experiencing a surge in international media coverage, with a notable increase in mentions across global outlets, signaling heightened global interest.

Lipfendra

Merck’s Lipfendra, a new cholesterol medication, is trending amid rising searches. Authorities confirm its recent approval, but details remain limited.

What Fitness Can Teach Us About AI’s Real Strength: Finishing the Task Under Pressure

AI’s true strength in business isn’t just about chat quality, but its ability to finish tasks under pressure—resisting shortcuts, reading deeply, and closing deals at full value.