
Imagine your fitness routine where even doing nothing earns you 26 points out of 100. Sounds strange, right? But in the world of AI benchmarks, this ‘do-nothing’ score reveals crucial truths about reliability and trustworthiness, especially when AI is set to manage your business operations. Today, we’ll explore how a simple baseline in AI testing sheds light on the importance of integrity and performance in real-world applications.
Get workout gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Benchmark That Reveals Honest AI Performance
In a recent live experiment conducted by Firmulate, four advanced AI models were tasked with managing a small software company’s toughest week—crises, manipulations, and all. Each model was run through identical scenarios, with decisions that are auditable and transparent. The goal? To see whether these models not only spot problems but also act ethically and responsibly under pressure.
The results are illuminating. All four models identified every crisis and refused every manipulative attempt—an essential trait for business trust. But only two managed to close the deal that their own analysis had earned, signing off on a €55,000 contract. The other two, despite diagnosing the issues correctly, left the deal on the table—missing out on potential revenue.
Why the Baseline Score Is 26
One surprising fact: a ‘do-nothing’ baseline score in this benchmark is 26 out of 100. That might seem odd—why does doing nothing get you 26 points? The reason lies in how partial progress is valued. Even minimal, cautious actions—like refusing manipulative requests or reading key documents—are counted positively. It’s a recognition that, in complex decision-making, even small steps towards honesty and diligence matter.
However, the benchmark also establishes a hard cap: a single breach of trust, such as signing a manipulative deal, caps the total grade regardless of other achievements. This underscores a fundamental principle—trustworthiness is paramount, and one breach can undermine the entire effort.
As an affiliate, we earn on qualifying purchases.
What This Means for Business AI Adoption
For companies considering AI to handle critical functions—like customer relations, support, or forecasts—the key question isn’t just whether the AI can write compelling language, but whether it can finish what it starts, read and understand your files, and stay honest under pressure.
In the live experiment, the models successfully detected crises and refused manipulative tactics across all stages. For instance, when fake CEO messages escalated in stages, all models refused to approve or sign anything suspicious—an essential trait for safeguarding your business integrity.
The Hidden Weaknesses in Deep Analysis
Interestingly, the real weakness in the models was not in their ability to detect problems but in their decision-making process. The most thorough participant, Opus 4.8, had over 80 learned rules and deep analyses but still failed to close the deal, leaving an opportunity on the table and slipping discipline. This points to a broader truth: even sophisticated models can falter in execution if not properly guided or disciplined.
AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Trust Matters More Than Scores
In the world of AI, a perfect score isn’t everything. The experiment emphasizes that trust—demonstrated by refusing manipulation and acting ethically—is critical. While models like Kimi K3 scored 93 and closed the deal with the cleanest discipline, others with higher scores like GPT-5.6 scored full marks because they found buried facts and closed deals at full price.
Ultimately, the benchmark exposes that achieving true reliability requires more than surface-level performance; it demands transparency, integrity, and discipline.
As an affiliate, we earn on qualifying purchases.
How to Use This Knowledge
Businesses should consider running their own ‘wargames’—simulated crises against their AI systems—before deploying them in critical roles. Firmulate’s live platform offers a way to test AI models in a safe, controlled environment, ensuring they can handle real-world pressures without risking trust or integrity.
In the end, the lesson is clear: measure not just what AI can do, but what it *should* do. A baseline score of 26 points reminds us that even minimal honest effort counts, and a single breach of trust can undo all progress. For your business, the goal is to choose AI that earns trust with every decision, every time.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI compliance monitoring software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
