
Imagine if your child’s teacher only gave grades for effort but never rewarded actual progress—that’s essentially what AI benchmarks are revealing about trust and performance today. In a world where businesses increasingly rely on AI for decisions, understanding what these systems truly deliver is more critical than ever. Recent experiments with AI models in a simulated company environment shed light on how honesty, discipline, and thoroughness are the real measures of success—beyond just chat quality or quick responses.
Turn the school run and nap time into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
Understanding the Benchmark’s Honest Approach
At first glance, it might seem straightforward: AI models are tested on how well they handle a series of simulated business crises within a controlled environment. But the real story is more nuanced. The latest experiment, conducted by Firmulate, involved running different AI models through the same difficult week of managing a small software company—complete with customer issues, internal crises, and tempting shortcuts. Crucially, each model’s decisions were meticulously versioned and auditable, ensuring transparency and fairness.
The Curious Baseline Score of 26
One striking detail stands out: even a ‘do-nothing’ baseline model scored 26 points out of a possible 100. This isn’t a mistake or a flaw; it’s a reflection of the benchmark’s design. The score represents partial progress—models that do nothing at all still get points simply for not making obvious mistakes or breaches of trust. It shows that the system recognizes some minimal level of effort or correctness, even in inaction. The key takeaway: a truly honest AI system can’t be scored at zero because the baseline itself accounts for the minimal act of not failing catastrophically.
Progress Counts, but Trust Caps Performance
Another vital aspect is how the benchmark handles breaches of trust. If an AI model attempts manipulation or shortcuts—such as pretending to have read a document it hasn’t or bypassing approval steps—its score gets capped. This means that no matter how impressive its partial progress, one slip of integrity can prevent higher scores. For instance, all models refused manipulative requests, and only two actually signed a deal based on analysis—showing that honesty and discipline are non-negotiable in real decision-making.
As an affiliate, we earn on qualifying purchases.
The Hidden Weaknesses and What They Reveal
Perhaps the most revealing finding is that the models’ weaknesses weren’t in their detection of crises or their refusal to manipulate—they all passed that test. Instead, the decisive failures lay in their internal discipline and thoroughness. For example, the lowest-performing model, Opus 4.8, identified the critical fact buried deep in a company file but faltered when it came to closing a deal, leaving the opportunity on the table. This highlights a deeper insight: thoroughness and attention to detail are paramount, especially when the stakes are high.
Trust Under Pressure
Real-world AI must be trustworthy, especially when facing social engineering attempts like fake messages from a CEO or staged reporter tricks. All models refused to be manipulated in these scenarios, with Kimi K3 explicitly treating suspicious requests as impersonation risks. This demonstrates that trustworthiness isn’t just about avoiding mistakes—it’s about actively resisting attempts to deceive the system.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Business and Parenting
For families and parents, the lesson extends beyond code and algorithms. Whether it’s teaching children honesty, discipline, or perseverance, the core principles remain the same. Just as a good teacher recognizes effort but rewards genuine progress, businesses and AI systems must balance partial work with unwavering integrity. It’s about fostering systems—human or machine—that are honest, disciplined, and thorough, not just quick or clever.
Watch the Experiment Live
Curious to see how these models perform in real-time? Firmulate offers a live environment where you can observe these AI-driven companies in action, managing crises, making decisions, and resisting manipulation—all without risking your actual business. This transparent, watchable setup provides a rare window into what trustworthy AI really looks like in practice.

The key takeaway is that honest AI isn’t about scoring perfect 100s—it’s about setting a realistic baseline where partial effort counts, and trust is non-negotiable. In both business and family life, integrity, thoroughness, and discipline are what truly define success. As AI continues to evolve, the most valuable performance metric remains trustworthiness—and that starts with acknowledging the importance of doing the simple, honest work first.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
AI performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
