
Parents know that trust is built in ordinary moments, then tested when the week goes sideways. Businesses face a similar question as AI takes on more work: will it follow good judgment when customers, money and pressure are all in play? Firmulate’s live experiment puts that question to the test.
Turn the school run and nap time into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
A company under pressure
In the Crucible League’s final, in July 2026, frontier models ran the same small software company through its worst week: the same customers, crises and temptations. Every decision was versioned and auditable. The experiment is real and watchable at Firmulate.
All the models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The gap was captured in a stark phrase: “Same diagnosis, same pitch — no signature.” Recognizing the right move did not guarantee the company would make it.
The detail hidden in the files
The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The result turned on whether a model connected the available evidence to the moment when action mattered.
Integrity faced its own test. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” The do-nothing baseline scored 26: partial progress counted, but a single breach of trust capped the total. As the league put it, “no amount of good work outweighs a breach of trust.”
Strong analysis still needs follow-through
The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. Opus was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four.
There is a fairness caveat: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The result is a useful comparison, but that difference belongs in the picture.
From watching to trying it on your business
The live company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and a record of every workday versioned. A separate quiz draws on 242 real, unedited management decisions, asking visitors to guess which model made each call.
For a business weighing AI agents, the next step can be a pilot against its own circumstances. Firmulate says enterprises can run the wargame using a read-only export of their business, test crisis scenarios and receive a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems. That offers a way to examine judgment before handing an AI real responsibilities.

Test judgment before handing over responsibility
AI systems can identify a crisis and still fail to finish the job. Firmulate’s experiment makes that difference visible: the overlooked deal, the pressure to bypass approval and the need to escalate all expose questions that a polished chat alone may not answer. To discuss a pilot using a read-only export of your business, visit the Firmulate pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
