AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Parents know that trust is built in ordinary moments, then tested when the week goes sideways. Businesses face a similar question as AI takes on more work: will it follow good judgment when customers, money and pressure are all in play? Firmulate’s live experiment puts that question to the test.

For listenersOffer from Amazon

Turn the school run and nap time into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

A company under pressure

In the Crucible League’s final, in July 2026, frontier models ran the same small software company through its worst week: the same customers, crises and temptations. Every decision was versioned and auditable. The experiment is real and watchable at Firmulate.

All the models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The gap was captured in a stark phrase: “Same diagnosis, same pitch — no signature.” Recognizing the right move did not guarantee the company would make it.

The detail hidden in the files

The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The result turned on whether a model connected the available evidence to the moment when action mattered.

Integrity faced its own test. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” The do-nothing baseline scored 26: partial progress counted, but a single breach of trust capped the total. As the league put it, “no amount of good work outweighs a breach of trust.”

Strong analysis still needs follow-through

The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. Opus was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four.

There is a fairness caveat: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The result is a useful comparison, but that difference belongs in the picture.

From watching to trying it on your business

The live company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and a record of every workday versioned. A separate quiz draws on 242 real, unedited management decisions, asking visitors to guess which model made each call.

For a business weighing AI agents, the next step can be a pilot against its own circumstances. Firmulate says enterprises can run the wargame using a read-only export of their business, test crisis scenarios and receive a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems. That offers a way to examine judgment before handing an AI real responsibilities.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Test judgment before handing over responsibility

AI systems can identify a crisis and still fail to finish the job. Firmulate’s experiment makes that difference visible: the overlooked deal, the pressure to bypass approval and the need to escalate all expose questions that a polished chat alone may not answer. To discuss a pilot using a read-only export of your business, visit the Firmulate pilot page or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Parenting content here is informational. For medical questions about your child, consult a pediatrician.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Robot Vacuums for Families: Features That Help With Baby Crumbs

Theoretically, a robot vacuum with smart features can make cleaning up baby crumbs easier—discover which features truly make a difference for busy families.

The Truth About BPA‑Free Plastic in Baby Products

Ineffective BPA-free labels may hide hidden risks and environmental concerns—discover what’s really inside your baby products before trusting the label.

Strollers Simplify Newborn Care

With the right stroller, your outings can become effortless and enjoyable, but do you know which features to prioritize for your newborn’s comfort?

Ducklings Early Learning Center Surges In Global Coverage

Ducklings Early Learning Center experiences a surge in international coverage, with 45 mentions in recent media monitoring, marking unprecedented global interest.