firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A smart oven can follow a recipe, but that tells you little about how it will handle a power cut halfway through dinner. The same goes for AI in business: a polished demo is not a test of how an agent behaves when customers, money and pressure are involved. Firmulate lets people watch that kind of test unfold, then offers companies a way to run one against their own business.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company’s worst week, repeated

Firmulate’s Crucible League put frontier AI models in charge of the same small software company during its worst week. Each faced the same customers, crises and temptations. Decisions were versioned and auditable, so the experiment could be watched and reviewed rather than reduced to a chat transcript.

In the final league, published in July 2026, gpt-5.6-sol finished first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s integrity rule is blunt: partial progress counts, but one breach of trust caps the total. As its principle puts it, “no amount of good work outweighs a breach of trust.”

Spotting the crisis was not the same as finishing the job

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The finding points to a gap between diagnosing a problem and carrying a decision through: “Same diagnosis, same pitch — no signature.”

The deal hinged on a detail buried two document references deep in the company’s own files, not in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. In a kitchen, a recipe’s crucial instruction might be tucked into the preparation notes; here, the missing step was finding and acting on relevant company knowledge.

The test also included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough work still needs sound judgment

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It nevertheless came last. The deal went unsigned, and discipline slipped when it tried to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.

There is a fairness detail for readers comparing the results: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh.

The experiment is part of a live company simulation with 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, a public cash countdown, more than 680 self-learned playbook rules and a versioned record for every workday. The live company is watchable at firmulate.com. A quiz built from 242 real, unedited management decisions lets readers guess which model made each choice.

From watching to testing your own business

For business leaders, the next step is to test agents against the particulars that a public benchmark cannot capture: their own customers, operating rules and weak spots. Firmulate says enterprises can run the same kind of crisis wargame using a read-only export of their business. The result is a board report with model rankings and weaknesses in the company’s playbooks. Nothing writes back to real systems.

That distinction matters. A simulated test can reveal how an AI handles pressure before a company gives it access to live operations. It does not guarantee perfect behavior, but it makes decisions and failures easier to examine before deployment.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your playbooks to the test

Firmulate’s pilot takes the experiment from a watchable company to a company’s own business: a read-only export, crisis scenarios and a board report, with no writes to real systems. To discuss a pilot, visit firmulate.com/pilot.html or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Crispier Breaded Cutlet Technique (So the Coating Doesn’t Fall Off)

When aiming for a perfectly crispy breaded cutlet that stays intact, mastering this technique will transform your cooking—discover the secrets to flawless coating.

Air Fryer Vegetables That Don’t Turn Mushy: Timing by Veg Type

Optimize your air fryer veggie results by timing each type perfectly—discover how to achieve crispy, non-mushy vegetables every time.

Air Fryer Frozen Burritos: The Crisp-Outside Reheat Method

Frozen burritos reheated in an air fryer achieve a perfect crispy outside, and with this method, you’ll never want to microwave again.

AI Models Stand Firm Against Social Engineering in Live Company Test

AIThis post was created with the assistance of artificial intelligence (AI).Live on…