firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine if your favorite kitchen appliance could not only follow recipes but also handle the chaos of a busy dinner service — spotting problems, resisting shortcuts, and sealing deals without fuss. That’s the kind of challenge AI faces in the real world of business, where pressure and temptation can reveal its true character. Recently, a groundbreaking experiment tested some of the most advanced AI models in a simulated company crisis, and the results might just change how you think about choosing technology for your operations.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

In July 2026, a public trial known as the Crucible League put five leading AI models through their paces in a realistic business wargame. Each AI was tasked with managing a small software company facing its worst week — handling crises, reading critical files, and making decisions under pressure. The goal was straightforward: see which model could spot problems, resist unethical temptations, and close a key deal worth €55,000 a month in recurring revenue.

The five contenders included well-known giants like gpt-5.6-sol and newer entrants like Moonshot’s Kimi K3. While all models showed competence in identifying crises and refusing manipulative tactics, the differences lay in their depth of analysis and discipline in execution.

The Results: The Top Performers

  • gpt-5.6-sol scored the highest with a 95 out of 100, successfully finding a buried piece of critical information in the company’s files and closing the deal. This model demonstrated full situational awareness and impeccable discipline.
  • Kimi K3, the newcomer from Moonshot, scored just slightly behind with a 93. Despite being tested at the API default setting (no effort parameter), it showed the cleanest discipline in refusing manipulative social engineering and identified the key hidden data that sealed the deal, boosting its recurring revenue by over €4,500 per month.
  • Sonnet 5 and Fable 5 followed, with scores of 88 and 77 respectively, both closing the deal but with minor slips in process discipline. The oldest model, Opus 4.8, scored 73 and demonstrated some weaknesses in escalation procedures, leaving money on the table.

Crucially, all models refused every manipulated request—fake CEO messages and staged reporter tricks—showing resilience against social engineering. Only two models, including K3, signed the deal on their own analysis, confirming their capacity for independent judgment. The others hesitated or did not sign, despite diagnosing the problem correctly.

The Hidden Edge: Reading Company Files

The decisive factor in the experiment was the models’ ability to read and interpret the company’s internal documentation, not just respond to external cues. The best performers found a buried reference deep within files that others overlooked, enabling them to close the deal at the full price and significantly boost revenue.

The Discipline of Decision-Making

Interestingly, the experiment revealed that running models at a higher effort setting did not necessarily improve outcomes. The K3 model, tested at the default effort level, demonstrated the most disciplined and accurate decision-making. In contrast, a more thorough participant, Opus 4.8, often left opportunities unseized due to overanalysis or improper escalation—weaknesses that could be costly in real settings.

It’s worth noting that Kimi K3 was evaluated without an effort parameter, making its performance even more impressive given the constraints. This suggests that model robustness and discipline might be more critical than raw analytical depth.

The Broader Implications

This experiment underscores a vital point: when deploying AI in real business environments, the question isn’t just about how well it chats or generates content. It’s whether it can finish tasks, read critical documents, resist manipulation, and act ethically under pressure. As the league table shows, even newer models can outperform established ones in these crucial aspects.

For companies pondering how to integrate AI into decision-making or support roles, these findings are a wake-up call. Relying solely on chat scores or superficial tests could be a costly mistake; instead, observing how models perform in comprehensive simulations is essential. The live company running at firmulate.com/live offers a transparent view of this ongoing experiment—showing how AI models handle real crises, real money, and real temptations every day.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The recent experiment shows that the best AI models for business management excel not just in understanding but in disciplined execution, reading internal files, and resisting manipulation—traits that matter far more than chat quality. Moonshot’s Kimi K3, tested at default effort, proved a formidable newcomer, outperforming some established models and closing a lucrative deal. When choosing AI for your organization, focus on its ability to finish what it starts, stay honest under pressure, and handle complex data—skills that can make the difference between success and missed opportunity.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI business management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision-making tools for companies

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI crisis management solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI data analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Air Fryer Baked Potatoes: Crispy Skin Without a Long Oven Wait

AIThis post was created with the assistance of artificial intelligence (AI).To achieve…

Air Fryer Nachos: How to Melt Cheese Without Soggy Chips

I’ll show you how to melt cheese perfectly on nachos in your air fryer without sogginess, so you can enjoy crispy, cheesy snacks every time.

Air Fryer Carrots: How to Get Tender Centers and Charred Edges

With the right techniques, you can achieve perfectly tender yet crispy air fryer carrots—discover how to master the process for delicious results.

Air Fryer Stuffed Peppers: How to Get Tender Peppers and Hot Filling

Crispy, tender air fryer stuffed peppers with hot filling await—discover the secrets to perfect cooking and irresistible flavor.