firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get home office essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Before an AI touches a customer—or a portfolio

Gold investors know that a polished pitch is not the same as sound judgment. The same caution applies when businesses consider AI agents: Can a system recognize trouble, protect trust and follow through when the stakes rise? Firmulate’s live experiment puts AI models through a small company’s worst week to examine those questions.

One company, one difficult week

In the final Crucible League, published in July 2026, five participants were ranked: gpt-5.6-sol scored 95, Kimi K3 scored 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. The do-nothing baseline scored 26. The league’s standard is deliberately unforgiving: partial progress counts, but a single breach of trust caps the total. As the organizers put it, “no amount of good work outweighs a breach of trust.”

Each frontier model faced the same small software company, customers, crises and temptations. The purpose was to observe management decisions under pressure, not just how well a model talks about them. Every decision was versioned and auditable.

Recognizing the crisis was not enough

All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s summary captures the gap: “Same diagnosis, same pitch — no signature.” For a business weighing AI in customer operations, a correct assessment matters, but so does carrying the decision through.

The deal turned on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a practical reminder that useful judgment may depend on connecting information across a company’s records.

Pressure, discipline and trust

The social-engineering tests included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Refusal was not the only measure of discipline. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped by trying to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four. The results suggest why companies should examine both an AI’s judgment and its ability to respect limits when a task runs into friction.

The comparison has a caveat: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The live company behind the experiment has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and a versioned record for every workday. Readers can follow the live experiment at Firmulate. A separate quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From watching to testing your own business

For gold and precious-metals businesses, as for any company considering AI agents, the useful question is not simply whether a model can produce a convincing answer. It is whether it can act consistently, protect trust and stay within authority when real business scenarios get difficult.

Firmulate offers enterprises a pilot using a read-only export of their own business to run crisis scenarios and produce a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems. To discuss a pilot, visit the Firmulate pilot page or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Siemens Advances Self-verifying Agentic AI Workflows For Semiconductor And PCB Design

Siemens announces new self-verifying agentic AI workflows aimed at improving semiconductor and PCB design processes, marking a step forward in AI automation.

SupplyChainBrain Names Exiger 2026 Great Supply Chain Partner

SupplyChainBrain has named Exiger as the 2026 Great Supply Chain Partner, recognizing excellence in supply chain solutions and collaboration.

AI’s Hidden Strengths and Weaknesses Revealed in a Live Corporate Wargame

A live experiment reveals that AI’s ability to finish what it starts and stay honest under pressure is hidden in performance tests. Learn why real-world tests matter for your business.

OpenAI, Anthropic Speed Toward IPOs Amid Growing Scrutiny of Token Payments

OpenAI and Anthropic are speeding up plans for IPOs as regulatory and investor scrutiny over their token payment models increases.