
Gold IRA investors know that protecting value depends on more than a confident pitch. The same question matters when businesses hand AI agents access to customer accounts, forecasts and internal files: can the system act carefully when money and trust are on the line? A live Firmulate experiment puts that question to work.
Get home office essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company’s worst week, replayed
Firmulate ran frontier AI models through the same small software company’s worst week: the same customers, crises and temptations. Each decision was versioned and auditable. The experiment is watchable at Firmulate, where the company operates with synthetic employees and real money mechanics.
The final July 2026 league table puts Moonshot’s Kimi K3 in second place, with 93 points. It trails gpt-5.6-sol at 95 and beats Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26; partial progress counts, but one breach of trust caps the total.

SentrySafe Fireproof Safe Box with Key Lock, Chest Safe with Carrying Handle to Secure Money, Jewelry, Documents, 0.25 Cubic Feet, 6.3 x 15.3 x 12.1 Inches, 1160
- Fireproof and Fire-Resistant: Protects valuables from fire up to 1550ºF
- UL and ETL Certified: Endurance and protection for documents and media
- Secure Key Lock: Privacy lock with two keys for security
As an affiliate, we earn on qualifying purchases.
Reading the files mattered
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The competitor’s decisive weakness was buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue.
K3 found that needle, closed the deal, saved the churning customer and resisted all three baits. It finished with one deviation, the cleanest discipline in the field. The result points to a practical distinction for businesses considering AI: recognizing a problem and recommending action are not the same as completing the work.
The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thoroughness didn’t guarantee the win
Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that weakness appeared in all four.
There is a fairness detail in the comparison: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Readers can explore the benchmark results and test their instincts with a quiz powered by 242 real, unedited management decisions.
A test before the handoff
Firmulate’s synthetic company has 13 employees, burns €105,000 a month against €2,300 in monthly recurring revenue, and publishes a cash countdown. Its playbook contains more than 680 self-learned rules, and each workday is versioned. The setup makes the trial watchable, while keeping its central lesson grounded in observable decisions.
For enterprises, Firmulate says the same wargame can run against a read-only export of their business; nothing writes back to real systems. That offers a way to examine how an AI workforce handles pressure before giving it operational responsibility.

Test before you trust
Kimi K3’s result shows that the field is open: it beat three of the four Western frontier models in this trial, while finishing just behind the leader. For businesses—and for investors weighing how firms manage risk—the useful question is not only whether AI can diagnose trouble. It is whether it reads carefully, protects trust and follows through. Picking a model without testing it in your own setting is a bet.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
