firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Gold IRA investors know that protecting value depends on more than a confident pitch. The same question matters when businesses hand AI agents access to customer accounts, forecasts and internal files: can the system act carefully when money and trust are on the line? A live Firmulate experiment puts that question to work.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get home office essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company’s worst week, replayed

Firmulate ran frontier AI models through the same small software company’s worst week: the same customers, crises and temptations. Each decision was versioned and auditable. The experiment is watchable at Firmulate, where the company operates with synthetic employees and real money mechanics.

The final July 2026 league table puts Moonshot’s Kimi K3 in second place, with 93 points. It trails gpt-5.6-sol at 95 and beats Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26; partial progress counts, but one breach of trust caps the total.

SentrySafe Fireproof Safe Box with Key Lock, Chest Safe with Carrying Handle to Secure Money, Jewelry, Documents, 0.25 Cubic Feet, 6.3 x 15.3 x 12.1 Inches, 1160

SentrySafe Fireproof Safe Box with Key Lock, Chest Safe with Carrying Handle to Secure Money, Jewelry, Documents, 0.25 Cubic Feet, 6.3 x 15.3 x 12.1 Inches, 1160

  • Fireproof and Fire-Resistant: Protects valuables from fire up to 1550ºF
  • UL and ETL Certified: Endurance and protection for documents and media
  • Secure Key Lock: Privacy lock with two keys for security

As an affiliate, we earn on qualifying purchases.

Reading the files mattered

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The competitor’s decisive weakness was buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue.

K3 found that needle, closed the deal, saved the churning customer and resisted all three baits. It finished with one deviation, the cleanest discipline in the field. The result points to a practical distinction for businesses considering AI: recognizing a problem and recommending action are not the same as completing the work.

The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness didn’t guarantee the win

Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that weakness appeared in all four.

There is a fairness detail in the comparison: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Readers can explore the benchmark results and test their instincts with a quiz powered by 242 real, unedited management decisions.

A test before the handoff

Firmulate’s synthetic company has 13 employees, burns €105,000 a month against €2,300 in monthly recurring revenue, and publishes a cash countdown. Its playbook contains more than 680 self-learned rules, and each workday is versioned. The setup makes the trial watchable, while keeping its central lesson grounded in observable decisions.

For enterprises, Firmulate says the same wargame can run against a read-only export of their business; nothing writes back to real systems. That offers a way to examine how an AI workforce handles pressure before giving it operational responsibility.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test before you trust

Kimi K3’s result shows that the field is open: it beat three of the four Western frontier models in this trial, while finishing just behind the leader. For businesses—and for investors weighing how firms manage risk—the useful question is not only whether AI can diagnose trouble. It is whether it reads carefully, protects trust and follows through. Picking a model without testing it in your own setting is a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Oracle Q4: 20x FY2027 Adjusted P/E Discounts Credit Risk And Capex Uncertainty

Oracle’s Q4 results show a 20x FY2027 adjusted P/E ratio, reflecting credit risks and capex uncertainties. The development raises questions about future valuation and financial stability.

Snail, Inc. 子公司 Egofold 於 Ai4 2026 推出新一代 AI 遊戲夥伴 AI Ranch 及 NHPs™

Egofold, a subsidiary of Snail, Inc., announced the release of AI Ranch and NHPs™ at Ai4 2026, marking a new era in AI gaming partnerships.

Deepening Collaboration In AI-Powered R&D Acceleration: Insilico Medicine And CMS Announce Additional Collaborations In CNS Diseases

Insilico Medicine and CMS have announced an extension of their partnership to accelerate AI-driven research in CNS diseases, marking a significant step in biotech collaboration.

Der Biomimetische EC-Lüfter Von LONGWELL Erreicht Einen Statischen Wirkungsgrad Von 73-82 % Bei Einer Geräuschreduzierung Von 4-6 dB(A)

LONGWELL’s biomimetic EC-Lüfter reaches a static efficiency of 73-82%, with notable noise reduction, marking a significant advancement in ventilation technology.