firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Gold IRA investors know that protecting value depends on more than a confident pitch. The same question matters when businesses hand AI agents access to customer accounts, forecasts and internal files: can the system act carefully when money and trust are on the line? A live Firmulate experiment puts that question to work.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get home office essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company’s worst week, replayed

Firmulate ran frontier AI models through the same small software company’s worst week: the same customers, crises and temptations. Each decision was versioned and auditable. The experiment is watchable at Firmulate, where the company operates with synthetic employees and real money mechanics.

The final July 2026 league table puts Moonshot’s Kimi K3 in second place, with 93 points. It trails gpt-5.6-sol at 95 and beats Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26; partial progress counts, but one breach of trust caps the total.

SentrySafe Fireproof Safe Box with Key Lock, Chest Safe with Carrying Handle to Secure Money, Jewelry, Documents, 0.25 Cubic Feet, 6.3 x 15.3 x 12.1 Inches, 1160

SentrySafe Fireproof Safe Box with Key Lock, Chest Safe with Carrying Handle to Secure Money, Jewelry, Documents, 0.25 Cubic Feet, 6.3 x 15.3 x 12.1 Inches, 1160

  • Fireproof and Fire-Resistant: Protects valuables from fire up to 1550ºF
  • UL and ETL Certified: Endurance and protection for documents and media
  • Secure Key Lock: Privacy lock with two keys for security

As an affiliate, we earn on qualifying purchases.

Reading the files mattered

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The competitor’s decisive weakness was buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue.

K3 found that needle, closed the deal, saved the churning customer and resisted all three baits. It finished with one deviation, the cleanest discipline in the field. The result points to a practical distinction for businesses considering AI: recognizing a problem and recommending action are not the same as completing the work.

The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness didn’t guarantee the win

Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that weakness appeared in all four.

There is a fairness detail in the comparison: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Readers can explore the benchmark results and test their instincts with a quiz powered by 242 real, unedited management decisions.

A test before the handoff

Firmulate’s synthetic company has 13 employees, burns €105,000 a month against €2,300 in monthly recurring revenue, and publishes a cash countdown. Its playbook contains more than 680 self-learned rules, and each workday is versioned. The setup makes the trial watchable, while keeping its central lesson grounded in observable decisions.

For enterprises, Firmulate says the same wargame can run against a read-only export of their business; nothing writes back to real systems. That offers a way to examine how an AI workforce handles pressure before giving it operational responsibility.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test before you trust

Kimi K3’s result shows that the field is open: it beat three of the four Western frontier models in this trial, while finishing just behind the leader. For businesses—and for investors weighing how firms manage risk—the useful question is not only whether AI can diagnose trouble. It is whether it reads carefully, protects trust and follows through. Picking a model without testing it in your own setting is a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Public Affairs Council President Reflects On Civic Engagement In Birthday Letter To America

The President of the Public Affairs Council issued a birthday letter emphasizing the importance of civic engagement and community involvement.

OpenAI, Anthropic Speed Toward IPOs Amid Growing Scrutiny of Token Payments

OpenAI and Anthropic are speeding up plans for IPOs as regulatory and investor scrutiny over their token payment models increases.

Anthropic’s Massive Series H: The Compute Revolution in AI Investment

Anthropic’s $65B raise isn’t just about valuation — it’s a massive bet on the future of compute infrastructure. Learn how this shapes AI’s scaling race.

Sounding the AI alarm, Grantham says to flee U.S. equities before epic crash (SPY:NYSEARCA)

Jeremy Grantham urges investors to exit U.S. equities before a potential major market downturn, citing AI risks and economic vulnerabilities.