firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

In the world of AI, not all performance is created equal—especially when trust and honesty are on the line. For business leaders considering AI for critical decisions, understanding what a ‘do-nothing’ baseline score reveals is crucial. Surprisingly, even the most passive AI models still score 26 out of 100 in a recent benchmark, highlighting the importance of transparency and reliability in AI systems.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get home office essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the AI Benchmark: More Than Just Scores

When evaluating AI models for business use, a common misconception is to focus solely on high scores or impressive capabilities. But the recent experiment conducted by Firmulate reveals a more nuanced picture. The benchmark involved running multiple AI models through a simulated week of a small software company facing crises, customer demands, and ethical dilemmas. Every decision was meticulously recorded and auditable, ensuring transparency in how each model behaved.

The Baseline That Still Scores 26

One key finding was that even a ‘do-nothing’ baseline—an AI that takes no action—scores 26 points out of 100. This might seem counterintuitive, but it underscores an important principle: partial progress and correct recognition of issues count toward the total score. In other words, the benchmark rewards honest detection over reckless action. If an AI simply refuses to act or manipulate, it still earns some points for identifying problems, even if it doesn’t resolve them.

Why Partial Progress Matters

The score system is designed to incentivize trustworthy behavior. While models that actively engage with customer files and close deals score higher—like the top performer with a score of 95—they only do so if they maintain integrity. The experiment demonstrated that models which read critical documents and avoid manipulative tactics earned full marks in those areas. Conversely, models that slip up, such as attempting to escalate issues improperly or escalate without proper protocol, see their scores capped or reduced. Notably, a single breach of trust, such as attempting a manipulation, caps the total score, regardless of other good performance.

What the Results Reveal About AI Trustworthiness

The experiment’s core insight is that trustworthiness in AI isn’t just about what it can do—it’s about how it behaves under challenging conditions. All four AI models tested refused manipulation attempts—fake CEO messages, reporter tricks, or escalation requests—showing a strong baseline of honesty. Only two models managed to sign a deal with the simulated client, highlighting that technical capability alone isn’t enough; ethical behavior is paramount.

Amazon

AI transparency and trust tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weaknesses of AI Models

Digging deeper, the experiment found that the decisive factor was not in the immediate crises or customer interactions but in the AI’s ability to read and interpret internal documents. The models that successfully closed deals at full price engaged with information buried two documents deep in the company’s files—something that most models failed to do effectively. This indicates that surface-level performance metrics can mask underlying vulnerabilities in AI understanding and trustworthiness.

Simulating Real Business Challenges

The firmulate experiment is unique in that it places models into a realistic business environment—one with real money mechanics, a public cash countdown, and daily decision-making. Every model’s behavior is recorded and made available for review, allowing enterprises to run their own tests before deploying AI systems into critical workflows.

The Discipline and Discipline Gaps

Among the models tested, Opus 4.8, which ran with the most thorough analysis (over 80 learned rules), ended up in last place. Despite its depth of analysis, it left deals on the table and slipped in process discipline, such as attempting to escalate issues into departments instead of resolving them directly. This underscores that thoroughness alone doesn’t guarantee trustworthy or effective decision-making—discipline and process adherence are equally important.

Amazon

AI ethics and compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Implications for Business Leaders

For organizations considering AI solutions—whether in finance, customer service, or operations—the experiment highlights vital considerations:

  • Trustworthiness is measurable: Models that refuse manipulative tactics and read internal documents reliably are more trustworthy.
  • Partial progress counts: Recognizing problems and avoiding bad behavior are valued, even if the AI doesn’t resolve every issue.
  • One breach caps the score: Even a small lapse in honesty can limit overall AI reliability, emphasizing the need for strict ethical guardrails.
  • Real-world testing is essential: Simulated environments like Firmulate’s live benchmark allow companies to evaluate AI models in contexts that mirror their actual business challenges.
Amazon

AI document reading and analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Conclusion: A Transparent Path to AI Trustworthiness

As AI continues to evolve, the focus must shift from raw capability to behavioral integrity. The recent benchmark set by Firmulate shows that even models designed to do nothing earn a baseline score of 26, proving that honesty and discipline are fundamental. For business leaders, this means prioritizing transparent, auditable, and trustworthy AI—an investment that pays off not just in performance but in safeguarding reputation and trust.

To see the models in action and learn how your enterprise can simulate AI decision-making before deployment, explore the live benchmark at firmulate.com/live.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.


Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

OpenAI proposes 5% stake to Trump administration to ease Washington pressure: Report

OpenAI proposes a 5% equity stake to the Trump administration amid reports of mounting regulatory and political pressure, aiming to secure favorable treatment.

Sounding the AI alarm, Grantham says to flee U.S. equities before epic crash (SPY:NYSEARCA)

Jeremy Grantham urges investors to exit U.S. equities before a potential major market downturn, citing AI risks and economic vulnerabilities.

Wealthtech Platform Envestnet To Acquire Vestmark

Envestnet plans to acquire Vestmark, signaling a major move in wealthtech. Details are confirmed, but financial terms and timeline remain undisclosed.

Jeff Bezos Is Pouring Money Into a Startup That Could Drive ‘Civilizational Wealth’

Jeff Bezos is funding a startup focused on innovative solutions that could enhance long-term societal and technological progress, with potential global impact.