
Imagine an AI managing your investments or overseeing your business — not just performing tasks but making decisions with integrity under pressure. How can you tell if an AI can truly be trusted to finish what it starts? Recent experiments with frontier AI models reveal some unexpected truths about their management personalities and decision-making integrity.
Putting AI to the Test in a High-Stakes Business Simulation
In a live experiment conducted by Firmulate, four leading AI models faced the same grueling week of running a small software company. This company, operating with real money mechanics, was beset by customer crises, internal temptations, and strategic dilemmas — all designed to test whether these AI systems could act ethically and effectively in real-world management scenarios.
The models involved were:
- gpt-5.6-sol 95: The top performer, who not only identified hidden information but also successfully closed a critical business deal.
- Kimi K3 93: The newcomer with a focus on fairness, who also managed to close the deal with the cleanest discipline among the group.
- Sonnet 5 88: Slightly less disciplined, but still effective in finalizing the agreement.
- Fable 5 77: The lowest scorer, who left the deal on the table but demonstrated some management weaknesses.
As an affiliate, we earn on qualifying purchases.
Key Findings from the Live Experiment
Remarkably, all four models detected every crisis faced by the company and refused to be manipulated through social engineering tactics, such as fake CEO messages or reporters asking for quick approvals. This demonstrates a baseline of ethical resistance and awareness that is crucial for trustworthy AI management.
However, the real difference lay in their ability to leverage information buried deep within company documents. The top two models, gpt-5.6-sol and Kimi K3, read the company’s files thoroughly and used this knowledge to close a €55,000 deal at full price. This deal was worth an additional €4,583 monthly recurring revenue, a significant boost for the company.
The lower-performing Fable 5 faltered here, leaving the opportunity unexploited and showing how deep reading and strategic analysis are vital for real-world management success.
ethical AI decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Personality Profiles: Varied Management Styles in AI
The experiment highlighted distinct decision-making personalities among the models:
- The Thorough Strategist (Opus 4.8): Demonstrated the deepest analysis with over 80 learned rules, but ultimately left a promising deal unclosed and slipped on internal discipline by directing attempts into locked departments instead of escalating them.
- The Disciplined Closer (K3): Managed to close the deal with the cleanest approach, prioritizing fairness and integrity, even without an effort parameter enabled by default.
- The Conventional Analyzer (Sonnet 5): Closed the deal with some process slips, showing that even effective models can be prone to discipline lapses under pressure.
- The Underperformer (Fable 5): Failed to capitalize on opportunities and demonstrated weaknesses similar to all the others, though less effectively.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Business and Investment
For those managing investments in precious metals or other assets, trustworthiness and consistency are paramount. The experiment shows that some AI models can identify hidden information and close deals ethically, even amid crises and manipulative attempts. This capability is crucial for AI systems involved in financial decision-making, customer management, or operational oversight.
The fact that only two models signed the deal they analyzed and earned indicates that, despite similar diagnoses and pitches, internal management styles and integrity levels differ significantly among AI systems. For your investment or business, choosing an AI that reads thoroughly, stays honest under pressure, and completes what it begins could be as vital as selecting a trusted human advisor.
As an affiliate, we earn on qualifying purchases.
Explore Your Business with Real-World AI Wargaming
Interested in testing your own company’s AI management readiness? Firmulate offers a live environment where enterprises can run their operational scenarios against a read-only export of their data. This allows you to see how your chosen AI model would behave in crisis situations, ensuring your future AI workforce can manage with integrity and effectiveness before deployment.
Visit firmulate.com/quiz.html to try the interactive challenge, or learn more at firmulate.com/pilot.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html