firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get gifts for the two of you delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Are AI models ready to run a company — and do they play fair?

In a world where AI promises to revolutionize everything from customer service to strategic decision-making, a recent experiment puts these claims to the test. Can a newcomer really beat seasoned models — especially when the stakes are real, and the pressure is high? Let’s explore how cutting-edge AI models tackled their toughest week yet, revealing surprising insights about trust, honesty, and performance in automated management.

Amazon

AI management decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Challenge: Simulating a Tough Week

Data from the latest Crucible League competition in July 2026 shows five top AI models put through their paces on a real software company facing its worst week. Every model received identical scenarios: the same customers, crises, and temptations. The goal? To see if these AI agents could diagnose issues, make honest decisions, and close a €55,000 deal based solely on their analysis and recommendations.

Amazon

business AI simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Models Performed

Out of the group, four models managed to spot all crises and refused every manipulation attempt, demonstrating strong integrity. The standout was gpt-5.6-sol, which scored 95 out of 100, narrowly beating the newcomer Kimi K3 with a 93. The latter also closed the deal, maintaining the cleanest discipline of the field, and found the buried security detail needed to secure the deal at full price — worth more than €4,500 in monthly recurring revenue.

Similarly, Sonnet 5 and Fable 5 also closed the deal, but with minor slips, scoring 88 and 77 respectively. Opus 4.8, despite a thorough approach with over 80 learned rules, lagged behind at 73, leaving some opportunities on the table and slipping in discipline. Intriguingly, all models refused a staged social engineering attack — fake CEO messages escalating over multiple stages — illustrating their resistance to manipulation.

Amazon

AI company management platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness

Despite these strong performances, a critical weakness was uncovered: the decisive difference was in reading company files more deeply. Kimi K3, for instance, found a buried reference in the company’s own documents that clinched the deal at full price. These hidden insights proved vital, yet such analysis isn’t apparent in simple chat demos; it’s the difference between winning and losing in real management decisions.

Amazon

AI business analytics software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real-World Company and Its AI Workforce

The experiment isn’t just theoretical. The AI models manage a real company with 13 synthetic employees, real money mechanics, and a public cash countdown. This live setup burns €105,000 monthly against a modest €2,300 in monthly revenue, highlighting the challenge of real business management under AI control. Every decision is versioned and auditable, providing transparency and accountability in this high-stakes environment.

Fairness and Testing Conditions

It’s important to note that Kimi K3 was run without an effort parameter (the API default), while the others operated at a high effort setting. This fairness note ensures comparisons are grounded in equivalent testing conditions.

Implications for Business and AI Adoption

What does this mean for companies considering AI automation? Simply put, the question isn’t just whether an AI model writes well or answers convincingly. It’s whether it can finish what it starts, stay honest under pressure, and uncover hidden insights that matter. The current league table shows a tight race, with the newcomer Kimi K3 showing it can rival and even beat established players like gpt-5.6-sol.

Watch and Decide

For those curious about how these models perform in real-time, the experiment is live at firmulate.com/live. You can see the company in action, observe the decision-making process, and even test your own management choices against the models through interactive quizzes and pilots. This transparent setup underscores a vital point: when it comes to AI in management, trust and performance are intertwined, and a model’s ability to stay honest may be its greatest asset.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Key Takeaway

The latest AI management experiment shows that newcomers can outperform established models in critical tasks, especially when deep analysis and honesty matter most. As AI continues to enter real business environments, choosing a model that can reliably finish what it starts — and resist manipulation — will be essential for success.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

New Research: Childhood Trauma’s Link to Sex and Porn Addiction

Discover how childhood trauma may influence sex and porn addiction and why understanding this link could be crucial to your healing journey.

After Lockdown: Did the Pandemic Make Sex and Porn Addictions Worse?

A deeper look reveals how pandemic-induced stress may have worsened sex and porn addictions, prompting important questions about recovery and control.

Headline: Politician Blames Sex Addiction for Scandal – Can It Be True?

I wonder how blaming sex addiction affects perceptions of accountability and the broader debate on mental health and morality.

Money Secrets and Dating Burnout: Unpacking Betrayal and Fatigue in Modern Relationships

Navigating modern love reveals that 40% hide financial secrets and dating fatigue is rising—discover how transparency and emotional connection can transform your relationships.