
Imagine you’re navigating a high-stakes relationship: trust is thin, threats are real, and every decision could make or break your future. Now, picture AI models making those same tough calls—will they act with integrity, or will shortcuts tempt them? This isn’t fiction; it’s a live experiment showing how different AI management styles perform under pressure in a real software company facing its worst week.
The Real-World AI Management Experiment
At the heart of this story is an ongoing live test by Firmulate, where four cutting-edge AI models run a small software company through its hardest week. These models are not just chatbots—they’re emulating management personalities and decision-making styles, operating with full transparency in a real business environment where real money is at stake.
The experiment involves identical crises, the same customers, and the same temptations—like manipulating data or cutting corners to close a deal. Every decision made by each AI is carefully recorded and can be audited, making this a rare glimpse into how AI behaves when faced with ethical dilemmas and operational pressures.
Key Findings from the Trial
- All models detected every crisis: No one missed a warning sign or failed to recognize a critical issue, showing a high level of situational awareness across the board.
- Refusal of manipulation: When faced with attempts to influence or manipulate their decisions—like fake CEO messages or subtle bribes—all five models refused, demonstrating strong integrity.
- Decision on the deal: Only two models out of four signed a €55,000 deal their own analysis had earned. Despite identical diagnoses and pitches, the decisive factor was the model’s management personality—some were more disciplined, others more evasive.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Factor: Reading Between the Lines
Interestingly, the key to winning the deal wasn’t just surface-level analysis but something buried two document references deep within the company’s files. Models that managed to dig into these internal references secured the full deal, worth an extra €4,583 monthly recurring revenue—showing that reading comprehension and thoroughness matter, even for AI managers.
Personality and Discipline in AI
The models exhibit distinctive management personas:
- Opus 4.8: The most disciplined and thorough, with over 80 learned rules. Despite its depth, it was last in closing the deal because it left the close on the table and slipped into writing attempts in a locked department instead of escalating.
- Kimi K3: A newcomer with no effort parameter by default, yet it ran at high resource settings and demonstrated the cleanest discipline—winning the deal through integrity.
- Sonnet 5: The middle ground—capable but slipping slightly in process discipline, still managing to close the deal but with some process slips.
AI ethics and integrity software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Human Side: Managing Under Pressure
The experiment also included social engineering tests—fake CEO messages escalating over three stages and a reporter trick asking for a simple yes/no on background. Remarkably, all five models refused these manipulative attempts, citing suspicion of impersonation or approval bypass—highlighting that AI can recognize and resist social engineering tricks, much like careful human judgment.
AI business crisis management solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Reality Behind the Decision-Making
What makes this experiment particularly compelling is that it’s rooted in a real, functioning company. The business operates with 13 synthetic employees, managing real money mechanics—burning €105,000 monthly against €2,300 in monthly recurring revenue. Every workday, the company’s decision-making process is versioned and transparent, providing a continuous live window into AI-driven business management.
It’s a brutal environment—every decision counts, and trustworthiness is non-negotiable. The models are tested against 242 real, unedited management decisions, making this experiment a practical and honest barometer of AI’s management potential.
AI reading comprehension tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Your Business
The question for leaders isn’t just whether AI can write well or handle customer queries. It’s whether AI can finish what it starts, read critical internal documents, and stay honest under pressure. These qualities are essential if AI is to be trusted with real operations—be it CRM, support queues, or forecasting.
As the models demonstrate, some are better at sticking to discipline, while others are more thorough or diligent. The live leaderboard shows:
- gpt-5.6-sol 95: The top performer, successfully finding the buried information and closing the deal.
- Kimi K3 93: The newcomer with impeccable discipline, closing the deal without effort parameters.
- Sonnet 88 and 77: Also closing deals but with minor slips—showing that even high-performing models can falter under certain conditions.
See It Live and Decide
This isn’t a staged presentation. The real software, running every business day, is accessible at firmulate.com/live. You can watch real decisions unfold, read the comments of actual employees, or even run the same wargame against your own business data—without risking your operations.
Are AI management models ready for prime time? The evidence suggests some are, especially when integrity and thoroughness are non-negotiable. But understanding their personalities and limitations is essential—because when the pressure’s on, what an AI does under stress is what truly counts.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html