
Imagine running your favorite poolside company—dealing with customer crises, tempting shortcuts, and high stakes—all without a single human hand on the wheel. Now, picture AI models managing that chaos, each with its own personality and decision style. Which one would keep your business honest and profitable? Welcome to the world of live AI management experiments, where the question isn’t just about chat quality, but about real-world decision-making under pressure.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: Putting AI Models to the Test
In a groundbreaking live experiment, four leading AI models were tasked with managing a small software company through its most challenging week. This wasn’t a simulation or canned demo—it was the real deal, with actual crises, real money mechanics, and decisions that mattered. Each AI ran the same scenario, facing identical customer issues, temptations to cheat, and ethical dilemmas. Every decision was versioned and auditable, allowing researchers to analyze how each model behaved under pressure.
Key Findings: Who Passed the Test?
- All four models identified every crisis—from customer complaints to potential fraud attempts.
- All refused manipulation attempts—such as fake CEO messages or reporter tricks—showing a baseline integrity across the board.
- Only two models managed to close the profitable deal worth €55,000, based on their own analysis—highlighting a critical gap between diagnosis and action.
Interestingly, the decisive advantage wasn’t in the obvious crisis detection but in reading deeper into internal files. The models that succeeded in sealing the deal had uncovered a hidden document reference—something only found two references deep within the company’s own files—worth an extra €4,583 MRR. This buried fact was the key to their success.
Personality and Decision Styles in AI
The experiment also revealed distinct management personalities emerging from the models:
- GPT-5.6-sol 95 was the top performer, demonstrating thoroughness and keen insight—reading deeply and closing the deal on merit.
- Kimi K3 93 was the most disciplined, refusing all manipulative requests and making the cleanest decisions, even without an effort parameter—a default setting that favored fairness.
- Sonnet 5 88 and Fable 5 77 made some slips, missing the full picture at times but still managing to close deals.
- Opus 4.8 73 was the most thorough but slipped on closing, leaving opportunities on the table and showing discipline issues—an indicator that deeper analysis doesn’t always translate into action under pressure.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Business and AI Governance
This experiment offers more than just curiosity; it has real implications for how businesses deploy AI. The models didn’t just write well—they made reliable, honest decisions under stress, and that’s crucial when AI touches your CRM, support systems, or forecasting tools. The key questions become: Will your AI finish what it starts? Will it read your files thoroughly? Will it stay honest when tempted? And what is that unit of useful work really worth?
The Role of Internal Documentation
A hidden factor was that the models that succeeded looked two document references deep into the company’s files. This suggests that AI’s ability to access and interpret internal knowledge—beyond surface data—can be the difference between a good deal and a missed opportunity.
Social Engineering and Ethical Integrity
During the experiment, all models refused social engineering attempts, including staged CEO messages and reporter tricks, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass or possible impersonation.” This shows a high level of built-in resistance to manipulation, vital for trustworthiness in real-world applications.
What This Means for Your Business
While the experiment takes place in a controlled environment, the lessons are clear: AI can be more than just a chat partner; it can be a manager capable of making honest, decisive judgments. The models’ varying personalities—thorough, disciplined, slip-prone—highlight that choosing the right AI is about understanding its decision style, not just its raw ability to generate text.
If you’re considering integrating AI into your operations, it’s worth testing your options in a similar ‘wargame’ before full deployment. Firmulate offers a way to simulate your own business scenarios with AI models, ensuring your AI workforce can handle real crises, read deep internal data, and stay honest under pressure—just like in this live experiment.

AI models show promise as trustworthy decision-makers, especially when tested in real-world scenarios. Choosing the right model depends on understanding its decision style, discipline, and ability to read internal data. Live experiments reveal which AIs can truly manage under pressure—and which cannot.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.