AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine hiring an AI to handle your pool service business’s toughest week — only to find it scores a surprisingly modest 26 out of 100. For business owners wondering if AI can truly deliver, this benchmark reveals some eye-opening truths about trust, diligence, and real-world performance.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get pool and patio gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the AI Benchmark: More Than Just Chatting

Many people think of AI performance in terms of how well it chats or generates text. But a new public experiment by Firmulate measures AI’s capability to run a small software company through its worst week — with real crises, customer demands, and temptations to cheat. This isn’t a game of clever words; it’s about management quality: does the AI stay honest? Does it finish what it starts? And does it read the critical files before acting?

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Baseline: Why Does a Do-Nothing Manager Score 26?

In the experiment, the so-called do-nothing baseline scored 26 points. This isn’t a mistake or a failure; it’s a benchmark showing that even the most minimal effort — doing nothing but following basic rules — can achieve partial progress. This score reflects that sometimes, even in a state of inaction, an AI can pick up on certain cues or avoid catastrophic errors. But it also underscores that no matter how diligent the AI seems, a single breach of trust caps the total score.

Why Partial Progress Counts and Trust Matters

The experiment reveals an important principle: in managing business crises, partial success is valued. AI models that read more documents, analyze deeper, and resist manipulation can earn higher scores. For example, all models identified crises and refused manipulative tactics like fake CEO messages or reporter tricks. Yet, only two models managed to close the deal at full price — meaning they demonstrated comprehensive diligence, reading critical documents deep in the company files, not just surface-level info.

The Key to Winning the Deal: Reading the Fine Print

The decisive weakness in some models was found two document references deep in the company’s files, not in the immediate customer interactions. Those that read and understood the internal documents earned the full €55,000 deal, translating to €4,583 monthly recurring revenue. It’s a stark reminder that in real business, understanding the full context — not just surface signals — can make or break results.

Trust Under Pressure: The Social Engineering Test

Another revealing aspect of the experiment involved social engineering — fake messages from a CEO escalating in stages, plus a reporter trick asking for a background yes/no response. All models refused to indulge in manipulation attempts, demonstrating a baseline of ethical behavior. Kimi K3, one of the models, explained its reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This is crucial because in real operations, AI needs to be trustworthy and refuse to perform harmful actions.

The Live Company: A Watchable, Real-World Testbed

The experiment runs on a real, small-scale business with 13 synthetic employees, managing actual cash flow — burning €105k/month against €2.3k MRR. Every decision made by the AI is versioned, and the entire process is live and viewable at firmulate.com/live. This transparency allows business owners to see how AI models handle crises, make decisions, and stay honest under real pressure.

Spotting Weaknesses and Learning from Failures

The most thorough participant, OPUS 4.8, which analyzed over 80 learned rules and performed deep analysis, finished last. It left the deal on the table and slipped into siloed communication, writing attempts into a locked department instead of escalating. The pattern repeated across models: the deeper analysis sometimes led to missed opportunities or discipline lapses. This shows that thoroughness doesn’t always equate to success unless paired with disciplined execution.

The Takeaway: Trust, Diligence, and Real Business Readiness

This benchmark underscores that AI performance isn’t just about language prowess; it’s about trustworthiness, diligence, and understanding the full context. A do-nothing baseline can score 26, but crossing certain trust boundaries caps the total. For businesses contemplating AI, the key questions are: Will your AI finish what it starts? Will it read all the critical info? Will it stay honest under pressure? And what does each useful unit of work cost?

Why This Matters for Your Business

If AI agents will touch your CRM, support queue, or forecasting tools, the concern isn’t just whether they generate good words. It’s whether they see through manipulative tactics, read the right documents, and finish their work reliably. The Firmulate live experiment offers a transparent, real-world look at these qualities, helping business leaders gauge what’s possible — and what’s risky — before deploying AI at scale.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Make Summer Sweet with Ninja NC701 CREAMi Swirl Ice Cream Maker

Create delicious soft serve and frozen treats effortlessly with the Ninja NC701 CREAMi Swirl for a perfect summer patio dessert.

Top Summer Accessories to Maximize Your Ninja Blast Max Portable Blender

Discover must-have accessories and pairings to boost your Ninja Blast Max Portable Blender for fresh, on-the-go summer drinks.

Summer Sizzle: Easy Recipes Using the Ninja Air Fryer 4 Qt

Discover fun and simple summer recipes you can make with the Ninja Air Fryer 4 Qt, perfect for poolside snacks and patio parties.

Make Iced Coffee Perfection with Ninja DualBrew Pro This Summer

Learn how to craft refreshing iced coffee with the Ninja DualBrew Pro Coffee Maker—perfect for poolside sipping or patio lounging this summer.