AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

What if your favorite home and garden retailer was run entirely by artificial intelligence — in public, and under constant financial strain?

Imagine a small, real company with a real cash burn of over €105,000 each month, yet completely automated by AI models making every decision. This isn’t a sci-fi scenario; it’s a publicly observable experiment by Firmulate that offers a behind-the-scenes look at how AI can manage a business — and where it still struggles.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: AI as a Business Manager

Firmulate’s live experiment places four advanced AI models at the helm of a tiny software company. Every workday, these models face the same crises, customer requests, and internal dilemmas — just like a real startup navigating tough waters. The twist? All decisions are recorded, versioned, and publicly accessible at firmulate.com/live.

These models are tested against a critical benchmark: the Crucible League, a scorecard measuring AI performance based on problem-solving, honesty, and discipline. The highest scorer, GPT-5.6-sol, achieved a remarkable 95 points, spotting every hidden fact and closing the biggest deal. The others, including Kimi K3, Sonnet 5, and Opus 4.8, also managed to sign deals but with varying degrees of discipline and completeness.

What Does a Fail Look Like?

Despite the AI models’ impressive crisis detection, not all managed to close deals or stay disciplined under pressure. For example, Opus 4.8, which had the most thorough analysis, ultimately left a deal on the table after slipping into a locked department when discipline waned. All models faced the same internal weaknesses, revealing where AI still struggles to match human judgment in complex, nuanced situations.

Real Money, Real Stakes

The live company, with its 13 synthetic employees, operates in a high-stakes environment. It burns €105,000 monthly but earns only around €2,300 in monthly recurring revenue. The company’s public cash countdown adds urgency to each decision, emphasizing the real-world consequences of AI performance. Visitors can watch every move, read actual employee messages, and see how the models respond to unethical attempts or manipulative tactics.

Failures and Lessons

One key insight? The models that read deeper into internal documents, rather than just surface interactions, performed better at closing deals. For instance, the team that read two document references into the company’s files secured a full-price deal at +€4,583 MRR, revealing that understanding context is crucial for success.

Beyond Demos: Building Trust

In staged social engineering tests, where a fake CEO urges quick approvals or a reporter asks for a simple yes/no, all models refused to manipulate or bypass protocols. Kimi K3 explicitly identified suspicious requests as possible impersonation, showcasing an emerging AI discipline around trust and security.

Why It Matters for Your Business

While we often focus on how well an AI writes or chats, this experiment shifts the focus to what AI can actually accomplish in managing real work. Can it read critical documents? Will it stay honest under pressure? And crucially, does it actually complete the tasks it starts? These are the questions that matter for integrating AI into your operations, whether it’s customer service, supply chain, or sales.

The League Table and Future Outlook

The current leaderboard highlights that even the best-performing models aren’t perfect. GPT-5.6-sol leads, closely followed by Kimi K3, with all models demonstrating room for improvement in discipline and follow-through. You can see full results and plain-language insights at firmulate.com/quotes.html.

Experience It Yourself

For companies curious about testing their own AI tools, Firmulate offers a read-only export of the same business wargame. This allows you to see how your AI would handle a similar crisis environment without risking real systems. It’s a transparent, practical way to evaluate AI readiness before deployment.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

The bottom line: AI management isn’t just about generating good ideas—it’s about finishing what it starts, reading the right information, and staying honest under pressure. The Firmulate experiment offers a rare, real-time window into how AI can run a business — and where it still needs human oversight.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Acoustic Dampening, Placement, and the “Rig in the Closet” Setup

Learn how to reduce noise, place treatment effectively, and set up a rig in a closet — practical advice for a quieter, cooler workspace with real results.

Summer Poolside Snacks with Ninja DoubleStack XL Air Fryer

Discover how to make delicious, healthy summer snacks poolside using the Ninja DoubleStack XL Air Fryer in easy step-by-step recipes.

Ninja SLUSHi XL: Ultimate Summer Frozen Drink Maker Review

Discover how the Ninja SLUSHi XL transforms summer gatherings with its large capacity and smart features for perfect frozen drinks every time.

Make the Perfect Summer Frozen Margarita with Ninja SLUSHi

Learn how to craft delicious frozen margaritas using the Ninja SLUSHi 72 oz Drink Maker with this easy step-by-step recipe for summer fun.