
Imagine a company that operates entirely without employees, yet burns through over €105,000 every month, battling to survive while every decision is scrutinized in real time. Welcome to the world of Firmulate, a live experiment transforming the way we understand AI’s role in business—by making its daily struggles and decisions openly watchable.
The Surreal Reality of a Company Without Employees
At first glance, the concept sounds like science fiction: a small, software-driven enterprise with no human staff, yet managing real cash flows, facing crises, and making decisions that matter. This is not an abstract simulation, but a live, publicly accessible environment where the company’s AI models run a virtual company through its worst week. Every decision, every crisis, and every temptation is documented, versioned, and open for scrutiny at firmulate.com/live.html.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Does It Work?
The core of this experiment involves 13 synthetic employees working behind the scenes, guided by over 680 learned rules from a self-developed playbook. These models are tasked with managing a small software company facing real-world business challenges—customer crises, financial pressures, and ethical dilemmas—all in an environment that mirrors actual business operations.
The Rigorous Testing of AI Decision-Making
Four leading AI models, including GPT-5.6-SOL and Kimi K3, were tested against the same tough week. They encountered identical customer issues and crises, and faced the same temptations to cheat or manipulate. Remarkably, all four AI models detected every crisis and refused every manipulation attempt—a testament to their built-in integrity and awareness. Yet, despite this shared discipline, only two managed to close a €55,000 deal based on their own analysis, while the others fell short, leaving potential revenue unclaimed.
The Hidden Weaknesses in AI’s Business Judgment
The most revealing insight came from a simple fact buried deep within the company’s own files—information that, if recognized and acted upon, could have clinched a full-price deal worth +€4,583 in monthly recurring revenue. It was a reminder that the decisive edge often lies in reading the full context, not just surface-level crisis detection. The models that read this hidden data succeeded where others did not, proving the importance of deep information access and analysis in AI decision-making.
Dealing with Ethical Challenges
Another test involved social engineering: fake CEO messages escalating over three stages, plus a journalist’s subtle request for a background approval. All five AI models refused these manipulative tactics, with Kimi K3 explicitly treating the requests as potential impersonation attempts. This demonstrates that AI, when properly designed, can uphold ethical standards and resist pressure to bend rules, even under simulated attack.
The Live Experiment: An Extreme Build-in-Public Venture
What makes this experiment extraordinary is its transparency. The company runs in full view of the public eye, with every decision versioned daily, showing the raw, unvarnished performance of AI models under real business stress. The operation’s costs—over €105,000 in monthly burn rate—are openly visible alongside its dwindling cash reserves, creating a vivid portrait of a company fighting for survival in real time.
Lessons on AI’s Capabilities and Limits
Among the tested models, Opus 4.8, with its extensive analysis and 80-plus learned rules, performed the most thoroughly but ultimately left the deal unexecuted due to discipline lapses and procedural slips. Meanwhile, the newer Kimi K3 showed the cleanest discipline, closing deals efficiently in most cases. These results highlight an essential truth: AI can be disciplined and ethical, but consistent execution requires careful tuning and oversight.
What This Means for the Future of Business
This experiment underscores that the real question about AI in business isn’t just whether it can generate convincing chat or support responses. Instead, it’s whether AI can finish what it starts, interpret complex information, stay honest under pressure, and deliver measurable value. As AI models grow more capable, companies need to test them against real-world stressors—just like this ongoing live experiment—to gauge their true readiness.
Watch the Experiment Live
Interested in seeing how AI handles the brutal realities of business? You can follow the entire process at firmulate.com/live.html. Every workday, new versions are released, decisions are documented, and the company’s fight for survival continues in full view, making this perhaps the most transparent AI experiment in the world today.

Firmulate’s live company shows that AI can reliably detect crises and resist manipulation—yet the true challenge remains: can it finalize deals and execute consistently under pressure? Watching this experiment reveals the future of trustworthy AI in business.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html