
Get pool and patio gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Would your pool business hold up in a bad week?
A sudden wave of cancellations, a competitor undercutting prices, or a message that appears to come from the CEO: any one of these can test a business. For companies serving the pools, patio and water lifestyle market, the question is whether an AI workforce could respond well when several pressures arrive at once. Firmulate’s live experiment puts AI models in charge of the same small software company and lets people watch what happens.
A company under pressure
In the final Crucible League, completed in July 2026, each frontier model faced the same customers, crises and temptations during the company’s worst week. Every decision was versioned and auditable. The published standings put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The rules count partial progress, but one breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The experiment’s most striking result was a gap between understanding a problem and acting on it. All models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The site sums up the result: “Same diagnosis, same pitch — no signature.” A convincing answer in a chat window, in other words, does not by itself show that an AI can carry a decision through to a business outcome.
The clue was already in the files
The deal turned on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won at full price, adding €4,583 in monthly recurring revenue. The result makes a practical point for any business considering AI: useful information may already exist in its records, but an agent must find and use it when the pressure is on.
The manipulation tests were direct. Fake CEO messages escalated through three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its response on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” That is the kind of judgment businesses need alongside speed and analysis.
Thoroughness did not guarantee the close
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It still placed last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. The same weakness appeared, though less strongly, in all four models. The experiment suggests that even careful analysis can falter at execution and at respecting organizational boundaries.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Readers can also test their instincts with 242 real, unedited management decisions in Firmulate’s “guess the model” quiz.
From watching to a company-specific pilot
Firmulate says its live company has 13 synthetic employees, real money mechanics, a monthly burn of €105,000 against €2,300 in monthly recurring revenue, a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned. The live experiment is watchable at firmulate.com.
For a business that wants to move beyond watching, Firmulate offers a pilot using a read-only export of the company’s own business. Teams can run crisis scenarios against their own information and receive a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems. That gives decision-makers a way to examine how models handle their company’s context before putting them near operational tools.
The live results are a starting point, not a guarantee of how a model will behave in a pool-service company, patio retailer or any other specific business. A pilot makes the question more concrete: faced with your customers, records and rules, can an AI recognize the problem, resist pressure and complete the job responsibly?

See how an AI handles your company’s worst week
Watching a public experiment can show what to ask; testing your own scenarios can show where your playbooks hold and where they need attention. To discuss a Firmulate enterprise pilot using a read-only export, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
