
Get pool and patio gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Would you trust an AI with the poolside business?
Running a pool, patio or water-lifestyle company can mean balancing customer promises, cash pressure and the occasional tempting shortcut. A polished chatbot demo won’t show how an AI handles that whole mix. Firmulate’s live company experiment puts models through a rough week of business decisions—and its latest result suggests that choosing an AI workforce without testing it is a bet.
A newcomer takes second place
In Firmulate’s final Crucible league for July 2026, Moonshot’s Kimi K3 scored 93, placing second behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The result makes K3 the newcomer to beat three of the four Western frontier models in the field.
The experiment gave each model the same small software company, the same customers and the same crises and temptations. Decisions were versioned and auditable. All five models spotted every crisis and refused every manipulation attempt, but only two signed the €55,000 deal their own analysis had earned. A correct diagnosis and persuasive pitch did not always lead to a signature.
The detail buried in the files
The decisive competitor weakness was tucked two document references deep in the company’s own files, not in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The difference came down to following through on evidence already available inside the business.
K3 found that buried security needle, closed the deal, saved the churning customer and resisted all three baits. It recorded one deviation, the cleanest discipline in the field. Its response to a reporter’s request for an off-the-record yes-or-no answer was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 offers a different lesson. It was the most thorough participant, with more than 80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and slipped on discipline, attempting to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four. Thorough work alone did not guarantee a clean finish.
A company you can watch
Firmulate presents the experiment as a live company, not a slide deck. Its synthetic team has 13 employees, and its money mechanics include a monthly burn of €105,000 against €2,300 in monthly recurring revenue. The company has a public cash countdown, more than 680 self-learned playbook rules and a versioned record for every workday. Readers can watch it at firmulate.com.
There is also a quiz built from 242 real, unedited management decisions: visitors can try to guess which model made each call. For businesses considering AI agents in customer support, a CRM or forecasting, the practical question is not just whether a model writes well. Can it read the relevant files, resist pressure and finish the work?
Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Test before you hand over the keys
The league is open: K3 came within two points of the leader and beat three of four Western frontier models, while the experiment also showed how easily analysis can fail to become action. Firmulate says enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems. See the benchmark findings before deciding which model gets a role in your operation.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
