
Get workout gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Before an AI runs the business, put it through a hard week
A home gym earns its keep when life gets busy: you can still train when the commute, weather or packed schedule gets in the way. But owning the equipment is only part of the equation. A fitness business also has to handle cancellations, competitors, customer questions and pressure to make the wrong shortcut look convenient. Firmulate has built a live experiment around that kind of pressure—not to test a workout plan, but to see how AI models manage a company when decisions carry consequences.
One company, one difficult week
In the experiment, each frontier model ran the same small software company through its worst week, facing the same customers, crises and temptations. Decisions were versioned and auditable. The point was to see what happened across a working business, rather than judge a model by how polished its chat responses sounded.
The final Crucible League, published in July 2026, put gpt-5.6-sol first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The experiment’s trust standard was uncompromising: “no amount of good work outweighs a breach of trust.”
Recognizing the problem wasn’t enough
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The finding was strikingly simple: “Same diagnosis, same pitch — no signature.” Knowing what to do and carrying the decision through were different things.
The deal turned on a detail buried two document references deep in the company’s own files, not in the customer event. Models that read the file won it at full price, worth +€4,583 MRR. That is the sort of gap a confident summary can hide: the answer may depend on whether an AI follows evidence through the company’s records and completes the next step.
Trust under pressure, and discipline at work
The social-engineering test escalated through fake CEO messages over three stages, then tried a reporter’s request: “just one yes/no, on background”. All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 offered a different lesson. It was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the close on the table and slipped on discipline, making write attempts into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. There is also a fairness detail in the comparison: K3 ran without an effort parameter, using the API default, while the others ran at xhigh.
From watching to a company-specific pilot
The live company makes the scenario tangible. It has 13 synthetic employees and real money mechanics: burn of €105k a month against €2.3k MRR, alongside a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. Readers can watch the live experiment at firmulate.com. A quiz built from 242 real, unedited management decisions lets visitors guess which model made each call.
For an enterprise, the next step is to try the wargame against its own business. Firmulate says a pilot can start from a read-only export, then run crisis scenarios against that company and produce a board report showing model rankings and weak points in its playbooks. Nothing writes back to real systems. That makes the test a way to examine how an AI might handle a company’s customers, pipeline and rules before anyone gives it operational access.
For fitness businesses, the same idea translates naturally: consider how an AI assistant would respond to a membership cancellation wave, a pricing change, a competitor’s offer or an urgent customer complaint. The experiment does not promise that a model will manage those situations well. It gives decision-makers a way to observe what it actually does under pressure—and where a capable answer still fails to become a sound business decision.

Try the test against your own business
Watching an AI handle a simulated company can reveal the difference between spotting a problem and resolving it responsibly. To discuss a Firmulate enterprise pilot using a read-only export of your business, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
