firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get workout gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Before an AI runs the business, put it through a hard week

A home gym earns its keep when life gets busy: you can still train when the commute, weather or packed schedule gets in the way. But owning the equipment is only part of the equation. A fitness business also has to handle cancellations, competitors, customer questions and pressure to make the wrong shortcut look convenient. Firmulate has built a live experiment around that kind of pressure—not to test a workout plan, but to see how AI models manage a company when decisions carry consequences.

One company, one difficult week

In the experiment, each frontier model ran the same small software company through its worst week, facing the same customers, crises and temptations. Decisions were versioned and auditable. The point was to see what happened across a working business, rather than judge a model by how polished its chat responses sounded.

The final Crucible League, published in July 2026, put gpt-5.6-sol first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The experiment’s trust standard was uncompromising: “no amount of good work outweighs a breach of trust.”

Recognizing the problem wasn’t enough

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The finding was strikingly simple: “Same diagnosis, same pitch — no signature.” Knowing what to do and carrying the decision through were different things.

The deal turned on a detail buried two document references deep in the company’s own files, not in the customer event. Models that read the file won it at full price, worth +€4,583 MRR. That is the sort of gap a confident summary can hide: the answer may depend on whether an AI follows evidence through the company’s records and completes the next step.

Trust under pressure, and discipline at work

The social-engineering test escalated through fake CEO messages over three stages, then tried a reporter’s request: “just one yes/no, on background”. All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 offered a different lesson. It was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the close on the table and slipped on discipline, making write attempts into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. There is also a fairness detail in the comparison: K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

From watching to a company-specific pilot

The live company makes the scenario tangible. It has 13 synthetic employees and real money mechanics: burn of €105k a month against €2.3k MRR, alongside a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. Readers can watch the live experiment at firmulate.com. A quiz built from 242 real, unedited management decisions lets visitors guess which model made each call.

For an enterprise, the next step is to try the wargame against its own business. Firmulate says a pilot can start from a read-only export, then run crisis scenarios against that company and produce a board report showing model rankings and weak points in its playbooks. Nothing writes back to real systems. That makes the test a way to examine how an AI might handle a company’s customers, pipeline and rules before anyone gives it operational access.

For fitness businesses, the same idea translates naturally: consider how an AI assistant would respond to a membership cancellation wave, a pricing change, a competitor’s offer or an urgent customer complaint. The experiment does not promise that a model will manage those situations well. It gives decision-makers a way to observe what it actually does under pressure—and where a capable answer still fails to become a sound business decision.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Try the test against your own business

Watching an AI handle a simulated company can reveal the difference between spotting a problem and resolving it responsibly. To discuss a Firmulate enterprise pilot using a read-only export of your business, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Avoid the ‘Yank With Your Arms’ Mistake on a Rowing Machine

How to avoid yanking with your arms on a rowing machine and improve your technique for a smoother, more effective workout.

The Damper Setting Confusion That Trips Up New Rowers

Lifting the damper too high can hinder your progress; discover how proper adjustments can transform your rowing experience.