
Imagine pushing through every set in your workout, hitting every rep with relentless discipline—but still walking away without reaching your fitness goal. In the world of AI-driven business decisions, the same paradox applies: relentless diligence isn’t enough. It’s not just about working hard; it’s about working smart, knowing which moves matter most. That’s the core lesson from a groundbreaking public experiment where AI models were tested like athletes in a high-stakes competition, revealing surprising truths about performance and impact.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Experiment: Putting AI Through Its Paces
In a real-world, transparent test, four leading AI models were tasked with managing a small software company through its worst week—crises, customer demands, temptations to cut corners. Every decision was tracked, versioned, and auditable, creating a detailed record of how each AI responded under pressure. The company faced typical challenges, from customer dissatisfaction to internal communications, all designed to simulate real business stress. The goal: see which AI could best navigate the chaos and close a crucial €55,000 deal based on their analysis.
As an affiliate, we earn on qualifying purchases.
The Results: Diligence Meets Its Limits
All four models demonstrated strong awareness: they identified every crisis and refused manipulation attempts, including fake CEO messages and reporter tricks. Despite this, only half of them managed to close the deal. The top performers—gpt-5.6-sol and Kimi K3—secured the agreement, with scores of 95 and 93 out of 100 respectively. The others, Sonnet 5 and Fable 5, fell short, scoring 88 and 77. Interestingly, the experiment revealed that the key weakness was not in recognizing problems but in execution—specifically, the failure to follow through and escalate critical issues appropriately.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: The Power of Context
Digging deeper, the decisive advantage was found not in surface-level analysis but in reading and understanding the company’s internal documents. The models that accessed two document references deep in the company files uncovered critical information that enabled closing the deal at full price—adding over €4,583 MRR. Conversely, models that did not review these files missed the vital insight, losing the opportunity. This underscores a vital point: diligence alone, even with extensive learned rules, is insufficient without effective prioritization and context comprehension.
As an affiliate, we earn on qualifying purchases.
Beyond the Test: Real Business, Real Money
The experiment’s setting was a live, functioning company with 13 synthetic employees and actual financial mechanics—burning €105k monthly against a mere €2.3k MRR, with a public cash countdown. Every workday, decisions are versioned, and rules are learned and refined. The models are not just chatbots but active, decision-making companies, tested publicly at firmulate.com/live. This transparency allows anyone to see how AI manages real crises, not just simulated conversations.
AI analysis software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Takeaways for Business Leaders and Fitness Enthusiasts Alike
The lessons from this AI experiment resonate beyond the tech sphere. For fitness fans, it’s akin to pushing through every crunch and squat but forgetting the importance of strategic rest or targeted exercises. Discipline is vital, but without prioritization—knowing which exercises target your goals—you won’t see the gains you want. Similarly, in AI-driven business processes, relentless work must be complemented by smart decision-making and context-aware focus.
What Truly Matters in AI Performance
Analysis from the experiment shows that the most thorough AI—Opus 4.8—despite its deep rule set and comprehensive approach, still finished last because it lacked disciplined escalation and missed critical insights buried in internal files. Meanwhile, models that prioritized reading relevant documents and validated their actions performed better in closing deals. This highlights a universal truth: volume of effort is secondary to focus and strategic prioritization.
Final Reflection: Diligence Is Not Enough
In sport, as in business and AI, the greatest success comes from knowing what to do—and doing it at the right moment. Overextending without clear priorities can lead to missed opportunities, even when effort is relentless. For those deploying AI in critical decision-making, the message is clear: train your models not just to work hard but to work smart, focusing on insights that matter, reading deeply, and escalating when needed.
Interested in seeing how your enterprise can prepare its AI workforce? Try the free pilot program and run your own wargame—without risking real systems or data. Because in the end, performance isn’t just about effort; it’s about impact.

In AI and fitness alike, relentless effort must be paired with strategic focus. Diligence alone doesn’t win; prioritization and context are key to closing the deal and reaching goals.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.