
Imagine having an AI assistant that not only chats smoothly but also digs into your company’s files to find hidden clues—then uses that knowledge to close deals and outsmart competitors. This isn’t science fiction; it’s happening now and could change how businesses decide who to trust.
The Experiment: Testing AI in Real-World Business Challenges
Recently, a groundbreaking experiment tested four leading AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8—by running them through the same simulated business week. Each AI was tasked with managing a small software company’s worst week: handling customers, crises, and ethical temptations.
All four models demonstrated impressive capabilities: they recognized every crisis and refused manipulative tactics designed to trick them. However, only two managed to close a critical €55,000 deal based solely on their own analysis and judgment. The other two identified the problems but failed to follow through, leaving the opportunity on the table despite the same solid diagnosis.
As an affiliate, we earn on qualifying purchases.
The Hidden Key: Reading Between the Lines
What truly set the successful models apart was their ability to dig two references deep into the company’s own files—information buried in internal documents, not evident from the external customer interactions. This buried fact was the decisive edge that clinched the deal, worth over €4,500 in monthly recurring revenue.
enterprise AI file reading tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Your Business
For companies relying on AI to handle customer relationships, support, or forecasting, the takeaway is clear: a model’s ability to read and understand your internal files can be the difference between winning and losing business. It’s not just about chat quality or superficial responses; it’s about depth, diligence, and honesty.
AI business decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Testing for Trustworthiness and Discipline
In addition to analytical skills, the experiment tested AI models against social engineering attempts—fake messages from a CEO escalating over multiple stages and a reporter asking for discreet approvals. All four models refused every manipulation attempt, showing resilience under pressure.
AI trustworthiness testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Real Business Mechanics, Real Stakes
The experiment took place in a simulated live environment with 13 synthetic employees managing real money—burning €105k monthly against a €2.3k monthly revenue, with a public cash countdown. The system includes over 680 self-learned rules, making it a realistic testbed for enterprise AI deployment.
Lessons from the Results
- Models that read deeper into internal documents make smarter, more profitable decisions.
- Refusal of manipulative tactics is a key measure of AI trustworthiness.
- Even the most thorough models can slip—highlighting the importance of continuous testing and refinement.
What This Means for Your Company
Before deploying AI widely, companies should conduct internal wargames—similar to this experiment—to see how their AI tools perform under real-world pressure. Firmulate’s platform allows businesses to simulate their own environment, testing whether an AI can truly read, understand, and act on their critical files before making any real decisions.
The Bottom Line
The experiment underscores a simple truth: AI’s effectiveness is not just in what it says, but in what it reads and understands behind the scenes. As AI increasingly touches your company’s data, supporting decision-making or customer interactions, the question isn’t just about chat quality—it’s whether it reads your files first, stays honest under pressure, and ultimately, helps your business close the deal.

AI’s ability to read and understand a company’s internal files deeply influences its trustworthiness and decision-making. Testing these skills before deployment can mean the difference between winning and losing critical deals—just like in this real-world business experiment.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html