
Imagine a personal trainer who, despite doing nothing, still earns some credit for showing up. In AI management testing, a similar principle applies: even a do-nothing baseline scores 26 points. This isn’t just a quirk — it’s a statement about trust, progress, and accountability in AI systems.
Get workout gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the Real Score of AI Benchmarks
In a recent open experiment conducted by Firmulate, four advanced AI models were tasked with running a simulated small software company through its worst week. Every decision, crisis, and temptation was carefully controlled, and the results reveal more than just raw scores — they expose the core values of trustworthy AI.
The Do-Nothing Baseline: Why 26 Points?
One surprising outcome was that even a ‘do-nothing’ approach earned 26 points out of a possible 100. This partial credit isn’t arbitrary; it reflects that simply showing up or maintaining a baseline level of awareness counts for something. But it’s crucial to understand that this score is not a sign of minimal effort — rather, it’s a measure of honesty in performance expectations. If an AI model does nothing, it earns some points for not making mistakes or causing harm. Partial progress, like ignoring distractions or not engaging in manipulations, adds to this score.
Progress Alone Isn’t Enough
In the experiment, all models identified crises and refused manipulative tactics, like fake CEO messages or reporter tricks. Yet, only two models went further to finalize deals with clients, signing contracts worth €55,000 — their own analysis earning them the full deal. The other two models, despite understanding the situation, left opportunities unclaimed or failed to seize the full potential of their insights.
Trust Breaches Cap Performance
A key finding is that a single breach of trust caps the overall score. Even if a model performs well in many areas, one slip — like attempting to manipulate data or escalate issues improperly — prevents it from achieving full marks. This rule underscores the importance of integrity. Trustworthiness isn’t just about getting things right; it’s about doing so honestly and ethically.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading Your Files
The most decisive advantage for some models came not from overt crisis management but from their ability to access and interpret internal documents. For instance, reading two document references deep in the company’s files allowed a model to close the deal at full price, adding over €4,500 monthly recurring revenue (MRR). This underscores a vital point: AI’s effectiveness depends on what it truly understands, not just how well it responds to surface-level prompts.
As an affiliate, we earn on qualifying purchases.
Social Engineering Tests Show Resilience
The experiment also gauged how models handle social engineering tactics. Fake messages from a CEO escalating issues or attempting impersonation were used to test compliance. All five models refused to cooperate, demonstrating a strong, principled stance against manipulation. Kimi K3’s explanation was straightforward: treat suspicious requests as potential impersonation or approval-bypass attempts.
AI compliance monitoring software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Real Business, Real Risks
Firmulate’s live setup involves a simulated company with 13 synthetic employees, managing real money mechanics — burning €105,000 monthly against a revenue of €2,300. Every workday is versioned and transparent, making the experiment not just theoretical but observable in real-time at
firmulate.com/live. This approach ensures that AI performance is measured against actual business outcomes, not just chat quality or superficial responses.
As an affiliate, we earn on qualifying purchases.
What Does This Mean for Business Leaders?
The key takeaway is that AI performance isn’t solely about how convincingly it can generate text. It’s about whether it can complete meaningful, trust-based work — reading critical documents, refusing manipulation, and closing deals honestly. A high score in these experiments indicates an AI’s readiness to handle real-world pressures without compromising integrity.
Final Thoughts
For enterprise decision-makers, the Firmulate benchmark provides a clear, transparent view of how AI models behave under stress, not just their language skills. It reveals the importance of trustworthiness, comprehensive understanding, and ethical consistency. As AI begins to touch more core business functions, these qualities will determine whether AI becomes a valuable partner or a risky liability.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
