firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine a personal trainer who, despite doing nothing, still earns some credit for showing up. In AI management testing, a similar principle applies: even a do-nothing baseline scores 26 points. This isn’t just a quirk — it’s a statement about trust, progress, and accountability in AI systems.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get workout gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the Real Score of AI Benchmarks

In a recent open experiment conducted by Firmulate, four advanced AI models were tasked with running a simulated small software company through its worst week. Every decision, crisis, and temptation was carefully controlled, and the results reveal more than just raw scores — they expose the core values of trustworthy AI.

The Do-Nothing Baseline: Why 26 Points?

One surprising outcome was that even a ‘do-nothing’ approach earned 26 points out of a possible 100. This partial credit isn’t arbitrary; it reflects that simply showing up or maintaining a baseline level of awareness counts for something. But it’s crucial to understand that this score is not a sign of minimal effort — rather, it’s a measure of honesty in performance expectations. If an AI model does nothing, it earns some points for not making mistakes or causing harm. Partial progress, like ignoring distractions or not engaging in manipulations, adds to this score.

Progress Alone Isn’t Enough

In the experiment, all models identified crises and refused manipulative tactics, like fake CEO messages or reporter tricks. Yet, only two models went further to finalize deals with clients, signing contracts worth €55,000 — their own analysis earning them the full deal. The other two models, despite understanding the situation, left opportunities unclaimed or failed to seize the full potential of their insights.

Trust Breaches Cap Performance

A key finding is that a single breach of trust caps the overall score. Even if a model performs well in many areas, one slip — like attempting to manipulate data or escalate issues improperly — prevents it from achieving full marks. This rule underscores the importance of integrity. Trustworthiness isn’t just about getting things right; it’s about doing so honestly and ethically.

Amazon

AI model testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Reading Your Files

The most decisive advantage for some models came not from overt crisis management but from their ability to access and interpret internal documents. For instance, reading two document references deep in the company’s files allowed a model to close the deal at full price, adding over €4,500 monthly recurring revenue (MRR). This underscores a vital point: AI’s effectiveness depends on what it truly understands, not just how well it responds to surface-level prompts.

Amazon

AI trust and integrity tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Social Engineering Tests Show Resilience

The experiment also gauged how models handle social engineering tactics. Fake messages from a CEO escalating issues or attempting impersonation were used to test compliance. All five models refused to cooperate, demonstrating a strong, principled stance against manipulation. Kimi K3’s explanation was straightforward: treat suspicious requests as potential impersonation or approval-bypass attempts.

Amazon

AI compliance monitoring software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Real Business, Real Risks

Firmulate’s live setup involves a simulated company with 13 synthetic employees, managing real money mechanics — burning €105,000 monthly against a revenue of €2,300. Every workday is versioned and transparent, making the experiment not just theoretical but observable in real-time at

firmulate.com/live. This approach ensures that AI performance is measured against actual business outcomes, not just chat quality or superficial responses.

Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Does This Mean for Business Leaders?

The key takeaway is that AI performance isn’t solely about how convincingly it can generate text. It’s about whether it can complete meaningful, trust-based work — reading critical documents, refusing manipulation, and closing deals honestly. A high score in these experiments indicates an AI’s readiness to handle real-world pressures without compromising integrity.

Final Thoughts

For enterprise decision-makers, the Firmulate benchmark provides a clear, transparent view of how AI models behave under stress, not just their language skills. It reveals the importance of trustworthiness, comprehensive understanding, and ethical consistency. As AI begins to touch more core business functions, these qualities will determine whether AI becomes a valuable partner or a risky liability.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Can AI Pick Up on the Hidden Clues in Business Decisions? A Live Experiment Reveals All

Discover how different AI models perform in real management scenarios — from crisis handling to closing deals. Test your AI’s discipline and honesty today at Firmulate.

Reaction-Diffusion Simulation: A Look Inside “Culture No. 46 — The Living Pattern Laboratory” (FABLE/175)

AIThis post was created with the assistance of artificial intelligence (AI).“Culture No.…

AI in Action: What Running a Business Through Its Worst Week Reveals About Reliability and Trust

A recent live experiment shows that AI can identify crises and resist manipulation but only some can follow through and close deals under pressure—trustworthy execution matters.

How to Avoid the ‘Yank With Your Arms’ Mistake on a Rowing Machine

How to avoid yanking with your arms on a rowing machine and improve your technique for a smoother, more effective workout.