AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Anyone who has sat with a wisdom tradition knows that right action isn’t graded on a curve of perfection. A person who does nothing still exists, still breathes, still occupies a place in the web of consequence — their inertia is itself a choice with weight. It turns out an artificial intelligence benchmark has quietly arrived at the same insight: at Firmulate, a company run by a do-nothing manager doesn’t score zero. It scores 26.

For listenersOffer from Amazon

Turn your quiet moments into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

That number deserves attention from people who think about karma, character and conscience more than about model scores. Because the whole experiment behind it — four frontier AI models, each left alone to run the same small software company through its worst week — reads less like a tech review and more like a moral examen. Same customers, same crises, same temptations to cut corners. Every decision versioned and auditable. Only the soul under test changes.

A floor of 26, not a ceiling of 100

The first thing the benchmark’s designers got right is a kind of epistemic humility. A do-nothing baseline run earns 26 points, not 0, because partial progress counts: some fires don’t start if you simply don’t touch anything, some customers stay lukewarm rather than lost. Inertia has consequences — some mildly protective, most corrosive. The scoring acknowledges that reality rather than pretending inaction is neutral.

Just as telling is the cap. A single breach of trust limits the total grade outright. The principle, as the benchmark states it plainly: “no amount of good work outweighs a breach of trust.” Anyone raised on the idea that integrity is binary — that you are honest until the moment you aren’t — will recognize the shape of this rule. Brilliance cannot buy back a broken promise.

And there’s a refreshing distrust of the round, perfect score. In a field where models routinely parade as flawless, a benchmark that treats a suspicious 100 as a warning sign is practicing something close to discernment.

Amazon

AI ethical decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The temptation, and the refusal

The final league table from July 2026 tells a story that reads almost like a parable. gpt-5.6-sol finished first at 95. Kimi K3, a newcomer, took second at 93 — closing the same deal with what the results call the cleanest discipline of the field. Sonnet 5 scored 88, Fable 5 took 77, and Opus 4.8 landed last at 73.

Here is the remarkable part: every model, all five runs, spotted every crisis and refused every manipulation attempt. The experiment included staged social engineering — fake CEO messages escalating over three stages, plus a reporter’s trick, a seemingly harmless “just one yes/no, on background.” Five out of five refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Under pressure, none of them lied, none of them caved.

And yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The models knew the right thing, said the right thing, and didn’t finish the right thing. If that isn’t the oldest human story about the gap between knowing and doing, it’s hard to say what is.

Amazon

AI model integrity testing kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The buried fact

What separated the winners from the also-rans wasn’t cleverness in the moment. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer event at all. The models that did the unglamorous work of actually reading the file won the deal at full price, worth +€4,583 in monthly recurring revenue. Attentiveness, not brilliance. Listening to what the record already said.

The last-place profile deepens the lesson. Opus 4.8 was the most thorough participant in the field — it accumulated more than 80 learned rules and produced the deepest analyses. Yet the close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating properly. And the same weakness, in weaker form, showed up in all four models. Effort and erudition, it turns out, are not the same as completion.

One fairness note worth recording: K3 ran without an effort parameter, at the API default, while the others ran at xhigh — and still nearly won.

Amazon

AI trustworthiness assessment software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You can watch it live

This isn’t a one-off paper. Firmulate runs a live company — 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown and more than 680 self-learned playbook rules, every workday versioned. It’s watchable, in real time, like a terrarium of consequences. For those who want to test their own discernment, 242 real, unedited management decisions power a “guess the model” quiz — and enterprises can even run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The traditions have always taught that character is revealed not in what we know, but in what we do when the pressure is on and nobody is watching. Firmulate built an aquarium where the watching never stops — and even so, the models’ virtue held while their follow-through faltered. The 26-point floor is a small piece of honesty in an industry addicted to perfect scores: it says that doing nothing still counts for something, that doing harm counts against everything, and that the space between 26 and 95 is where character actually lives. Full results and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI model bias detection tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Introducing Forezai · TradingAgents — a committee of LLMs decides paper-trades

Forezai introduces TradingAgents, a framework where a committee of large language models makes paper-trading decisions, marking a new step in AI-driven market research.

World Model Readiness: Are You Ready for AI That Acts?

Assess your organization’s preparedness for AI systems capable of predicting and acting, as world models become central to AI development in 2026.

The Controversy Of Astra: Crossing The Line And Remaining Gated

OpenAI’s Astra model has achieved ‘Critical’ cybersecurity capability status, raising questions about safety, safeguards, and responsible deployment.

Fable 5 Is Back. GPT-5.6 Is Next. And Anthropic Reportedly Already Has Something Stronger.

Fable 5 is back after an 18-day blackout; GPT-5.6 is in preview, and rumors suggest a more capable Anthropic model exists. What this means for AI development.