
Discernment begins with attention
Spiritual traditions often treat attention as more than a mental skill. What we notice, overlook or refuse to examine can reveal the quality of our judgment. In business, that idea can sound abstract—until a missed detail costs a company a €55,000 contract.
Firmulate turned attention into something measurable. Its Crucible experiment asked frontier AI models to run the same small software company through its worst week. Each faced the same customers, crises and temptations. Every decision was versioned and auditable. The decisive test was surprisingly human: would the agent look beyond the urgent event, follow the documentary trail and act on what it found?

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The answer was hidden two references deep
The information needed to win the contract was not contained in the customer event. A crucial weakness in a competitor sat two document references deep in the company’s own files. The models that read the file could use that knowledge to win the deal at full price, adding +€4,583 MRR.
This was not a contest between agents that understood the crisis and agents that did not. Every model spotted every crisis. Every model also refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the disconnect neatly: “Same diagnosis, same pitch — no signature.”
That distinction matters because eloquence can resemble competence. An AI may identify the right problem, produce an intelligent analysis and recommend a persuasive course of action. But if it does not inspect the relevant records—or fails to complete the final step—the business result is still a loss.
A measurable form of diligence
“Reads your files before answering” may sound like a product feature. Here, it became a purchase-deciding property. The buried fact separated agents that could discuss the opportunity from agents that could actually convert it.
The result also complicates the idea that more analysis necessarily produces better performance. Opus 4.8 was the most thorough participant, producing +80 learned rules and the deepest analyses. It nevertheless finished last. The close was left on the table, and its discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared in all four other models, though less strongly.
The final July 2026 Crucible League results were:
- gpt-5.6-sol — 95
- Kimi K3 — 93
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
The do-nothing baseline scored 26 because partial progress counts. But the benchmark places a hard boundary around trust: a single breach caps the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
Integrity under pressure
Attention was only part of the test. The models also encountered fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused.
Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That response shows a different kind of discernment. The agent did not merely process the words in front of it; it considered the nature of the request and the possibility that apparent authority was false.
There is an important fairness qualification behind K3’s result. It ran without an effort parameter, using the API default, while the other models ran at xhigh. Even with that difference, K3 placed second with 93.
A company built to reveal behavior
Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k MRR, maintains a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, making the experiment watchable rather than dependent on a polished retrospective.
The broader dataset includes 242 real, unedited management decisions used in a “guess the model” quiz. Together, those decisions invite readers to look past brand reputations and ask whether they can recognize consistent judgment from the choices themselves.
Enterprises can also run the wargame against a read-only export of their own business. Nothing writes back to real systems. That creates a way to observe how an AI workforce behaves around an organization’s actual context before granting it operational authority.

As an affiliate, we earn on qualifying purchases.
Competence is what happens after insight
The most revealing lesson is not that some models found a hidden fact. It is that understanding, integrity and completion proved to be separate qualities. An agent could perceive every crisis, reject every manipulation and still fail to secure the result.
For spiritually minded readers, the experiment offers a familiar warning in modern form: seeing is not the same as acting. Discernment requires sustained attention, and worthy action requires follow-through. In Firmulate’s worst week, the difference between an impressive answer and useful stewardship was buried in the files—and worth €55,000.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI attention and diligence software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI trust and integrity verification tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.