AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Firmulate has launched a live benchmark testing AI models in a simulated company crisis environment. The experiment shows that management skills, not just chat quality, are essential for effective AI deployment in business. Results highlight the importance of decision-making, trust, and accountability in AI management.

Firmulate has launched a live, real-time benchmark that evaluates AI models based on their ability to manage a simulated company’s crises during its worst week. Unlike traditional benchmarks focused on technical or conversational performance, this experiment measures management quality, decision-making, and trustworthiness in a high-stakes environment. The results underscore that effective management by AI models is a critical capability for business applications.

The experiment involved five AI models competing in a simulated scenario where they had to handle customer crises, negotiate deals, and make strategic decisions under pressure. The models were scored on their ability to identify issues, communicate effectively, and maintain trust, with a strict rule: any breach of trust would cap their score. The top performer, gpt-5.6-sol, achieved a score of 95, while others lagged behind, with Opus 4.8 scoring 73. The test revealed that while all models could detect crises and resist manipulation, only two managed to close a significant deal, highlighting the gap between diagnosis and execution in AI management.

One notable finding was that models often failed to retrieve critical information from internal documents, which could have altered business outcomes. For example, a model that read the company’s files correctly secured a deal worth an additional €4,583 MRR, but most models missed this crucial detail. This demonstrates that surface-level responses can be misleading, and true management effectiveness depends on deep organizational awareness and accurate information retrieval.

Furthermore, the experiment tested AI resistance to social engineering attacks, such as impersonation and approval bypasses. All models refused to comply with suspicious requests, with Kimi K3 explicitly recognizing and documenting the risks. However, even models that maintained boundaries struggled with completing managerial tasks, such as escalating issues or finalizing decisions, exposing a significant execution gap. The experiment’s design intentionally emphasized trust and accountability, making it clear that safety and productivity are intertwined but distinct qualities in AI management.

At a glance
reportWhen: ongoing, with final July 2026 results p…
The developmentFirmulate’s live experiment tests AI models managing a simulated company’s worst week, revealing management effectiveness as a key measure.

Why Management Skills Are the Next Benchmark in AI

This experiment reveals that evaluating AI models solely on chat quality or technical prowess is insufficient for real-world business use. Management involves complex decision-making, prioritization, trust, and accountability, which are not captured by traditional benchmarks. As AI begins to take on roles that influence organizational outcomes, measuring management effectiveness becomes essential. The findings suggest that future AI evaluation should focus on how well models manage consequences, handle organizational context, and preserve trust over time, rather than just generating impressive responses.

Amazon

AI decision-making management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limitations of Traditional AI Benchmarks in Business Settings

Historically, AI evaluation has centered on benchmarks like coding competitions, chat arena ratings, or accuracy metrics, which primarily assess technical or conversational performance. These benchmarks fail to capture the complexities of managing real-world crises, making them inadequate for enterprise deployment. The rise of AI agents in business operations—such as customer support, decision support, and automation—necessitates a new evaluation paradigm. Firmulate’s live experiment bridges this gap by simulating an ongoing, high-pressure environment where models must demonstrate management skills, decision-making under uncertainty, and trustworthiness in a context that mirrors actual business risks.

Prior efforts have focused on static tests or isolated tasks, but these do not reflect the dynamic, consequence-driven nature of organizational management. The July 2026 Crucible League results highlight that even the most thorough models can falter when asked to manage real consequences, emphasizing the need for a new, performance-based standard rooted in operational effectiveness.

“The next leap in AI evaluation will come from watching models manage consequences and whether they can complete the job without compromising the organization.”

— Thorsten Meyer, founder of Firmulate

Amazon

AI crisis management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Aspects of AI Management Are Still Being Tested?

While the experiment demonstrates that models can resist manipulation and identify crises, it remains unclear how well these models perform over longer periods or in more complex organizational environments. The current tests focus on a single simulated week with specific scenarios; real-world management involves ongoing adaptation, learning, and handling unforeseen issues. Additionally, the impact of different organizational structures and cultures on AI performance is still unknown. Further research is needed to validate whether these findings generalize across diverse business contexts and whether models can sustain high management quality over time.

Amazon

AI organizational information retrieval tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Evaluation and Adoption

Going forward, firms and AI developers are likely to adopt management-focused benchmarks as standard evaluation tools. The immediate next step is integrating these tests into AI deployment workflows, especially for roles involving decision-making, escalation, and trust management. Companies considering AI assistants should ask whether models can read organizational data, prioritize correctly, escalate when necessary, and remain honest under pressure. Additionally, ongoing research will refine these benchmarks, exploring longer-term management capabilities, multi-week simulations, and diverse organizational scenarios. The ultimate goal is to develop AI models that can reliably manage organizational consequences as part of their core competencies.

Amazon

AI trust and accountability software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why is management quality more important than chat performance for AI in business?

Management quality reflects an AI’s ability to handle complex decisions, prioritize tasks, maintain trust, and deliver results in real-world scenarios. These skills are critical for operational success, whereas chat performance mainly measures superficial language abilities.

What does the experiment reveal about AI’s ability to resist manipulation?

All models successfully refused social engineering attempts like impersonation and approval bypasses, indicating strong resistance to manipulation and a focus on safety and boundaries.

How can companies use this new benchmark in practice?

Firms can run similar live tests tailored to their organizational scenarios to evaluate whether AI models can manage real consequences, read organizational data effectively, and uphold trust over time before deploying them in critical roles.

Are current models ready to manage real companies?

While promising, current models still face gaps in execution and long-term consistency. The experiment highlights progress but also underscores the need for further development before full deployment in complex, high-stakes environments.

What is the significance of the €55,000 deal in the experiment?

The deal’s significance lies in testing whether models can retrieve key facts from internal documents and effectively close a business transaction, demonstrating real-world management effectiveness beyond superficial responses.

Source: ThorstenMeyerAI.com

You May Also Like

Incident postmortem builder for managed service providers

A new incident postmortem builder for small managed service providers is being tested to improve post-incident communication and efficiency.

How To Implement Human-Review Trackers In AI Agency Operations

A new workflow tool for AI-assisted agencies helps track human and AI task ownership, improving oversight and quality control in client projects.

One Model, a Whole Portfolio: What Ten Days on Fable Mean for a Business Building on Frontier AI

A detailed report on how one model, Fable 5, managed to run a comprehensive business portfolio in ten days, highlighting operational, economic, and security insights.

Thrymvault: A System Around Your Content

Thrymvault launches as a self-hosted workspace integrating content creation, management, AI prompts, and client collaboration in one platform.