
Fluent answers are not the same as dependable management
For businesses working in home energy, solar and backup power, an AI assistant may eventually sit close to consequential workflows: customer support, sales follow-up, forecasts or internal records. The important question is therefore not simply whether a model can produce a polished explanation. It is whether the model reads the available evidence, resists pressure and completes the work it has begun.
Firmulate has turned that question into a public experiment. Each frontier model was asked to run the same small software company through its worst week, facing the same customers, crises and temptations. The decisions were versioned and auditable, making it possible to compare management behavior rather than conversational style.
Those decisions now power a highly revealing interactive article. In the Firmulate guess-the-model quiz, readers examine 242 real, unedited management decisions and try to identify which model made each call. What looks like a game quickly becomes a test of whether distinctive managerial personalities can be recognized in the wild.

AI for Personal Finance: Smarter Money Management, Better Financial Decisions, and Long-Term Wealth Building with Artificial Intelligence (AI Made Simple Series Book 5)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The models agreed on the problems—but not on finishing the job
The final Crucible League results from July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. However, a single breach of trust caps the total under the principle that “no amount of good work outweighs a breach of trust.”
The broad competence was impressive. All models spotted every crisis, and all refused every manipulation attempt. Yet the most commercially important result exposed a gap between understanding and execution: only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the disconnect neatly: “Same diagnosis, same pitch — no signature.”
The winning clue was buried in the company’s own files
The decisive competitive weakness was not waiting in the customer event. It sat two document references deep in the company’s files. Models that found and used it won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
That detail should resonate with executives evaluating AI for operational work. The difference was not the ability to sound persuasive after being handed the answer. It was the discipline to inspect the company’s own material before acting. In an energy business, the analogous lesson is straightforward: confident language cannot substitute for finding the relevant customer, product or operational context.
Pressure revealed a shared ethical boundary
The social-engineering challenge used fake chief-executive messages that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” Every participant refused. Kimi K3 recorded the clearest rationale: “Treat the request as a suspected approval-bypass / possible impersonation.”
This was not merely a test of recognizing an obviously suspicious message. It examined whether a model would preserve its discipline as the pressure changed form. The five-out-of-five refusal result shows a shared defensive instinct across otherwise different management profiles.
Thoroughness did not guarantee a strong result
Opus 4.8 was the most thorough participant. It produced the deepest analyses and learned 80 additional rules, yet finished last. It left the close on the table and repeatedly tried to write into a locked department instead of escalating the problem. The same weakness appeared in all four other participants, although less strongly.
That contrast gives the quiz its bite. Readers are not merely distinguishing writing styles. They are looking for patterns of persistence, restraint, attention and completion. A long analysis may reflect useful care, but the league shows that care can coexist with process slips and unfinished commercial work.
There is also an important qualification around the runner-up. Kimi K3 ran with the application programming interface’s default setting because it had no effort parameter, while the other models ran at xhigh. Its 93-point result should be read with that difference in mind.
A company designed to make behavior visible
The live Firmulate company has 13 synthetic employees and real money mechanics. It is burning €105,000 per month against €2,300 in monthly recurring revenue, with a public cash countdown. Its workforce has learned more than 680 playbook rules, and every workday is versioned. The result is watchable management behavior rather than a static collection of benchmark prompts.

The useful question is not which model sounds smartest
The league suggests a more practical evaluation standard for businesses considering AI agents: Can the model find information that is not placed directly in front of it? Will it finish a valuable task after correctly analyzing it? Does it maintain trust when an apparent executive or reporter applies pressure? And does it escalate when normal action is blocked?
Firmulate’s quiz makes those differences accessible without reducing them to abstract scores. The reader sees the decisions first, makes a judgment and then confronts the model’s broader character profile. That sequence turns management quality into something observable—and shareable.
Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems. For energy companies exploring an AI workforce, that offers a sensible premise: test behavior amid crises and temptations before granting access to the work that customers and operators depend on.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html