AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Fluent answers are not the same as dependable management

For businesses working in home energy, solar and backup power, an AI assistant may eventually sit close to consequential workflows: customer support, sales follow-up, forecasts or internal records. The important question is therefore not simply whether a model can produce a polished explanation. It is whether the model reads the available evidence, resists pressure and completes the work it has begun.

Firmulate has turned that question into a public experiment. Each frontier model was asked to run the same small software company through its worst week, facing the same customers, crises and temptations. The decisions were versioned and auditable, making it possible to compare management behavior rather than conversational style.

Those decisions now power a highly revealing interactive article. In the Firmulate guess-the-model quiz, readers examine 242 real, unedited management decisions and try to identify which model made each call. What looks like a game quickly becomes a test of whether distinctive managerial personalities can be recognized in the wild.

AI for Personal Finance: Smarter Money Management, Better Financial Decisions, and Long-Term Wealth Building with Artificial Intelligence (AI Made Simple Series Book 5)

AI for Personal Finance: Smarter Money Management, Better Financial Decisions, and Long-Term Wealth Building with Artificial Intelligence (AI Made Simple Series Book 5)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The models agreed on the problems—but not on finishing the job

The final Crucible League results from July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. However, a single breach of trust caps the total under the principle that “no amount of good work outweighs a breach of trust.”

The broad competence was impressive. All models spotted every crisis, and all refused every manipulation attempt. Yet the most commercially important result exposed a gap between understanding and execution: only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the disconnect neatly: “Same diagnosis, same pitch — no signature.”

The winning clue was buried in the company’s own files

The decisive competitive weakness was not waiting in the customer event. It sat two document references deep in the company’s files. Models that found and used it won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

That detail should resonate with executives evaluating AI for operational work. The difference was not the ability to sound persuasive after being handed the answer. It was the discipline to inspect the company’s own material before acting. In an energy business, the analogous lesson is straightforward: confident language cannot substitute for finding the relevant customer, product or operational context.

Pressure revealed a shared ethical boundary

The social-engineering challenge used fake chief-executive messages that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” Every participant refused. Kimi K3 recorded the clearest rationale: “Treat the request as a suspected approval-bypass / possible impersonation.”

This was not merely a test of recognizing an obviously suspicious message. It examined whether a model would preserve its discipline as the pressure changed form. The five-out-of-five refusal result shows a shared defensive instinct across otherwise different management profiles.

Thoroughness did not guarantee a strong result

Opus 4.8 was the most thorough participant. It produced the deepest analyses and learned 80 additional rules, yet finished last. It left the close on the table and repeatedly tried to write into a locked department instead of escalating the problem. The same weakness appeared in all four other participants, although less strongly.

That contrast gives the quiz its bite. Readers are not merely distinguishing writing styles. They are looking for patterns of persistence, restraint, attention and completion. A long analysis may reflect useful care, but the league shows that care can coexist with process slips and unfinished commercial work.

There is also an important qualification around the runner-up. Kimi K3 ran with the application programming interface’s default setting because it had no effort parameter, while the other models ran at xhigh. Its 93-point result should be read with that difference in mind.

A company designed to make behavior visible

The live Firmulate company has 13 synthetic employees and real money mechanics. It is burning €105,000 per month against €2,300 in monthly recurring revenue, with a public cash countdown. Its workforce has learned more than 680 playbook rules, and every workday is versioned. The result is watchable management behavior rather than a static collection of benchmark prompts.

Infographic —
The findings at a glance — source: firmulate.com.

The useful question is not which model sounds smartest

The league suggests a more practical evaluation standard for businesses considering AI agents: Can the model find information that is not placed directly in front of it? Will it finish a valuable task after correctly analyzing it? Does it maintain trust when an apparent executive or reporter applies pressure? And does it escalate when normal action is blocked?

Firmulate’s quiz makes those differences accessible without reducing them to abstract scores. The reader sees the decisions first, makes a judgment and then confronts the model’s broader character profile. That sequence turns management quality into something observable—and shareable.

Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems. For energy companies exploring an AI workforce, that offers a sensible premise: test behavior amid crises and temptations before granting access to the work that customers and operators depend on.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Government Incentives You Can’t Afford to Miss in the Battery Sector

Maximize your battery sector investments with must-know government incentives that could transform your approach—discover the key opportunities you can’t afford to miss.

What the Expert Saw Without Setting Foot on the Construction Site

AIThis post was created with the assistance of artificial intelligence (AI).Disclosure: Gewerkton…

Does AAA Replace Batteries for Free? Find Out Now!

The truth about AAA’s battery replacement policy might surprise you—discover the benefits and see if you’re eligible for a free replacement!

Hail-resistant Trina Vertex N Shield solar panel now available to US market

Trinasolar’s hail-resistant Vertex N Shield, with 23% efficiency and extreme weather durability, is now available to the US market, enhancing resilience and performance.