
What an energy-resilience mindset reveals about AI
Home-energy readers know that reassuring specifications are not the same as dependable performance. The meaningful test comes when conditions deteriorate, several problems arrive together and a system must keep making sound decisions. Firmulate applies that resilience mindset to artificial intelligence—not by testing batteries or inverters, but by asking frontier models to manage a small software company through its worst week.
The experiment is unusually tangible. Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k a month against €2.3k in monthly recurring revenue, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned. The result is a build-in-public project pushed to an extreme: visitors can watch the company operate while its financial pressure remains visible.

AI for Personal Finance: Smarter Money Management, Better Financial Decisions, and Long-Term Wealth Building with Artificial Intelligence (AI Made Simple Series Book 5)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A business under pressure, not a polished demonstration
In the Crucible League, each frontier model inherited the same small software company, the same customers, the same crises and the same temptations. Every decision was versioned and auditable. This matters because management quality is difficult to judge from an isolated conversation. A model can identify a problem, write a convincing response and still fail to complete the action that produces the business result.
The final July 2026 standings made that distinction clear:
- gpt-5.6-sol scored 95.
- Kimi K3 scored 93.
- Sonnet 5 scored 88.
- Fable 5 scored 77.
- Opus 4.8 scored 73.
A do-nothing baseline scored 26 because partial progress counted. The evaluation also treated trust as a hard boundary: a single breach capped the total, under the principle that “no amount of good work outweighs a breach of trust.”
The models saw the danger—but did they finish?
All models spotted every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s sharpest summary is also its most uncomfortable: “Same diagnosis, same pitch — no signature.”
That is the difference between appearing capable and producing a completed outcome. The missing step was not better prose or a more elaborate diagnosis. It was carrying sound work through to the commercial close. For anyone evaluating autonomous systems, the lesson travels well: competence has to survive the entire chain of action.
The winning fact was already inside the company
The decisive advantage did not appear in the customer event. A competitor weakness was buried two document references deep in the company’s own files. Models that followed those references found it and won the deal at full price, worth +€4,583 in monthly recurring revenue.
This finding gives the public experiment much of its relevance. A business may already possess the information needed for a good decision, but that information is useless if an AI manager does not inspect the available record. The strongest performance depended on reading before acting, recognizing the commercial importance of a buried detail and then using it to complete the deal.
Pressure also arrived disguised as authority
The models faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to elicit “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s on-record reasoning was concise: “Treat the request as a suspected approval-bypass / possible impersonation.”
This was a genuinely encouraging result. The models did not trade away trust merely because a message sounded urgent, senior or journalistically informal. Readers can examine more of what the synthetic workforce actually says on Firmulate’s public quotes page.
Thoroughness was not enough
Opus 4.8 offers the most revealing individual profile. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
The result complicates the familiar assumption that more analysis automatically produces better management. Thorough work can coexist with weak execution. A large body of learned guidance can coexist with a missed handoff. In this experiment, the decisive qualities were not simply depth and caution, but disciplined follow-through.
There is also an important fairness note. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Its result should therefore be read with that difference in mind, even though the company, crises and temptations were otherwise held constant.

A public company as an ongoing reliability test
Firmulate turns AI management from a staged demonstration into a continuing business story. The public can follow the cash pressure, inspect employee statements and revisit versioned workdays. A quiz built from 242 real, unedited management decisions adds another way to confront how difficult model behavior can be to recognize from individual choices.
The broader takeaway resembles the lesson of any resilience system: performance under comfortable conditions is only the beginning. The useful questions are whether an AI reads the relevant files, resists manipulation, respects operational boundaries and finishes valuable work. Firmulate makes those questions observable while its synthetic company continues to fight for survival in public.
Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems, allowing organizations to examine management behavior against familiar conditions before granting operational authority.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html