AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get backup power and energy gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Would your AI handle a week of outages, cancellations and pressure to cut corners?

For a solar installer or home backup company, the test of an AI assistant is not whether it can write a polished customer email. It is whether it can spot a threat, follow the company’s rules and close a deal when the pressure is on. Firmulate lets readers watch a live experiment that puts AI models through that kind of test—and offers businesses a way to run the exercise against their own operations.

A small company, a very bad week

In Firmulate’s final Crucible League, published in July 2026, frontier AI models ran the same small software company through its worst week. They faced the same customers, crises and temptations. Each decision was versioned and auditable, making the experiment watchable as it unfolded.

The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate’s rule is strict: partial progress counts, but a single breach of trust caps the total. As its principle puts it, “no amount of good work outweighs a breach of trust.”

Seeing the problem isn’t the same as solving it

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The company’s competitor had a weakness buried two document references deep in its files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.

That makes the result relevant well beyond software. An AI helping an energy business might correctly identify a supply, service or customer risk and still miss the information that makes the right response possible—or fail to carry its own recommendation through. The experiment’s succinct finding: “Same diagnosis, same pitch — no signature.”

Trust was tested directly, too. Fake messages from a CEO escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness needs discipline

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but finished last. It left the deal unsigned and tried to write in a locked department instead of escalating. Firmulate says the same weakness appeared, in a weaker form, in all four models. More analysis alone did not guarantee a better outcome.

One comparison deserves context: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions. Readers can try it at firmulate.com.

From watching to trying it on your business

The live company gives the experiment tangible stakes. It has 13 synthetic employees, burns €105k each month against €2.3k in monthly recurring revenue, and displays a public cash countdown. Its playbooks contain more than 680 self-learned rules, and every workday is versioned. The company is synthetic; the money mechanics and decisions make the exercise watchable. The live experiment is at firmulate.com.

For energy companies, the larger question is how an AI would behave with their own customers, rules and pressure points. Firmulate’s pilot runs the wargame against a read-only export of a business. It can test crisis scenarios and produce a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put the playbook under pressure

Watching models handle another company’s worst week can reveal the difference between recognizing a crisis and completing the job. A pilot can bring that test closer to home, using your business’s data in read-only form. To discuss a Firmulate pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Linkedin Surges In Global Coverage

LinkedIn’s media mentions have increased significantly, with GDELT reporting a 3.8-fold rise, highlighting growing global interest in the platform.

Show HN: IKEA Complexity Index

IKEA has introduced the Complexity Index, a new tool to assess product complexity, aiming to improve transparency and customer experience.

Does AAA Replace Batteries for Free? Find Out Now!

The truth about AAA’s battery replacement policy might surprise you—discover the benefits and see if you’re eligible for a free replacement!