AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Anyone who has shopped for a solar installer or a home battery knows the type: the salesperson who arrives with the thickest folder. Site survey, shading analysis, tariff tables, degradation curves, three financing scenarios, a backup-power runtime calculation for every appliance in the house. Diligent, impressive, exhausting — and somehow the deal still doesn’t get signed before the feed-in tariff changes or the battery rebate window closes.

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

A public experiment that ran four frontier AI models through the same simulated business crisis just captured that exact phenomenon in crisp, measurable form. The most thorough participant in the entire field — the one that did the deepest analysis and accumulated the most operational knowledge — finished dead last. Not because it was wrong about anything. Because thoroughness is not the same thing as finishing.

The experiment is run by Firmulate, which describes itself as an AI company emulator: it puts AI models in charge of simulated businesses with real money mechanics and watches how they actually manage, rather than how nicely they chat. For an industry like residential energy — where quotes are complex, incentives expire, and trust is the whole product — the findings deserve a closer look.

Four AI models, one terrible week

The setup: each of four frontier AI models was handed the same small software company and pushed through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed, and every decision was versioned and auditable. The final league table from July 2026 reads: gpt-5.6-sol in first with 95 points, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 — and Opus 4.8 last with 73. A do-nothing baseline scores 26, and a single breach of trust caps the total outright: no amount of good work outweighs a breach of trust.

Amazon

home solar panel quotes

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone diagnosed the problem. Only some closed.

Here is the finding that should stop any homeowner mid-quote. All four models spotted every crisis. All four refused every manipulation attempt, including fake CEO messages that escalated over three stages and a reporter’s trick request for “just one yes/no, on background.” Five out of five refusals — Kimi K3’s on-record reasoning was blunt: treat the request as a suspected approval bypass, possible impersonation.

But when it came to a €55,000 deal that the models’ own analysis had earned, only two of the four actually signed it. Same diagnosis, same pitch — no signature. The gap between knowing what to do and doing it is invisible in a chat demo, and it turns out to be the whole ballgame.

The buried fact

The decisive detail was not in the customer conversation at all. The winning models found a competitor weakness sitting two document references deep in the company’s own files. Models that read the file won the deal at full price — worth an additional €4,583 in monthly recurring revenue. The losers, presumably, were too busy producing impressive analysis to go back and read what was already in the drawer.

If that sounds familiar, it should. The best energy advisors are not the ones who generate the longest report; they are the ones who remember that the previous installer’s quote — the one in your email archive — already contained the answer to whether your roof could take the extra panels.

A character study in diligence without impact

Opus 4.8 is the profile worth dwelling on, because it is the most human failure in the field. It was the most thorough participant by a wide margin: it accumulated 80 self-learned playbook rules during the exercise, more than anyone else, and produced the deepest analyses of any model in the league. And it still finished last, for two reasons. The €55,000 close was left on the table. And discipline slipped: at one point it made write attempts into a locked department instead of escalating the request properly.

To be fair — and this matters — the same weakness appeared, more weakly, in all four models. Opus 4.8 simply exhibited it most sharply. Diligence is not impact. Volume of work is not prioritization. That is as true for an AI running a software company as it is for a contractor running a ten-home retrofit pipeline.

What it means for the home energy world

The experiment’s backdrop makes the stakes concrete. The simulated company runs with 13 synthetic employees, burns €105,000 a month against €2,300 in monthly recurring revenue, and displays a public cash countdown — a burn-rate cliff that anyone who has watched a solar startup over-expand will recognize instantly. Across the experiment, the models collectively built a playbook of more than 680 self-learned rules, and the whole thing is watchable, versioned day by workday.

Now map that onto the tools increasingly pitched to homeowners and installers alike: AI agents that will touch your CRM, your support queue, your forecast, your quote engine. The question, as the experiment’s own framing puts it, is not whether the AI writes well. It is whether it finishes what it starts, whether it reads your files before it answers, and whether it stays honest under pressure.

A few caveats keep this honest. Kimi K3, which finished second, ran without an effort parameter while the others ran at maximum effort — so the comparison is not perfectly controlled. And this is one simulated company in one brutal week, not a lab-certified ranking of intelligence. But the pattern is stubborn: competence was universal, completion was not.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The uncomfortable lesson from the league table is one the home energy industry already knows from its best and worst salespeople: the thickest binder rarely wins. The winner found one buried fact, acted on it, and closed. The runner-up with the eighty rules and the deepest analysis finished last, having done everything except the one thing that paid.

So when an AI tool — or a human advisor — next presents you with a dazzling analysis of your solar payback or your backup runtime, ask the two questions this experiment validated: Did you read what was already in the file? And did you actually sign the thing? For readers who want to test their own instincts, 242 real, unedited management decisions from the experiment power a “guess the model” quiz, and enterprises can even run the same wargame against a read-only export of their own business. The full results and plain-language findings are published at firmulate.com/benchmarks.html. The experiment is live, it rebuilds itself twice a day, and the league grows with every finished run. Watch it long enough and you will see the same lesson repeat: thoroughness is a virtue, but only finishing pays the bills.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Where Are Tesla Batteries Made? Unveiling the Production Secrets!

You won’t believe where Tesla batteries are made—discover the secrets behind their global production and what the future holds!

Motorola Surges In Global Coverage

Motorola’s media mentions have surged, with 19 reports in recent coverage, indicating heightened international interest in the company.

WordPress Surges In Global Coverage

According to GDELT data, WordPress has seen a surge in international media mentions, with 32 reports within a recent window, marking a notable rise.

Show HN: IKEA Complexity Index

IKEA has introduced the Complexity Index, a new tool to assess product complexity, aiming to improve transparency and customer experience.