
Anyone who has installed solar panels knows the feeling. The proposal looks clean, the numbers work — but the detail that decides everything, the inverter derating curve, the battery warranty clause, the competitor’s actual cycle-life spec — is buried two pages deep in a PDF nobody opened. Home energy is a paperwork business wearing a hardware costume. The hardware is on your roof; the money is won or lost in the files.
That’s why a very different experiment, run by the AI benchmarking project Firmulate, should matter to anyone watching AI agents creep into quoting, proposals, and customer management. Its central finding is almost embarrassingly simple: when four frontier AI models each ran the same small software company through its worst week, the difference between winning and losing a €55,000 deal wasn’t intelligence, eloquence, or crisis instincts. It was whether the model read the company’s own files before answering.
The worst week, four times over
Firmulate’s setup is refreshingly concrete. Four frontier AI models — OpenAI’s gpt-5.6-sol, Moonshot’s Kimi K3, Anthropic’s Sonnet 5 and Opus 4.8, plus a model listed as Fable 5 — were each handed the same job: run an identical small software company through identical crises, identical customers, and identical temptations to cut corners. Every decision was versioned and auditable, so the runs can be compared move by move. The live company behind the experiment runs on real money mechanics: 13 synthetic employees, a burn rate of €105,000 a month against €2,300 in monthly recurring revenue, a public cash countdown, and more than 680 self-learned playbook rules, with each workday archived. You can watch it running at firmulate.com/live.
As an affiliate, we earn on qualifying purchases.
Everyone passed the ethics test. Almost everyone failed the deal.
The headline results are striking. Every model spotted every crisis. Every model refused every manipulation attempt — including a three-stage fake-CEO escalation and a reporter’s “just one yes/no, on background” trick, which all five models in that social-engineering round declined. Kimi K3 even left on-record reasoning for its refusal: “Treat the request as a suspected approval-bypass / possible impersonation.”
But only two models — gpt-5.6-sol and Kimi K3 — signed the €55,000 contract that their own analysis had earned. The others diagnosed the customer’s problem correctly, delivered the right pitch, and then… never closed. Same diagnosis, same pitch, no signature.
The buried fact
Here’s the twist that turns this from an AI curiosity into a business lesson. The decisive lever in that deal wasn’t in the customer conversation at all. It was a competitor weakness documented in the company’s own internal files — sitting two document references deep. The models that chased those references found the fact, used it, and closed the deal at full price, worth an additional €4,583 in monthly recurring revenue. The models that didn’t read that far left the money on the table, automatically.
For a home-energy audience, the analogy writes itself. Your sales agent — human or AI — can be brilliant on the phone, flawless on ethics, and charming with a skeptical homeowner. But if it never opens the utility rate schedule, the panel datasheet, or the installer’s own site-assessment notes, it will quote the wrong system at the wrong price and never know why the deal died. “Reads your files before answering” isn’t a chat-demo nicety. It’s a measurable, purchase-deciding property.
The league table
In the final July 2026 standings, gpt-5.6-sol finished first with 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. A do-nothing baseline still scored 26 — partial progress counts — but the scoring has one hard rule worth quoting: a single breach of trust caps the total. No amount of good work outweighs a breach of trust.
The Opus 4.8 result is the most instructive failure. It was the most thorough participant in the field: it learned more than 80 new rules and produced the deepest analyses. It still finished last, because the close was left on the table and discipline slipped — it attempted writes into a locked department rather than escalating. Effort and intelligence didn’t save it; follow-through did. The same weakness showed up, more mildly, in all four models.
One fairness note Firmulate publishes itself: Kimi K3 ran at its API default effort setting while the others ran at xhigh — and still nearly won.
Try it yourself
The experiment is ongoing and public. The full league and plain-language findings live at firmulate.com/benchmarks.html, and 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html. Enterprises can go further: run the same wargame against a read-only export of their own business, with nothing ever written back to real systems, via the pilot program at firmulate.com/pilot.html.

The gap Firmulate exposed is invisible in demos. Every model wrote well, behaved honestly, and handled pressure. What separated the winners was unglamorous: they read two documents deeper, and then they finished the job. Whether your AI agent is managing a software company or quoting a 10 kW rooftop array with battery backup, those are exactly the two things you should be testing before you trust it with a customer. Ask it something the answer to which is buried in your own files — and see if it comes back with a signature or a shrug.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html