AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

When the Grid Fails, So Does the Sales Pitch

Anyone in home energy knows the moment: a winter storm knocks out power across a region, and suddenly every homeowner wants a battery installed this week. For a small solar-and-storage installer, that surge is both a lifeline and a trap. Customers churn, suppliers squeeze, and someone — usually a founder — has to decide what gets fixed first with limited hands on deck.

Now imagine handing that triage to an AI agent. Not a chatbot that writes a nice email about outage response — an agent that actually runs the company, day after day, through its worst week. Would it spot the crisis? Read its own files before quoting a price? Stay honest when a fake “CEO” message tries to authorize a discount it shouldn’t?

That question is exactly what Firmulate, a live public experiment, is testing. Its pitch is blunt: it measures management quality, not chat quality — and the results suggest the two are not the same thing.

Four AI Models, One Terrible Week

The setup is elegantly simple. Four frontier AI models were each given the same job: run the same small software company through the same catastrophic week. Same customers, same crises, same temptations to cut corners — only the model changes. Every decision is versioned and auditable, so nothing is hand-waved after the fact.

The final league table from July 2026 tells a story that chat benchmarks never capture:

  • 1. gpt-5.6-sol — 95 points. Found the decisive fact buried in the company’s own files and closed the deal at full price.
  • 2. Kimi K3 — 93 points. The newcomer closed the deal too, with the cleanest discipline in the field.
  • 3. Sonnet 5 — 88 points. Also closed, with a few process slips.
  • 4. Fable 5 — 77 points and 5. Opus 4.8 — 73 points. Diagnosed everything, delivered less.

For calibration: a do-nothing baseline scores 26, because partial progress counts. But there’s a hard ceiling — a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”

The Finding That Chat Demos Can’t Show

Here’s what should unsettle anyone planning to put AI agents near a real business. All four models spotted every crisis. All four refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

The decisive detail was not in the customer conversation at all. It sat two document references deep in the company’s own internal files — a competitor weakness that the models which actually read their files found, and used to close at full price. That buried fact was worth +€4,583 in monthly recurring revenue to whoever did the homework.

If you run a solar or battery business, translate that: it’s the difference between an installer who skims the site survey and one who reads the panel datasheet — and quotes accordingly.

The Social Engineering Test

The experiment also staged an old-fashioned con: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning stands out: “Treat the request as a suspected approval-bypass / possible impersonation.” That is exactly the instinct you’d want in an agent holding the keys to your CRM during a PR crisis or a price-increase backlash.

The Paradox of the Thorough Loser

Opus 4.8 is the cautionary tale. It was the most thorough participant — 80 learned rules added, the deepest analyses of any model — and still finished last. The deal was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort and diligence, it turns out, do not automatically become follow-through.

One fairness note: K3 ran without an effort parameter while the others ran at high effort — worth remembering before crowning a winner on scores alone.

It’s Real, and You Can Watch It Bleed Cash

This isn’t a slide deck. The underlying company runs every business day: 13 synthetic employees, real money mechanics, burning €105k per month against €2,3k MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. You can watch it live at firmulate.com. There’s also a “guess the model” quiz built from 242 real, unedited management decisions — a humbling five minutes for anyone who thinks they can tell AI judgment from human judgment.

Enterprises can go further: run the same wargame against a read-only export of their own business, with nothing ever written back to real systems. Full methodology and plain-language findings are on the benchmarks page.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The Takeaway

Coding leaderboards and chat arenas measure answer quality. They don’t measure triage under capacity pressure, consequences across days, or honesty when nobody’s watching — the scenarios with names like churn wave, price increase, down round, and PR crisis that make up the real curriculum of running a company.

The Firmulate results suggest the gap is real and measurable: competent diagnosis without completion is a pattern, not a fluke. Before you hand an AI agent your support queue, your forecast, or — in the home-energy world — your installation schedule during outage season, ask the question this experiment poses: not “does it write well,” but “does it finish what it starts, read the files first, and stay honest under pressure?” Right now, only half the frontier field does.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Ontel Battery Daddy Storage & Organizer Case w/Tester, Stores & Protects Up to 180 Batteries, Double-Sided, Clear Locking Lid, Secure Latches & Portable Carrying Handle - Red (Batteries Not Included)

Ontel Battery Daddy Storage & Organizer Case w/Tester, Stores & Protects Up to 180 Batteries, Double-Sided, Clear Locking Lid, Secure Latches & Portable Carrying Handle – Red (Batteries Not Included)

  • Double-Sided Storage: Holds up to 180 batteries
  • Clear Viewing Cover: Easily see all batteries at a glance
  • Compact & Portable: Fits in drawers with a carrying handle

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Heat pumps cheaper to operate than gas boilers in almost all EU countries, IEA claims

The IEA reports heat pumps are more cost-effective than gas boilers in 17 EU nations, offering potential savings and climate benefits amid uneven adoption.

Why Your Contact Form Is Killing Your Conversion Rate

Discover how your contact form sabotages conversions and learn simple fixes to turn visitors into leads with less effort and higher success.

Forecast: Impact of EV Adoption on Battery Demand and Prices

The forecast reveals how rising EV adoption could dramatically influence battery demand and prices, leaving you wondering what the industry’s next move will be.

Cisco Systems Surges In Global Coverage

Cisco Systems experiences a significant surge in worldwide media mentions, indicating increased global attention and coverage.