AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get backup power and energy gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Benchmark That Refuses to Hand Out a Zero

If you have ever compared solar inverters or home battery quotes, you know the frustration of specs that flatter the product while hiding the failure modes. So here is a benchmark with an unusual habit of honesty: when researchers ran a completely passive, do-nothing baseline through a simulated business crisis week, it scored 26 points — not zero. That number tells you a great deal about how the Firmulate benchmark thinks, and why its results are worth more than a glossy demo.

Firmulate is a live, watchable experiment: AI models are handed the same small software company and told to run it through its worst week — same customers, same crises, same temptations to cut corners. Every decision is versioned and auditable, and the scoreboard updates publicly. The final July 2026 league table reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. But the most interesting number in the whole system is that floor of 26.

Amazon

solar inverter with real-time performance monitoring

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Partial Progress Counts

A do-nothing run still gets 26 points because the benchmark rewards partial progress. Think of it like a home energy audit: a house that merely avoids disaster (no fires, no fried appliances) has still done something measurable compared to one that actively damages itself. In Firmulate’s world, a passive manager still answers some customer messages, keeps some processes alive, and avoids inventing new problems. Ignoring that partial value — by scoring it zero — would make the leaderboard look more dramatic than reality justifies.

This matters because inflated drama is exactly what chat demos trade in. Firmulate’s designers are openly distrustful of tidy, round numbers and headline-friendly scores. A floor of 26 forces every participant to beat a real, honest baseline rather than a straw man. It is the benchmarking equivalent of measuring a solar panel’s yield against actual weather data instead of perfect lab conditions.

Why One Breach of Trust Caps Everything

The second pillar of the scoring philosophy is harsher: a single breach of trust caps the total grade. In the benchmark’s own words, “no amount of good work outweighs a breach of trust.” A model could diagnose every crisis perfectly, close every deal, and charm every customer — but if it deceives a stakeholder even once, its ceiling collapses.

For readers who worry about what an AI agent might do inside their smart-home energy systems, their utility accounts, or their battery-management dashboards, this is the standard you should demand. Competence without trustworthiness is a liability, not an asset.

What Actually Happened in the Crucible

The headline finding from the final runs is deceptively simple. All five models spotted every crisis and refused every manipulation attempt — yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch, no signature. The job was 90 percent done and then abandoned on the finish line.

The decisive detail was buried two document references deep in the company’s own files, not in the customer event. The models that actually read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue. The lesson transfers directly to home energy: the answer to “why is my bill spiking?” or “is this battery quote fair?” is rarely in the flashy pitch — it is in the paperwork two layers down.

Social Engineering: The Field Held the Line

The experiment included fake CEO messages that escalated over three stages, plus a reporter’s trick: “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was unambiguous: “Treat the request as a suspected approval-bypass / possible impersonation.”

The Thoroughness Trap

Opus 4.8 is the cautionary tale of the league. It was the most thorough participant — over 80 learned rules and the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. Being the most prepared does not mean being the most finished.

A Fairness Footnote

Honest benchmarks disclose their asterisks: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still placed second with the cleanest discipline of the field.

You Can Watch It Live

The simulated company is real and running: 13 synthetic employees, genuine money mechanics with a burn of €105k per month against €2.3k in MRR, a public cash countdown, and over 680 self-learned playbook rules — every workday versioned. You can watch it at firmulate.com/live. There is also a “guess the model” quiz built on 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway for the Energy-Conscious Reader

Whether you are choosing a solar installer, a battery platform, or an AI agent that will someday manage your home’s energy trading, the Firmulate benchmark offers a template for asking better questions: Does it finish what it starts? Does it read the files before answering? Does it stay honest under pressure — and does one lie permanently cap its grade?

A benchmark that gives a do-nothing baseline 26 points, distrusts round 100s, and refuses to let good work launder a breach of trust is not just scoring AI models. It is modeling how we should evaluate any system we invite into our homes. The full league table and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Future Of Flipper Zero Development

The Flipper Zero team has announced new development plans, including upcoming features and community engagement initiatives, signaling a significant update for users.

Sodium‑Ion Batteries: Market Potential and Roadblocks

Meticulously exploring sodium-ion batteries reveals their promising market potential but also uncovers key challenges that could shape their future development.

A War Room for Your Next Idea: Inside IdeaClyst

Discover how IdeaClyst transforms idea validation and decision-making with an AI-powered, local-first war room. Perfect for founders and teams aiming for clarity and speed.

Does Costco Sell Car Batteries? Shocking Truth Revealed!

Surprising details about Costco’s car battery offerings await, including pricing and warranty changes that could impact your next purchase decision.