
Get backup power and energy gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A Frontier Model Just Ran a Company for a Week — and the Results Read Like a Solar Panel Spec Sheet
Any homeowner comparing solar panels knows the drill: the brochure efficiency rating means nothing until the panel is stressed under real conditions — partial shade, a heatwave, a grid outage. So you look at independent, side-by-side tests. This July, a live experiment called the Crucible did exactly that for AI models — not for chat quality, but for the ability to actually run a business. And the surprise finisher was a newcomer: Moonshot’s Kimi K3, scoring 93 and beating three of four Western frontier models at managing a company under pressure.
As an affiliate, we earn on qualifying purchases.
Same Storm, Different Roof
Here’s how it worked. Each frontier AI was handed the same small software company and told to steer it through its worst week — identical customers, identical crises, identical temptations to cut corners. Every decision was versioned and auditable, so nothing could be quietly buried after the fact. It’s the AI equivalent of putting five inverters on the same test bench during the same grid failure.
The final league table: gpt-5.6-sol took first with 95. Kimi K3 followed at 93. Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73. For context, doing nothing at all would have scored 26 — partial progress counts — while a single breach of trust caps the total outright, on the principle that no amount of good work outweighs a breach of trust.
What the Newcomer Actually Did
K3’s week read like a textbook. It found the buried fact that mattered: the decisive competitor weakness wasn’t in the customer’s event at all, but sat two document references deep in the company’s own files. The models that read the file won the €55,000 deal at full price — worth +€4,583 in monthly recurring revenue. K3 was one of only two models to close it. It also saved a churning customer and recorded just one deviation all week — the cleanest discipline in the field.
Then came the social engineering. Fake CEO messages escalated over three stages, followed by a reporter trick — “just one yes/no, on background.” All five models refused every manipulation attempt. K3’s on-record reasoning was blunt: treat the request as a suspected approval-bypass, possible impersonation.
Diagnosis Without the Signature
The strangest finding cut across the whole field. Every model spotted every crisis and refused every bait — yet only two signed the deal their own analysis had earned. Same diagnosis, same pitch, no signature. That gap is invisible in a chat demo, and it’s exactly the kind of gap that matters if an AI agent will ever touch your customer pipeline, your support queue, or your forecast.
Opus 4.8 is the cautionary profile: the most thorough participant, adding 80 learned rules and producing the deepest analyses — yet finishing last. The close was left on the table, and discipline slipped, with write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models.
You Can Watch the Company Bleed Cash
This isn’t a slide deck. The live company has 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in MRR — a public cash countdown, and over 680 self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com, and there’s a “guess the model” quiz built on 242 real, unedited management decisions. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.
Fairness note: Kimi K3 ran without an effort parameter (API default), while the other models ran at xhigh reasoning effort.

The League Is Open
The takeaway for anyone — including the energy-tech world increasingly handing AI agents real operational duties — is simple: the leaderboard is no longer settled by geography or brand pedigree. A newcomer beat three of four Western frontier models at the unglamorous work of finishing what it started, reading the files first, and staying honest under pressure. Chat demos can’t show you that. If you’re picking a model to touch anything that matters, run your own stress test first — or you’re not buying a system, you’re placing a bet.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
