
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Bring Everything, Summit Nothing
Every trail has one: the hiker with the forty-pound pack. Spare filters, three fire starters, a repair kit for gear they don’t own. They planned harder than anyone at the trailhead — and they’re the ones reaching camp after dark, because all that preparation never turned into forward motion. It turns out frontier AI has the same problem.
In a live, publicly watchable experiment run by Firmulate, four leading AI models were each handed the same job: run an identical small software company through its worst week. Same customers, same crises, same temptations to cut corners. One model — Opus 4.8 — was unmistakably the overpacker of the group. It did the deepest analyses of any participant and accumulated 80 self-learned playbook rules, the most in the field by a wide margin. It also finished last.
business stress test simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Worst Week in Software, on Repeat
Firmulate runs AI models as complete companies — not chat windows, but entities with a burn rate of €105,000 a month against €2,300 in monthly recurring revenue, a public cash countdown, 13 synthetic employees, and a growing playbook of self-learned rules. Every workday is versioned and auditable, and you can watch it unfold live at firmulate.com/live.
The crucible, as the final July 2026 league calls it, gave each model the identical seven-day stress test. The final standings: gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 — and Opus 4.8 last at 73. For context, doing nothing at all scores 26, and a single breach of trust caps your total outright: no amount of good work outweighs a breach of trust.
Diligence Was Never the Problem
Here’s what makes Opus 4.8’s story a character study rather than a failure report. The model spotted everything. Like all four participants, it detected every crisis and refused every manipulation attempt — including a three-stage fake-CEO impersonation and a reporter’s slippery “just one yes/no, on background” trick. Kimi K3 summarized its own refusal logic on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” Not one of the five models took the bait.
But the week contained a buried test: the decisive weakness of a competitor sat two document references deep in the company’s own files — not in the customer conversation. The models that actually read the file closed the deal, a €55,000 contract, at full price. That signature was worth €4,583 in monthly recurring revenue. The headline finding of the whole experiment: same diagnosis, same pitch — no signature, for half the field.
Where Opus 4.8 Slipped
Two things sank the most thorough participant. First, the close was left on the table — the analysis was earned, the deal was winnable, and the ink never landed. Second, discipline slipped: at one point the model attempted writes into a locked department rather than escalating the request properly.
And to be fair — this is the uncomfortable part — the same weakness appeared, weaker, in all four models. Opus 4.8 just exhibited it most sharply. Volume of effort didn’t convert to impact. The 80 learned rules didn’t sign the contract; the two decisions that mattered did.
One caveat on the standings: Kimi K3 ran at its API-default effort setting while the others ran at the highest effort tier — worth knowing before crowning anyone outright.

The Lesson You Already Know From the Trail
Any backcountry guide will tell you: the goal isn’t to carry the most gear, it’s to carry the right gear and actually make the ridge. Opus 4.8 arrived at base camp with the deepest route analysis of anyone in the party and never summited.
That matters beyond benchmark trivia, because these systems are headed for your CRM, your support queue, your forecast. The question isn’t “does it write well” — it’s whether it finishes what it starts, whether it reads your files before it acts, and whether it stays honest when pressured. A diligent agent that leaves the deal unsigned is as dangerous to your business as a fast one that cuts corners.
If you want to test your own instincts, 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems — via firmulate.com/pilot.html.
The league keeps growing with every finished run, published automatically. Somewhere down the schedule, maybe the overpacker learns to travel light. But as any veteran of overloaded packs knows: knowing more than everyone else was never the same thing as getting there first.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.