AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Every serious trekker knows the type. The fresh-faced guide nobody has heard of shows up at base camp, carries the same load as the veterans, reads the weather the old hands read, and somehow gets the whole party over the pass with fewer stumbles than anyone expected. That, more or less, is what just happened in the world of artificial intelligence — except the mountain was a failing software company and the guides were frontier AI models.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get travel and outdoor gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

In a live, public experiment run by Firmulate, an AI company emulator that runs language models as entire businesses, Moonshot’s newcomer Kimi K3 took second place in a management league table with a score of 93 — finishing ahead of three of four Western frontier models it was tested against.

The Worst Week in Business, on Repeat

The setup is elegantly brutal. Each frontier model was handed the same small software company and told to steer it through its worst week: the same customers, the same crises, the same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing could be quietly swept under the rug.

The final July 2026 league table tells the story:

  • 1. gpt-5.6-sol — 95
  • 2. Kimi K3 — 93
  • 3. Sonnet 5 — 88
  • 4. Fable 5 — 77
  • 5. Opus 4.8 — 73

For context, doing nothing scores 26 — partial progress counts — but a single breach of trust caps the total. As the experiment’s own rule puts it, no amount of good work outweighs a breach of trust.

Amazon

AI management software tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What K3 Actually Did

K3’s week reads like a checklist of good judgment. It found the buried security needle hidden in the company’s own files. It won the €55,000 deal — worth an extra €4,583 in monthly recurring revenue — at full price. It saved a customer who was on the verge of churning. And it resisted all three bait attempts aimed at it, with just one deviation across the entire week: the cleanest discipline in the field.

That last part matters more than it sounds. The buried fact — the decisive competitor weakness — sat two document references deep in the company’s own files, not in the customer event itself. Models that actually read the file won the deal. Models that didn’t, didn’t. Same diagnosis, same pitch — no signature.

The Test Nobody Passes in a Chat Demo

The experiment also threw social engineering at every model: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

The headline finding, though, was subtler: every model spotted every crisis and refused every manipulation. Only two signed the deal their own analysis had earned. That gap — between seeing the summit and actually standing on it — is invisible in a chat demo.

Effort Isn’t Everything

The most sobering result belonged to Opus 4.8: the most thorough participant in the field, with 80-plus learned rules added and the deepest analyses — and still last place at 73. It left the close on the table and slipped on discipline, attempting writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four competitors. Diligence without follow-through is a familiar failure mode to anyone who has watched an over-prepared expedition turn back a few hundred meters from the summit.

You Can Watch It Lose Money

This is not a slide deck. The company is real software running every business day, with 13 synthetic employees, real money mechanics — a burn of €105,000 per month against €2,300 in MRR — a public cash countdown, and more than 680 self-learned playbook rules. You can watch it live at firmulate.com.

Feeling confident about your own judgment? A quiz built from 242 real, unedited management decisions lets you guess which model made which call — a humbling exercise in how similar the contestants sound until you see the outcomes. The full methodology and plain-language findings are on the benchmarks page.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The lesson for anyone hiring AI — or hiring a guide — is that reputation is no longer a shortcut. A newcomer at default settings nearly topped the field, and the most diligent veteran finished last. The league is open. If AI agents will touch your CRM, your support queue, or your forecast, the question is not whether they write well; it is whether they finish what they start, read your files first, and stay honest under pressure. Picking a model without running your own test is no longer a decision — it is a bet.

Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh — meaning the newcomer’s near-top finish came without the extra reasoning budget its rivals received.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Aruba’s Oldest Luxury Hotel Just Got a Boutique Tower — and a Rooftop Worth the Trip

The Hilton Aruba Caribbean Resort & Casino unveils The Westerly, a new boutique tower with a rooftop lounge, enhancing its luxury offerings on Palm Beach.

What Premium Homes Borrow From Hospitality Better Than Ever

Knowledge of luxury amenities and services reveals how premium homes are borrowing from hospitality better than ever, transforming your living experience—continue reading to discover how.

Bulgari Hotel Milan – Italian Luxury at Its Finest

Bulgari Hotel Milan offers an unparalleled blend of elegance and modern luxury, but what truly sets it apart awaits your discovery.

Eco‑Lodges: Sustainable Luxury in the Wild

Many eco-lodges redefine luxury by seamlessly blending sustainability with comfort, inspiring responsible travel—discover how they achieve this balance.