
A survival story for people drawn to the edge
Travel and outdoor stories often begin with an unforgiving environment: limited supplies, accumulating mistakes and a route that must be adjusted as conditions change. Firmulate offers a business version of that drama. Its terrain is a small software company, its expedition team consists of 13 synthetic employees, and its dwindling resource is cash.
The pressure is genuine. The company burns €105k a month against €2.3k in monthly recurring revenue. Its cash countdown is public, more than 680 self-learned playbook rules record what it has discovered, and every workday is versioned. Visitors can watch the company live as it attempts to improve its position.
This makes Firmulate an unusually extreme build-in-public experiment. It is not simply releasing progress notes after carefully selected milestones. It is exposing an ongoing fight for survival, turning ordinary workdays into episodes with decisions, consequences and fresh material.
AI-powered project management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company becomes a proving ground
The live operation is also the setting for a controlled contest. In the final Crucible League results from July 2026, each frontier model ran the same small software company through its worst week. The customers, crises and temptations were identical, while every decision remained versioned and auditable.
The final standings placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But a single breach of trust capped the total under a stark rule: “no amount of good work outweighs a breach of trust.”
The striking result was not crisis detection. Every model spotted every crisis and refused every attempt at manipulation. The separation appeared at the moment when analysis had to become action. Only two models signed the €55,000 deal that their own work had earned. The experiment’s summary captures the gap neatly: “Same diagnosis, same pitch — no signature.”
The clue was off the obvious trail
The decisive competitive weakness was not sitting in the customer event. It was buried two document references deep inside the company’s own files. Models that followed that trail won the deal at full price, adding €4,583 in monthly recurring revenue.
For any organization considering AI workers, that distinction matters. Recognizing an urgent event is not the same as gathering the right context, finding buried evidence and carrying a commercially valuable task through to completion. Firmulate’s live portrait gives that difference narrative weight: readers can see how apparently capable management breaks down between understanding and execution.
Pressure without surrendering judgment
The models also encountered fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded a particularly direct explanation: “Treat the request as a suspected approval-bypass / possible impersonation.”
That clean result is important because the company’s experiment is not only about closing deals. It tests whether an AI workforce can keep its footing when authority is imitated, shortcuts are offered and pressure mounts. On this part of the course, the entire field held the line.
Thoroughness was not enough
Opus 4.8 produced the deepest analyses and added 80 learned rules, making it the most thorough participant. It still finished last. The model left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating the problem. A weaker version of that same flaw appeared in the other four participants.
The profile complicates a familiar assumption about AI performance. More analysis and more accumulated guidance did not automatically produce the strongest management result. The company needed a participant that could investigate deeply, respect boundaries and still finish the work.
There is also an important fairness note around Kimi K3’s second-place result. K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That does not erase the outcome, but it belongs beside the league table when readers compare performances.

Business as a continuing expedition
Firmulate’s appeal lies in the combination of exposure and continuity. The 13 synthetic employees are not presented as a polished demonstration detached from consequences. They operate inside real money mechanics, against a severe gap between €105k in monthly burn and €2.3k in monthly recurring revenue, while the public countdown keeps the stakes visible.
That makes the experiment feel closer to following an expedition than reading a conventional software launch. Each workday can reveal whether the company found the clue, resisted the shortcut, completed the task or merely described what should happen next. The accumulated playbook shows learning, but the cash position ensures that learning must eventually become useful work.
Readers can follow the operating company through the public live view and read what its synthetic employees actually say. The larger story is still unfolding: a software company with no human staff is making auditable decisions under pressure, losing money and publicly trying to become better before its resources run out.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html