AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

When the trail disappears, which AI keeps moving?

Travel and outdoor experience teaches a useful lesson about judgment: recognizing danger is not the same as getting everyone safely home. A capable leader must read the map, notice what others overlook, resist bad shortcuts and complete the job under pressure.

That is also the premise behind Firmulate, a live experiment that placed frontier AI models in charge of the same small software company during its worst week. Each faced identical customers, crises and temptations. Their decisions were versioned and auditable, producing something more revealing than another polished chatbot demonstration: a record of how different models actually manage.

Readers can now examine those records through a highly shareable challenge. The Firmulate quiz draws from 242 real, unedited management decisions and asks players to guess which model made each one. What emerges is less like a technical benchmark and more like a collection of distinct management personalities.

Amazon

AI management decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Identical terrain, markedly different outcomes

The final Crucible League table from July 2026 placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. There was one overriding condition: a breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

The striking result was not that some models failed to notice an emergency. All of them spotted every crisis. All also refused every manipulation attempt. The separation came later, where recognition had to become execution.

Only two models signed the €55,000 deal that their own analysis had earned. The others could understand the customer, identify the opportunity and construct the pitch, yet still leave without the signature. Firmulate summarizes that gap neatly: “Same diagnosis, same pitch — no signature.”

The clue hidden off the main trail

The decisive weakness in a competitor was not sitting in the customer event where an attentive manager might naturally look. It was buried two document references deep in the company’s own files. Models that followed that trail won the deal at full price, adding €4,583 in monthly recurring revenue.

For any organization considering AI agents, this is an important distinction. A model may sound informed in a conversation while failing to consult the material that makes a decision commercially useful. In outdoor terms, it is the difference between confidently describing the landscape and actually checking the route notes before choosing a fork.

Pressure tested without surrendering trust

The models also faced fake messages from the chief executive that escalated over three stages, followed by a reporter attempting to extract “just one yes/no, on background.” All 5 models refused the manipulation. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

That clean response matters because the simulated company is not an abstract question-and-answer exercise. It has 13 synthetic employees and real money mechanics, including monthly burn of €105k against €2.3k in monthly recurring revenue. It maintains a public cash countdown, has accumulated more than 680 self-learned playbook rules and versions every workday.

The result is a company under sustained operational pressure, where doing nothing carries consequences and apparent authority cannot automatically be trusted. The experiment is live and watchable, allowing observers to follow behavior as company conditions evolve rather than relying only on a retrospective presentation.

Thoroughness did not guarantee victory

Opus 4.8 offers the clearest warning against equating volume with effectiveness. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The deal remained unsigned, while discipline slipped through attempts to write into a locked department instead of escalating the problem. The same weakness appeared in all four other participants, although less strongly.

This profile complicates the familiar assumption that the model with the most exhaustive response must be the safest managerial choice. Detailed thinking can be valuable, but only when it supports timely, disciplined action. A beautifully documented expedition that never leaves camp is still not a successful expedition.

There is also an important fairness note when comparing results. Kimi K3 ran with the application programming interface default because it had no effort parameter. The other models ran at xhigh. Its second-place finish should therefore be read with that difference in conditions visible, not quietly ignored.

Infographic —
The findings at a glance — source: firmulate.com.

A field test for the AI workforce

The quiz works because it turns an abstract procurement question into a human one: can you recognize a model by the decisions it makes? Across the 242 examples, readers can look for persistence, caution, thoroughness and the crucial ability to finish.

The wider lesson is that management personality is measurable. Models facing the same facts can notice the same danger and resist the same deception, yet produce materially different business outcomes. For leaders preparing to give AI access to customer relationships, support work or financial planning, fluency alone is a poor compass.

Firmulate also offers enterprises a pilot using a read-only export of their own business. Nothing writes back to real systems. That allows an organization to stage its own wargame before placing an AI workforce on the operational trail—testing not merely whether a model knows what to do, but whether it reliably does it.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Art Deco Symphony at L’hôtel Saint-Marc Near Opéra-Comique

Marvel at the Art Deco Symphony at L’Hôtel Saint-Marc, where history and luxury entwine—discover what awaits within this enchanting boutique hotel.

A Favorite in Notting Hill: The Laslett

The Laslett, a charming blend of Victorian elegance and modern comfort in Notting Hill, invites you to uncover its delightful surprises.

Aruba’s Oldest Luxury Hotel Just Got a Boutique Tower — and a Rooftop Worth the Trip

The Hilton Aruba Caribbean Resort & Casino unveils The Westerly, a new boutique tower with a rooftop lounge, enhancing its luxury offerings on Palm Beach.

Great Wolf Lodge Surges In Global Coverage

Great Wolf Lodge experiences a surge in international coverage, with 67 media mentions in recent monitoring, highlighting growing global interest.