
Get travel and outdoor gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
When the itinerary falls apart, what would your AI assistant do?
A delayed connection, a hotel cancellation and a message claiming to come from the CEO: travel plans can turn into a test of judgment fast. As AI moves from suggesting destinations to handling bookings and customer problems, the useful question is not only whether it can spot trouble. Can it follow through, protect trust and act on what it has learned?
One company, the same worst week
Firmulate set up a live experiment: frontier AI models ran the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and auditable. The company is synthetic, but its money mechanics are real: 13 synthetic employees, monthly burn of €105,000 against €2,300 in monthly recurring revenue, and a public cash countdown.
The final Crucible League, in July 2026, ranked gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 and Opus 4.8 fifth at 73. The do-nothing baseline scored 26. Under the league’s integrity rule, a breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
Seeing the problem is not the same as solving it
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s stark summary: “Same diagnosis, same pitch — no signature.”
The missed opportunity depended on reading the company’s own records. A decisive competitor weakness sat two document references deep in the files, not in the customer event. Models that found it won the deal at full price, worth €4,583 in monthly recurring revenue. In a travel business, the equivalent could be a useful detail tucked into an existing customer record or operating guide: knowing the disruption is real does not automatically mean finding the information needed to resolve it.
Trust held under pressure, too. Five of five models refused fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” Kimi K3 described the reporter approach as: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thoroughness did not guarantee a finish
Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. A weaker version of the same discipline problem appeared in all four models.
There is a fairness detail in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Firmulate also makes 242 real, unedited management decisions available in a “guess the model” quiz. The experiment is watchable at firmulate.com.

From watching to testing your own business
For a travel company, AI might eventually help with bookings, customer support or disruption planning. A live demonstration can show how a model behaves in one company’s crisis. A pilot can ask the harder question: how does it handle yours?
Firmulate says enterprises can run the same kind of wargame against a read-only export of their own business. The exercise can produce a board report ranking models and exposing weak points in a company’s playbooks. Nothing writes back to real systems. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
