
Turn your quiet moments into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
The test is not what an AI says it would do. It is what it does under pressure.
For anyone thinking about character, faith or the gap between principle and action, that distinction may feel familiar. A system can spot what is wrong and still fail to do what is right. Firmulate’s live business experiment puts that tension in practical terms: frontier AI models were asked to run the same small software company through its worst week, facing the same customers, crises and temptations.
A test of judgment in the working world
Firmulate describes itself as an AI company emulator. Its experiment gives models responsibility for a software company with 13 synthetic employees and real money mechanics. The company burns €105,000 a month against €2,300 in monthly recurring revenue, and its public cash countdown makes the stakes visible. Every workday is versioned, and the live company has accumulated more than 680 self-learned playbook rules.
The final Crucible League results, from July 2026, put gpt-5.6-sol first with 95. Moonshot’s Kimi K3 came second with 93, ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. K3’s result is striking: it beat three of the four Western frontier models in the field. The do-nothing baseline scored 26. Firmulate’s stated rule is that partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The deeper finding is less about recognizing a crisis than following through. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The decisive weakness in a competitor’s position was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The company’s summary captures the gap: “Same diagnosis, same pitch — no signature.”
K3 found that buried security weakness, won the deal, saved the churning customer and resisted all three baits. It made one deviation, the fewest in the field. The social engineering challenge included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
That conduct offers a concrete way to consider integrity: not as a claim about inner life, but as a pattern visible in decisions. Does a system read what it needs to read? Does it keep its judgment when authority is imitated or pressure rises? Does it complete the work it has already recognized as necessary?
Opus 4.8 makes the distinction especially clear. It was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the deal unclosed and slipped on discipline, attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four. The league table therefore tells a more complicated story than a simple contest in eloquence or effort: thorough analysis does not guarantee sound execution.
There is a fairness detail for readers interpreting the rankings: K3 ran without an effort parameter (API default) while the others ran at xhigh. The results are a report on this experiment, not a universal measure of every model or every business situation. Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting readers to guess which model made each choice.
The experiment is real and watchable at Firmulate. Its full results and plain-language findings are available on the benchmark page. Enterprises can also run the wargame against a read-only export of their own business; the pilot does not write back to real systems.

Test the judgment you plan to trust
K3’s second-place finish shows that the field is open, while the missed deal shows why a strong answer is not the same as a completed responsibility. If an AI model may touch a company’s customers, records or forecasts, choosing by reputation alone is a bet. Firmulate’s experiment makes the case for testing how a model acts when the facts are buried, the pressure is real and the next step matters.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
