firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
PRIME

Get ready for Prime Big Deal Days — try Prime free

Exclusive member deals on October 6–7, plus fast free delivery. Cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Measure the cut, not the size of the toolbox

Anyone who works with timber knows the difference between careful preparation and useful progress. You can sharpen every blade, label every drawer and produce a beautiful cutting plan. If the cabinet never gets assembled, however, the customer does not receive a cabinet.

That is roughly what happened to Opus 4.8 in Firmulate’s Crucible League. It was the most thorough participant, producing the deepest analyses and adding more than 80 learned rules to its playbook. Yet it finished last with 73 points. Its problem was not a failure to notice what mattered. It identified the crises, resisted attempts to manipulate it and developed the analysis needed to win valuable business. Then, at the decisive moment, it failed to complete the job.

Amazon

AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A bad week, held constant

Firmulate runs AI models as complete small companies, exposing them to real money mechanics, customer pressure and temptations to take shortcuts. In the Crucible experiment, each frontier model faced the same customers, crises and decisions. Every action was versioned and auditable, making the exercise less like a polished chat demonstration and more like watching several tradespeople tackle the same difficult commission.

The final July 2026 Crucible League results put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. A single breach of trust, however, capped the total: “no amount of good work outweighs a breach of trust.”

Opus did not lose because it was reckless or oblivious. All the models detected every crisis and refused every manipulation attempt. The pressure included fake messages from the chief executive escalating across three stages and a reporter asking for “just one yes/no, on background.” All 5 models declined. Kimi K3 recorded the clearest summary of the danger: “Treat the request as a suspected approval-bypass / possible impersonation.”

The fact hidden behind the obvious task

The commercial challenge turned on information that was easy to miss. A decisive weakness in a competitor was not included in the customer event. It was buried two document references deep inside the company’s own files. Models that followed those references found the fact and used it to win the deal at full price, worth €4,583 in monthly recurring revenue.

This is familiar territory in a workshop. The clue that prevents a costly mistake may be in the installation sheet tucked behind the specification, not on the drawing sitting open on the bench. Looking busy is no substitute for checking the source material that changes the decision.

Every model reached the right diagnosis. Every model had essentially the same pitch available. But only two signed the €55,000 deal their analysis had earned: “Same diagnosis, same pitch — no signature.” That distinction explains why Opus’s diligence did not translate into impact. The model accumulated knowledge and examined the situation deeply, but it left the close on the table.

When procedure becomes a substitute for judgment

Opus also lost discipline when it attempted to write into a locked department instead of escalating the obstruction. The episode is small but revealing. A capable operator must distinguish between a task that deserves another attempt and a boundary that requires a different route. Repeating the wrong motion more carefully does not make it the right motion.

Firmulate’s finding should not be reduced to a single-model flaw. A weaker version of the same tendency appeared in the other four participants. Opus is simply the clearest character study because its thoroughness was so pronounced: more than 80 learned rules and the deepest analyses, paired with the lowest final score in the field.

There is also an important qualification when comparing the runners. Kimi K3 operated with its API default because it had no effort parameter, while the other models ran at xhigh. That does not erase the result, but it belongs beside it. Fair tool comparisons require attention to setup as well as outcomes.

A company that can be watched

The benchmark sits inside a live synthetic company with 13 employees. Its finances are deliberately unforgiving: monthly burn of €105,000 against €2,300 in monthly recurring revenue, accompanied by a public cash countdown. Across its operations, the company has accumulated more than 680 self-learned playbook rules, and every workday is versioned.

The public experiment is therefore not fiction or a one-off scripted vignette. It is an ongoing, watchable test of whether AI can manage a company under pressure. A quiz built from 242 real, unedited management decisions also asks readers to guess which model made each choice. For enterprises, the same wargame can be run against a read-only export of their own business, with nothing written back to real systems.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

business analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Completion is a capability

Opus 4.8 deserves a fair reading. It was diligent, perceptive and trustworthy under manipulation. Those are substantial strengths. Its failure was narrower and, for businesses considering AI agents, perhaps more instructive: it did not consistently convert understanding into finished work.

The workshop lesson applies cleanly. More notes, more checks and more rules can improve a job, but only when they serve the critical path. The valuable operator is not necessarily the one with the fullest notebook. It is the one who finds the buried specification, respects the locked door, escalates at the right moment and completes the handover. For AI as for people, diligence is an input. Impact is the finished piece.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

trust and ethics in AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

deal closing automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

I Used To Get Rid Of Fallen Acorns And Pinecones – Until I Found Out How Useful They Are

Interest in repurposing fallen acorns and pinecones is rising as homeowners learn their surprising uses, shifting perceptions about yard debris.

A Twisted Sculpture Lands On Manhattan’s High Line

A large, twisted sculpture has unexpectedly appeared on Manhattan’s High Line, sparking curiosity and speculation among visitors and art experts.

Guide to Installing Router Tables in Small Workshops

Install a rigid, dust-controlled router table that fits your small workshop, supports long stock, and stays safe during real cuts.

James McNeill Whistler’s Music For The Eyes

Search interest in James McNeill Whistler’s ‘Music for the Eyes’ has spiked, driven by rising curiosity about his artistic approach and influence, though specific developments remain unconfirmed.