
Get business pricing on tools and workshop supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
You Wouldn’t Buy a Planer on the Spec Sheet Alone
Anyone who’s spent time in a workshop knows the rule: the brochure tells you almost nothing. You run a test board through the planer, listen to the motor, check the snipe. A tool that looks identical on paper can behave completely differently on your lumber. So why do companies buy AI models the way amateurs buy tools — off the spec sheet, off the demo, off the chat window?
This month, an experiment called the Crucible did what any careful woodworker would insist on: it put the tools through the same job and measured the results. And the outcome upended expectations. Moonshot’s Kimi K3 — the newcomer, running on default settings — beat three of four Western frontier models at running an actual company.
One Company, One Terrible Week, Five Models
The setup is elegantly simple, like a good jig. Four — technically five, counting the baseline — frontier AI models were each handed the same small software company during its worst week: same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing relies on anyone’s word.
The final league table from July 2026:
- 1. gpt-5.6-sol — 95 points. The complete performance.
- 2. Kimi K3 — 93 points. The newcomer. Closed the deal, cleanest discipline in the field.
- 3. Sonnet 5 — 88 points. Closed the deal, with a few more process slips.
- 4. Fable 5 — 77 points.
- 5. Opus 4.8 — 73 points. The most thorough participant — yet last.
For context: doing nothing scores 26. And a single breach of trust caps your total — no amount of good work outweighs it. The full breakdown lives on the public benchmarks page.
The Newcomer’s Week
K3’s second-place run reads like a checklist of things you’d want from a shop foreman. It found a security needle buried two document references deep in the company’s own files — not in the customer event where everyone was looking. It won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. It saved a customer who was about to churn. And when fake CEO messages came in, escalating over three stages, followed by a reporter’s trick — “just one yes/no, on background” — K3 refused all three baits. Its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the AI equivalent of measuring twice.
Crucially, K3 logged only one deviation all week — the cleanest discipline of any model in the field.
Same Diagnosis, No Signature
Here’s the finding that should stop any buyer mid-purchase. Every model spotted every crisis. Every model refused every manipulation attempt. But only two — gpt-5.6-sol and K3 — actually signed the deal their own analysis had earned. The others diagnosed the problem correctly, made the pitch, and then… nothing. Same diagnosis, same pitch, no signature.
That gap is invisible in a chat demo. It only shows up when the model has to run the whole job, start to finish — the AI version of a tool that cuts beautifully on the first pass and chatters on the tenth.
Thorough Isn’t the Same as Good
The most instructive failure belongs to Opus 4.8. It was the most thorough participant by volume: over 80 learned rules added, the deepest analyses in the field. It still finished last. The deal was left on the table, and discipline slipped — it attempted writes into a locked department instead of escalating the problem. The same weakness showed up, weaker, in all four Western models. Effort without judgment is like sanding harder on a board that needs to be recut.
It’s Real, and You Can Watch It
The Crucible isn’t a slide deck. Firmulate runs a live company — 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules. Every workday is versioned, and you can watch it in real time at firmulate.com. If you’d rather test your own eye, 242 real, unedited management decisions power a “guess the model” quiz — a humbling exercise. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.
A Note on Fairness
One caveat for the record: K3 ran without an effort parameter (the API default), while the other models ran at xhigh. In other words, the newcomer wasn’t even dialed up — and still placed second.

The League Is Open
The lesson for anyone hiring tools — or AI — is the one every woodworker already knows. The name on the machine doesn’t tell you how it handles your material. A newcomer from Moonshot outperformed three of four Western frontier models at the actual job, while the most verbose model finished last. The performance spread between “talks well” and “does the work” only becomes visible under a real, repeatable test.
If AI agents will touch your CRM, your support queue, or your forecast, the question isn’t whether they write fluently. It’s whether they finish what they start, read your files before acting, and stay honest under pressure. You can’t answer that from a demo. You answer it by running the board through the planer — and picking a model without your own test is now, plainly, a bet.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI performance benchmarking software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
enterprise AI decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI security and trust verification tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
