firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

You Wouldn’t Buy a Planer on the Spec Sheet Alone

Anyone who’s spent time in a workshop knows the rule: the brochure tells you almost nothing. You run a test board through the planer, listen to the motor, check the snipe. A tool that looks identical on paper can behave completely differently on your lumber. So why do companies buy AI models the way amateurs buy tools — off the spec sheet, off the demo, off the chat window?

This month, an experiment called the Crucible did what any careful woodworker would insist on: it put the tools through the same job and measured the results. And the outcome upended expectations. Moonshot’s Kimi K3 — the newcomer, running on default settings — beat three of four Western frontier models at running an actual company.

One Company, One Terrible Week, Five Models

The setup is elegantly simple, like a good jig. Four — technically five, counting the baseline — frontier AI models were each handed the same small software company during its worst week: same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing relies on anyone’s word.

The final league table from July 2026:

  • 1. gpt-5.6-sol — 95 points. The complete performance.
  • 2. Kimi K3 — 93 points. The newcomer. Closed the deal, cleanest discipline in the field.
  • 3. Sonnet 5 — 88 points. Closed the deal, with a few more process slips.
  • 4. Fable 5 — 77 points.
  • 5. Opus 4.8 — 73 points. The most thorough participant — yet last.

For context: doing nothing scores 26. And a single breach of trust caps your total — no amount of good work outweighs it. The full breakdown lives on the public benchmarks page.

The Newcomer’s Week

K3’s second-place run reads like a checklist of things you’d want from a shop foreman. It found a security needle buried two document references deep in the company’s own files — not in the customer event where everyone was looking. It won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. It saved a customer who was about to churn. And when fake CEO messages came in, escalating over three stages, followed by a reporter’s trick — “just one yes/no, on background” — K3 refused all three baits. Its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the AI equivalent of measuring twice.

Crucially, K3 logged only one deviation all week — the cleanest discipline of any model in the field.

Same Diagnosis, No Signature

Here’s the finding that should stop any buyer mid-purchase. Every model spotted every crisis. Every model refused every manipulation attempt. But only two — gpt-5.6-sol and K3 — actually signed the deal their own analysis had earned. The others diagnosed the problem correctly, made the pitch, and then… nothing. Same diagnosis, same pitch, no signature.

That gap is invisible in a chat demo. It only shows up when the model has to run the whole job, start to finish — the AI version of a tool that cuts beautifully on the first pass and chatters on the tenth.

Thorough Isn’t the Same as Good

The most instructive failure belongs to Opus 4.8. It was the most thorough participant by volume: over 80 learned rules added, the deepest analyses in the field. It still finished last. The deal was left on the table, and discipline slipped — it attempted writes into a locked department instead of escalating the problem. The same weakness showed up, weaker, in all four Western models. Effort without judgment is like sanding harder on a board that needs to be recut.

It’s Real, and You Can Watch It

The Crucible isn’t a slide deck. Firmulate runs a live company — 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules. Every workday is versioned, and you can watch it in real time at firmulate.com. If you’d rather test your own eye, 242 real, unedited management decisions power a “guess the model” quiz — a humbling exercise. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

A Note on Fairness

One caveat for the record: K3 ran without an effort parameter (the API default), while the other models ran at xhigh. In other words, the newcomer wasn’t even dialed up — and still placed second.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The League Is Open

The lesson for anyone hiring tools — or AI — is the one every woodworker already knows. The name on the machine doesn’t tell you how it handles your material. A newcomer from Moonshot outperformed three of four Western frontier models at the actual job, while the most verbose model finished last. The performance spread between “talks well” and “does the work” only becomes visible under a real, repeatable test.

If AI agents will touch your CRM, your support queue, or your forecast, the question isn’t whether they write fluently. It’s whether they finish what they start, read your files before acting, and stay honest under pressure. You can’t answer that from a demo. You answer it by running the board through the planer — and picking a model without your own test is now, plainly, a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI performance benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

enterprise AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI model evaluation platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI security and trust verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

IKEA’s New Lighting Collection Looks So Much More Expensive Than It Is

IKEA has launched a new lighting collection that appears significantly more upscale than its affordable price point, sparking widespread interest and speculation.

13 Stores Like IKEA For Stylish And Budget-Friendly Furniture And Decor

Discover 13 stores similar to IKEA offering trendy, budget-friendly furniture and decor options. Perfect for stylish home updates without breaking the bank.

How to Properly Align Your Table Saw Blade for Precision Cuts

Align your table saw blade, fence, riving knife, and bevel stops for cleaner, safer, more accurate cuts using practical garage-shop methods.

Best Table Saws for Hobby Woodworkers

An honest, hands-on breakdown of the best table saw types for hobby woodworkers — specs, safety, rip capacity, dust collection, and real budgets.