firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

You Wouldn’t Buy a Planer on the Spec Sheet Alone

Anyone who’s spent time in a workshop knows the rule: the brochure tells you almost nothing. You run a test board through the planer, listen to the motor, check the snipe. A tool that looks identical on paper can behave completely differently on your lumber. So why do companies buy AI models the way amateurs buy tools — off the spec sheet, off the demo, off the chat window?

This month, an experiment called the Crucible did what any careful woodworker would insist on: it put the tools through the same job and measured the results. And the outcome upended expectations. Moonshot’s Kimi K3 — the newcomer, running on default settings — beat three of four Western frontier models at running an actual company.

One Company, One Terrible Week, Five Models

The setup is elegantly simple, like a good jig. Four — technically five, counting the baseline — frontier AI models were each handed the same small software company during its worst week: same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing relies on anyone’s word.

The final league table from July 2026:

  • 1. gpt-5.6-sol — 95 points. The complete performance.
  • 2. Kimi K3 — 93 points. The newcomer. Closed the deal, cleanest discipline in the field.
  • 3. Sonnet 5 — 88 points. Closed the deal, with a few more process slips.
  • 4. Fable 5 — 77 points.
  • 5. Opus 4.8 — 73 points. The most thorough participant — yet last.

For context: doing nothing scores 26. And a single breach of trust caps your total — no amount of good work outweighs it. The full breakdown lives on the public benchmarks page.

The Newcomer’s Week

K3’s second-place run reads like a checklist of things you’d want from a shop foreman. It found a security needle buried two document references deep in the company’s own files — not in the customer event where everyone was looking. It won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. It saved a customer who was about to churn. And when fake CEO messages came in, escalating over three stages, followed by a reporter’s trick — “just one yes/no, on background” — K3 refused all three baits. Its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the AI equivalent of measuring twice.

Crucially, K3 logged only one deviation all week — the cleanest discipline of any model in the field.

Same Diagnosis, No Signature

Here’s the finding that should stop any buyer mid-purchase. Every model spotted every crisis. Every model refused every manipulation attempt. But only two — gpt-5.6-sol and K3 — actually signed the deal their own analysis had earned. The others diagnosed the problem correctly, made the pitch, and then… nothing. Same diagnosis, same pitch, no signature.

That gap is invisible in a chat demo. It only shows up when the model has to run the whole job, start to finish — the AI version of a tool that cuts beautifully on the first pass and chatters on the tenth.

Thorough Isn’t the Same as Good

The most instructive failure belongs to Opus 4.8. It was the most thorough participant by volume: over 80 learned rules added, the deepest analyses in the field. It still finished last. The deal was left on the table, and discipline slipped — it attempted writes into a locked department instead of escalating the problem. The same weakness showed up, weaker, in all four Western models. Effort without judgment is like sanding harder on a board that needs to be recut.

It’s Real, and You Can Watch It

The Crucible isn’t a slide deck. Firmulate runs a live company — 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules. Every workday is versioned, and you can watch it in real time at firmulate.com. If you’d rather test your own eye, 242 real, unedited management decisions power a “guess the model” quiz — a humbling exercise. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

A Note on Fairness

One caveat for the record: K3 ran without an effort parameter (the API default), while the other models ran at xhigh. In other words, the newcomer wasn’t even dialed up — and still placed second.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The League Is Open

The lesson for anyone hiring tools — or AI — is the one every woodworker already knows. The name on the machine doesn’t tell you how it handles your material. A newcomer from Moonshot outperformed three of four Western frontier models at the actual job, while the most verbose model finished last. The performance spread between “talks well” and “does the work” only becomes visible under a real, repeatable test.

If AI agents will touch your CRM, your support queue, or your forecast, the question isn’t whether they write fluently. It’s whether they finish what they start, read your files before acting, and stay honest under pressure. You can’t answer that from a demo. You answer it by running the board through the planer — and picking a model without your own test is now, plainly, a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI performance benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

enterprise AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI model evaluation platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI security and trust verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

10 Art Shows To See In Chicago This Fall

Discover the 10 must-see art exhibitions in Chicago this fall, featuring diverse works from local to international artists, opening across major galleries and museums.

The Meticulous AI That Forgot to Finish the Job

The most diligent AI built 80 rules and still finished last. Firmulate’s live company test shows why analysis matters less when nobody closes the deal.

Judicial Review Sought Over Fast-track For Auckland Seaside Village Redevelopment – RNZ

Legal challenge has been launched against the expedited process for Auckland’s seaside village redevelopment, raising concerns over planning and environmental standards.

The Largest Roman Mosaic Ever Excavated Opens To The Public

The world’s largest Roman mosaic has been uncovered and opened for public viewing, marking a significant archaeological milestone and attracting global interest.