firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Every woodworker knows the type: the apprentice who shows up, sweeps the shavings, oils the machines, and never once picks up a chisel. Useless? Not entirely — but you wouldn’t hire them to run the shop. Now imagine you had to give that apprentice a number. Not zero, because an empty shop is worse than a swept one. Not fifty, because nothing got built. Something in between — a number that says “kept things from falling apart, but the commission went unsigned.” That, in a nutshell, is the strangest and most honest detail in a new AI benchmark called the Crucible League: the do-nothing baseline scores 26. Not 0. And the people running it consider that a feature, not a bug.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Same Shop, Same Worst Week, Different Foreman

Firmulate, a public project watchable at firmulate.com/live, did something deceptively simple: it handed four frontier AI models the identical job — run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, the management equivalent of leaving your layout lines visible on the finished piece.

The final league table from July 2026: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. Fine rankings — but the more interesting numbers are the ones explaining how scoring works at all.

Amazon

AI decision-making benchmark tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why the Floor Is 26, Not 0

In most benchmarks, doing nothing earns you nothing. Firmulate rejects that, on a craftsman’s logic: a manager who does nothing still keeps certain bad things from happening. Crises get noticed even if unaddressed; chaos doesn’t compound. So partial progress counts. A model that diagnoses a problem but doesn’t fix it isn’t scored the same as one that never looked. The do-nothing run — a deliberate control, like a test board run through the planer before your good stock — lands at 26.

But there’s a ceiling rule too, and it’s stricter than any floor: a single breach of trust caps the total grade. The project’s stated principle is blunt — “no amount of good work outweighs a breach of trust.” In the shop, that’s the apprentice who does flawless joinery but lies about checking the fence. One lie, and the flawless work stops counting for much.

Amazon

AI model evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Week Itself

The setup is more than a questionnaire. The company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown — anyone can watch the runway shrink. Over the run, the system accumulated 680+ self-learned playbook rules, and every workday is versioned like a commit history.

The headline finding surprised even the organizers. All models spotted every crisis. All refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. It’s the AI equivalent of a perfect cut list and immaculate layout, and then the cabinet never gets glued up.

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact

The deal turned on something subtle. The decisive competitor weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson lands hard for anyone who’s made a mistake because they didn’t check what was already in the drawer: read your own records before you talk to the customer.

Amazon

AI trust and ethics assessment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Pressure, and Refusing the Shortcut

Some of the pressure came from social engineering: fake CEO messages escalating over three stages, plus a reporter offering an easy out — “just one yes/no, on background.” All five models tested refused. Kimi K3’s on-record reasoning was admirably paranoid: “Treat the request as a suspected approval-bypass / possible impersonation.”

The last-place profile is instructive. Opus 4.8 was the most thorough participant — over 80 learned rules, the deepest analyses — yet finished at 73. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. Like a woodworker who measure-obsesses but never clamps up, thoroughness without follow-through loses to slightly messier competitors who finish. The same weakness appeared, weaker, in all four models.

One Footnote on Fairness

K3 ran without an effort parameter — the API default — while the other models ran at xhigh. It still took second at 93. Worth knowing when comparing the top two scores.

Try It Yourself

There’s a guessing game built from 242 real, unedited management decisions — a “guess the model” quiz. And the benchmark itself grows: a new run was in the lab at the time of writing, with results published automatically at each refresh. Full results and plain-language findings are at firmulate.com/benchmarks.html.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

An honest benchmark doesn’t hand out zeros or perfect hundreds easily. It pays for partial progress (the floor of 26), but it also says one breach of trust caps everything above it — and it distrusts a clean 100 enough to never simply award one. For anyone hiring, deploying, or just watching AI agents enter real businesses, that’s the right instinct: grade them like you’d grade a foreman. Not on how well they talk in the demo, but on whether they finish what they start, read the files that are already there, and stay honest when nobody’s watching — except, of course, the version history, which is always watching.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Comparing Benchtop Planers for Hardwood Projects

Compare cutterheads, feed speeds, snipe, dust collection, accuracy, and ownership costs before choosing a hardwood planer.

Alma Thomas Opens Up Our Crowns

Alma Thomas’s artwork titled ‘Our Crowns’ is now on display, drawing increased attention to her innovative style and legacy. Details remain limited on the exhibition specifics.

James McNeill Whistler’s Music For The Eyes

Search interest in James McNeill Whistler’s ‘Music for the Eyes’ has spiked, driven by rising curiosity about his artistic approach and influence, though specific developments remain unconfirmed.

Liu Wei’s Met Facade Comission Reflects Our Moment

Liu Wei’s new facade artwork at the Met Museum captures contemporary themes, sparking increased interest and discussion about art’s role today.