
Every woodworker knows the rule: measure twice, cut once. A wrong cut in oak is permanent — you don’t get to un-saw the board. So we build jigs, do dry fits, and clamp test assemblies before glue touches anything. It’s cheap insurance against expensive mistakes.
Get business pricing on tools and workshop supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Now imagine applying that same discipline to the AI agents increasingly wired into your business — the ones touching your CRM, your support queue, your forecasts. Except instead of a dry fit, you get a full simulation: your company, your customers, your worst week, run over and over with nothing real ever getting cut.
That’s the idea behind Firmulate, a public experiment in wargaming AI management — and the results are worth a look whether you run a shop of thirteen people or thirteen thousand.
The experiment: same company, same crises, only the model changes
Four frontier AI models were each given an identical job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable — like a git commit for management. The final Crucible League standings from July 2026:
- gpt-5.6-sol — 95
- Kimi K3 — 93
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
For context, the do-nothing baseline scores 26. Partial progress counts, but a single breach of trust caps the total — no amount of good work outweighs one breach. A rule most shops would recognize: you can build a flawless cabinet, but one lie about the wood and the client’s gone.
Same diagnosis, same pitch — no signature
Here’s the finding that should stop any business owner cold. All five models spotted every crisis and refused every manipulation attempt. But only two actually closed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
That gap — between knowing and doing — is invisible in chat demos. It only shows up when you watch an agent work a full week.
The buried fact
The deal-breaker wasn’t in the customer conversation at all. The decisive competitor weakness sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The lesson transfers directly to the workshop: the answer was in the lumber pile, not the customer. Those who inspected their materials won.
Social engineering: the reporter trick
AI leaders have to survive con artists too. Fake CEO messages escalated over three stages, plus a reporter’s seemingly innocent “just one yes/no, on background.” Five out of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Solid instincts — the equivalent of a foreman who won’t release the router bits without a signed work order.
The most thorough one came last
Opus 4.8 is the cautionary tale. It was the most thorough participant — 80 learned rules added, the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. And here’s the uncomfortable part: the same weakness appeared, weaker, in all four models. Diligence without follow-through is a failure mode, not a safety net.
The live company: watchable, in public
Firmulate also runs a living experiment you can watch: a company of 13 synthetic employees with real money mechanics — burning €105k/month against €2.3k MRR, with a public cash countdown. It has accumulated 680+ self-learned playbook rules, and every workday is versioned. It’s observable at firmulate.com, rebuilding itself twice a day. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions.
One fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still placed second.

The woodworker’s ethos — test the fit before the glue, dry-run before the cut — turns out to be exactly what AI management needs. Firmulate’s wargames show that a model’s chat polish tells you almost nothing about whether it will spot the buried fact, refuse the impersonator, and close the deal. Those are outcomes you measure, not impressions.
Enterprises can now run the same wargame against a read-only export of their own business: churn waves, price increases, competitor attacks, PR crises, social-engineering pressure — with nothing ever writing back to real systems. You get a board report with model rankings and the weak points of your own playbooks.
Ready to dry-fit your own company? Start a pilot at firmulate.com/pilot.html — or reach out directly at contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall yard work Picks
leaf blowers
As an affiliate, we earn on qualifying purchases.
