firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine trying to run a busy restaurant during its worst week — managing crises, dodging scams, and closing deals—all while ensuring the staff stays honest. Now, replace that restaurant with a software company, and your staff with AI models. That’s exactly what a groundbreaking live experiment has revealed about the future of AI-driven management.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen staples and gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Recently, a live, real-world test pitted five advanced AI models against each other, simulating a tough week in a small software business. The goal? To see which AI could best handle crises, avoid manipulation, and even close a crucial deal. The results? All models identified every crisis and refused manipulation attempts. Yet, only two managed to close the deal and sign it at full price — a critical measure of their practical competence.

The League Table of AI Performance

  • gpt-5.6-sol: Scored the highest with a 95, found the buried fact, and secured the €55,000 deal.
  • Kimi K3: A newcomer from Moonshot, scored just behind at 93, and demonstrated the cleanest discipline in the field, also closing the deal.
  • Sonnet 5: Achieved an 88, closing the deal but with some process slips.
  • Fable 5: Scored 77, also closing, but with weaker discipline.
  • Opus 4.8: Scored 73, closing the deal but showing similar weaknesses as others.

The Hidden Weaknesses

Interestingly, the decisive advantage for the winning models lay not in superficial chat skills but in their ability to read deeply into company files. Both winning models dug two document references deep into the company’s own files—an insight that proved crucial for closing the deal at full price. Conversely, the lower-performing models missed this buried information, leaving money on the table.

Handling Social Engineering and Trust

During the test, all models faced simulated social engineering attacks—fake CEO messages escalating in urgency and a reporter’s attempt to coerce approval. All refused these attempts, with Kimi K3 explicitly reasoning: ‘Treat the request as a suspected approval-bypass / possible impersonation.’ This shows a commendable level of caution and discernment, vital for real-world application where trust can be exploited.

The Live Business Environment

The experiment was conducted within a real, functioning business setup with 13 synthetic employees and actual financial mechanics — burning €105k monthly against €2.3k MRR. The system is live, constantly analyzing decisions based on over 680 self-learned rules, and is openly accessible for watching at firmulate.com/live. This transparency underscores the experiment’s seriousness and its relevance for actual enterprise management.

Insights from the Models

The most thorough participant, Opus 4.8, with over 80 learned rules, demonstrated deep analytical capabilities but ultimately left the close on the table, showing a lapse in discipline by diverting work into a locked department instead of escalating. This highlights that even the most detailed analysis can falter if discipline slips—an important lesson for deploying AI in management.

Fairness and Testing Conditions

It’s worth noting that Kimi K3 was evaluated without an effort parameter (the API default), while the other models ran at xhigh. This ensures a fair comparison, revealing K3’s robustness under typical operational settings.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The live experiment demonstrates that AI models can do more than just produce convincing chat—they can manage real business situations, identify critical hidden facts, refuse manipulation, and even close deals at full value. The winner, Kimi K3 from Moonshot, proved that a newcomer could outperform established Western models in discipline and effectiveness. For enterprises considering AI management tools, the key takeaway is clear: selecting a model isn’t just about chat quality, but about whether it can complete critical work under pressure, read deeply into data, and stay honest—traits that matter more than ever in a data-driven world. The league is wide open, and as these models improve, the choice of AI partner becomes less about reputation and more about proven results.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI business management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

When AI Fails to Close: Lessons from a Live Business Experiment in Trust and Discipline

Live AI business test reveals that closing deals, reading internal files, and resisting manipulation are invisible in demos but crucial for success—only two models succeeded.

The Real Reason Your Toast Never Tastes Like a Café’s

Inefficient home equipment and techniques often prevent your toast from matching café quality, but the key to improvement lies in understanding why.

Watch an AI-Run Business Struggle and Survive in Real-Time 

Discover how a live, AI-driven company faces crises, refuses manipulation, and fights to stay afloat, revealing what trustworthy AI in business truly looks like.

13 Fall Groceries The Trader Joe’s Employees Are Most Excited About (Out Of Hundreds!)

Trader Joe’s staff highlight their favorite 13 fall groceries, sparking increased interest in seasonal products amid rising consumer curiosity.