
Get home appliances delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
AI’s New Role: Running a Business in Real-Time
Imagine AI not just answering questions or managing smart home devices, but actively steering a company through its toughest week. That’s exactly what the latest live experiment from Firmulate demonstrates—turning AI models into decision-makers in a high-stakes business environment, with real money and real crises at play.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Testing AI in a Business Crisis
In a groundbreaking live experiment, four advanced AI models were tasked with managing a small software company during its worst week. The same set of customers, crises, and temptations were faced by each model—only the AI’s decision-making engine changed. Every choice was documented and auditable, providing a transparent lens into their behavior under pressure.
The League Table of Performance
- gpt-5.6-sol: Achieved the highest score of 95, uncovering buried critical information and sealing a €55,000 deal.
- Kimi K3 (Moonshot): The surprise runner-up with a score of 93, demonstrating exceptional discipline and winning the deal on its own merits.
- Sonnet 5: Scored 88, successfully closing the deal but with some process slips.
- Fable 5: Scored 77, also closing the deal but showing more signs of process slips.
- Opus 4.8: Scored 73, the lowest but still managing to close the deal.
All models refused manipulative tactics such as social engineering or fake CEO messages, demonstrating a robust ethical stance during the test. Notably, the key to winning the biggest deal was reading two documents deep into the company’s own files—an area where the winning models excelled, reflecting their ability to uncover hidden insights crucial for decision-making.
As an affiliate, we earn on qualifying purchases.
Discipline Matters More Than Chat Quality
In these tests, superficial chat finesse takes a backseat. The models’ real strength lies in their ability to finish what they start, read relevant documents, and maintain honesty under pressure. For example, the Firmulate experiment found that the models refused all manipulation attempts, such as escalating fake CEO requests, which were staged in a three-stage social engineering scenario.
What About the Real Business?
The experiment wasn’t just theoretical. The company involved is real, with 13 synthetic employees managing actual money mechanics—burning €105,000 monthly while earning only €2,300 in monthly recurring revenue. Every move, every decision, is versioned daily, allowing observers to see exactly how each AI navigates real-world pressures.
As an affiliate, we earn on qualifying purchases.
What We Learn for Home and Smart Devices
This experiment underscores a vital point for the smart home industry: AI isn’t just about clever replies or automations. It’s about creating systems that can handle complex, pressure-filled situations with discipline and honesty. Whether managing security alerts, optimizing energy use, or handling support requests, the ability to finish tasks reliably and ethically is paramount.
AI business crisis management systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Bigger Picture
As AI models continue to advance, the ability to make disciplined, honest decisions under stress becomes a key differentiator. The leaderboard from the Firmulate live test shows that even newcomers like Kimi K3 can outperform established models, provided they are tested rigorously in real-world scenarios. This emphasizes that choosing AI tools based solely on superficial chat scores might be risky—performance in critical decision-making under pressure is what truly matters.
The Fairness and Transparency of the Test
It’s important to note that K3 ran without an effort parameter (the default API setting), while the other models operated at xhigh, making K3’s performance even more impressive. The live experiment is accessible and transparent, demonstrating what AI can and cannot do in complex, decision-heavy environments.
Conclusion: AI as a Business Partner You Can Trust
This experiment from Firmulate reveals that AI models are capable of disciplined decision-making in real-world, pressure-filled business situations. For companies and even smart home systems, this means more reliable, ethical, and effective AI-driven management—if tested properly, like in this live wargame.

Key Takeaway
The live experiment shows that AI models like Kimi K3 can outperform expectations by demonstrating discipline, honesty, and problem-solving skills under pressure. Choosing AI based on rigorous real-world testing is crucial for trustworthy automation in business and smart home environments.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
