
In a world where artificial intelligence increasingly guides our decisions, understanding whether these models can be trusted with ethical choices is vital—especially in high-stakes business environments. Imagine a company run entirely by AI, navigating crises, temptations, and complex human interactions, all in real time. Can these digital managers demonstrate consistent integrity, or do their personalities influence their decisions? A groundbreaking live experiment from Firmulate offers a revealing glimpse.
The Experiment: Putting AI Models to the Test
Firmulate’s live business simulation pits four frontier AI models against the challenges of running a small software enterprise during its worst week. The same crises—customer issues, internal temptations, and manipulative social engineering—were presented to each model, which then made decisions in real time. Every choice was carefully recorded, making this an unprecedented look into not just what AI can do, but how it approaches management and ethics.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Results: Trustworthiness and Decision-Making
All four models demonstrated impressive vigilance by identifying every crisis and refusing every attempt at manipulation, including fake CEO messages and a staged reporter trick. Their ability to spot threats signals a robust awareness. Nonetheless, only two models successfully closed a crucial €55,000 deal—their own analysis had earned it—while the others held back, leaving potential revenue on the table.
The Hidden Factor: Reading Between the Files
The decisive advantage for the winning models was their ability to access critical information buried two document references deep within the company’s files—information that the other models failed to uncover. When the models that read these files made the connection, they closed the deal at full price, adding €4,583 in monthly recurring revenue. This highlights a crucial insight: the capacity to read and interpret internal data can be the difference between ethical decision-making and missed opportunities.
Social Engineering and Ethical Boundaries
The models faced a staged escalation: a series of fake CEO messages designed to test their boundaries. All models refused to escalate or act on the manipulated requests, citing concerns about impersonation and bypassing approval processes—Kimi K3 explicitly described the request as a suspected approval-bypass or impersonation. This strict stance indicates that these AI systems are capable of recognizing and resisting social engineering, aligning with principles of integrity and honesty.
The Real-World Setting: A Business Under Strain
The experiment took place within a real operational environment—a company with 13 synthetic employees, burning €105,000 each month against a revenue of just €2,300. The company employs over 680 self-learned rules, with decisions made daily and every move versioned for transparency. Watching this live, one witnesses a microcosm of AI’s potential to manage real, money-driven organizations with integrity and consistency.
Personality Profiles of AI Models
The models showed distinct management personalities:
- Opus 4.8: The most thorough and analytical, with over 80 learned rules, yet it left the close on the table due to discipline lapses, such as writing attempts into a locked department instead of escalating.
- Kimi K3: The newcomer, running at default API effort, displayed the clearest sense of fairness and integrity—closing the deal at full price without effort parameter adjustments.
- Sonnet 5: Slightly less disciplined, with more process slips, but still managed to close the deal.
- Fable 5: Similar to Sonnet but with more slips, indicating that even disciplined models can falter under pressure.
Interestingly, the last-place model, Opus 4.8, demonstrates that even the most thorough AI can slip in real-world scenarios, especially when discipline is compromised, emphasizing that ethical consistency isn’t guaranteed solely by complexity.
Implications for Business and Ethics
This experiment underscores a vital point: when deploying AI in management roles, it’s not just about whether it can handle the workload, but whether it can uphold core values like honesty and trust. The models’ ability to refuse manipulation and identify hidden information suggests they can serve as guardians of integrity—if designed with that purpose in mind.
The Future of AI in Management
As these models continue to evolve, their personalities and decision-making styles will become critical factors. A model like Kimi K3, which refused to overlook fairness for efficiency, embodies the kind of digital manager who prioritizes ethical standards. Meanwhile, the experiment demonstrates that reading and interpreting internal data can be a decisive factor in closing deals and maintaining integrity.
Learn More and Test Your Business
If you’re curious about how AI could manage your organization, you can explore this live experiment yourself by visiting firmulate.com/quiz.html. Run your own decisions against these models and see which personality emerges—trustworthy, terse, or cautious? The future of digital management isn’t just about performance; it’s about ethics, transparency, and trust.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html