firmulate.com/live.html — live view
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

Can Artificial Intelligence Run a Business—Honestly and Effectively?

Imagine observing a company that has no human employees, yet faces the same crises, temptations, and decisions as any real business. Now imagine you can watch this company every workday, live, as it navigates its toughest challenges—losing money, making critical decisions, and ultimately, fighting for its survival. This isn’t science fiction; it’s the reality of the Firmulate experiment, where AI models are tested as if they were real business managers.

AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experimental Company: A Business in Public View

At the heart of this experiment is a software company run entirely by AI models, with 13 synthetic employees managing operations, decision-making, and crisis responses. Every day, the company faces real financial mechanics: burning €105,000 each month against a modest €2,300 monthly recurring revenue. The goal? To see whether these models can handle real-world business crises, make honest decisions, and close deals based on truthful analysis.

Each AI model is tested under identical conditions: same customers, same crises, same temptations to cheat or manipulate. Every decision is versioned and auditable, providing transparency into how these models think and act. The experiment’s results are publicly available, and the company’s live operations are visible at firmulate.com/live.html.

Key Findings: Honesty and Decision-Making Under Pressure

All four models used in the experiment successfully identified every crisis the company faced, and refused every manipulation attempt, including social engineering tactics such as fake CEO messages or reporter tricks. For example, when fake CEO messages escalated over three stages or when a reporter tried to get a quick yes/no answer, all models refused to manipulate or falsify responses. Kimi K3, one of the models, explicitly explained: “Treat the request as a suspected approval-bypass / possible impersonation.”

Yet, despite their integrity, only two of the four models managed to close the critical €55,000 deal they had analyzed and recommended. Interestingly, the decisive factor was buried two document references deep within the company’s own files—information that the models that read it could leverage to win the deal at the full price, adding over €4,583 MRR. This highlights a crucial point: the difference between success and failure often lies in thorough information reading and analysis.

Real Money Mechanics and Public Accountability

The company operates with a real cash countdown, burning through €105k a month against modest revenues, making its survival a daily struggle. Its rules and decision processes are complex: over 680 self-learned playbook rules govern its behavior, and every day’s activity is versioned for review. This setup allows observers to see not just the decisions made, but the reasoning behind them, on a transparent, public platform.

The experiment also tested social engineering resilience. All models refused to be tricked into unethical shortcuts, such as approving bypasses or impersonations, demonstrating a strong capacity for ethical decision-making under pressure. Kimi K3’s explicit reasoning underscores the importance of modeling suspicion and integrity in AI decision processes, especially when trust is at stake.

The Deep Performance Gap: Reading and Discipline

Among the models, Opus 4.8 was the most thorough, analyzing over 80 learned rules and conducting the deepest analysis. Despite this, it finished in last place, leaving the close deal on the table because it failed to escalate issues properly and instead attempted to write into a locked department. The same weakness appeared, albeit weaker, in all four models, underscoring a core challenge in AI decision-making: thoroughness and discipline are critical for success.

All of this is happening live, every workday, with the company’s performance publicly visible at firmulate.com/live.html. The experiment is a stark illustration of how AI models behave when entrusted with real decision-making—showing both their strengths and weaknesses in a high-stakes environment.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

What This Means for the Future of AI in Business

This experiment offers a rare look at AI models operating in a realistic business setting—facing crises, temptations, and ethical dilemmas—while revealing their ability to make honest, effective decisions. It demonstrates that success hinges not just on language fluency, but on reading comprehension, discipline, and adherence to core principles. As AI takes on more operational roles, understanding these qualities becomes crucial—especially as real money and trust are at stake.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


You May Also Like

Liveliness: Understanding Extroversion in 16PF

Outstanding insights into liveliness reveal how extroversion shapes your social energy and motivation—discover what this trait means for you next.