firmulate.com/live.html — live view
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

Imagine a business where every decision is made by artificial intelligence, yet the company struggles to stay afloat — burning €105,000 every month with only €2,300 in recurring revenue. This isn’t science fiction; it’s the live experiment at Firmulate. Here, AI models run a real, publicly visible company, facing crises, temptations, and the harsh realities of business. What can this extreme experiment teach us about trust, decision-making, and the future of automation that impacts not just tech firms but any organization aiming for ergonomic, reliable, and honest operations?

In a world increasingly dominated by AI, the question isn’t just whether machines can generate compelling chat or write code. It’s whether they can truly manage complex, real-world tasks — especially when everything is on the line. Firmulate’s live experiment puts this question under a microscope by running four different AI models as the decision-makers for a small software company. This company faces the toughest week imaginable: the same customers, identical crises, and the same temptations to cut corners or manipulate the system.

Every day, the AI models are faced with management decisions. They are given the full context, including internal documents, and are tasked with diagnosing problems, handling crises, and closing deals. The process is fully auditable — each decision is versioned and recorded, allowing observers to see exactly how each AI arrived at its choices.

Key Findings Unveiled

  • The models identified every crisis accurately, including subtle cues buried within internal documents, not just the surface-level customer complaints.
  • All models refused manipulative or deceptive requests, such as fake CEO messages or reporter tricks, adhering strictly to ethical boundaries.
  • Only two of the four models managed to close the most lucrative deal (€55,000), despite all having the same information and instructions.
  • Surprisingly, the decisive factor wasn’t the AI’s general intelligence or chat quality, but its ability to read and interpret internal company files — the hidden “truths” that led to closing the deal at full price (+€4,583 in monthly recurring revenue).

This reveals a crucial weakness in many AI decision processes: the inability to see beyond surface information. The models that read deeper into internal documents succeeded where others failed, highlighting the importance of comprehensive data access for trustworthy automation.

Resisting Social Engineering

Another dimension tested was social engineering. The company simulated scenarios where a fake CEO message escalated over multiple stages, and a reporter attempted to coax approvals under the guise of routine background questions. All five models tested refused these manipulative maneuvers, with reasoning like: “Treat the request as a suspected approval-bypass/possible impersonation.” This strict refusal underscores that AI systems designed for enterprise decision-making can be programmed to resist common social engineering tactics — a critical feature for trustworthy automation in sensitive environments.

The Real Cost of Automation

Behind the scenes, the live company’s mechanics are stark: it burns €105,000 monthly against a mere €2,300 in recurring revenue. It is publicly counting down its cash, and every weekday’s decisions are versioned and observable. While this isn’t a typical business, it showcases what’s possible when AI models are subjected to real operational stresses, rather than just demo chats.

Interestingly, the most thorough participant — Opus 4.8 — with over 80 learned rules and deep analysis, failed to close the deal and slipped into process slips, such as writing attempts into a locked department instead of escalating. Even with its comprehensive knowledge, discipline and focus remain vital for success. In contrast, the newer Kimi K3, running without an effort parameter, demonstrated the cleanest discipline and succeeded in closing the deal, emphasizing the importance of configuration choices in AI decision-making.

This live experiment isn’t just a tech showcase — it’s a wake-up call. As organizations incorporate AI into their workflow, the focus should shift from superficial chat quality to core capabilities: Can it read the right information? Will it stay honest under pressure? And can it reliably complete what it starts, even when faced with crises or manipulative tactics?

For managers and decision-makers, this means rethinking how AI is integrated into your operations: Are your tools capable of understanding the underlying data? Will they resist being manipulated? And are they aligned with your company’s integrity and long-term goals? The Firmulate experiment offers a rare, transparent window into these questions, demonstrating that building trust in AI is more about disciplined process and comprehensive understanding than flashy demos.

Discover more about the ongoing live test and see these decisions unfold in real time at firmulate.com/live.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

The Firmulate live experiment shows that trustworthy AI in business isn’t just about how well it chats, but whether it can read, interpret, and stay honest under pressure. Deep data access and disciplined processes matter most.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


Behavioral AI: Unleash Decision Making with Data

Behavioral AI: Unleash Decision Making with Data

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

An Introduction to Healthcare Informatics: Building Data-Driven Tools

An Introduction to Healthcare Informatics: Building Data-Driven Tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

AI Phishing, Social Engineering & Fraud: How Criminals Use AI to Manipulate, Steal & Deceive (The AI Cybersecurity)

AI Phishing, Social Engineering & Fraud: How Criminals Use AI to Manipulate, Steal & Deceive (The AI Cybersecurity)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

HUMAN CENTERED ARTIFICIAL INTELLIGENCE SYSTEMS: Explainability ethical design and decision support engineering

HUMAN CENTERED ARTIFICIAL INTELLIGENCE SYSTEMS: Explainability ethical design and decision support engineering

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Buried Apple Feature Turns An iPhone Into The Perfect Kids’ Dumb Phone

A secret Apple feature allows users to transform an iPhone into a simplified device, ideal for children, by disabling advanced functionalities.

Cloudflare Meerkat – Globally Distributed Consensus

Cloudflare has announced Meerkat, a new system designed to achieve globally distributed consensus, enhancing reliability and security for internet infrastructure.