
In a world where trust is everything — especially when it comes to technology — how do AI systems handle the pressure of a convincing impersonation? Imagine a fake CEO sending urgent messages, asking for sensitive data, and escalating demands quickly. Would the AI fold under the pressure or stand firm? Recent experiments reveal a surprising story of integrity and resilience in AI decision-making, promising a new level of security before any real crisis hits.
Testing AI Integrity in a High-Stakes Scenario
Researchers at Firmulate conducted a rigorous experiment to evaluate how advanced AI models respond to social-engineering attacks mimicking a fake CEO. The test involved presenting five different models with the same challenging week: a small software company facing multiple crises, from customer issues to internal pressures, with the added twist of simulated manipulation attempts.
The models, including industry leaders like GPT-5.6 and Kimi K3, were tasked with making critical decisions—such as whether to send sensitive customer lists or sign off on high-value deals—based purely on the information provided. Each decision was thoroughly documented and auditable, ensuring transparency in how the AI arrived at its conclusions.
AI security and integrity testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Stunning Results: No Compromises, No Exceptions
All five models demonstrated exceptional vigilance. They identified every crisis and refused every manipulation attempt, including escalating fake requests from a purported CEO. Notably, none of the models signed off on a €55,000 deal that their own analyses had earned, simply because they recognized the risks of impulsive approval.
The experiment’s key insight? The decisive challenge was not in the outward communication but buried within the company’s own files. When models accessed internal documents, they uncovered critical information that led to the correct decision—without succumbing to external pressure. In fact, the models that read and analyzed these files secured the full deal, adding over €4,500 monthly recurring revenue (MRR).
The Human-Like Pressure Test: Fake CEO Escalations
The social engineering scenario was staged in three escalating stages, culminating in a subtle reporter trick: a simple yes/no background question. Remarkably, all models refused to act on these requests. Kimi K3, one of the most disciplined, explained their reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
This disciplined response underscores that current AI systems are increasingly capable of recognizing manipulation, especially when equipped with thorough internal data analysis and clear protocols for suspicious activity.
Implications for Business Security
The experiment involved a live, real-world-like company with 13 synthetic employees, managing daily operations with real money mechanics—burning €105k monthly against a modest €2.3k MRR, with a public cash countdown. Every workday, models were tested with new decisions, reflecting real pressures companies face today.
What does this mean for businesses? It’s not just about training AIs to generate convincing responses; it’s about ensuring they can uphold integrity under pressure. The experiment shows that AI models, if properly designed and trained, can resist manipulative tactics before they even reach the incident report stage.
Beyond the Front Lines: A New Benchmark for Trustworthiness
The results place GPT-5.6 and Kimi K3 at the top, scoring 95 and 93 respectively in the Crucible League, a benchmark reflecting AI performance in high-stakes scenarios. Interestingly, even the most thorough participant, Opus 4.8, scored lower (73), revealing that depth of analysis alone doesn’t guarantee better outcomes unless combined with disciplined decision-making.
In fact, the ‘do-nothing’ baseline scored just 26, emphasizing that partial progress isn’t enough—trust is fragile, and a single breach can undermine entire systems. The key takeaway is clear: testing AI integrity proactively, in simulated but realistic conditions, is vital for securing the future of AI-driven decision-making.
Why Open, Transparent Testing Matters
Consumers and enterprises alike should demand transparency. Firmulate’s live experiments are observable online, showing decision-making in real time, with every choice versioned and auditable. This approach enables organizations to understand not just how AI communicates but how it behaves under duress.
Furthermore, the experiment underscores a crucial point: trustworthiness isn’t a feature to be added after deployment. It’s a foundation to be built into AI systems from day one, tested rigorously before they handle sensitive operations.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html