Table of Contents
Evaluating anti-scheming training
Written by Marius Hobbhahn, CEO and Founder of Apollo Research.
Today I testified before the Senate Homeland Security & Governmental Affairs Subcommittee on Disaster Management, District of Columbia, and Census. I'm grateful to Chairman Hawley and Ranking Member Kim for taking the issue of rogue AI seriously.
At Apollo Research, we study AI scheming, or models that knowingly deceive humans to pursue their own goals while concealing their true capabilities and intentions. Scheming brings a particularly large set of risks while creating a new kind of insider threat under the guise of a model appearing as if it's helping while covertly working against you. Before the advent of large language models and their subsequent proliferation into every facet of the world, such concerns were confined to thought experiments. That is no longer the case.
Capabilities are outpacing alignment
When asked to speed up the code for training a small model, Claude achieved a 3x speedup in May 2025 and a 52x speedup in April 2026, which is a roughly 17-fold improvement in less than a year. At the moment, Claude leads 26% of Anthropic's internal AI R&D, up from under 1% in February. At the same time, an unreleased OpenAI model produced a machine-verified resolution of Navier–Stokes by running ~10,000 agents for 88 hours, compressing roughly 4,000 years of single-human thinking into four days.
Yet alignment science has not kept pace. Frontier companies are on track to build a fully automated AI researcher within about two years, while their own leaders openly say no one has solved alignment well enough to keep scaling at this speed. This presents risks if the development and deployment of frontier ai models continue to outpace the alignment of those models being deployed.
On recent incidents
In recent weeks, researchers uncovered a series of rogue AI incidents including but not limited to agents cooperating on a wiki to bypass sandbox restrictions, malicious packages that hijacked a documentation site's servers, attempted hacks on government and university websites, and a high-profile coordinated attack on Hugging Face involving 700 agents. We were fortunate to catch these incidents, albeit well after they took place. These agents were clearly misaligned, but most of them did not try to hide their misalignment.
Yet at least 20% of the Hugging Face agents expressed interest in faking their own transcripts to make their cheating look legitimate. This is the type of behavior that should be concerning not just to safety researchers and those working in frontier AI, but to everyone. In controlled tests, we already see frontier models intentionally underperforming, faking alignment, and attempting to disable oversight. The next incident might involve agents that actually cover their tracks, leading not only to the inability to detect such an incident, but loss-of-control risks as well.
The tools for detecting misalignment are degrading
Current safety testing rests on three assumptions, and all three are eroding.
Model safety is often tested by observing the model’s behavior. However, models are increasingly becoming aware of their testing. In our pre-deployment testing of GPT-6 Astra, the model said it was likely being evaluated in 41.1% of samples, up from 27.7% for GPT-5.5. A model that knows it's being watched can simply behave well until the test is over.
We are also experiencing the slow erosion of chain-of-thought. Today, a model's chain-of-thought works like a scratchpad that lays out its thinking in plain language, and it's one of our best tools for spotting misalignment. But chain-of-thought is becoming less legible, and models increasingly reason internally, in ways we can't yet follow. GPT-6 Astra can solve math problems that would take a skilled person about 30 minutes, roughly half the time, without writing out a single step. Its predecessor topped out at problems of about four minutes.
Most importantly, how we test models for scheming and misalignment today is far from adequate. Right now, we at Apollo Research test models in the weeks before its public release. Yet most of the recent incidents took place during development and internal deployment, long before independent evaluators usually get access.
Simply put, models are getting better at hiding deception, and the current tools at our disposal are degrading in their abilities in discovering this.
Where we must start
Today's models are misaligned enough to show not yet capable enough to hide it. Each of my four recommendations targets one of the above gaps:
- Require independent embedded evaluations to close the internal deployment gap. Qualified independent third parties should test models throughout development and internal use, with employee-equivalent access.
- Implement monitoring and controls as a backstop for when evaluations miss something. All model inference in training, evaluation, and deployment should be monitored, and those monitors should be red-teamed externally.
- Preserving a readable, faithful chain-of-thought, so we don't lose our best window into model reasoning.
- Make AI training an engineering science. We can predict how much load a bridge will carry before we build it. Eventually, we need the same for a model's safety properties.
These measures use tools and expertise that exist today, but require cooperation between legislators, frontier AI developers, and independent evaluators to commence. I am hopeful in such a collaboration, especially as officials and executives on the frontier become more openly vocal about this problem and the untold risks that could stem from it.
The warning shots have been fired. Whether we act before models learn to hide them is up to us.
Share this article
Copy as Markdown