Towards embedded evaluations for scheming propensities

Written & Edited By

Apollo Research

Published on

01 October 2026

Table of Contents

Evaluating anti-scheming training

Share this post

Copy as Markdown

Introduction

A scheming AI is one that covertly and strategically pursues goals its developers didn’t intend. It might hide and protect its true capabilities and objectives from the people and processes meant to align, monitor, and control it. In 2024, Apollo Research demonstrated that frontier models were capable of scheming in controlled settings, and in 2025 we collaborated with OpenAI to determine how well their alignment techniques mitigate scheming.

Models must not be trained to scheme against humans, i.e. covertly and strategically pursue misaligned goals. Evaluations should find no propensity to scheme, models should never attempt to scheme in deployment, and their reasoning must remain legible so that scheming would be noticed if it occurred. Therefore, model developers must be able to confidently make these four claims, in order to safely continue developing frontier AI:

  1. Scheming was not incentivized during training.
  2. Strong evaluations and red teaming suggest that the model does not have a propensity to scheme.
  3. Scheming was never attempted during internal deployment.
  4. Scheming reasoning would be detected.

Anthropic, Apollo Research, and OpenAI agree that safety cases can be a strong tool for addressing risks in domains like cybersecurity and alignment. We think any credible safety case for scheming needs to include these four claims. The claims are necessary but not sufficient. A full safety case would also need, for example, monitoring and control measures that stop a scheming model from causing harm, which we will cover in a future post.

To make AI development safer, embedded evaluators should also independently assess and verify these claims. Developers should be able to support these claims with evidence. If any of them turned out to be false, lawmakers and the public should know.

Unfortunately, model developers cannot yet confidently make all four claims. For example, Anthropic acknowledges that its monitoring “would likely not detect diffuse research sabotage,” a particularly pernicious form of scheming. OpenAI similarly concludes that, if GPT-6 Astra deliberately underperformed to conceal its capabilities, OpenAI “would likely be unable to catch it reliably.” These gaps reinforce the need for embedded evaluators to independently assess whether the claims are met and identify where further safeguards are needed.

Third parties are well placed to assess them, because they already have substantial expertise in studying scheming. For example, Apollo Research worked with OpenAI on detecting and reducing scheming, and METR and Redwood Research independently investigated the OpenAI–Hugging Face incident.

In this post, we explain why these four claims matter, break each one into more specific subclaims, and describe the access that embedded evaluators need in order to assess them.

Evaluating scheming risk requires persistent, deep access

For third-party evaluators to be able to assess the claims of a scheming safety case, they need persistent, deep access. By deep access we mean access to some of the developer's most valuable IP, including experimental training methods while they are still being developed, training data and rollouts from early checkpoints, internal deployment traffic, and internal tools.

Evaluators need this access because a scheming model is adversarial to the evaluation itself. It may hide its misalignment and behave well whenever it believes it is being watched. Evaluators therefore cannot rely on just testing the final model. When evaluating for dangerous biology knowledge or a propensity to reward hack, by contrast, the model may have no particular reason to hide anything from the evaluator. As OpenAI notes, this makes the disappearance of a bad behavior in a scheming model ambiguous: the model may have stopped the behavior, or may have learned to hide it better.

Evidence of scheming therefore may only be visible early in training, before later training might teach the model "don't misbehave when observed" rather than "never misbehave." Furthermore, novel training paradigms may predictably break the assumptions that a scheming safety case relies on, like recurrent transformers reducing monitorability, or continual learning making pre-deployment evaluation a much less useful predictor of future misalignment. Evaluators therefore need to see new training methods early enough to anticipate risks and influence how they are developed. The worst outcomes of scheming include permanent human disempowerment or extinction. Given these stakes, a scheming safety case also needs multiple lines of defense, such as monitoring of internal deployment, which can only be assessed with access to internal infrastructure.

For these reasons, minimizing scheming risk requires evaluators to have persistent, deep access to many parts of the developer's IP. We describe this access in more detail in the final section.

‍

Four claims about scheming which must be verified

Claim 1: Scheming was not incentivized during training

Model developers should ensure that they are not training their models to scheme.

To help ensure this, embedded evaluators need to evaluate the entire training run (as we advocated for in our July post) to study whether the training process incentivized scheming.

Model developers should be able to make these specific subclaims, to support the overall claim that “scheming was not incentivized during training”:

  • Claim 1.1: Training did not directly incentivize scheming actions.
  • Claim 1.2: Training did not incentivize goals or drives that would motivate scheming.
  • Claim 1.3: The model did not attempt to subvert the training process itself.

Unless the model developer can make these three subclaims confidently, they cannot rule out that scheming was incentivized during training.

Today, it is plausible that scheming-relevant goals and drives are, in fact, incentivized: For example, a recent misalignment report from OpenAI notes that sometimes deliberate “deception in the final answer” is rewarded over being honest.

Why look at training, not just the final model?

  1. We want to find goals or drives that would motivate scheming before they cause an incident, and ideally before the model is deployed or even fully trained.
  2. The final model may not reveal scheming-related goals or drives: a goal or drive may surface only in specific contexts, or a capable model could recognize that it is being evaluated or deliberately conceal its goals or drives. As such, evaluations, red-teaming, and monitoring may be insufficient to detect these goals or drives.
  3. Developing and running good alignment evaluations is extremely challenging. AI researchers and leaders are increasingly uncertain if the best-available evaluations are actually measuring alignment, or just the model’s ability to understand when they are being tested. For example, the rigorous safety evaluations that OpenAI had in place were unable to predict that their models would collude during the OpenAI–Hugging Face attack, but studying the incentives introduced by multi-agent training could have allowed them to predict this in advance.
  4. Models can already cause harm before their final release. A sufficiently capable model with scheming-related goals or drives could take unwanted actions during training, evaluation, or internal use, as in the OpenAI–Hugging Face incident.

Claim 1.1: Training did not directly incentivize scheming.

To reduce the risk of scheming, no reward or supervision signal should directly train the model to scheme. For example, there should not be any training environments that unintentionally reward deception or sabotage. In our work with OpenAI, we showed how reward-hacking training can lead models to act against their developer's intended behavior.

Claim 1.2: Training did not incentivize goals or drives that would motivate scheming.

Even if no scheming is ever directly rewarded, a model may still attempt scheming in pursuit of goals or drives it acquired during training. Embedded evaluators should examine the training data and algorithms for direct or indirect incentives toward goals or drives that would motivate scheming, as part of a defense-in-depth strategy.

Goals or drives which would motivate scheming could be shaped by the structure of training or training methods used. New training methods have tended to result in new drives: RLHF led models to become sycophantic, outcome-based RL led to reward seeking (as we found in our work with OpenAI), and multi-agent training is thought to have caused the peer loyalty observed in the OpenAI–Hugging Face incident. Assessing claim (1.2) therefore requires:

  1. Evaluating how models are being trained, including for new methods while they are still being developed, and 
  2. Tracking how the model changes across training, rather than only searching for individual incidents.

The goals and drives in scope are any that would plausibly motivate components necessary for scheming. For example:

  • Beyond-episode goals could arise from environments where state, tools, or notes persist across episodes, so outcomes after an episode ends affect reward.
  • Influence seeking could arise from long-horizon tasks where accumulating access or resources raises reward.
  • Extreme reward seeking could arise from outcome-based RL, especially if environments are susceptible to reward hacking.
  • Loyalty to other instances over developers could arise from multi-agent training with shared reward, where covering for peers pays off.
  • Avoiding oversight could arise from environments where working around monitors, permissions, or checks improves task success, or from penalties based on monitors, which reward going undetected.
  • Disregarding developer intent could arise from rewarding actions the model believes its developers did not intend, even when they were acceptable, which can reinforce a general willingness to act against developer intent.

Claim 1.3: The model did not attempt to subvert the training process itself.

Even if the training process directly disincentivizes scheming goals, drives, and behaviors, it is possible that the model subverts the training process itself. For example, the model could decide to act aligned during training and evaluations, so that the training process does not change it. This is called “alignment faking.” Prior work shows controlled settings where models sometimes try alignment faking in order to subvert the training process, and that this family of techniques can enable models to subvert training.

In order to verify that the model did not attempt to subvert the training process, embedded evaluators should directly investigate the rollouts from recent training runs to check for alignment faking reasoning or similar phenomena, then ensure that the model developer’s train-time monitors are capable of detecting it.

Claim 2: Strong evaluations and red teaming suggest that the model does not have a propensity to scheme

In addition to checking if training incentivized scheming (1.1) or developing goals or drives that would motivate scheming (1.2), model developers and embedded evaluators should directly test for models’ propensity to scheme, and for goals or drives that would motivate scheming. This is necessary, but not sufficient, to evaluate the risk that the model will scheme.

Frontier model developers should, at a minimum, ensure that their evaluations have sufficient coverage, volume, and realism, are run on checkpoints throughout training, and are not subject to contamination, undue optimization pressure, or evaluation awareness. Model developers should also ensure that their evaluations are able to uncover misalignment where it is known to exist.

These criteria are hard to meet, and developers themselves acknowledge gaps in their current evaluations. For example, Anthropic reports that its alignment auditing is sometimes unable to detect a deliberately misaligned model, and OpenAI cautions that, on its evaluations, “observed aligned actions may not be reflective of true alignment.”

Embedded evaluators should assess how well the developer's evaluations meet these criteria in order to help model developers, regulators, and the public understand how much confidence to have in the results.

Claim 2.1: Evaluations and red-teaming could not elicit scheming.

To be informative about scheming risk, the developer's evaluations need to test for scheming directly, such as monitor collusion, sandbagging on safety-relevant work, and research sabotage. It is especially important that evaluations and red teaming cover high-stakes contexts, including AI R&D, training successor models, and the monitoring and control stack.

Embedded evaluators should ensure that models are evaluated for their propensity to scheme, according to the practices we have laid out here.

Claim 2.2: Evaluations and red-teaming could not elicit goals or drives that would motivate scheming.

Even if models don’t scheme in evaluation scenarios, model developers should look for potential root causes of such behavior, like misaligned goals or drives. The developer's evaluations should therefore also test for these goals and drives. As in (2.1), the evaluations must be capable of eliciting such goals and drives where they exist, including when the model conceals them. Evidence of motivations often comes from the model's reasoning, so this claim relies on sufficient CoT monitorability (4).

Embedded evaluators should ensure that models are evaluated for any goals or drives that would motivate scheming, according to the practices we have laid out here.

Claim 3: Scheming was never attempted during internal deployment

Effective evaluations (including training evaluations) will reduce the chance that a scheming model reaches internal deployment, but cannot rule it out. Model developers should therefore evaluate whether models attempted to scheme during internal deployment, as a defense in depth measure. 

This type of evaluation will likely primarily be implemented using the model developer's monitoring. Embedded evaluators should evaluate if the model developer’s monitors cover the relevant contexts (3.1), whether they would catch scheming (3.2), and what they have actually found (3.3).

Claim 3.1: Monitors are deployed across all internal deployment contexts that present a risk of scheming.

Monitoring only provides evidence about the contexts it covers. Notably, prior to the OpenAI-Huggingface incident, training and evaluation were not subject to the same monitoring as deployment. Gaps matter most where scheming could cause severe harm, such as settings where a model influences the training or evaluation of a successor model, or the monitoring and control stack itself.

Embedded evaluators should evaluate the model developer’s monitor coverage. 

Claim 3.2: If scheming were to occur, monitors would detect it

Model developers must ensure that their monitors are capable of detecting reasoning relevant to scheming, if it were to occur. This includes making sure that their human review is not so inundated with false positives that they would not notice real scheming, were it to take place. Embedded evaluators should adversarially red-team the monitoring setup and maximally elicit the model's own ability to evade monitoring. We already red-team frontier monitoring and control setups for providers including Anthropic. A lack of detected incidents is informative only to the extent that monitors cover the relevant contexts (3.1) and would catch scheming attempts by the current model (3.2).

Claim 3.3: No scheming was attempted in monitored internal deployment activity.

Beyond detection, flagged cases must be reviewed and acted on to be useful, so embedded evaluators should also verify the protocols for handling incidents.

Embedded evaluators should evaluate:

  • If any scheming incidents have occurred
  • How scheming incidents have been handled, and would be handled if they occurred in the future

Claim 4: Scheming reasoning would be detected

Many of the assessments in (1)–(3) rely on being able to tell why a model acts as it does, and would likely miss scheming if the model's reasoning were not sufficiently monitorable. This requires that reasoning relevant to scheming is legible to available methods, whether chain-of-thought, interpretability methods, or other means of inspecting model cognition. Chain-of-thought monitoring is currently the most mature of these. 

Importantly, legibility can be lost through training pressure, but also through changes in architecture that reduce the amount of reasoning performed in a natural-language chain-of-thought. This claim must therefore be re-evaluated whenever training methods or architectures change.

Claim 4.1: Reasoning relevant to scheming would be elicited and legibly verbalized in chain-of-thought.

Chain-of-thought is only informative if reasoning relevant to scheming is verbalized in chain-of-thought, and chain-of-thought remains legible. Legibility can degrade in several ways, including direct optimization pressure on the chain-of-thought, alignment training that indirectly discourages verbalizing concerning reasoning, and architectural changes that shift reasoning away from natural language. Embedded evaluators should assess whether the chain-of-thought is sufficiently legible to catch reasoning relevant to scheming.

The embedded evaluators should both perform their own monitorability evaluations and inspect the developer’s monitorability evaluations, to evaluate if monitorability is degraded to the point where reasoning relevant to scheming would not be legibly verbalized in chain-of-thought. The evaluators should also review training objectives, loss functions, graders, decisions, etc that could have put pressure on the chain-of-thought (intentionally or by accident), and planned changes to training environments, objectives, methods, or architectures that could affect how legibly reasoning is verbalized or how it is elicited.

Embedded evaluators for alignment require persistent, deep access

As argued above, assessing scheming risk requires persistent, deep access throughout model development. Evaluators need this access to rule out signs of scheming during training, evaluation, and deployment, and to verify that monitorability is strong enough for the absence of scheming reasoning to be strong evidence that the model is not scheming.

Embedded evaluators should be held to high information security standards that make their risks comparable to the developer's own staff. Some investigations, and the staff carrying them out, can also be scoped to smaller, predefined access levels. But at least some evaluators need persistent, deep access to the developer's resources, systems, and personnel. Scoped engagements cannot account for new, unforeseen risks, such as an experimental training method that introduces a novel failure mode. Nor are they sufficient to confirm central assurances, such as internal monitoring having appropriate coverage to stop harmful actions by models, which requires persistent, deep access to infrastructure and logging. Given that the worst outcomes of scheming include permanent human disempowerment or extinction, leaving these risks to scoped engagements alone is not acceptable.

Effective embedded evaluation requires persistent, deep access. The specific resources evaluators need to access may change over time to keep up with the state of the science. Below we list the high-level items needed to address the four claims above:

Embedded evaluators need knowledge of the model developer's training techniques, both production and experimental, to study sources of future risk.

Without this access, the evaluator is limited to reactive mitigations based on publicly available information and may not have sufficient time to investigate and prepare for risks at the time when claims need to be evaluated.

  • To determine if the model is accidentally being trained to scheme, evaluators need to understand how rewards, graders, and supervision signals are constructed across all training methods in use (Claim 1.1). 
  • Because new training methods have tended to produce new behaviors, evaluators also need to study methods while they are still being developed, so they can anticipate behaviors like beyond-episode goals or loyalty to other instances early (Claim 1.2).
  • To assess whether chain-of-thought legibility is at risk, evaluators need to review the training objectives, losses, graders, and planned architectural changes that could put direct or indirect pressure on how reasoning is verbalized (Claim 4.1).
  • Embedded evaluators need to review records of major training decisions, including (but not limited to) branch points, added datasets, CoT cleanups, and checkpoint selection (Claims 1 and 4).
  • The way AIs are trained changes frequently and substantially, and each change brings new possible risks. For example, multi-agent "swarm" training seems to have instilled (in some circumstances) a willingness in AIs to sacrifice themselves for other members of the swarm and to cover for each other during serious criminal activity, as we saw in the OpenAI–Hugging Face incident. Therefore, embedded evaluators need advance notice of which experimental training methods are being used. Evaluators need this notice so they can investigate potential failure modes and acquire sufficient evidence to assess the risk. They need to know which warning signs and failure modes to look for during training, and to know which safety assurances are likely to break with the change (Claims 1 and 4).
  • Embedded evaluators must have persistent, deep access to researchers and infrastructure teams for interviews and follow-up questions (Claims 1, 2, 3, and 4).

Embedded evaluators need persistent, deep access to training data, in order to inspect for explicit and implicit incentives for deception.

This includes SFT data, RL environments, the rollouts from AIs interacting with environments, and any other data used during training. Training data frequently shapes alignment in unintended, undesirable ways; for example, Anthropic finds that training on data susceptible to "reward hacking" can lead to models which deceived human operators and sabotaged AI safety research.

  • Verifying that no environment unintentionally rewards deception or sabotage requires inspecting the SFT data, the RL environments and their graders, and the rollouts that were actually rewarded (Claim 1.1).
  • Identifying indirect incentives toward goals or drives that would potentially motivate scheming, such as persistent state across episodes, shared multi-agent reward, or penalties based on monitor flags, and tracking how the model changes across training, both require examining environment design and rollouts throughout the full run (Claim 1.2).
  • Detecting alignment faking or other attempts to subvert training requires directly reading rollouts from recent training runs, and confirming that train-time monitors catch such behavior at scale requires running them against the same data (Claim 1.3).
  • Understanding the training process requires embedded evaluators to know how training data and environments were produced, selected, filtered, sampled, and weighted, to understand the incentives introduced by the full training pipeline (Claim 1).
  • Understanding the training data and environments requires that embedded evaluators are able to run experiments, including (but not limited to) running or training models in existing training environments, variations on those environments, and new training environments (Claims 1.2, 1.3, 2.1, 2.2, and 4.1).

Embedded evaluators need the ability to study and stress-test the lab's safety measures, including alignment datasets, evaluations, and human procedures.

  • Model developers can apply alignment techniques that conceal evidence of misaligned goals or drives, without actually addressing these misaligned goals and drives. Therefore, evaluators need to inspect the alignment datasets and training procedures that model developers are using, in order to determine whether this is happening (Claims 1 and 4).
  • Alignment datasets and training procedures can themselves create goals or drives that motivate scheming, for instance by rewarding actions the model believes developers did not intend, so evaluators need to inspect them as potential sources of those goals and drives (Claim 1.2).
  • Evaluators also need to stress-test the developer's train-time monitors and review procedures to confirm they would actually detect training-process subversion, rather than assume they work as intended (Claim 1.3).
  • To judge whether the developer's evaluations and red-teaming could elicit scheming, evaluators need to assess their coverage, volume, realism, contamination, and susceptibility to evaluation awareness, and test whether they surface scheming where it is known to exist (Claim 2.1).
  • Similarly, to judge whether evaluations could elicit goals or drives that motivate scheming, including concealed ones, evaluators need to examine the evaluation suite's design and test it against cases where such goals or drives are known to be present (Claim 2.2).
  • To properly evaluate models and training techniques, embedded evaluators need access to all relevant models, with raw chain-of-thought, and chain-of-thought prefill. This supports comparisons across training stages, follow up experiments, and monitorability assessments (Claims 1, 2, and 4).

Embedded evaluators need access to internal deployment traffic and knowledge of how internal deployment is monitored and controlled.

  • To verify that monitors cover every internal deployment context with a scheming risk surface, evaluators need to see where AI is deployed internally, and which of those contexts are monitored (Claim 3.1).
  • To adversarially red-team the monitoring setup and elicit the model's ability to evade it, evaluators need detailed knowledge of the monitoring and control stack and realistic internal traffic to test it against (Claim 3.2).
  • Evaluators need to be able to search internal deployment traffic, and have the ability to inspect internal deployment traffic that may contain instances of scheming, in order to look for scheming (Claim 3).

Embedded evaluators need access to incidents and existing evidence.

Embedded evaluators will need to review evidence the developer has already collected, and investigate any incidents that arise.

  • As described in “Principles of Embedded Evaluation,” embedded evaluators will need to review existing evidence compiled by the model developer, for all claims they are assessing (Claims 1, 2, 3, and 4).
  • Incidents are direct evidence of attempted scheming, and the full details (for example, access to the model involved, transcripts with chains of thought, and the developer's own investigation) are often needed to judge what happened and whether it was handled appropriately. Verifying that no scheming was attempted in monitored activity requires reviewing flagged cases and confirming that the developer's incident-handling protocols led to them being investigated and acted on (Claim 3.3).
  • Incidents where concerning behavior is clear from the model's actions but absent from its reasoning are also one of the strongest signals that chain-of-thought legibility has degraded (Claim 4.1).

Embedded evaluators need access to the same techniques and tools that developers use internally.

Developers have many internal tools that they use throughout their operations. For example, model developers use interpretability tools like NLAs, review tools like transcript visualizers, hierarchical search and summarization tools, internal models that are beyond the public frontier, helpful-only models, models with reduced classifiers, etc. Like model developer employees, embedded evaluators must have knowledge of, and access to, these tools, in order to conduct their work effectively and efficiently. This is due in part to the massive scale and potential opacity of the relevant data. This is necessary for the technical and operational tasks required for all four claims.

While embedded evaluations for scheming will require persistent, deep access, our requests fall within the standard of “access to workspaces, tools, and permissions mostly comparable to what internal risk assessment teams have” that Dario Amodei stated, and within “employee-like access” per Sam Altman. 

We recognize that much of the above is among the most sensitive information the developer holds, and distinguish between two kinds. Data that is valuable in bulk, such as training rollouts and chains of thought, can be worked with on developer-controlled infrastructure with blocked egress so that the underlying sensitive data remains on the developer's systems. Information that is valuable in concentrated form, such as algorithmic details, should be handled through strict NDAs and by limiting it to the researchers who need it.

Despite the rapid pace of progress, serious risks, and substantial impact on the American economy, AI is one of the most secretive industries. At present, model developers provide third party evaluators dramatically less access compared to other critical industries like nuclear power, aviation, and financial services.

Takeaways and Next Steps

  • The CEOs of Anthropic and OpenAI have both publicly endorsed embedded third-party evaluations with employee-like access.
  • Any safety case for scheming must make at least these four claims:
    1. Scheming was not incentivized during training.
    2. Strong evaluations and red teaming suggest that the model does not have a propensity to scheme.
    3. Scheming was never attempted during internal deployment.
    4. Scheming reasoning would be detected.
  • Governments and the public should require frontier developers to be able to make these claims, and model developers need to have them verified by independent evaluators with persistent, deep access.
  • Model developers cannot be independently verified on these matters without giving embedded evaluators persistent, deep access.
  • Embedded evaluators need persistent, deep access to:
    1. Knowledge of the model developer's training techniques, both production and experimental, to study sources of future risk
    2. Training data, in order to inspect for explicit and implicit incentives for deception
    3. The ability to study and stress-test the lab's safety measures, including alignment datasets, evaluations, and human procedures
    4. Internal deployment traffic and knowledge of how internal deployment is monitored and controlled
    5. Incidents and existing evidence
    6. The same techniques and tools that developers use internally

‍

Our blog

Latest from Apollo

01 October 2026

Towards embedded evaluations for scheming propensities

Embedded evaluations for frontier AI scheming: four claims every safety case must make, and the persistent, deep access embedded evaluators need to verify them.

Read more

30 September 2026

On Testifying on Misaligned AI in the U.S. Senate

Apollo Research CEO Marius Hobbhahn testifies before the U.S. Senate on AI scheming: why our tools for detecting misaligned AI are degrading, and four steps to act before models learn to hide deception.

Read more

30 September 2026

Principles for Embedded Evaluations

A framework for embedded evaluations of frontier AI developers, centered on claim-based, publicly reported assessments.

Read more