From Research to Product

AI Agent Monitoring

Our monitoring research aims to translate compute into scalable security. We are building Watcher, a monitoring tool for coding agents, to bring this frontier research into production.

highlights

Latest updates

July 23, 2026

What makes a good monitoring prompt?

We want to understand which principles work well for high-quality monitoring prompts. We use a 1-10 severity scale and full-trajectory scoring, as opposed to real-time binary prompts.

Read more

Product

July 13, 2026

Red-teaming auto mode: lessons from our first external monitor campaign with Anthropic

Apollo Research ran a pilot monitor campaign for auto mode, Anthropic's monitoring system for Claude Code agents that decides whether an agent's next action should be allowed or blocked.

Read more

Research

July 7, 2026

Evaluating LLM Calibration for Coding-Agent Monitoring

We evaluate 16 LLMs scoring coding-agent trajectories on a 1-10 severity scale across three datasets and five failure modes to assess their capabilities and calibration as coding agent monitors.

Read more

Research

May 8, 2026

A scalable monitoring research agenda

The goal of this agenda is to translate compute into security at scale. In the best case, we figure out how to spend on the order of $10-100M to produce extremely good monitors for coding agents.

Read more

Research

our findings

All research




Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

A scalable monitoring research agenda

Research

compute, security at scale, coding agents, monitors, research agenda, scalable oversight, AI control, monitor training, roadmap, agent monitoring, Watcher

What makes a good monitoring prompt?

Product

severity scale, rubric, reasoning structure, trajectory scoring, ablations, calibration, failure modes, worked example, grading rubric, compression, chain of thought, MAE, Spearman, prompt engineering, prompt design, monitor design, LLM judge, best practices

Red-teaming auto mode: lessons from our first external monitor campaign with Anthropic

Research

Anthropic, Claude Code, auto mode, automode, red-teaming, red team, monitor campaign, monitoring system, allowed, blocked, agent action, pilot, adversarial testing, bypass, agent safety, external campaign

Evaluating LLM Calibration for Coding-Agent Monitoring

Research

severity scale, trajectories, failure modes, calibration, capabilities, datasets, 16 LLMs, monitors, scoring, benchmark, LLM judge, model comparison, grading, monitor models, evaluation

Apollo x Tailscale: Introducing “Watcher” for AI Oversight & Control

Product

Watcher, Tailscale, oversight layer, AI agents, agent monitoring, safety failures, security failures, liabilities, flags, product launch, partnership, security tool, monitoring tool, AI control, announcement, coding agents

Apollo’s product vision

Product

agent monitoring, coding agents, Claude Code, Codex, Cursor, real-time interventions, multi-agent teams, AGI, agent logs, Deep Research, remote worker, safety tools, autonomous agents, product strategy, roadmap, Watcher, agent security, vision

watcher

Your AI agent monitoring tool by Apollo Research

Watcher catches dangerous coding agent behavior before it becomes an incident.

Meet Watcher