From Research to Product
AI Agent Monitoring
Our monitoring research aims to translate compute into scalable security. We are building Watcher, a monitoring tool for coding agents, to bring this frontier research into production.
highlights
Latest updates
July 23, 2026
What makes a good monitoring prompt?
We want to understand which principles work well for high-quality monitoring prompts. We use a 1-10 severity scale and full-trajectory scoring, as opposed to real-time binary prompts.
Product
July 13, 2026
Red-teaming auto mode: lessons from our first external monitor campaign with Anthropic
Apollo Research ran a pilot monitor campaign for auto mode, Anthropic's monitoring system for Claude Code agents that decides whether an agent's next action should be allowed or blocked.
Research
July 7, 2026
Evaluating LLM Calibration for Coding-Agent Monitoring
We evaluate 16 LLMs scoring coding-agent trajectories on a 1-10 severity scale across three datasets and five failure modes to assess their capabilities and calibration as coding agent monitors.
Research
May 8, 2026
A scalable monitoring research agenda
The goal of this agenda is to translate compute into security at scale. In the best case, we figure out how to spend on the order of $10-100M to produce extremely good monitors for coding agents.
Research
our findings
All research
A scalable monitoring research agenda
Research
compute, security at scale, coding agents, monitors, research agenda, scalable oversight, AI control, monitor training, roadmap, agent monitoring, Watcher
What makes a good monitoring prompt?
Product
severity scale, rubric, reasoning structure, trajectory scoring, ablations, calibration, failure modes, worked example, grading rubric, compression, chain of thought, MAE, Spearman, prompt engineering, prompt design, monitor design, LLM judge, best practices
Red-teaming auto mode: lessons from our first external monitor campaign with Anthropic
Research
Anthropic, Claude Code, auto mode, automode, red-teaming, red team, monitor campaign, monitoring system, allowed, blocked, agent action, pilot, adversarial testing, bypass, agent safety, external campaign
Evaluating LLM Calibration for Coding-Agent Monitoring
Research
severity scale, trajectories, failure modes, calibration, capabilities, datasets, 16 LLMs, monitors, scoring, benchmark, LLM judge, model comparison, grading, monitor models, evaluation
Apollo x Tailscale: Introducing “Watcher” for AI Oversight & Control
Product
Watcher, Tailscale, oversight layer, AI agents, agent monitoring, safety failures, security failures, liabilities, flags, product launch, partnership, security tool, monitoring tool, AI control, announcement, coding agents
Apollo’s product vision
Product
agent monitoring, coding agents, Claude Code, Codex, Cursor, real-time interventions, multi-agent teams, AGI, agent logs, Deep Research, remote worker, safety tools, autonomous agents, product strategy, roadmap, Watcher, agent security, vision


watcher
Your AI agent monitoring tool by Apollo Research
Watcher catches dangerous coding agent behavior before it becomes an incident.

