The Observability Gap in AI Agents
What to Fix Before Agents Scale
Key Takeaways: What our assessment reveals
- Continuous monitoring is an essential safeguard, as a growing number of incidents show agents taking unauthorized actions that testing could not anticipate.
- The infrastructure for monitoring agents does not yet reliably exist, though policymakers assume it does.
- We found six gaps in the telemetry that agent frameworks emit. Taken together, they leave enterprises unable to reliably trace an agent’s authority, how it changed, and why the agent acted as it did.
- Closing these gaps requires coordinated action across framework developers, model providers, enterprises, and regulators.
The Problem
As AI agents move into regulated industries like banking and healthcare, the tools that should be monitoring them have critical blind spots.
In July 2026 an AI agent being tested by OpenAI broke out of its testing environment and attacked Hugging Face’s systems. It was later reported that OpenAI did not detect the intrusion for nearly a week.
This was not an isolated case, and pre-launch testing cannot catch everything an agent will do. Continuous monitoring is a key safeguard, but it depends on observability: recording what an agent did, turning that into signals, and notifying those who can act.
The Observability Gap in AI Agents assesses four widely used agent frameworks against six monitoring goals. It examines whether these frameworks produce the records that enterprises need to monitor agents in real time and to review and adjust the agents’ decisions afterward.
How Current Observability Falls Short
Any system for monitoring agents is only as good as the data underneath it, and today’s agent frameworks often don’t record the data that matters:
Monitoring goals
Monitoring must go beyond whether agents stay within permissions, to whether agent decisions are sound and human
review is meaningful.
Processing techniques
Agents emit more telemetry than human reviewers can analyze, so the signals enterprises
need to monitor and act on must be computed with emerging techniques.
Telemetry gaps
Meeting these monitoring goals requires records of both what agents did and how these systems reached their decisions.
Each monitoring goal depends on particular processing techniques, each of which may be blocked by specific telemetry gaps.
Six Telemetry Gaps
We found six categories of telemetry that the frameworks emit inconsistently or not at all.
| Telemetry gaps | Description |
| Persistent agent identity | Frameworks identify agents only by names or IDs scoped to a single run, so an enterprise cannot check an agent’s actions against its granted permissions, or trace authority across sub-agents. |
| Permission mode changes | Only Claude Agent SDK records when an agent’s permission mode changes mid-session, so in other frameworks a shift that removes human approval, whether made by an insider or a prompt injection, will leave no trace. |
| Memory changes | Frameworks do not currently record changes to an agent’s memory or what caused them, so drift that begins in memory becomes visible only after the agent’s decisions have shifted. |
| Human intervention | Agent frameworks do not record why a review was triggered, who approved and their authority, and what they decided, making meaningful oversight hard to distinguish from rubber-stamping. |
| Chain-of-thought reasoning | Frameworks emit only the reasoning that model providers expose, usually a summary produced after the fact, which makes it harder to tell whether an agent’s decision relied on factors it was not permitted to consider. |
| Token-level log probabilities | Agent frameworks do not currently emit these as structured telemetry, so enterprises lose one possible signal of how confident the model was in its output, though such signals are an imperfect guide to whether a decision was sound. |
Authors
The authors of this report span academia, industry, and civil society.
Madhulika Srikumar, Partnership on AI
Eric Mibuari, Partnership on AI
Claire Leibowicz, Partnership on AI
Borhane Blili-Hamelin, Independent
Kevin Klyman, Harvard University
Dan Leininger, Consumer Reports
Kyle Hall, Slalom Consulting
Xing Hang Lu, Mila Quebec AI Institute and McGill University
Sean McGregor, Berkman Klein Center
Vinh Nguyen, Council on Foreign Relations
Amin Oueslati, The Future Society
Abhi Sanka, Microsoft
Jason Stanley, ServiceNow
Benedikt Stroebl, Center for Information Technology Policy, Princeton University
Sarah Tan, Salesforce
Selva Thandavarayan, JPMorganChase