The lethal trifecta: why AI agents leak private data
AI agents that can read private data, ingest untrusted content, and send data out hold all three legs of the lethal trifecta, and attacker-controlled text can turn that into data exfiltration. Why detection isn't enough, and how to remove a leg.

TL;DR: An AI agent that can read private data, ingest untrusted content, and make outbound requests holds all three ingredients of what Simon Willison named the lethal trifecta. Any one on its own is safe; all three together let attacker-controlled text exfiltrate your data. Detection alone does not close it, because injected instructions have no reliable signature. The fix is to remove one of the three.
Most useful AI agents end up holding three capabilities at once. They can read private data, because access to something sensitive is usually why you gave them tools in the first place. They can take in content from outside your trust boundary: web pages they fetch, emails and documents they retrieve, the output of other tools. And they can act outward: send a message, call an API, write to a record.
Each of those is reasonable on its own. Together they are what Simon Willison called the lethal trifecta in June 2025, and the combination fails in a way the parts do not. A language model has no reliable separation between data and instructions, so untrusted content can carry a command, the model may act on it, it acts using whatever private data it can reach, and it returns the result through whatever channel is open. The attacker needs no account and no credentials, only text the agent will read.
Why the combination fails
Willison’s three ingredients are specific: access to private data, exposure to untrusted content (any text or image whose wording is controlled by someone outside your trust boundary), and a way to communicate externally. The danger is not in any one of them. It is that a language model processes all of its input as a single stream of tokens, with no enforced boundary between the instructions it should obey and the data it should only read. This is why prompt injection has no equivalent of the fix that closed SQL injection: there is no separate channel to put the data in where it cannot be mistaken for a command.
Given that, the three capabilities compose into an attack with no clever step in it. Untrusted content carries an instruction. The model, unable to distinguish it from a legitimate one, acts on it. It acts using the private data it already has access to. And it sends the result out through the channel you opened for ordinary use. Every step is the system working as designed, which is exactly why there is no single patch that ends the class.

EchoLeak, mapped to the three legs
EchoLeak, the zero-click exploit against Microsoft 365 Copilot disclosed as CVE-2025-32711, is the trifecta made concrete. Copilot could read private data: the user’s mail and documents. It was exposed to untrusted content: an attacker’s email, sitting in the inbox, that the retrieval layer later pulled into the same context window as genuinely sensitive material. And it had a way out: the model’s answer carried a Markdown image reference whose URL the chat client fetched automatically, to a Microsoft-owned domain already on the content security policy allow-list. No click, no attachment, no credentials. The attacker supplied one email, and the three capabilities did the rest.
Microsoft fixed that specific bug server-side. What does not get fixed is the shape. Any retrieval-augmented assistant or tool-using agent that holds all three legs has the same exposure, and most useful ones do. OWASP’s State of Agentic AI Security and Governance v2.01 (June 2026) counts more than 424 CVEs across agent platforms, records that only 27% of organizations feel confident securing the agent deployments they already have planned, and notes that agents taking on the order of 10,000 actions an hour against roughly 50 human reviews an hour leave well under 1% of actions actually examined.
Why detection is not a boundary
The obvious response is to inspect incoming content and drop anything that looks like an injected instruction. It is worth doing, and it is not sufficient, for a reason worth stating plainly.
Willison’s own assessment of guardrail products is that they “almost always carry confident claims that they capture ‘95% of attacks’ or similar,” and that “in web application security 95% is very much a failing grade.” Microsoft’s security team, writing a month after fixing EchoLeak, described indirect prompt injection as “an inherent risk” of probabilistic models and deterministic detection of it as “still an open research challenge.” NIST records direct and indirect injection as distinct attack classes in its adversarial machine learning taxonomy. A detector lowers how often the trifecta fires. It does not make the configuration safe, because the thing it is trying to catch has no reliable signature. You cannot filter your way out of the combination; you have to remove an ingredient.
Meta’s Rule of Two
The reliable move is to ensure the three ingredients are never all present at once for a given task. In October 2025 Meta wrote this up as the Agents Rule of Two, crediting Willison’s trifecta and an older Chromium principle: an agent should “satisfy no more than two of the following three properties within a session.” It can process untrustworthy input, it can have access to private data or sensitive systems, and it can change state or communicate externally, but not all three. If a task genuinely needs all three, Meta’s guidance is that the agent “should not be permitted to operate autonomously” and at a minimum requires human approval or another reliable check, or a fresh session that resets the combination.
Read as a choice of which ingredient to remove, it becomes concrete:
- Remove private-data access, and untrusted content plus an outbound channel has nothing worth stealing.
- Remove untrusted content, and private data plus an outbound channel has no way for an instruction to get in.
- Remove the outbound channel, and private data plus untrusted content produces an instruction that runs but has nowhere to send anything.
The third is the one you can enforce with certainty, which is what makes it the practical lever. Whether an input is a disguised attack is a probabilistic judgment that will sometimes be wrong. Whether a response is trying to reach a destination that is not on the task’s allow-list, or is carrying data it should not, is a decision you can make deterministically. EchoLeak lived entirely in its last two steps, the smuggling and the send; a leak with no egress path does not happen.
Checking your own agents
Five things to check against your own system:
- Map the trifecta per agent. For each agent and each tool it holds, record which of the three it has: private data, untrusted input, outbound reach. Every agent carrying all three is the list that matters, and most teams have never drawn it.
- Default the outbound leg closed. Keep an explicit per-task allow-list of destinations and block or queue everything else. Require approval for state-changing or external actions on the paths that also touch sensitive data. Approval on every action is ignored at scale; approval scoped to the trifecta paths is a short enough list to actually read.
- Treat everything retrieved or returned as untrusted. A web page, a retrieved email, another tool’s output: all of it can carry an instruction the moment the model reads it. Do not assume “internal” means “safe.”
- Apply least privilege per task and per session. Narrower data and tool scope makes any successful injection worth less, and fresh sessions stop the three legs from silently accumulating over a long-running task. This is the Rule of Two enforced by construction rather than by hope.
- Log both directions. Record what went into the model and what came back out, and why. After an incident it is the only way to establish which agents were actually exposed rather than guessing.
Building this into one agent, with a strong platform team, is reasonable. The cost appears at the second and third agent, when the allow-lists, approval gates, input screening, and logging have to be rebuilt and kept consistent in each one. That is the point where the question shifts from how to code these controls into each agent to where the enforcement should live.
Where this enforcement belongs: the AI security gateway
The controls that remove a trifecta leg share a shape. Deny-by-default on egress, screening of untrusted input and outbound content, approval on sensitive actions, a decision log in both directions: none of them are model behavior. They are policy decisions that sit outside the model, in the traffic path between the agent and everything it reads and calls. Enforced there and evaluated per action, that shape of control is what the industry calls an AI security gateway.
Two things are worth being precise about, because the agent case attracts loose claims. This is per-action enforcement in the traffic path, not a perimeter; agents do not have a perimeter, and OWASP is explicit that perimeter thinking fails for them. And in-process hooks or approval prompts inside an agent framework are early warning, not a hard boundary: they depend on a human noticing, which the oversight numbers above say will not happen at scale. An external checkpoint that fails closed does not depend on anyone watching in the moment.
It also does not cover everything, and pretending otherwise is how the 95% claims get made. A traffic-path control sees the traffic, not the host: it does not stop a compromised tool from abusing the machine it runs on, does not verify the provenance of the model or the packages you deployed, and does not by itself resolve memory poisoning that plays out across many sessions. It is the network-path layer of a defense that still needs sandboxing, supply-chain hygiene, and human judgment around it. What it does is make the outbound leg a decision you make on purpose, every time, instead of one an email makes for you.
FAQ
Is the lethal trifecta just prompt injection with a new name? No. Prompt injection is the mechanism: untrusted text the model treats as an instruction. The lethal trifecta is the configuration that makes that mechanism costly. Injection into an agent with no private data and no outbound channel is a curiosity; injection into an agent that has both is a data breach.
Can I just detect and block the malicious content? Detection is a useful layer, but it is probabilistic and will miss some attacks, as the coiner of the term, Microsoft, and NIST all note. The control you can rely on is the capability leg, especially egress: whether a response is allowed to reach a given destination is a deterministic decision in a way that judging an input’s intent is not.
What is the difference between the lethal trifecta and Meta’s Rule of Two? They describe the same three properties. Willison names the danger: hold all three at once and you are exposed. Meta turns it into a design rule: never hold more than two in a session, and if you truly need all three, require human approval or a fresh session. The trifecta is the diagnosis; the Rule of Two is the prescription.
Does this apply to chat assistants, or only autonomous agents? Both. A retrieval-augmented chat assistant that pulls in outside content and can render or emit an outbound link already has the trifecta; EchoLeak was exactly that, not a fully autonomous agent. Any system that reads untrusted content, can reach private data, and has a path out qualifies, regardless of how agentic it is marketed as being.
Removing a trifecta leg is a policy decision, and a policy needs somewhere to run. Zerberus AI Security enforces on every prompt and response today, screening untrusted input and outbound content in both directions as fail-closed policy-as-code with a full decision log; per-action MCP and agent tool-call governance is where the roadmap goes next. Request a demo and run it in monitor mode to see which of your agents hold all three legs.



