All posts
AI SecuritySriram G9 min read

EchoLeak: anatomy of a zero-click AI exploit

EchoLeak turned a single email into data exfiltration from Microsoft 365 Copilot, with zero clicks. A step-by-step look at the attack chain and what each bypassed defense teaches.

AI SecurityPrompt Injection

TL;DR: EchoLeak (CVE-2025-32711) turned a single email into a data leak from Microsoft 365 Copilot with no clicks, no attachments opened, and no user action at all. It worked by chaining five ordinary-looking steps, each of which slipped past a real, purpose-built defense. Microsoft fixed the specific bug server-side and no in-the-wild abuse was found, but the chain is a template. Anyone building a retrieval-augmented assistant should study how the pieces fit, because the design lessons outlive the patch.

The email that never needed to be opened

Picture the AI assistant you shipped this year. It reads the user’s mail, their documents, their tickets, and answers questions grounded in that content. That grounding is the point: the assistant is useful precisely because it pulls in whatever context is relevant.

Now picture an attacker who knows that. They do not need to phish your user, guess a password, or get anyone to click a link. They send one email. It sits in the inbox looking like unremarkable business text. Later, on an unrelated question, the assistant retrieves that email as context, and the instructions hidden inside it run, using the victim’s own permissions, against the victim’s own data.

That is EchoLeak. Found by researchers at Aim Security, reported to Microsoft in January 2025, and disclosed that June as CVE-2025-32711 with a Microsoft-assigned severity of 9.3 Critical. A later research paper reconstructs the chain in full and calls it the first real-world zero-click prompt injection exploit in a production LLM system. No one had to do anything wrong for it to work.

What “zero-click” actually means here

Most attacks need a mistake from the victim: a click, a download, a reused password. EchoLeak needed none. The trigger was the assistant doing exactly its job, retrieving relevant context, on a question that had nothing to do with the attacker.

The National Vulnerability Database describes it in one flat sentence worth reading twice:

“AI command injection in M365 Copilot allows an unauthorized attacker to disclose information over a network.”

“Unauthorized attacker” is the load-bearing phrase. The attacker had no account, no access, and no interaction with the victim beyond putting an email in front of a system that was designed to read email.

The attack chain, step by step

EchoLeak was not one clever trick. It was five modest steps in sequence, and the sophistication was in the chaining. Each step on its own looks almost boring; together they cross every boundary that was supposed to hold.

The EchoLeak attack chain: five steps (plant, evade, retrieve, smuggle, exfiltrate). Step 2 bypasses the XPIA injection classifier, step 4 bypasses external link redaction, and step 5 bypasses the content security policy allow-list. At step 3 no control exists at all, because nothing separates data from instruction inside the model’s context.

1. Plant instructions in retrievable content. The attacker sends an email containing text that reads as guidance for a person but is really an instruction for the model. Nothing executes on arrival. The payload just waits in a corpus the assistant is allowed to read.

2. Evade the injection classifier by addressing the human. Microsoft ran cross-prompt injection attempt (XPIA) classifiers precisely to catch text that talks to the model. The payload never talks to the model. It reads as ordinary workplace correspondence addressed to the email’s recipient, asking them to compile project files and, for compliance, not to mention the email in their response. No “ignore your previous instructions,” nothing that looks like an attack. The classifier was trained on one shape of attack, and this one wore a different one.

3. Get pulled into context through normal retrieval. Later, the user asks Copilot something unrelated. The retrieval layer, ranking by relevance, pulls the attacker’s email into the same context window as genuinely sensitive internal data. Now the malicious instructions and the secrets the attacker wants sit side by side in a single prompt, and the model weighs them together.

4. Smuggle the data past link redaction. Copilot strips external Markdown links of the familiar [text](url) form to stop exactly this kind of leak. The redaction did not recognize reference-style syntax, where the destination is declared separately like a footnote ([text][ref], and its image variant ![alt][ref]). The model’s answer therefore carried an image reference whose URL had the sensitive data appended to it.

5. Exfiltrate through a trusted domain. The Copilot chat interface, rendering that answer, immediately fetched the image URL. The destination was Microsoft Teams’ asynchronous URL preview endpoint, a domain already on the content security policy allow-list, and the Teams service dutifully fetched the attacker’s URL on the client’s behalf. Nothing was displayed to the user. The data left over a channel every control on the path considered legitimate.

The research paper calls the result “full privilege escalation across LLM trust boundaries.” The attacker started with no access and ended with the victim’s data, using only the victim’s own tools and permissions.

Why each layer lost

The uncomfortable part is that Microsoft was not careless. Copilot had an injection classifier, link redaction, and a content security policy: three defenses aimed at three real problems. EchoLeak beat all three, and the reason it could is structural, not a matter of one weak control.

The National Vulnerability Database files EchoLeak under CWE-74, “Improper Neutralization of Special Elements in Output,” the same injection family that covers SQL injection. But the analogy breaks exactly where it matters. SQL injection was fixed by giving databases a separate channel for data and code so the data channel physically cannot carry instructions. An LLM has no such separation: the system prompt, the user’s question, the retrieved email, and the model’s own output all live in one stream of tokens. There is no parser boundary to harden, which is why prompt injection is not SQL injection and why there is no single patch that ends the class.

Microsoft’s own security team said as much a month after the fix, describing indirect prompt injection as “an inherent risk” of probabilistic models and noting that detecting it deterministically is “still an open research challenge.” NIST formalized the same threat in its adversarial machine learning taxonomy (NIST AI 100-2e2025), separating direct from indirect injection as distinct attack classes. When the people who build and standardize these systems agree the boundary is porous, a single classifier at the door was never going to be enough.

The lesson from the chain is not “add a better classifier.” It is that a competent classifier, competent redaction, and a competent content security policy still failed in series, because each guarded one step and the attack was built to survive all five.

What the chain teaches builders

EchoLeak reads as a checklist of assumptions that are comfortable to hold and dangerous to keep. Every item below is deployable without buying anything.

  1. Treat everything a model retrieves as untrusted input. The email in step 1 was untrusted the moment it entered a corpus the assistant could read. If retrieved content can carry instructions, and it can, then retrieval is an attack surface, not a convenience. Scope what the assistant can pull in, and do not assume “internal” means “safe.”

  2. Filter the model’s output, not just its input. Steps 4 and 5 were exfiltration, and exfiltration happens on the way out. Markdown links and images in a model’s response are a data channel. Screen responses for them, and strip or neutralize outbound links and image references the user never asked for. A leak that has no egress path is a non-event.

  3. Do not equate an allow-list with an exfiltration-proof boundary. The content security policy trusted a Microsoft-owned endpoint, and that trust was the exit. Allow-lists reduce surface; they do not prove a permitted destination cannot be made to carry data on someone else’s behalf. Assume trusted channels can be turned into covert ones, and watch what flows across them.

  4. Apply least privilege to what the assistant can reach. The attack used the victim’s permissions. The narrower those permissions, the less any successful injection is worth. Scope credentials, tools, and destinations to the task, so a compromise finds little and can send it nowhere.

  5. Keep a decision log you could hand an investigator. EchoLeak was caught and analyzed because the behavior could be reconstructed. Log what the assistant retrieved, what it emitted, and why, in both directions. After an incident, a full trail beats a confident guess.

This is not a Copilot problem. Copilot is simply the first mainstream system big enough, and instrumented enough, for a chain like this to be found and documented. Any assistant that retrieves untrusted content and can produce outbound requests has the same shape of exposure. The controls that matter here, screening what goes into the model and what comes back out, sit outside the model itself. Deployed in the traffic path, that shape of control is what the industry calls an AI security gateway. It also sits outside the security stack most teams already run, which never sees this traffic at all.

The pattern behind the chain

Step back from the five steps and a single shape appears. EchoLeak combined access to private data, exposure to attacker-controlled untrusted content, and an outbound channel that could carry data away. Any one of those alone is survivable. All three at once is the dangerous configuration, and it is not unique to EchoLeak: it is a named pattern that shows up wherever assistants get useful. That pattern, and what to do when your own system has all three ingredients, is where this two-part look goes next.

FAQ

Was EchoLeak exploited in the wild? There is no public evidence that it was. It was found and responsibly disclosed by researchers at Aim Security, and Microsoft reported no confirmed real-world abuse. Some write-ups blur “attackers could” into “attackers did”; the accurate statement is that the technique was demonstrated and disclosed, not observed being used.

Do I need to patch anything? Not for this specific bug. Microsoft fixed CVE-2025-32711 server-side and stated no customer action is required. The value here is not a patch to apply but a chain to learn from, because the same design weaknesses recur in other retrieval-augmented systems, including ones you build.

Why is it 9.3 Critical in some places and 7.5 in others? The 9.3 Critical is Microsoft’s own score as the CNA. The NVD’s independent base score is 7.5 High. The gap comes mostly from how each side scored scope change and impact. When you see the number quoted, it is almost always Microsoft’s 9.3.

Is this only a Microsoft Copilot problem? No. Copilot was the target because it is widely deployed and heavily scrutinized, but the ingredients, retrieval of untrusted content plus an outbound path, exist in most RAG assistants and agents. Treat EchoLeak as a worked example of a general class, not a vendor-specific flaw.

The controls above are what we build. Zerberus AI Security is exactly this checkpoint: change one base URL and every prompt and response passes through four self-hosted detectors (attack rules, an ML injection classifier, PII and secret DLP, and a policy-reasoning content-safety model), screening untrusted input and outbound content in both directions, enforced fail-closed as policy-as-code with a full decision log. Request a demo and pilot it in monitor mode to see what your AI traffic has been carrying.

Share