All posts
AI SecuritySriram G7 min read

Your firewall can't read prompts: the blind spot in every AI deployment

Firewalls, WAFs, and DLP never see what your apps send to LLMs, or what comes back. What NIST, OWASP, NCSC and NSA guidance says to do about the AI traffic blind spot.

AI SecurityLLM Security

TL;DR: Every request your applications send to an LLM, and every response that comes back, passes through a gap in the security stack. Firewalls see approved HTTPS. WAFs find no signatures. DLP never inspects a model’s reply. Meanwhile that traffic carries customer data outbound and model-generated content inbound, and attackers have already shown they can hijack it. Security guidance from OWASP, NIST, NCSC, and the NSA converges on the same answer: inspect and control AI traffic at a point you own, outside the model.

Trace one request through your stack

You shipped an AI feature this quarter. A support copilot, a summarizer, a chat interface: something that takes what users type, wraps it in a prompt, and sends it to a model API.

Now trace that request through your security stack. Your firewall saw HTTPS to an approved endpoint and waved it through. Your WAF checked for SQL injection and XSS signatures and found none, because natural language doesn’t have signatures. Your DLP watches email and file shares; it has never parsed a model response. Your SIEM shows nothing, because nothing in that path emits an event.

If that prompt carried a customer’s account history, or an API key someone pasted into a ticket, or if the reply came back carrying something it never should have said, the honest answer is: nothing you run today would have noticed.

This is not a “you” problem. The whole stack was built for an older kind of traffic. AI traffic is different in one fundamental way: the payload is natural language, that language acts like instructions, and your sensitive data and an attacker’s instructions travel in the same channel.

This gap is already being exploited

In 2025, researchers disclosed EchoLeak (CVE-2025-32711), which the National Vulnerability Database describes plainly: “AI command injection in M365 Copilot allows an unauthorized attacker to disclose information over a network.” Microsoft assigned it a CVSS score of 9.3 (Critical). A single crafted email caused Copilot to hand over internal data, with zero clicks from the victim. The full attack chain is worth reading: five steps, each one past a defense built to stop it.

The lesson is not that Microsoft made a mistake; they patched it server-side. The lesson is that AI traffic is a new traffic class, one that carries both your sensitive data and attacker-controlled instructions in the same channel, and in most organizations nothing stands in its path. That combination of private data, untrusted content and a way out has a name: the lethal trifecta.

And the exposure is measurably common. IBM’s Cost of a Data Breach 2025 report found that 13% of organizations had already suffered a breach of an AI model or application, and 97% of those lacked proper AI access controls.

What the institutions recommend

None of what follows is vendor opinion. Over the past three years, the major security bodies have converged on the same set of practices for AI systems:

  • The OWASP Top 10 for LLM Applications (2025) puts prompt injection at #1 (LLM01) and sensitive information disclosure at #2 (LLM02), and its mitigations lean on input/output filtering, least privilege, and monitoring rather than trusting the model.
  • The Guidelines for Secure AI System Development (UK NCSC and CISA, co-sealed by agencies across 18 countries) make logging, monitoring, and secure operation first-class requirements across the AI lifecycle. You cannot secure what you cannot see.
  • NSA and CISA’s joint guidance Deploying AI Systems Securely tells operators to treat AI systems as untrusted by default: validate and continuously monitor them rather than assume they behave.
  • The NIST AI Risk Management Framework and its Generative AI Profile name information security and data privacy among the core generative-AI risks, and expect controls that are measurable and managed, not aspirational.
  • NIST’s adversarial ML taxonomy (NIST AI 100-2e2025) formalizes direct and indirect prompt injection as attack classes, giving security teams shared vocabulary for threat modeling.
  • Most recently, the NSA’s AI Security Center published security design considerations for the Model Context Protocol (May 2026), recommending that external agent connections pass through a “filtering outgoing proxy” with tightly pinned access. Even the guidance for the agent era points to inspection in the traffic path.

Different bodies, one theme: the model cannot be the control point. Something outside it, something you own, has to inspect and enforce.

What good looks like

Here is that guidance translated into a practical checklist. All of it is achievable without buying anything, if you have the engineering time:

  1. Know your AI traffic. Inventory which applications call which models through which providers, including agents and anything unofficial.
  2. Screen requests before they reach the model. Pattern rules catch known injection attacks cheaply; an ML classifier catches novel phrasings that rules miss. You want both layers.
  3. Screen responses before they reach your users. Data leakage and unsafe content travel outbound too; a request-only control covers half the problem.
  4. Prefer masking over blocking for data findings. Redacting an SSN in flight keeps the application working; a hard block teaches teams to route around your controls.
  5. Write your policy down as code, not as a PDF. Versioned, reviewable configuration that says what is allowed, per application. Auditors ask for exactly this.
  6. Decide the failure mode in advance. If your screening layer goes down, does traffic flow unfiltered? The safe answer is no: fail closed, so an outage never silently disables security.
  7. Log every decision, and start in monitor mode. Watch what would have been blocked on real traffic before you enforce anything. Zero production risk, and the logs themselves are the evidence regulators and customers increasingly ask for.

Two ways to implement it

You can build this into each application: a filtering library, a redaction function, logging wired up per team. For a single application with a capable platform team, building it yourself is a reasonable choice, and everything in the checklist above is buildable from open components.

The trade-off shows up with the second and third app. Each ends up running its own version of the rules, updated on its own schedule, with its own gaps, and there is no single audit trail across them. At that point most teams consolidate: move the screening into the traffic path once, so every request and response passes one point that applies one policy and writes one log, regardless of which app or provider sits on either side.

This pattern has a name

The industry calls that second option an AI security gateway: a proxy between your applications and your LLM providers that inspects the traffic itself and applies graded verdicts (allow, mask, or block) under a policy you control. Adjacent tools differ mainly in where the enforcement decision runs: routing-focused AI gateways manage cost and rate limits with light guardrails attached, guardrail SDKs return classifications for your own code to act on, and a security gateway applies the policy in the path itself.

FAQ

Our LLM provider already has safety filters. Do we still need this? Provider filters run the provider’s policy, not yours. They do not mask your customers’ PII on the way out, produce no audit trail you can access, and do not extend to the second provider you will inevitably add. They are a complement, not a control.

Will putting security in the path break our application? Not if you roll out properly: start in monitor mode (log-only, nothing enforced) and watch would-block decisions on real traffic first. For data findings, in-flight masking keeps responses flowing rather than failing them.

Doesn’t this just route our prompts through yet another vendor’s cloud? It can, and that is the first question to ask of any tool in this space: where does detection run? If prompts are shipped to a third-party SaaS for scoring, you have traded one exposure for another. Whatever you evaluate (or build), know exactly what leaves your environment and for whom.

We’re a small team with no security engineers. Is this realistic? Keep it to three moves: put whatever screening you adopt in the path rather than in each app, keep the policy in one reviewable place, and roll out in monitor mode so there is no production risk while you tune. Building and maintaining the full checklist yourself is roughly a quarter of platform work; weigh that against your roadmap before choosing to build.

The checkpoint described above is what we build. VANGUARD is Zerberus’s AI security gateway: change one base URL and every prompt and response passes through four self-hosted detectors (attack rules, an ML injection classifier, PII and secret DLP, and a policy-reasoning content-safety model), enforced as fail-closed policy-as-code with a full audit trail. Request a demo and pilot it in monitor mode. More writing on AI security on our blog.

Share