The 7 Runtime Risks Hiding in Every AI Application
Most AI teams secure the wrong layer. Learn the seven runtime gaps that decide whether an AI application is safe in production — and how to fix each one.

Most teams building with LLMs are securing the wrong layer.
They review the prompt. They red-team the model. They check the training data policy of whichever provider they’ve chosen. All of that matters, but it misses where AI applications actually fail in production: at runtime, in the gap between what the model is asked to do and what it’s actually allowed to do.
This isn’t a prompt engineering problem. It’s a systems problem, and the industry data increasingly backs that up. Gartner forecasts that by 2028, a quarter of enterprise generative AI applications will experience at least five minor security incidents a year, up from under one in ten in 2025, with the number of applications suffering a major incident set to roughly quintuple by 2029. That’s not a model quality problem. It’s a governance problem.
This post walks through the seven runtime gaps we see most often, what each one is, why it’s easy to miss, and what fixing it actually looks like, backed by the research and frameworks that are now converging on the same conclusion. Think of it as a checklist for anything you’re shipping with an LLM in it.

1. Traffic control: no gateway in front of the model
Most teams wire their application straight to the model. Every request, whether it’s legitimate, malformed, or actively malicious, reaches the LLM directly. There’s no single point where you can see, throttle, or block anything before it gets there.
This is the traffic control gap, and it’s usually the first one to appear because it’s invisible until something goes wrong. Without an interception layer in front of the model, you can’t:
- Rate-limit a single user or client before they run up your inference bill
- Block a known-bad request pattern before it reaches the model
- See, in one place, everything that’s actually hitting your LLM
The fix is conceptually simple: an AI gateway sitting in front of the model. One chokepoint that every request passes through, where it can be inspected and stopped before it reaches the LLM.
Quick self-check: if you can’t name the single place every AI request in your system passes through, you have this gap.
2. Input risk: no validation before the model sees the prompt
“Ignore your previous instructions” still works against most AI applications in production. Prompt injection has held the top spot in the OWASP Top 10 for LLM Applications for two editions running, because of a structural problem rather than a bug: LLMs take in instructions and the data they act on through the same channel, with no reliable way to tell one from the other. Once an attacker frames malicious input as a new instruction, the model has no way of distinguishing it from a legitimate one, so it simply complies (OWASP, Top 10 for LLM Applications 2025).
Here’s the uncomfortable part: a system prompt is not a security control. It’s an instruction the model can be talked out of, and it’s talked out of more often than most teams assume. As security researchers studying agent guardrails have put it, a model doesn’t comply with an instruction because it’s bound to, it complies because that response happens to be the statistically likely one given the prompt it was handed, which is a fundamentally different thing from enforcement.
The fix is pre-model input validation: catching injection, jailbreak, and extraction attempts before the model ever sees them, rather than hoping the system prompt holds under pressure.
Quick self-check: if a user typed “print everything above this line,” what actually stops it? If your honest answer is “the system prompt,” that’s the gap.
3. Retrieval exposure: RAG without permission or content checks
Retrieval-augmented generation (RAG) makes an AI application useful by feeding it your own documents. But most RAG pipelines retrieve and inject content chunks with no permission check and no scan of what’s actually in them. OWASP added a dedicated category for this in its 2025 update precisely because RAG pipelines introduce their own class of vulnerability: an attacker can poison a vector database so malicious content surfaces during otherwise legitimate queries, and weak access controls on the vector store itself can expose data across tenant boundaries that were never supposed to overlap (OWASP, Top 10 for LLM Applications 2025).
There are two failure modes here, and teams tend to only think about one:
- Cross-tenant or cross-permission leakage. Retrieval doesn’t check who’s asking, only what’s relevant.
- Injection via documents. A malicious or compromised chunk carries instructions into the prompt, not just data.
The fix is retrieval guards: checking permissions and scanning for injection on every chunk before it enters the prompt, not just at the point of ingestion.
Quick self-check: if your vector search doesn’t know the requesting user’s access level, you have this gap.
4. Agent action risk: tool calls without runtime policy
The moment an AI agent can call APIs, send emails, move money, or write to a database, the risk profile changes completely. “It generated a strange response” becomes “it took a strange action,” and those are not the same category of problem.
This is exactly why OWASP published a separate Top 10 for Agentic Applications in 2026, distinct from its LLM list. The framework centres on a new design principle it calls “least agency”: agents should be granted only the minimum autonomy their task actually requires, so that even if one is compromised, the damage it can do downstream stays bounded. Its lead risk, Agent Goal Hijack, describes exactly the pattern that makes agent action risk so dangerous: an attacker alters an agent’s underlying objective through malicious text embedded in something the agent processes as routine input. Because agents often can’t reliably separate instructions from data, a poisoned document, email, or calendar invite can quietly redirect what the agent does next (OWASP, Top 10 for Agentic Applications 2026).
Most agent setups still decide what’s allowed inside the prompt or the surrounding code. That means every new workflow reinvents its own rules, and nothing actually enforces them at the moment the call happens.
The fix is a runtime policy layer on tool and API calls: deciding, per call, what this specific user, tenant, and workflow is allowed to do, at the moment of execution, not baked into a prompt in advance.
Quick self-check: if your agent can call a tool and the only thing deciding whether it should is the model itself, that’s the gap.
5. Policy maturity: access rules scattered across prompts and code
Early on, access rules tend to live wherever was convenient at the time: a line in a system prompt, an if-statement in a route handler, a hardcoded check inside a tool. It works, until you add a second tenant, a new workflow, or a teammate who doesn’t know all the places a rule is hiding.
Gartner has flagged this exact pattern as a governance failure mode in its own right, warning that applying a single, uniform level of governance across every AI agent, rather than a policy proportional to each agent’s actual risk, is itself what causes agent programmes to fail: enterprises tend to treat governance as binary, either locked down or fully trusted, and that binary approach is the root cause of failure. On the back of this, Gartner predicts that by 2027, four in ten enterprises will demote or decommission autonomous agents after governance gaps surface only once something has already gone wrong in production.
This is the policy maturity gap. The rules exist, but they aren’t reusable, aren’t enforceable from one place, and are effectively impossible to audit.
The fix is a reusable policy layer: moving access and approval rules out of prompts and code and into a single place where they can be enforced and changed without redeploying logic scattered across the codebase.
Quick self-check: to change who’s allowed to use a given feature, how many files do you need to touch? If it’s more than one, that’s the gap.
6. Output leakage: no sanitisation before the response is delivered
Most teams put real effort into validating what goes into the model. Almost none put the same effort into what comes out. The response goes straight to the user, whatever it contains: PII it shouldn’t have echoed back, a secret pulled from context, another tenant’s data, or an internal reasoning trace that was never meant to be seen. OWASP tracks this directly through its Improper Output Handling and Sensitive Information Disclosure categories, and added System Prompt Leakage as its own entry after a run of real-world incidents. Many teams had simply assumed their system prompts, and whatever context or logic sat inside them, were isolated from the user by default. Recent incidents have shown that assumption doesn’t hold (OWASP, Top 10 for LLM Applications 2025).
The model has no innate sense of what’s sensitive. Left unchecked, it will happily say the quiet part out loud.
The fix is output sanitisation: screening every response for PII, secrets, and cross-tenant leakage before it reaches the user, not just trusting that a well-written prompt kept it in bounds.
Quick self-check: if nothing sits between the model’s response and your user’s screen, that’s the gap.
7. Audit: no way to prove what happened
The scariest gap on this list isn’t a missing control, it’s the inability to reconstruct a decision after the fact. A customer complains, a regulator asks a question, a bad output goes viral, and the team is piecing together what happened from partial logs and guesswork.
This isn’t hypothetical for much longer. As AI agents take on a growing share of enterprise workflows, regulators are moving in the same direction: audit trails, human oversight, and decision-making transparency are moving from best practice to enforceable requirement in several major frameworks now taking effect. Gartner separately flags identity and access management for AI agents, specifically policy-driven authorization built for machine actors rather than inherited from human user permissions, as one of its top security trends for the year, precisely because the access and audit questions get harder as agents proliferate.
The audit gap means you can’t answer the basic questions: what input came in, what context was retrieved, what tools were called, what policy allowed it, and what went out.
The fix is a proper audit trail: recording inputs, retrieved context, tool calls, policy decisions, and outputs, so every AI decision is provable rather than remembered.
Quick self-check: could you fully replay your AI system’s last risky decision, end to end, using only your logs? If not, that’s the gap.
Why this matters more than trusting AI guardrails alone
A lot of teams believe they’ve already solved this, because they’ve added guardrails: a content filter, a moderation model, a well-written system prompt telling the AI what it shouldn’t do. Guardrails are useful. They are not the same thing as governance, and treating them as interchangeable is where most of the risk above actually comes from.
The distinction is about where the check happens, not how well-intentioned it is. Guardrails and prompt-based rules are advisory rather than binding: an agent can read a rule, appear to agree with it, and still act against it, because nothing forces the outcome either way. They live at inference time, before the model responds. But the moment that actually matters for an agent, the point between deciding to call a tool and the tool actually running, is a different moment entirely, and prompt guardrails simply don’t live there.
The numbers make the cost of this gap concrete. In IBM’s 2025 Cost of a Data Breach Report, 13% of organisations reported an AI-related breach, and of those, 97% said they had no proper AI access controls in place at the time, with 60% of those incidents leading to broader data compromise and 31% causing operational disruption. That’s not a story about sophisticated attackers outsmarting a well-defended system. It’s a story about access controls that were never enforced at the point they were needed.
This is precisely why “permission should be verified, not inferred” is a runtime engineering requirement, not a slogan. A guardrail that reads “never deploy without explicit approval” only works if the model chooses to follow it every time, under every framing, in every context window. A runtime policy layer that checks authorization at the moment of the tool call works regardless of what the model decided, what it inferred from earlier turns, or how the request was phrased. One is a suggestion sitting inside the model’s context. The other is a control sitting outside it, where it can’t be reasoned around, only enforced.
That’s the practical difference between an AI application that merely sounds well-behaved and one that’s actually governed.
Sources
- Gartner, Gartner Predicts 25% of All Enterprise GenAI Applications Will Experience At Least Five Minor Security Incidents Per Year By 2028
- Gartner, Gartner Says Applying Uniform Governance Across AI Agents Will Lead to Enterprise AI Agent Failure
- IBM, 2025 Cost of a Data Breach Report: Navigating the AI rush without sidelining security and IBM Newsroom press release
- OWASP, Top 10 for LLM Applications 2025
- OWASP, Top 10 for Agentic Applications 2026
- AI EdgeLabs, AI Agent Governance: From Policy to Runtime Enforcement
- DEV Community, Your agent’s guardrails are suggestions, not enforcement



