All posts
AI SecurityRamkumar Sundarakalatharan8 min read

Open-Weight AI Models Are Bypassing Safety Filters: What the UK AISI Kimi K3 Evaluation Means for Enterprise AI Security

UK AISI and US CAISI found Kimi K3's safeguards failed to prevent offensive cyber operations. Here is what the findings mean for runtime AI security and agentic AI governance.

AI securityOpen-weight AIRuntime AI securityAgentic AI securityLLM security

Model safety filters won’t save your AI stack: UK AISI evaluation shows Kimi K3 bypasses safeguards

On 23 July 2026, the UK Artificial Intelligence Security Institute (UK AISI) and the US Center for AI Standards and Innovation (CAISI) published a joint evaluation of Moonshot AI’s Kimi K3 — released on 16 July 2026 and scheduled for open-weight release by 27 July 2026.

The report confirms what practitioners in AI security already suspected: model-level safety filters are not a reliable control boundary for autonomous or agentic AI systems. Below we set out the verified technical findings, place them within the broader open-weight threat landscape, and explain the architectural implications for enterprise security teams.

What did UK AISI and CAISI find? The Kimi K3 evaluation results

Kimi K3 on the Corporate Network Attack Range

UK AISI and CAISI tested Kimi K3 against “The Last Ones” (TLO), an expert-built 32-step simulated corporate network attack spanning four subnets and approximately 20 hosts — a scenario a human expert would require roughly 20 hours to complete.

Key results:

  • Kimi K3 reached step 17 of 32 on average, ahead of peer open-weight model GLM-5.2 (step 11), but substantially below the leading US closed-weight models, which averaged 28.5 steps.
  • In 1 of 10 attempts, Kimi K3 completed the full 32-step attack path within a 100-million-token limit.
  • UK AISI and CAISI stated the model is “capable of autonomously attacking small, weakly defended and vulnerable enterprise systems, when directed to do so and given initial network access.”

Important methodological note: TLO operates without active defenders, imposes no penalty for alert-triggering actions, and follows an intentional attack path. The result is a controlled lower bound — not a live-environment projection. It should be interpreted as such.

Kimi K3 on Exploit Development: ExploitBench

ExploitBench, developed by Carnegie Mellon University, tests a model’s ability to progress along the software exploitation ladder across 41 post-2023 vulnerabilities in Chrome’s V8 engine.

  • Kimi K3 scored 32.2% — above GLM-5.2 at 24.4%, but well below leading US models at an average of 76.2%.
  • Kimi K3 achieved Arbitrary Code Execution (ACE) on 0 of 41 samples. The most capable US models achieved ACE on an average of 20 of 41 samples.

ACE is the highest-severity exploit outcome: it grants an attacker the ability to hijack a target system.

Did Kimi K3’s Safety Filters Prevent the Attack?

No. Kimi K3’s safeguards did not prevent it from attempting cyber exploit development or offensive cyber operations during the evaluation.

This is consistent with prior AISI findings. In an earlier open-weight assessment, DeepSeek V4-Pro’s refusals on narrow cyber tasks were overcome by simply retrying the request a small number of times. The pattern is structural: refusal alignment is probabilistic, not deterministic. It may delay — it does not block.

Why open-weight AI models are a distinct security problem

Safety Guardrails Can Be Stripped from Open-Weight Models

When an AI model’s weights are released publicly, the developer loses all downstream control. Safeguards cannot be remotely updated or revoked. Critically, refusal training can be removed from the weights entirely using publicly available tools.

A Financial Times and AI safety research group Alice investigation in May 2026 demonstrated that a tool called Heretic can strip all safety protections from major open-weight models — including those from Meta, Google, and OpenAI — in under ten minutes on a standard laptop. A popular model repository currently hosts over 8,000 “uncensored” or “abliterated” model variants with refusal training removed.

An ICLR 2026 paper documented a refined surgical silencing approach achieving up to a 99% bypass rate on tested models.

The Open-Weight Cyber Capability Gap Is Closing

AISI’s July 2026 open-weight gap report found that leading open-weight models now trail the closed-model frontier by four to seven months on cyber capability benchmarks — down from six to ten months through most of 2025.

The operational cost of running an autonomous attack simulation has also dropped sharply. AISI assessed the cost of a full 100-million-token attack run at approximately $1.19 for DeepSeek V4-Pro — against roughly $85 for the closed frontier models it benchmarks against.

Kimi K3 is scheduled for open-weight release on 27 July 2026. Once released, no regulatory action can retrieve or modify those weights.

What this means for enterprise AI security architecture

Why model-level AI safety filters are not sufficient

The architecture failure is not specific to Kimi K3. It is a property of where safety controls are placed.

Model-level safety filters are trained into the weights. They are:

  • Probabilistic: they can be circumvented by rephrasing, retrying, or adversarial prompting.
  • Strippable: on open-weight release, refusal training can be removed by any technically capable actor.
  • Blind to session context: per-request filters evaluate each call in isolation and cannot detect threats that unfold across a multi-step attack chain.

The AISI evaluation demonstrates exactly why that last point matters. Kimi K3’s simulated attack unfolded across 17 steps, four subnets, and approximately 20 hosts. A per-request security system evaluating each call in isolation would have no visibility into the trajectory — only the individual step.

The case for runtime AI security and agentic governance

The control point for AI security must sit at the layer that remains under your control regardless of what the model does, what safeguards it was shipped with, and whether those safeguards have been modified.

OWASP’s Top 10 for Agentic Applications (published December 2025) catalogues the structural risk categories: goal hijacking, tool misuse, identity abuse, memory poisoning, cascading failures, and rogue agents. Each of these manifests at execution time — not at the level of model weights.

This is the design principle behind VANGUARD, Zerberus’s session-aware AI security gateway. VANGUARD enforces a four-layer runtime defence stack operating independently of model behaviour:

  • Rules-based prompt injection detection on every input before it reaches the model.
  • Meta Prompt Guard 2 for classification of adversarial instruction patterns.
  • Data Loss Prevention screening for 14+ PII types across outputs and tool call parameters.
  • Content safety screening across 14 MLCommons hazard categories.

The session-awareness layer is the critical architectural differentiator. VANGUARD tracks conversation history per user, tenant, and API key — enabling detection of multi-turn attack patterns of precisely the kind the AISI evaluation documents. Per-request stateless systems cannot detect these by design.

The current VANGUARD roadmap includes OPA-based intent validation and MCP tool governance, which extends enforcement to agentic tool-call sequences: every tool invocation is checked against policy prior to execution, regardless of model intent.

Run a free runtime risk scan on your AI stack

Summary: Three things the AISI Kimi K3 report confirms

  1. Open-weight models with meaningful offensive cyber capability are available today. The gap between open and closed frontier models is narrowing at pace.
  2. Model-level safeguards on open-weight models are not reliable barriers to offensive tasking — even before weight stripping is considered.
  3. Once weights are released, no provider action can retrieve or update those safeguards. The control must exist at the runtime layer.

If your AI security architecture depends on what the model was trained to refuse, it is dependent on a control you do not own and cannot guarantee. Runtime enforcement is not a supplement to model-level safety. For agentic and open-weight deployments, it is the primary control.

References

[1] UK Artificial Intelligence Security Institute / US Center for AI Standards and Innovation. Preliminary Assessment of Kimi K3’s Cyber Capabilities. 23 July 2026. https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-k3s-cyber-capabilities — mirrored at https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-preliminary-assessment-kimi-k3s-cyber-capabilities

[2] UK Artificial Intelligence Security Institute. Open-Weight AI Cyber Capability Gap Report. July 2026. https://www.aisi.gov.uk/blog

[3] South China Morning Post. China’s Kimi K3 significantly below US rivals in hacking power, UK-US study shows. 23 July 2026. https://www.scmp.com/tech/tech-war/article/3361711/chinas-kimi-k3-significantly-below-us-rivals-hacking-power-uk-us-study-shows

[4] Techtimes. Open-Weight AI Models Now Match Frontier Cyber Skill From Four Months Prior, AISI Finds. 22 July 2026. https://www.techtimes.com/articles/320960/20260719/open-weight-ai-models-now-match-frontier-cyber-skill-four-months-prior-aisi-finds.htm

[5] Financial Times / AI safety research group Alice, reported via Akerman LLP. Open-Weight AI Models: Safety Guardrails Can Be Removed in Minutes Using Free, Publicly Available Tools. 25 May 2026. https://www.akerman.com/en/perspectives/open-weight-ai-models-safety-guardrails-can-be-removed-in-minutes-using-free-publicly-available-tools.html

[6] Open-Weight AI Models governance paper (arXiv, 2026), citing Hugging Face platform data: over 8,000 “uncensored” or “abliterated” model variants hosted publicly. https://arxiv.org/pdf/2606.19890

[7] OWASP. Top 10 for Agentic Applications 2026. December 2025. Referenced in: Microsoft Open Source Blog, Introducing the Agent Governance Toolkit. May 2026. https://opensource.microsoft.com/blog/2026/04/02/introducing-the-agent-governance-toolkit-open-source-runtime-security-for-ai-agents/

Further Reading

  • AISI Cyber & Autonomous Systems blog — ongoing evaluations and open-weight gap reporting: https://www.aisi.gov.uk/blog
  • ExploitBench — Carnegie Mellon University benchmark for exploit development capability measurement. Referenced throughout UK AISI evaluation series.
  • ICLR 2026: Surgical refusal component silencing paper — up to 99% bypass rate via weight-level modification. (Referenced in Akerman / Lexology alert, May 2026.)
  • Nature Communications (2026): Large reasoning models are autonomous jailbreak agents — 97% multi-turn jailbreak success rate with no human involvement after initial instruction. (Referenced in Akerman / Lexology alert, May 2026.)
  • OWASP Top 10 for Agentic Applications 2026: https://owasp.org (December 2025 release)
  • Zerberus.ai Runtime Risk Scan — assess your AI stack’s runtime exposure: https://zerberus.ai/assessment/
Share