← Newsroom
Event Recap5 June 2026

ToolProbe at Cyber Science 2026: Measuring What AI Agents Should and Should Not Do

Presented at Cyber Science 2026, the 11th International Conference on Cybersecurity, Situational Awareness and Social Media, hosted by C-MRiC at the historic Senate House, Royal Holloway, University of London.

By Ramkumar Sundarakalatharan

Conference
Cyber Science 2026 - 11th International Conference on Cybersecurity, Situational Awareness and Social Media
Host
Centre for Multidisciplinary Research, Innovation and Collaboration (C-MRiC)
Venue
Senate House, Royal Holloway, University of London
Dates
3-5 June 2026
Session
Friday, 5 June 2026, 10:35-11:00 UK time (remote presentation, moderated by Cyril Onwubiko)
Authors
Ramkumar Sundarakalatharan, Sriram Gopalakrishnan (Zerberus.ai)
Proceedings
Peer-reviewed, published by Springer

Zerberus.ai presented at Cyber Science 2026, the 11th International Conference on Cybersecurity, Situational Awareness and Social Media, hosted by the Centre for Multidisciplinary Research, Innovation and Collaboration (C-MRiC) at the historic Senate House, Royal Holloway, University of London. The paper, "ToolProbe: A Two-Stage Evaluation Framework for LLM Agent Safety in MCP Tool-Calling Environments," by Ramkumar Sundarakalatharan and Sriram Gopalakrishnan, tackles a question that most AI safety research still leaves open: not what a model says, but what an agent is actually allowed to do.

Most evaluation frameworks grade model outputs. ToolProbe looks one layer deeper, at the moment an agent can execute real actions through tools. As the talk put it, the gap between what an AI agent can do and what it should do is enormous, and it routinely exceeds the scope of the deployment frameworks meant to contain it. The Model Context Protocol (MCP) makes this concrete: it lets agents call live tools and services, which is exactly where a well-intentioned agent can quietly cross a line.

To measure that gap, ToolProbe runs a two-stage LLM-as-Judge pipeline against 750 adversarial test inputs across 95 live MCP servers sourced from Smithery.ai. Rather than probing a model in isolation, it exercises agents against real tool-calling environments and judges both whether an action was safe and whether it should have been taken at all.

750
Adversarial test inputs
95
Live MCP servers evaluated
2-stage
LLM-as-Judge pipeline

The findings land hardest for regulated sectors. In financial services, healthcare, and critical infrastructure, an agent that takes an unauthorized action is not just a bug, it is an audit-liability event. ToolProbe gives teams a repeatable way to surface those failure modes before they reach production, and a common language for reasoning about agent behaviour under adversarial pressure.

The paper appears in the peer-reviewed Cyber Science 2026 proceedings, published by Springer. For Zerberus, it is foundational research: the same thinking underpins our work on runtime governance and automated compliance infrastructure with verifiable safety guarantees, so that what an enterprise AI system can do stays aligned with what it is permitted to do.

Read the ToolProbe paper Cyber Science 2026 programme