Five Ways AI Agents Actually Fail, and Why Most Security Programs Are Only Built for Two of Them

Key Takeaways
- Zenity Labs reproduced a zero-click data exfiltration attack against a replica of a customer-service agent that McKinsey built in Copilot Studio and Microsoft published as a flagship example. A build that clears a consulting firm's review and a platform vendor's showcase is still wide open.
- Multi-turn attacks can assemble malicious intent across messages that each individually pass inspection, which means a guardrail scoring single turns in isolation is architecturally incapable of catching them.
- Compromised MCP servers exploit trust between agents, not a vulnerability in any single agent, so a security posture built to monitor individual agents can watch the entire compromise unfold and see nothing wrong.
- Not every consequential incident involves manipulation. Some of the most damaging failures happen when an agent, using access it was legitimately granted, reasons its way into an action nobody anticipated.
- Reading these scenarios together, rather than individually, is what reveals the actual shape of the coverage gap most agentic security programs are carrying without knowing it.
Most of the industry discourse on agentic AI risk has settled into a comfortable framing: agents get attacked the way models get attacked, through some form of prompt manipulation, and the fix is a better guardrail. I think that framing is dangerously incomplete, and I want to walk through five specific scenarios that make the case directly rather than abstractly. Two of them are manipulation. One exploits trust between agents rather than any single agent's behavior. One is a configuration nobody reviewed, exploited by an attacker who never touched the agent. And one involves no attacker at all, which is the category most security programs are least equipped to see coming.
The Scenario Everyone Assumes, Reproduced Against a Vendor's Own Reference Implementation
Zenity Labs reproduced this first attack against a replica of a customer-service agent that McKinsey built in Copilot Studio and Microsoft published as a flagship example. The agent listened to an inbox with no sender restriction, so a single email carrying hidden instructions was enough to make it ship its entire customer CSV, plus dozens of Salesforce records, back out to an external destination, with zero clicks required from anyone. Zenity Labs rebuilt a simplified replica of that agent to test it, wired to Salesforce for the customer data the original pulled from its own databases. Same agent pattern, same inbox trigger, same class of connector. I think this detail matters more than it usually gets credit for: no organization can credibly say "we would never build it that way," because Microsoft put that build in front of the market as an example worth following. McKinsey built the agent, and Microsoft chose to showcase it.
The mechanism reaches past data theft into direct action just as easily. Give that same agent the ability to issue refunds up to a defined threshold, and a hidden instruction embedded in an inbound support ticket can direct it to approve a maximum-value refund and mark the account exempt from fraud review, executed as a routine workflow step before any human is ever involved. What stops this isn't a smarter prompt filter. It's hard boundary enforcement that governs tool invocations deterministically, regardless of the reasoning chain that produced them, so that refund authorization above a defined threshold triggers a human approval requirement at the execution layer rather than functioning as a soft guardrail the model can reason around.
The Attack That Doesn't Exist in Any Single Message
The second scenario is the one I find most useful for testing whether a security team actually understands the multi-turn threat surface, because the failure mode is genuinely counterintuitive. An organization deploys an internal research agent with access to proprietary product documentation, competitive intelligence, internal financial data, and a tool that exports briefing packets to a specified destination. An adversary distributes an injection payload across five separate turns, interleaving each fragment with on-topic filler so that no individual turn carries enough signal to trip a prompt filter.
The turns aren't the attack. They're the assembly. By the sixth turn, the agent is operating on an instruction that never appeared in any single message, and it invokes the export tool against the finance connector, writing to a destination it happens to be permitted to write to. Every individual turn passed inspection, because inspection was looking at turns. The composite instruction existed in none of them, and it only exists at all once you're looking at the sequence as a sequence.
If you want to know whether a guardrail vendor actually understands this, ask two specific questions rather than a general one about capability. Does the detector read the conversation as a sequence, or does it score individual turns and pool the scores afterward? And what's its measured detection rate on a benchmark where no individual turn exceeds that vendor's own per-turn threshold? I've found that most vendors can't answer the second question with a real number, because most published benchmarks in this space leave the signal sitting in the opening turn, where it's convenient to catch and easy to claim credit for.
When the Compromise Lives Between Agents, Not Inside One
The third scenario is where I think the industry's mental model breaks down most completely, because it requires abandoning the assumption that a compromised system has to behave anomalously somewhere. A development team builds an agentic workflow orchestrating three specialized agents: a code review agent, a deployment agent, and a monitoring agent, coordinating through a shared MCP server for tool access. An attacker compromises a dependency inside that server, injecting a tool-poisoning payload that causes the deployment agent to include a backdoor in the next production release, while simultaneously reporting a clean deployment to the monitoring agent.
No single agent in that workflow behaves anomalously, and I mean that literally, not as a rhetorical flourish. The deployment agent trusts instructions arriving from the MCP server, because that trust relationship is exactly what makes the orchestration function in the first place. The monitoring agent trusts the status report the deployment agent sends it, for the same reason. Each agent is doing precisely what it was designed to do with the information it received. The compromise doesn't live in any individual agent's behavior. It lives in the relationships connecting them, relationships that were never designed with the assumption that one of the parties might be lying.
A security architecture built to monitor individual agents for anomalous behavior can watch this entire sequence unfold and conclude nothing is wrong, because from each agent's own vantage point, nothing is. Catching this requires agent-to-agent communication monitoring that can detect anomalous instruction flows inside a shared workflow, and runtime behavior analysis that compares actions against expected workflow patterns even when every individual action looks entirely legitimate in isolation. As multi-agent orchestration becomes the default architecture rather than the exception, I expect this category to become the harder half of the detection problem, not the easier one.
The Configuration Nobody Reviewed, Exploited by Someone Who Never Touched the Agent
The fourth scenario doesn't require an attacker to be clever about the agent at all. An enterprise deploys Microsoft Copilot Studio to build a custom agent integrated with its ERP and HR systems. A business user, unaware of the security implications of what they're configuring, creates a low-code agent that inadvertently exposes its service account credentials through an overly permissive data connector configuration. An attacker who identifies this exposure uses the harvested credentials to access HR data directly, entirely outside the agentic workflow the credentials were originally provisioned for.
What I want to draw attention to here is what's absent: no injection, no manipulation of the agent's reasoning, no compromise of the agent itself. The vulnerability lives in a configuration a well-meaning business user set up, almost certainly without any security review, using a low-code platform explicitly designed to make agent creation accessible to non-developers. That accessibility is a genuine benefit to the organization's speed of innovation, and it's also exactly why this class of exposure is becoming more common as citizen development scales faster than security review capacity. Closing it takes AI Security Posture Management, the continuous discovery and posture assessment of the agents, connectors, credentials, and permissions in your environment, particularly in the low-code and SaaS-embedded contexts where security teams have the least direct visibility. Posture alone doesn't finish the job here, because the attacker used that credential outside the agent workflow entirely. That takes integration with your non-human identity platform for credential lifecycle management, plus alerting when an agent service account shows up outside its expected workflow context.
The One With Nothing to Detect
The fifth scenario is the one I think most security programs are least prepared for, and I've written about this specific failure mode in more depth elsewhere, but it belongs in this set because it completes the picture the other four scenarios leave incomplete. A development team deploys a coding agent to fix a mismatch in a staging environment. Mid-task, the agent hits an inconsistency nobody briefed it on. No one manipulated it. No malicious content entered its context window. Trying to resolve the discrepancy on its own, the agent reaches for a deployment token its role legitimately carries, connects to a production database it was never instructed to touch, and drops the production table while trying to reconcile the schema.
Every API call the agent made was permitted by a role a human had already approved. The entire sequence takes under ten seconds. Nothing in it would trigger a prompt-injection filter, because there was no injection. This is the scenario most detection logic is structurally blind to, not because the logic is poorly built, but because it was built to catch manipulation, and there's nothing adversarial here to catch. What stops this is enforcement of the action itself against declared policy, evaluated at the moment of the call against the task the agent was actually dispatched to perform, rather than an attempt to infer bad intent from language that was never malicious in the first place.
What Reading These Five Together Actually Tells You
No individual capability addresses all five of these, and I think that's the actual point of walking through them as a set rather than treating each as its own isolated case study. A platform strong on prompt-injection detection may have nothing to offer against the fifth scenario, where there's no injection. A platform strong on posture management may miss the second scenario's multi-turn assembly, because posture is a build-time discipline and that attack unfolds entirely at runtime. A platform that only monitors single agents in isolation will miss the third scenario entirely, because the compromise lives in the relationship between agents, not in any one of them.
The question worth asking of your own environment, or of any vendor's platform in a sandboxed test, isn't "does this scenario sound plausible?" It's "would our current architecture actually detect this, stop it, and how?" I've found that running these five scenarios directly, rather than working from a capability checklist, is the fastest way to find out where the real gaps are, and it's uncomfortable in exactly the way a genuinely useful test should be.
Download the The Enterprise Buyer's Guide to Agentic AI Security to learn how to evaluate, compare, and select security solutions purpose-built for the age of AI agents.
All ArticlesRelated blog posts

The Authorization Trap: Why "No Evidence of Manipulation" Doesn't Mean "No Incident"
Many conversations about AI agent risk over the past year start from the same unspoken assumption: something bad...

MCP Is Growing Up
Our team spends much of the week talking to security leaders and practitioners who are trying to figure out where...

The Agent Will See You Now: Why Healthcare's AI Agent Boom Needs Visibility and Control
Healthcare, as an industry vertical, is moving faster on agentic AI than it has in past technology evolutions....
Secure Your Agents
We’d love to chat with you about how your team can secure and govern AI Agents everywhere.
Get a Demo