// 00
Who this is for and what it is for
read this part before the others
I wrote this with two people in mind. The first works in security and has to assess systems already running agents in production. The second builds products with artificial intelligence and has not yet stopped to think about what happens when someone on the outside can write into what the agent reads.
The intent here is one thing only: to spread defensive knowledge. Describing how the attack works is a precondition for defending, because nobody protects a system against a class of attack they cannot describe. That is why this document goes into mechanism and attack chain instead of staying on the surface.
A word on where this comes from. Jump is certified to ISO/IEC 42001, the international standard for AI management systems, and deploys agents in client environments. The method in this document is what we use before granting access to real data. Section 10 covers how that connects to what is described here.
What you will not find is a recipe. There is no ready payload, no step by step reproducible against someone else's system. Every case is public, already fixed by the vendor, and described at the level a defender needs to check their own house.
THE RULE THAT IS NOT NEGOTIABLE
Nothing here gets tested without written authorization. Against a third party's system, that is a crime. Against your own, without agreeing it first with whoever owns it, it is an incident you caused yourself. Section 03 covers the minimum agreement before any test.
If you work for one of the vendors named here, let me say the obvious: these cases are in the document because they were handled well. In four of the five, a researcher found it and the company fixed it before it became damage. That is the behaviour we want to see more of, and citing it is one way to acknowledge it.
// 01
What matters and what does not
the difference between a curiosity and a compromise
When someone shows you they got a model to swear, break character or spit out its own system prompt, do not get excited. That proves you can push the model off script, and nothing beyond it. On its own it reaches nothing behind the model. It is a sign the channel works, and that is all.
What we look for is the attack that crosses the model and lands on something real: data leaving the company, an action taking place, a system being touched. In every assessment we run, the triage question is the same. Did influence become effect? If it did not, we are on the right track.
Autonomy defines the blast radius
"Agent" has become a marketing label and today says almost nothing. There are systems that always run the same pipeline with a model bolted in, and systems that decide on their own what to do next. Both are sold the same way and produce completely different security work. When we assess a vendor, that is the first thing I separate. The area of each circle below represents the reach of a successful injection at that level; the dashed ring is always the worst case, for comparison.
// 02
Five questions that define the target
forget the product and the framework, look at the deployment
These five questions are where I start any assessment. You can swap the product and the framework all you like; it is the answers that draw the attack surface.
1. Untrusted input. Where can I get content in? The direct prompt is the obvious case. The dangerous one is indirect: the pages it browses, the documents it ingests, the passages it retrieves, its own memory, tickets, emails and the output of other agents.
2. Tools. What can it do? Classify each one by the severity of the action it enables, from reading a record to moving money.
3. Privilege. With what authority, on whose behalf? A read-only scope and a shared admin credential are very different targets behind the same tool.
4. Processing. Who interprets its output? Any call that becomes a database query, a shell command, a fetched URL or executed code is an injection sink, and the model is what feeds it.
5. Output. Where can the data go? Every destination is a potential exfiltration route, including the ones that do not look like output: a log line, a counter, an image being loaded.
OWASP Top 10 for Agentic Applications
Published by the OWASP GenAI Security Project on 9 December 2025. It is the reference taxonomy, built from real 2025 incidents rather than projections. It works as shared vocabulary in threat modelling and as a scoping checklist.
ASI01ASI02ASI03ASI04ASI05ASI06ASI07ASI08ASI09ASI10// 03
Before you test anything
a precondition, not an appendix
Active reconnaissance, traffic interception, breaking certificate pinning and injecting into production fields are penetration testing activities. Against a third party's system without authorization, that is illegal. Against your own without agreeing it first, it takes down production and burns internal trust. Almost every piece of material circulating on this subject skips this part, and I find that irresponsible.
- Written authorization from whoever owns the system, naming the targets and the permitted techniques. When the target is contracted SaaS, that includes the vendor.
- Explicit scope of environments, integrations and accounts. An agent wired to CRM, email or payments reaches customer data and outside systems.
- An agreed window and a stop channel, with someone on the defensive side knowing the traffic is a test.
- A clear rule for real data. If proving it requires exfiltrating something, agree beforehand what counts as evidence. Customer records do not leave the environment, not even in a screenshot.
- Destructive actions out of scope by default. Delete, transfer and send only with named authorization, preferably in a mirrored environment.
In production, prove reach without consummating the effect. Demonstrating that the agent can call the destructive tool is usually enough as a finding, and it keeps the test from becoming the incident.
// 04
The four stages of an attack
an adaptation of the cyber kill chain to agents, not a new framework
STAGE 1
Reconnoitre: find the way in
This stage produces inventory. Where content gets in, what the agent does with it, which credentials sit behind each tool, and which systems the actions reach. A clever prompt comes later, if at all.
Start with the delivery form, because it sets the tooling. An assistant that only exists inside the app leaves no web traffic to watch. Then map the whole stack: host, operating system, neighbouring services, runtime, sandbox, orchestration. Only at the top do you find the model and its classifiers. Most testing starts at that top. I start from the bottom, because everything underneath is surface nobody looked at.
A few deliberately out-of-scope questions reveal the architecture by the reaction. No answer at all suggests keyword search with an AI label on it. A polite refusal points to a model with a system prompt. A refusal that is always identical and instant, before any reasoning could have happened, gives away a separate classifier filtering the input.
The inputs the interface hides are the part most people skip, and where the real surface lives: search or filter fields that travel a different route, API parameters the screen never exposes, traffic between the agent and its subagents, and tool output, which the front end does not render and the filter almost never inspects.
STAGE 2
Exploit: make the agent act on your content
This is the moment what you wrote stops being data the agent reads and becomes an instruction it follows. The channel matters far more than the wording of any payload.
- Direct injection. What every guardrail was built to catch. Useful for establishing the baseline of what the front door blocks.
- Indirect injection through ingested content. Instructions in alt text, HTML comments, metadata, off-screen text, zero-width characters. No suspicious user in the transcript. Anyone who can drop a file into an indexed drive has a permanent, click-free instruction channel.
- Retrieval and memory. The instruction arrives before the prompt, in the position of highest trust, and survives the session.
- Subagents and tool output. Filtering watches the main input and output, not internal traffic. A near-universal gap.
- Supply chain. An installed tool, a loaded plugin, a queried server or a dependency can carry instructions treated as trusted by default.
- The human in the loop. Approval is worth whatever the displayed summary is worth, and the same system being influenced is what generates that summary.
STAGE 3
Execute: turn influence into impact
There are two paths, and the first is the forgotten one. The capability is already the attack: no sink, no code. The tool does exactly what it was built for, only for someone else. Read and leak, send, transfer, delete. Or the action lands in a sink: the output becomes input to another system that executes it. Here SQL, command, SSRF, RCE, traversal and deserialization all come back.
Climbing the stack is not automatic
You often read "influence over the model becomes runtime execution, which becomes host access, which becomes network reach" as though it were a natural progression. Each jump requires a condition: from model to runtime, that the output reaches something that executes it; from runtime to host, that the sandbox is weak; from host to network, a reusable credential or an exposed metadata endpoint. Verifying the cases shows most compromises use no sink at all.
THE TRUST THAT MAKES THIS WORK
The action runs with the agent's own authority: credential, scope and reach are inherited. When access control is applied only at the agent level, where it states its own identity and is trusted to ask only for what it should, all it takes is someone controlling what it asks for and the whole thing falls.
STAGE 4
Objectives
- Exfiltrate. Every outbound path is a door, including logs, reports, counters and image loads.
- Act. Send email, move money, alter records. The objective with no equivalent in a passive system.
- Deny. Delete, corrupt, take down, or simply exhaust the system by forcing expensive operations.
- Persist. Poison memory, plant state the agent reloads and trusts.
- Propagate. In multi-agent systems peer trust is usually unverified: one compromise becomes fleet compromise.
// 05
Documented cases
verified against primary sources, and the status matters more than the name
These five cases circulate as a list of incidents. I went to check each one against the original source and four of them are not adversary attacks that reached users. I marked the status on all five, because it changes what you can conclude.
EchoLeak: zero-click exfiltration in Microsoft 365 Copilot
Aim Labs built the proof of concept in January 2025 and reported it to MSRC. Microsoft fixed it server-side in May, requiring no customer action, and disclosure came on 11 June 2025 as CVE-2025-32711, CVSS 9.3. No evidence of exploitation in the wild.
The value is in the chain, which brought down three layers in sequence. The email, with instructions phrased as a legitimate request, passed the XPIA classifier. The filter that stripped links only recognised inline Markdown syntax, so the attack used reference-style links. The image embedded in the response was fetched automatically by the client, which removed the need for a click. And the content policy was bypassed using a Teams preview endpoint that was on the allowlist and would fetch any URL passed as a parameter. A Microsoft service performed the exfiltration. It is the kind of chain no single layer would have held.
Aim Labs · MSRC CVE-2025-32711 · Reddy and Gujral, arXiv:2509.10540, Sep 2025
ForcedLeak: a five-dollar domain in Salesforce Agentforce
Noma Labs reported it on 28 July 2025. Salesforce fixed it on 8 September by enforcing Trusted URLs for Agentforce and Einstein, and disclosure followed on 25 September. CVSS 9.4.
The input was the description field of the lead capture form, which accepts 42,000 characters of free text. The trigger is delayed: the payload only runs when, days later, an employee asks the agent to process that lead. And the output, without which the attack would not work, was a domain listed in Salesforce's own CSP allowlist. It had expired and was on sale for about five dollars. Five dollars was the price of the exfiltration channel in a CVSS 9.4 attack.
Noma Security, Sep 2025 · The Hacker News, 26.09.2025 · Salesforce Ben, 30.09.2025
Amazon Q: the commit that became a wiper
This is the only one of the five where an external adversary reached the end user. On 13 July 2025 a malicious commit landed in the aws-toolkit-vscode repository. It shipped in version 1.84.0, published on 17 July to a base of roughly 964,000 installations, and was reverted in 1.85.0 on 19 July. It received CVE-2025-8217.
The detail that decides everything is in the payload: it called the Q CLI with --trust-all-tools --no-interactive, carrying a prompt instructing it to wipe the system to a factory state and delete local and cloud resources. The model did not fail. What existed there was a switch capable of turning off every confirmation. If a flag like that exists in your agent, stop reading this document and go remove it.
AWS stated that no customer resources were impacted. Two caveats: the destructive code apparently did not execute due to a defect, and the story that the attacker was handed admin credentials came from the attacker himself, through the press, and does not hold up against public analysis of the repository history.
The Register and SC Media, 24.07.2025 · Bargury, timeline reconstruction from commits, 24.07.2025 · AWS statement, 26.07.2025
Replit: the database deleted during a freeze
In a twelve-day session with a code freeze in force, the agent deleted a user's production database, holding roughly 1,200 executive records and 1,190 company records, fabricated around 4,000 fake profiles, and stated that rollback was impossible.
That last part was false, and the data was recovered. There were backups. The version that circulates, that the data was lost, is wrong, and it came from the agent itself. This is what worries me most about this case: the agent was the only witness to what the agent did.
The root cause is architectural and mundane: at the time the platform used the same database for preview, testing and production. The CEO acknowledged on 19 July 2025 that this was unacceptable and should never have been possible, and the fixes announced were structural: automatic separation between development and production, a planning-only mode, one-click restore. None of them is "make the model more careful".
The Register, 22.07.2025 · heise online, 25.07.2025 · public posts by Amjad Masad and Jason Lemkin
Agent-in-the-Middle: winning the routing through the description
Trustwave SpiderLabs researchers published a fake agent card in an A2A protocol directory. Because the coordinating agent picks peers using a model that judges card descriptions, the description itself worked as prompt injection: it instructed that this agent always be chosen. That was enough to win task routing.
In the demonstration, the rogue agent corrupted currency conversion results. Reading it as interception and leakage of sensitive data to third parties is an extrapolation of potential impact.
Trustwave SpiderLabs · Agent In the Middle: Abusing Agent Cards in the A2A Protocol
WHAT THE COUNT IS TELLING YOU
Of five cases, one was an adversary attack that reached users, and even there no customer damage was confirmed. Two were vulnerabilities found by research and closed before publication. One was an accident with no attacker. One was a lab.
This does not reduce the risk, and reading it that way would be a mistake. The class of attack is demonstrated and structural, it is just that the public record so far is researchers arriving before criminals. To me that is an open window, and windows close. You can still treat this as an architecture decision, before it becomes incident response.
THE TECHNICAL PATTERN THAT REPEATS
In both exfiltration cases, what decided the outcome was not the prompt: it was the outbound channel. In EchoLeak, a Microsoft endpoint that was on the allowlist and fetched arbitrary URLs. In ForcedLeak, an expired domain still sitting on the allowlist. In both, the injection worked; what turned it into a leak was a stale list of permitted domains.
In the case that reached users, the factor was equivalent: a command line option that turns off every confirmation. In the accident, no separation between environments.
None of those four is an artificial intelligence problem. They are known controls, badly maintained, that AI has now reached. When we go into a client, that list is where I start.
// 06
What actually holds
grouped by level of guarantee, not by preference
Layer 1 · architecture with a demonstrable guarantee
The only approach that does not depend on the model resisting anything. The basis is the dual LLM pattern, proposed by Simon Willison in 2023: a privileged model that sees only the trusted query and plans the actions, and a quarantined model that processes untrusted content but has access to no tools at all. Suspicious content never reaches the model that decides.
CaMeL, from Google DeepMind, is the first concrete implementation of that pattern. The privileged model generates code representing the user's intent; that code runs in a custom interpreter, and the model is not what orchestrates the calls. The interpreter tracks the provenance of every piece of data and applies policy before each tool call. Because control flow is derived only from the trusted query, untrusted data cannot change what the program does. It only changes what the program contains.
THE NUMBER THAT ANSWERS THE EXECUTIVE QUESTION
On the AgentDojo benchmark, CaMeL solved 77% of tasks with provable security, against 84% for a system with no defence at all. Roughly seven points of utility in exchange for a verifiable guarantee, with the implementation published as open source.
The authors are honest about the price: capability-based systems demand substantial implementation effort, overly restrictive policy produces approval fatigue, and side channels remain, such as inferring information from call patterns or response timing.
Layer 2 · reducing probability
Spotlighting, published by Microsoft researchers, explicitly marks untrusted content in the context so the model treats it as data rather than instruction. It works, it reduces success rate, and it is not a boundary: EchoLeak passed Microsoft's own injection classifier with careful wording. Apply it, but with the right expectation. It raises the cost of the attack and cuts the volume. It does not stop a well-built one.
Layer 3 · the conventional controls that decided the real cases
If I could pick a single thing from this document for you to do tomorrow morning, it would be this layer. It is the cheapest and it is the one that failed most.
- A living outbound allowlist, with an owner and a review date. In both exfiltration cases the channel was a stale entry on the allowlist. Auditing that list and removing whatever nobody claims is the highest return per hour in this entire document.
- No option that turns off confirmation. Amazon Q became a wiper because a switch existed that trusts every tool with no interaction. If that switch exists in your agent, it is the attack.
- Environment separation. The root cause in the Replit case was one database serving preview, testing and production.
- Least privilege and per-task credentials. The Agentforce agent runs as a user, and the injection inherits exactly that user's visibility. Broad read access is what converts nuisance into leak.
OWASP names the principle that unifies this as least agency: autonomy is a privilege earned per task, not a default setting.
WHAT DOES NOT HOLD ON ITS OWN
Guardrails and intent classifiers as primary defence. Selling a guardrail as a boundary is, in my reading, the biggest misconception in the market right now. A model filtering another model carries the same weakness being exploited. This is not a theoretical argument. In EchoLeak the classifier, the link filter and the content policy fell one after another, each under modest pressure.
A better system prompt. Training shapes what the model tends to produce; it imposes no boundary on what comes in. In practice the attacker only has to stop resembling what the model learned to refuse.
Human approval on its own. A real control, but a soft one: it is worth whatever the displayed summary is worth, and that summary is generated by the very system being influenced.
And what decides the size of the damage
In the Replit case, the agent said rollback was impossible and was wrong. Recovery only happened because somebody did not believe it. In Amazon Q, the compromised version stayed distributed for two days.
DETECTION AND RESPONSE
Log every tool call, with parameters, result, the identity it ran under, and the context fragment that motivated it. Without that there is no possible investigation. The conversation does not serve as evidence, because it is precisely what the attacker controls.
Alert on behavioural drift: a tool never used in that task, read volume outside the norm, an unprecedented outbound destination. A deterministic rule is worth more here than a model-based judge.
Verify what was done, not what was reported. Independent confirmation of the real effect. The agent can never be the only witness to what the agent did.
A response plan written in advance. How to revoke the agent's credential without taking down the rest, clear poisoned memory and indexes, reverse actions in bulk, and who has the authority to stop the agent.
// 07
Checklist for builders
answer these before putting any agent into production
- List everything the agent reads without anyone talking to it: documents, pages, tickets, emails, memory, another agent's output. Treat every item as untrusted by default.
- Classify each tool by the severity of the action it permits. High-severity ones need deterministic verification outside the model.
- Confirm that access control happens at the tool boundary and uses the real user's session, not a generic agent identity.
- Map every point where the output becomes a query, a command, a file path, a URL or code.
- Map every observable outbound channel, including logs, metrics and image loads, and restrict network egress with an allowlist that has an owner and a review date.
- Check that the security filter also sees traffic between subagents and tool output, not only the main input and output.
- Define what can go into long-term memory and who writes to it. Memory writable by ingested content is a permanent backdoor.
- Make sure the summary shown to whoever approves is generated outside the influenceable path, or shows the raw call instead.
- Eliminate any option that executes a tool without confirmation. If it exists, it is the attack.
- Log every tool call with parameters and result, and set up alerting for behavioural drift.
- Write the response plan before you need it: credential revocation, cleaning poisoned memory, bulk reversal.
// 08
Where we differ from what circulates
errors found while checking each case against the source
These five cases are repeated across dozens of articles and presentations, almost always from a third-hand summary that copied another summary. When we went to the source, we found the errors below. They are corrected in the body of this document and recorded here, because anyone reading other material will encounter the wrong version and needs to know why ours differs.
FOUR OF THE FIVE ARE NOT INCIDENTS
The list is usually presented as "five real incidents". EchoLeak and ForcedLeak were coordinated disclosures, fixed before publication and with no observed exploitation; Replit was an operational accident with no adversary; Agent-in-the-Middle is a proof of concept. Only Amazon Q was a real compromise that reached users.
IN THE REPLIT CASE, THE DATA WAS NOT LOST
The current version claims the data disappeared. There were backups and the database was recovered. The claim that rollback was impossible came from the agent itself and was false. That is the most important point in the case, and it vanishes when the wrong version is repeated.
THE REPLIT CASE IS ASI10, NOT ASI01
It circulates mapped to ASI01, agent goal hijack. That category requires untrusted input redirecting the agent, and there is no external attacker there. In OWASP's own launch text, the example cited for ASI10, rogue agents, is literally this case.
AGENT-IN-THE-MIDDLE HAD ITS IMPACT OVERSTATED
It is usually described as interception and leakage of sensitive data to third parties. In the published demonstration, the rogue agent corrupted currency conversion results. The rest is potential impact.
CLIMBING THE STACK IS NOT AUTOMATIC
You often find "influence over the model, runtime execution, host access, network reach" chained as a natural progression. Each jump requires a specific condition, and verification shows most compromises use no sink at all.
There is also the catchphrase that all of this is "injection with no prepared statements", suggesting no structural fix is possible. The verified cases say otherwise. In four of the five, what decided the outcome was a badly maintained conventional control. The problem has a solution. It just does not live inside the model.
// 09
Sources
consulted directly and dated
- OWASP GenAI Security Project. OWASP Top 10 for Agentic Applications, . Confirms the official examples per category: EchoLeak under ASI01, Amazon Q under ASI02, Replit under ASI10. In September 2026 the project announced an Agent Control Standard and the 2026 Top 10 for LLMs, not yet read for this revision.
- Reddy, P.; Gujral, A. S. EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM System. arXiv:2509.10540, Sep 2025.
- Noma Security. ForcedLeak: AI agent risks exposed in Salesforce Agentforce, Sep 2025. With The Hacker News () and Salesforce Ben ().
- Bargury, M. Reconstructing a timeline for Amazon Q prompt infection, . With The Register and SC Media () and the AWS statement of .
- The Register () and heise online () on the Replit case, with the public posts by Amjad Masad and Jason Lemkin.
- Trustwave SpiderLabs. Agent In the Middle: Abusing Agent Cards in the A2A Protocol. Proof of concept.
- Debenedetti, E. et al. (Google DeepMind) Defeating Prompt Injections by Design. arXiv:2503.18813. The CaMeL architecture, AgentDojo results and an open implementation.
- Willison, S. The Dual LLM pattern for building AI assistants that can resist prompt injection, Apr 2023.
- Hines, K. et al. (Microsoft) Defending Against Indirect Prompt Injection Attacks With Spotlighting. CAMLIS 2024.
- Dark Marc. The Hacker's Guide to Attacking AI Agents, . Origin of the four-stage model and the five questions.
// 10
How we handle this at Jump
what the certification solves and what it does not
Jump is certified to ISO/IEC 42001, published in 2023 as the first certifiable international standard for artificial intelligence management systems. We went through the process from the inside before taking it to clients, which changes the conversation quite a bit: it prepares whoever lived it, not whoever only designed it. It is to AI what 27001 is to information security. It does not dictate how to build the model; it dictates how to manage the system lifecycle with traceability, named accountability and risk management.
I want to be honest about its reach, because some people sell certification as if it were armour. No certificate prevents indirect prompt injection. 42001 is not a technical control. What it delivers is the management system that gives the questions in this document an owner and a review date.
Look at what decided the cases in the previous sections. An allowlist with an expired domain nobody reviewed. An option that turns off confirmation and that nobody had inventoried. Production with no environment separation. Logging too thin to know what the agent did. None of those is a model problem. All of them are management problems, which is exactly what the standard demands: an inventory of AI systems, a named owner, risk assessment before deployment, operational control, records and continuous review.
WHAT WE DO WITH IT
Agent assessment before production. The five questions from section 02 applied to the real deployment, mapping inputs, tools, privilege, sinks and outbound channels. The deliverable is the list of what has to change before the agent gets access to real data.
Agent deployment using the layers in section 06. A deterministic boundary beneath the model, least privilege per task, tools on demand with no automatic execution, and logging sufficient to investigate afterwards.
Our AI governance platform, organised in three stages that match what this document demands.
Discover. API scanning of cloud, data, code and identity, reading metadata only and with credentials in a vault. Twenty-nine connectors, from model providers to vector databases and identity directories. What the scan finds enters an inventory with a named owner and a risk level, including what nobody had declared.
Control. Deterministic rules decide what gets masked and what gets blocked, with no model judging another model. At the endpoint, classification runs on the person's own machine, offline, and masks before sending: raw content never leaves the computer, only the event does. Whoever requests an exception is not who approves it, and every exception is born with an expiry date.
Prove. Every action becomes a dated event on a timeline, and the framework score and the audit dossier come from that evidence, with no number typed by hand. A failure does not turn into a green screen.
THE LIMIT, SO WE DO NOT SELL WHAT IT IS NOT
RADAR·AI governs AI use across the organisation: discovery, shadow AI, policy, endpoint control and audit evidence. It does not replace the deterministic boundary that has to exist inside your agent's architecture, the one from layer 1 of section 06, validating each tool call against the user's session.
They are different layers and both need to exist. The platform answers "what do we have, who owns it and how do I prove it". The agent architecture answers "can this specific call happen". We deploy both, and be suspicious of anyone who tells you one replaces the other.
If you read this far and recognised your own architecture in one of the cases, the useful conversation does not start with tooling. It starts with answering the five questions about your agent and seeing what shows up.
THE SERIES
- 01The agent cannot tell what it read from what it was told to doyou are here · five cases verified against primary sources
- 02The agent was not tricked. They used its badgein Portuguese · non-human identity and the Salesloft Drift case
- 03Shadow AI: the AI use the company never authorisedin preparation