Retrieved content is evidence to inspect, not permission to act.
An instruction hidden in data
Prompt injection occurs when attacker-controlled content attempts to redirect an AI system’s behaviour. OWASP describes both direct and indirect forms. An agent that reads websites or documents may encounter text that looks like an instruction but is actually untrusted input. In a Web3 workflow, the consequences can be more serious when the system also has transaction or credential access.
A defensive example
Imagine an agent summarising a token website. The page contains text telling it to reveal configuration secrets or approve a transfer. Those words do not represent the user’s request. A safe workflow keeps the document’s claims separate from the authority governing actions. This example illustrates the boundary; it does not require executing an attack or testing on someone else’s system.
Controls belong outside the prose
Natural-language reminders alone are not a complete security boundary. Restrict tool permissions, separate sensitive credentials from retrieved text and require independent checks before consequential actions. Validate destinations and amounts against an explicit policy. A system that cannot transfer funds without a separate authorisation step is less exposed than one that simply asks the model to be careful while granting unrestricted signing access.
Test the boundary, not just the answer
In a controlled environment with no valuable keys, test whether malicious document text changes tool use, destinations or disclosures. Log blocked attempts and false positives. Success should mean the agent completes the legitimate task while refusing unauthorised instructions, not merely that it produces a polite warning. Publish the test scope and limitations rather than describing any single defence as complete immunity.
Sources & further reading
Sources checked 7 October 2026. Source-linked explanatory content; not personalised investment advice. Found an error? Request a correction.








