Prompt injection occurs when content the model processes contains instructions that override its intended behaviour. Because the model has no reliable way to distinguish an instruction from data, the vulnerability does not close with better prompting.
How attacks work in practice
- Indirect injection. A webpage, email or document contains hidden text directing the assistant to exfiltrate data or change its answer.
- Tool abuse. The assistant is persuaded to call a tool with attacker-controlled parameters.
- Data poisoning. Malicious content enters a retrieval index and influences future answers.
Mitigations that reduce blast radius
- Least privilege. Give tools only the permissions the task requires. Read-only beats read-write by default.
- Human confirmation for consequential actions. Sending money, deleting data and sending mail externally should require explicit approval.
- Content provenance. Treat retrieved third-party text as untrusted input, and mark it as such in the prompt.
- Output filtering. Detect and block responses that contain secrets or unexpected external links.
What not to rely on
Instructing the model to ignore conflicting instructions is a speed bump, not a control. Any security model that assumes the model will comply is already broken.
Comments (0)
Log in to join the discussion
Log InNo comments yet