Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week News AMD Goes All-In on the 192GB "Agentic PC" Two Days Before the Nvidia RTX Spark Event: 300-Billion-Parameter Models, No Cloud Required ChatGPT OpenAI Will Put Sponsored Images Inside ChatGPT Image Generation — Testing in the US This Month for Its 1.2 Billion Weekly Users Opinion Hinton, Bengio and 20 Other Top Researchers Warn a Year of AI Progress Could Soon Take Five Weeks Business Schneider Electric to Buy PTC for $22.6 Billion in Its Largest-Ever Deal — and Its Stock Dropped 9% on the Price Tag Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week News AMD Goes All-In on the 192GB "Agentic PC" Two Days Before the Nvidia RTX Spark Event: 300-Billion-Parameter Models, No Cloud Required ChatGPT OpenAI Will Put Sponsored Images Inside ChatGPT Image Generation — Testing in the US This Month for Its 1.2 Billion Weekly Users Opinion Hinton, Bengio and 20 Other Top Researchers Warn a Year of AI Progress Could Soon Take Five Weeks Business Schneider Electric to Buy PTC for $22.6 Billion in Its Largest-Ever Deal — and Its Stock Dropped 9% on the Price Tag

Prompt Injection: The Vulnerability Without a Clean Fix

Prompt Injection: The Vulnerability Without a Clean Fix

Any text the model reads can contain instructions. Separating data from instructions is the core unsolved problem.

Prompt injection occurs when content the model processes contains instructions that override its intended behaviour. Because the model has no reliable way to distinguish an instruction from data, the vulnerability does not close with better prompting.

How attacks work in practice

  • Indirect injection. A webpage, email or document contains hidden text directing the assistant to exfiltrate data or change its answer.
  • Tool abuse. The assistant is persuaded to call a tool with attacker-controlled parameters.
  • Data poisoning. Malicious content enters a retrieval index and influences future answers.

Mitigations that reduce blast radius

  • Least privilege. Give tools only the permissions the task requires. Read-only beats read-write by default.
  • Human confirmation for consequential actions. Sending money, deleting data and sending mail externally should require explicit approval.
  • Content provenance. Treat retrieved third-party text as untrusted input, and mark it as such in the prompt.
  • Output filtering. Detect and block responses that contain secrets or unexpected external links.

What not to rely on

Instructing the model to ignore conflicting instructions is a speed bump, not a control. Any security model that assumes the model will comply is already broken.

Comments (0)

Log in to join the discussion

Log In

No comments yet