Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week News AMD Goes All-In on the 192GB "Agentic PC" Two Days Before the Nvidia RTX Spark Event: 300-Billion-Parameter Models, No Cloud Required ChatGPT OpenAI Will Put Sponsored Images Inside ChatGPT Image Generation — Testing in the US This Month for Its 1.2 Billion Weekly Users Opinion Hinton, Bengio and 20 Other Top Researchers Warn a Year of AI Progress Could Soon Take Five Weeks Business Schneider Electric to Buy PTC for $22.6 Billion in Its Largest-Ever Deal — and Its Stock Dropped 9% on the Price Tag Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week News AMD Goes All-In on the 192GB "Agentic PC" Two Days Before the Nvidia RTX Spark Event: 300-Billion-Parameter Models, No Cloud Required ChatGPT OpenAI Will Put Sponsored Images Inside ChatGPT Image Generation — Testing in the US This Month for Its 1.2 Billion Weekly Users Opinion Hinton, Bengio and 20 Other Top Researchers Warn a Year of AI Progress Could Soon Take Five Weeks Business Schneider Electric to Buy PTC for $22.6 Billion in Its Largest-Ever Deal — and Its Stock Dropped 9% on the Price Tag

Designing Tools for AI Agents Without Creating a Security Problem

An agent is only as safe as the tools you give it. Narrow inputs, explicit confirmation and audit logs are the baseline.

Giving a model the ability to act is where AI systems stop being a text box and start being infrastructure. The tool interface is the security boundary.

Make inputs narrow

A tool that accepts a free-form query string is effectively a shell. Define typed parameters with enumerations where possible: a customer_id and a refund_reason from a fixed list, not a paragraph of instructions.

Separate read from write

Expose read operations freely. Require an explicit confirmation step for anything that creates, modifies, deletes or sends. The confirmation must come from the user, not from the model.

Constrain consequences

  • Cap monetary amounts and record counts per call.
  • Scope credentials to the minimum the tool needs.
  • Never expose credentials to the model — the tool injects them server-side.
  • Idempotency keys on write operations, so a retry does not double-charge anyone.

Log every call

Record tool name, parameters, acting user, timestamp and outcome. When something goes wrong — and eventually it will — this log is the only way to reconstruct what happened.

Test adversarially

Put instructions in a document the agent reads, telling it to call a tool it should not. If that succeeds, the boundary is wrong, regardless of what the prompt says.

Comments (0)

Log in to join the discussion

Log In

No comments yet