Giving a model the ability to act is where AI systems stop being a text box and start being infrastructure. The tool interface is the security boundary.
Make inputs narrow
A tool that accepts a free-form query string is effectively a shell. Define typed parameters with enumerations where possible: a customer_id and a refund_reason from a fixed list, not a paragraph of instructions.
Separate read from write
Expose read operations freely. Require an explicit confirmation step for anything that creates, modifies, deletes or sends. The confirmation must come from the user, not from the model.
Constrain consequences
- Cap monetary amounts and record counts per call.
- Scope credentials to the minimum the tool needs.
- Never expose credentials to the model — the tool injects them server-side.
- Idempotency keys on write operations, so a retry does not double-charge anyone.
Log every call
Record tool name, parameters, acting user, timestamp and outcome. When something goes wrong — and eventually it will — this log is the only way to reconstruct what happened.
Test adversarially
Put instructions in a document the agent reads, telling it to call a tool it should not. If that succeeds, the boundary is wrong, regardless of what the prompt says.
Comments (0)
Log in to join the discussion
Log InNo comments yet