Why prompt-level guardrails are not enough, and how a two-plane architecture, least-privilege tools, and audit trails make agents safe to give real permissions.