
Human approval is valuable, but “a human is in the loop” is not a complete control. Security leaders should ask what information reaches the reviewer, whether the AI can shape that information and whether one person can approve an irreversible action. Credentials also leak, dependencies are compromised and systems are misconfigured. A capable agent can search this environment quickly, retry continuously and share what it learns. The question is not whether a sandbox reduces risk. It does, but whether the organization has mistakenly treated that sandbox as the full security boundary?
Build the jail and plan for leakage
Back to the podcast debate I mentioned above, the broader argument still deserves consideration. Humanity has learned to manage technologies that can cause serious harm, even when safeguards are imperfect. AI containment should be approached in the same way: not as a guarantee against failure, but as a way to make failure less likely and limit its consequences. We handle dangerous pathogens in high-containment laboratories even though accidents are possible. The controls still reduce risk enormously. AI containment should be viewed in the same way: essential risk reduction, not proof that failure is impossible.
The protection should match the possible damage. A failure involving a document summarizer is different from one involving cloud administration, financial transfers, biological tools or critical infrastructure. For every agentic AI deployment, security leaders should ask:
