1. Objective
To prevent AI agents from being turned against the organization, by limiting what each agent can do and by finding the specific combinations of capability that make an agent exploitable.
2. Key terms
Capability profile. A record of what an agent can take in, read, write and send, the tools it can call, and how autonomously it acts.
Exposure leg. One of four capabilities that together form an exploitable path: (1) accepts untrusted input, (2) reads private or confidential data, (3) writes to shared systems other people or agents rely on, (4) sends data outside the organization.
Exposure path. An agent, or a chain of connected agents, that holds a dangerous combination of exposure legs. An agent holding all four is a complete exposure path.
3. Requirements
4. Evidence
- Capability profiles and exposure-path findings with treatment decisions.
- Tool action classification, including the unmeasured list.
- Egress allow-lists and reviews.
- Non-human identity inventory and rotation records.
- Tool server registry and change approvals.
- Threat coverage map with stated limits.
5. Metrics
| Metric | Definition |
|---|---|
| Complete exposure paths | Number of agents or chains holding all four exposure legs. |
| Unmeasured actions | Tool actions not yet classified, divided by all tool actions available to agents. |
| Egress control | Externally sending agents with an approved allow-list, divided by externally sending agents. |
6. Basis for conclusions
Prompt injection turns any agent that reads untrusted content into a possible insider. Whether that matters depends on what else the agent can do. Reviewing components one at a time misses the risk, because no single leg fails its own review. The draft therefore requires analysis of combinations and chains. Requirement 500.3 addresses a common error: treating unknown tool actions as harmless reads.
7. Questions for respondents
- Are the four exposure legs the right decomposition? Should others be added?
- Should agents holding three legs also be rated high risk by default?
- Is human approval for material actions practical for high-volume agents?