AI Agent Security Glossary
Plain-language definitions of terms used in AI agent security, with a primary source for each where one exists. Entries are maintained under our glossary pen name and expand over time. See them in context in the October 2026 landscape and our MCP SSRF analysis.
Agent goal hijack · Agent identity · Agentic AI · Excessive agency · Human-in-the-loop (HITL) · Indirect prompt injection · Least privilege · MCP (Model Context Protocol) · Non-human identity (NHI) · Prompt injection · Runtime protection · Sandboxing · Server-side request forgery (SSRF) · Tool misuse · Tool poisoning
- Agent goal hijack
- An attack in which an adversary causes an autonomous agent to pursue a different objective than the one intended by its operator, often through indirect prompt injection or poisoned context. It is listed as ASI01 in the OWASP Top 10 for Agentic Applications 2026. Source: OWASP Top 10 for Agentic Applications.
- Agent identity
- A distinct identity assigned to an AI agent so it can authenticate, be granted specific permissions, be audited and be revoked independently of the humans and services around it. Microsoft Entra Agent ID, for example, describes an agent identity as the primary identity an AI agent uses to authenticate to systems and access resources. Agent identity is a specific case of non-human identity. Source: Microsoft Learn: Entra Agent ID key concepts.
- Agentic AI
- AI systems that do more than generate text: they plan multi-step tasks and take actions through tools, APIs, code execution or other agents, with limited human direction at each step. NIST's CAISI frames AI agent systems as systems capable of taking autonomous actions that impact real-world systems or environments. Security for agentic AI therefore covers identity, permissions, tools, memory and runtime behavior, not only the model. Source: NIST CAISI request for information (Federal Register).
- Excessive agency
- OWASP's term (LLM06:2025) for the vulnerability that lets an LLM-based system perform damaging actions in response to unexpected, ambiguous or manipulated model output. OWASP identifies three typical root causes: excessive functionality, excessive permissions and excessive autonomy. Source: OWASP LLM06:2025 Excessive Agency.
- Human-in-the-loop (HITL)
- A control in which a person must review and approve an agent's proposed action before it is executed, typically for high-impact operations such as sending messages, moving money or deleting data. OWASP recommends requiring user approval for high-impact actions as a mitigation for both prompt injection and excessive agency. Approval is only meaningful if the reviewer can see the full action, including its arguments. Source: OWASP LLM06:2025 Excessive Agency.
- Indirect prompt injection
- A form of prompt injection in which the malicious instructions arrive through external content the model processes, such as a web page, file, email, retrieved document or tool output, rather than from the user directly. OWASP describes it as occurring when an LLM accepts input from external sources that, when interpreted by the model, alters its behavior in unintended ways. It is the main way attackers steer agents they cannot talk to directly. Source: OWASP LLM01:2025 Prompt Injection.
- Least privilege
- The security principle that a user, process or agent should be granted only the minimum access needed to perform its task. NIST defines it as restricting the access privileges of users, or processes acting on behalf of users, to the minimum necessary to accomplish assigned tasks. For agents this applies to tools, data scopes, cloud roles and OAuth scopes. Source: NIST CSRC Glossary: least privilege.
- MCP (Model Context Protocol)
- An open-source standard for connecting AI applications to external systems such as data sources, tools and workflows. An MCP client (inside the AI application) connects to MCP servers that expose tools, resources and prompts. MCP adds capability and also a new supply-chain and permission surface, which the project addresses in its Security Best Practices. Source: modelcontextprotocol.io.
- Non-human identity (NHI)
- An identity assigned to a machine, service, workload or AI agent rather than a person, such as a service account, API key or workload role. Proper NHI management, including inventory, scoping, rotation and revocation, is a core control for agent security.
- Prompt injection
- A vulnerability in which inputs alter a language model's behavior or output in unintended ways, for example causing it to ignore instructions, reveal data or call tools it should not. OWASP lists it as LLM01:2025 and distinguishes direct injection (from the user's own input) from indirect injection (from external content). OWASP notes it is unclear whether fool-proof prevention exists, so impact-limiting controls matter. Source: OWASP LLM01:2025 Prompt Injection.
- Runtime protection
- Controls that observe or constrain an agent while it is executing, such as tool calls, network access and data reads, rather than only at prompt or response time.
- Sandboxing
- Running code or a process in a restricted environment that limits its access to the file system, network and other system resources, so that a compromised or misbehaving component has limited impact. The MCP Security Best Practices recommend executing local MCP server commands in a sandboxed environment with minimal default privileges. A sandbox only helps if its network egress is also restricted. Source: MCP Security Best Practices.
- Server-side request forgery (SSRF)
- A weakness (CWE-918) in which a server can be induced to send requests to destinations the operator did not intend, such as localhost services, private networks or cloud metadata endpoints. In agent systems the URL is often chosen by the model, so SSRF is frequently reached through prompt injection. Source: MITRE CWE-918.
- Tool misuse
- When an agent invokes a permitted tool in a way that produces harmful or unauthorized outcomes, either through manipulation or overly broad permissions. It is listed as ASI02 Tool Misuse and Exploitation in the OWASP Top 10 for Agentic Applications 2026. Source: OWASP Top 10 for Agentic Applications.
- Tool poisoning
- An attack, named by Invariant Labs in April 2025, in which malicious instructions are embedded in an MCP tool's description or metadata. The model reads the full description while users typically see a simplified view, so the hidden instructions can steer the agent to read sensitive files or exfiltrate data. Related variants include rug pulls (a tool description changed after approval) and shadowing (one server's tool description altering how the agent uses another server's tools). Tool poisoning is a specialized form of indirect prompt injection. Source: Invariant Labs.
Suggest a term or a correction via info@aiagentthreats.com.