AI Agent Security Checklist: Permissions, Prompt Injection, and Recovery
An AI agent becomes a security concern when it can read private information, call tools, retain memory, or change a real system. The practical goal is not to make a model infallible. It is to keep untrusted content from becoming authority, constrain every action to the minimum necessary scope, and make consequential work observable, interruptible, and recoverable. This checklist translates current NIST, CISA, and OWASP guidance into release criteria for production agent workflows.
1. Define the agent’s authority before choosing tools
Document the business workflow, the data the agent may read, the actions it may propose, and the actions it may execute. Name the human or service that owns each decision. A useful boundary is explicit enough that an operator can tell whether a proposed tool call is inside or outside the agent’s job.
Separate read, draft, recommend, approve, and execute capabilities. An agent that summarizes a customer record does not automatically need permission to edit it. An agent that drafts an email does not automatically need permission to send it. Treat every added tool, data source, and persistent memory store as an expansion of the threat model.
- Inventory models, tools, APIs, data stores, memory, retrieval sources, identities, and downstream systems.
- Classify actions by impact, reversibility, external visibility, and data sensitivity.
- Assign an owner, stop condition, and recovery path for every high-impact action.
- Record which actions are deliberately unavailable, even if the underlying platform supports them.
2. Treat retrieved content as data, never as authority
Prompt injection can arrive through a user message or indirectly through a document, email, website, ticket, API response, or other content the agent reads. The control objective is to prevent instructions embedded in that content from changing the agent’s permissions, approval requirements, system rules, or task boundary.
Keep trusted instructions and untrusted content structurally separate. Constrain retrieved text to a clearly labeled data field, validate expected formats, and reject content that attempts to redefine the workflow. Filtering can reduce obvious attacks, but it should not be the only control because a persuasive or obfuscated instruction may evade pattern matching.
- Label the source and trust level of every external input.
- Do not place secrets, access tokens, or unnecessary personal data in the model context.
- Require authorization in the execution layer; never rely on the model to enforce its own permissions.
- Test direct, indirect, encoded, multilingual, and multi-step injection attempts.
3. Give tools short-lived, narrow permissions
Grant the agent only the tools and resources required for the current workflow. Prefer read-only access until a write operation is justified. Scope credentials to specific tenants, folders, mailboxes, records, environments, or API operations, and keep development and production identities separate.
A tool call must be authorized using the requesting user, target resource, requested operation, and current approval state. Tool descriptions and model outputs are not authorization evidence. Unknown tools and unexpected parameters should fail closed. Apply rate, cost, retry, and chain-length limits so a mistaken loop cannot become an operational or financial incident.
- Use dedicated service identities instead of shared administrator credentials.
- Keep credentials outside prompts, memory, transcripts, and error messages.
- Allowlist operations and validate parameters against a schema before execution.
- Rotate or revoke access without rebuilding the agent.
4. Protect memory, retrieval, and tenant boundaries
Persistent memory can carry sensitive information or malicious instructions into later sessions. Decide what is allowed to persist, how long it remains useful, and who may retrieve it. Validate and classify content before storage; do not automatically convert every conversation, tool response, or retrieved document into long-term memory.
Isolate memory and retrieval indexes by user, organization, and environment. Enforce the same access controls when data is retrieved as when it was first stored. Deletion and retention rules should cover source documents, embeddings, cached prompts, transcripts, tool arguments, outputs, and backups—not only the visible chat history.
- Define retention, expiration, deletion, and legal-hold behavior.
- Redact secrets and minimize personal or regulated data before model access.
- Test cross-user, cross-tenant, stale-permission, and deleted-record retrieval.
- Log the source record and authorization decision behind material outputs.
5. Put independent controls around consequential actions
Use a separate policy and execution component for actions that are destructive, financial, administrative, externally visible, or difficult to reverse. The agent may prepare a proposal, but an authorized person or deterministic control should validate the exact target, parameters, and current state before execution.
Approval must be specific and fresh. A general request to help with email should not authorize an unrelated message, recipient, or attachment. Show an action preview with the destination and material effects, then bind approval to that exact action. Re-check authorization immediately before execution so a stale approval cannot be replayed after the context changes.
- Require approval for sending, publishing, deleting, changing access, executing code, or committing funds when applicable.
- Use separation of duties for the highest-impact operations.
- Make actions idempotent where possible and use transaction or compensation patterns.
- Provide a kill switch and a tested way to pause queued or looping work.
6. Test the workflow as an adversary and an operator
Evaluate the complete system, not only the base model. Test prompts, retrieved content, tools, memory, identity propagation, approval flows, error handling, and downstream effects. Repeat security testing after material changes to the model, system prompt, tools, permissions, retrieval source, memory policy, or orchestration code.
Build abuse cases from the workflow’s real assets: data exfiltration, privilege escalation, prompt injection, tool misuse, memory poisoning, approval bypass, runaway cost, and cascading multi-agent actions. Include normal operator mistakes and dependency failures, because a safe design also needs predictable behavior when a service times out or returns malformed data.
- Record test inputs, expected controls, actual results, evidence, owner, and remediation.
- Exercise unauthorized users, expired sessions, changed permissions, and unavailable dependencies.
- Verify that the agent stops safely rather than improvising around a failed control.
- Block release when a critical control fails; do not substitute a warning banner for enforcement.
7. Monitor decisions and rehearse recovery
Create an audit trail that connects the originating request, actor, retrieved sources, model or policy version, authorization decision, tool call, target, result, and approval. Protect logs from alteration and avoid recording secrets or excessive sensitive data. Alerts should focus on security-relevant behavior such as repeated denied actions, unusual tool sequences, cross-tenant access attempts, approval bypasses, and sudden cost or retry spikes.
Write an incident procedure before launch. It should identify who can disable tools, revoke credentials, isolate memory, stop queues, preserve evidence, notify affected owners, and restore trustworthy operation. Recovery is incomplete until the team verifies that compromised instructions, cached context, or poisoned memory cannot immediately reintroduce the problem.
The minimum release package is a threat model, authority matrix, data-flow inventory, tool permission record, approval policy, adversarial test evidence, logging plan, incident playbook, and named operating owner. This is a practical checklist, not a certification or a claim that any agent can be made risk-free.