An AI agent that can act on your systems is useful precisely because it can do things. That is also why it needs limits. The practical question is not whether to keep a person involved, but where: which actions an agent can take on its own, which need approval, and which it should never take. This guide offers a simple way to decide.
What the guidance says
Anthropic's guide to building effective agents notes that agents can pause for human feedback at checkpoints or when they hit blockers, and it recommends starting with the simplest solution and adding complexity only when needed. OWASP's Top 10 for LLM Applications (2025) lists excessive agency (LLM06) as a risk: an LLM-based system is often granted a degree of agency beyond what is necessary. In our view, these ideas point to the same design: give an agent the minimum access and autonomy a task requires, and place people where the stakes are higher. The steps below are our practical framework, not a standard.
Step 1: List what the agent can do
Write down every action: read data, draft text, create or change records, send messages, spend money, change production systems. If you cannot list them, the agent probably has more access than you realize.
Step 2: Classify each action by risk
- Read-only. The agent looks at information. Risk is mainly about what data it can see.
- Reversible. The agent changes something that can be undone easily, such as a draft or a labeled record.
- Hard to reverse. The agent sends an email to a customer, changes a record others depend on, or touches production.
- High stakes. Money, legal commitments, personal or sensitive data, safety.
Step 3: Decide the gate for each class
- Read-only and reversible, low risk: the agent can act, with logging and periodic review of a sample.
- Hard to reverse: require approval before the action, with the context the reviewer needs to decide.
- High stakes: a person makes the decision and the agent prepares the information, or the action is excluded entirely.
Step 4: Limit access (least privilege)
Give the agent only the data, tools, and permissions the workflow needs, and scope them to the engagement. Use separate credentials for agents, and remove access when the work ends. Limiting capability reduces the damage of any mistake or manipulation, including prompt injection (OWASP LLM01).
Step 5: Make approvals meaningful
Approval gates fail when they become rubber stamps. To keep them useful:
- Show the reviewer what matters. The proposed action, the reason, the inputs, and what will change.
- Keep the volume manageable. If a reviewer sees hundreds of items, quality drops. Automate the low-risk ones and review the rest.
- Allow rejection and correction, and feed corrections back into the instructions or tests.
- Sample the automated actions to check that low-risk really is low-risk.
Step 6: Log and be able to undo
Keep a record of what the agent did, what it was asked, and who approved. Define how to roll back an action and who is responsible when something goes wrong.
Step 7: Revisit as trust is earned
Start with more review than you think you need. As evidence builds that a class of actions is reliable, you can reduce approvals for that class deliberately and keep them where errors are costly. Evidence means measured quality over a meaningful sample, not a good first impression.
A hypothetical example
An agent triages incoming customer emails. Reading and labeling are allowed automatically. Drafting a reply is allowed, but a person approves anything that is sent. Refunds are never issued by the agent; it prepares the case for a person. Access is limited to the support inbox, and a sample of labels is reviewed each week.
That is the design behind an AI Pod: agents execute defined work, and people hold the decisions that matter. To plan your first automation with these limits in mind, read how to choose your first workflow to automate with AI.