An AI agent that can act on your systems is useful precisely because it can do things. That is also why it needs limits. The practical question is not whether to keep a person involved, but where: which actions an agent can take on its own, which need approval, and which it should never take. This guide offers a simple way to decide.

What the guidance says

Anthropic's guide to building effective agents notes that agents can pause for human feedback at checkpoints or when they hit blockers, and it recommends starting with the simplest solution and adding complexity only when needed. OWASP's Top 10 for LLM Applications (2025) lists excessive agency (LLM06) as a risk: an LLM-based system is often granted a degree of agency beyond what is necessary. In our view, these ideas point to the same design: give an agent the minimum access and autonomy a task requires, and place people where the stakes are higher. The steps below are our practical framework, not a standard.

Step 1: List what the agent can do

Write down every action: read data, draft text, create or change records, send messages, spend money, change production systems. If you cannot list them, the agent probably has more access than you realize.

Step 2: Classify each action by risk

Step 3: Decide the gate for each class

Step 4: Limit access (least privilege)

Give the agent only the data, tools, and permissions the workflow needs, and scope them to the engagement. Use separate credentials for agents, and remove access when the work ends. Limiting capability reduces the damage of any mistake or manipulation, including prompt injection (OWASP LLM01).

Step 5: Make approvals meaningful

Approval gates fail when they become rubber stamps. To keep them useful:

Step 6: Log and be able to undo

Keep a record of what the agent did, what it was asked, and who approved. Define how to roll back an action and who is responsible when something goes wrong.

Step 7: Revisit as trust is earned

Start with more review than you think you need. As evidence builds that a class of actions is reliable, you can reduce approvals for that class deliberately and keep them where errors are costly. Evidence means measured quality over a meaningful sample, not a good first impression.

A hypothetical example

An agent triages incoming customer emails. Reading and labeling are allowed automatically. Drafting a reply is allowed, but a person approves anything that is sent. Refunds are never issued by the agent; it prepares the case for a person. Access is limited to the support inbox, and a sample of labels is reviewed each week.

That is the design behind an AI Pod: agents execute defined work, and people hold the decisions that matter. To plan your first automation with these limits in mind, read how to choose your first workflow to automate with AI.

Sources