Back to Insights
Agents ·27 September 2026 ·7 min read

When does a human need to approve? A decision tree for AI agent autonomy

Not every action an AI agent can take deserves the same amount of freedom. A practical framework based on reversibility, impact, scale and input trustworthiness — with a concrete decision tree for when a human should approve first.

An AI agent can act autonomously when the consequences are limited and reversible, the financial, legal and human impact is low, and the information driving the decision is sufficiently trustworthy. Once one of those conditions stops being true, human approval should become part of the workflow.

The practical question is therefore not do we trust the agent? It is what happens if this specific action is wrong? An agent that labels an internal document presents a very different risk from one that sends a contract, pays a supplier or emails hundreds of customers. Autonomy should therefore be assigned per action, not per agent.

Start with the action, not the intelligence of the agent

It is tempting to give more autonomy to more capable models. A stronger model would then receive broader permissions than a smaller or cheaper one. That is the wrong basis for access control.

Even a highly capable model can receive incorrect information, misunderstand an instruction or be influenced by prompt injection. Conversely, a fairly simple agent can safely operate autonomously when the actions available to it have limited consequences.

Define the autonomy level for each tool or operation instead. For example:

  • create a draft document: autonomous;
  • classify an internal support ticket: autonomous;
  • propose a calendar appointment: autonomous;
  • delete an existing appointment: approval may be required;
  • pay an invoice: human approval;
  • send a contract on behalf of the company: human approval.

Autonomy is therefore a property of the combination agent + action + context, not of the model alone.

Question 1: can the action easily be reversed?

Reversibility is often the most useful first dividing line.

If an error can be corrected without significant consequences, an agent can usually be given more freedom. Examples include adding metadata, drafting a reply or moving an internal file to the wrong folder.

Actions that are difficult or impossible to reverse require tighter control. Examples include:

  • making a payment;
  • permanently deleting data;
  • sending a legal statement;
  • blocking an account;
  • sharing confidential information externally;
  • placing a binding order.

Technical reversibility is not the only consideration. You can technically follow up an incorrect email with a correction, but a mistaken message to a customer or regulator may have reputational or legal consequences that cannot truly be undone.

A useful rule is: the harder it is to restore the situation that existed before the action, the stronger the case for human approval.

Question 2: what is the impact if the action is wrong?

Not every mistake has the same consequences. An agent misclassifying one internal document has a different risk profile from an agent changing prices or disabling customer accounts. At minimum, consider three categories of impact.

Financial impact. Can the action directly spend money, create a financial obligation or affect revenue? An agent may safely prepare an invoice or purchase request while the final payment or order still requires approval.

Legal or contractual impact. Can the action create obligations, alter rights or trigger formal communication? Examples include accepting terms, sending a contract, submitting an official declaration or changing an employee record.

Operational impact. Could an error interrupt business processes, damage data or prevent employees or customers from accessing a service? An incorrect production configuration or bulk account change can have much greater consequences than a badly worded draft message.

The greater the potential impact, the weaker the case for unrestricted autonomous execution.

Question 3: how many people or systems can be affected?

Scale can turn a small error into a serious incident.

An agent drafting a single email has a limited blast radius. An agent automatically sending that same message to the entire customer database does not. The same principle applies to technical operations: updating one record may be low risk, while applying the same operation across a complete database requires a different level of control.

Ask: if this decision is wrong, how many people, records, accounts or systems are affected?

A practical pattern is to reduce autonomy as scale increases:

  • individual internal action: often suitable for autonomy;
  • small group or limited system component: autonomy with safeguards;
  • large group, production environment or external communication: approval or additional verification;
  • organisation-wide or difficult-to-recover change: almost always human review.

This prevents automation from multiplying one local mistake across an entire organisation.

Question 4: how trustworthy is the input?

An agent's decision can only be as reliable as the information used to make it. An action based on structured data from a controlled internal system is generally more predictable than one based on an incoming email, web page or arbitrary document.

That distinction matters because external and unstructured inputs can be incomplete, outdated or malicious. An email may contain a misleading instruction. A document may no longer reflect current policy. A web page may contain text designed to manipulate the agent into performing an unrelated action.

It is therefore useful to distinguish between:

  • controlled internal data;
  • data from known systems with validation;
  • human free-text input;
  • external documents and emails;
  • the public internet;
  • input whose origin cannot be established with confidence.

The less trustworthy the source, the less sensible it is to connect that input directly to an irreversible action.

The practical decision tree

For every action an agent can perform, run through the same questions:

  1. Is the action fully and easily reversible? Yes: continue. No: require human approval unless the potential impact is demonstrably negligible.
  2. Could an error have financial, legal, privacy or security consequences? No: continue. Yes: add approval or a strong policy control.
  3. Could a single mistake affect many people, accounts, records or systems? No: continue. Yes: reduce the permitted scope or require approval.
  4. Is the decision based on trustworthy, controlled input? Yes: autonomous execution may be appropriate. No: validate, escalate or require approval first.
  5. Can the agent verify that the intended outcome stays within explicit boundaries? Yes: allow autonomy within those boundaries. No: have a human review the decision.

The outcome does not have to be a binary choice between "autonomous" and "human decides." There is an important middle ground: bounded autonomy. An agent might independently prepare purchase requests but not place orders. It might schedule meetings but be unable to delete existing appointments. It might send customer messages from pre-approved templates while presenting newly generated free text as a draft.

Use approval as a safety mechanism, not as the default

Requiring human approval for everything may sound safe, but it removes much of the value of an agent system. When employees have to approve dozens of low-risk actions, approval can also deteriorate into a routine click rather than meaningful oversight.

Human-in-the-loop controls are most useful when humans are reserved for decisions that genuinely require judgement. Automate actions that have low impact and clear boundaries. Escalate exceptions, uncertainty and high-risk actions.

An agent that requires confirmation for every step is little more than an advanced assistant. An agent that can execute everything independently is difficult to control. Useful agent architectures usually operate somewhere between those extremes.

Combine autonomy with least privilege and observability

An approval model should not operate in isolation. Autonomous actions should still be restricted to the permissions the agent genuinely needs. An agent that schedules meetings has no reason to delete customer records — that is least privilege applied to AI agents.

You should also be able to reconstruct what happened afterwards: which input the agent used, which tool it called, which action it executed and whether approval was required or granted — exactly what agent observability is for.

Least privilege limits what can go wrong. Observability shows what actually happened. Human approval prevents selected high-impact actions from being executed without review. Together, these controls provide a much stronger safety model than any one of them on its own — the core idea behind AI agents with the right guardrails for SMEs.

Conclusion

An AI agent should require human approval when an action is difficult to reverse, can have substantial financial, legal or operational consequences, affects many people or systems, or relies on insufficiently trustworthy input. Low-impact actions with limited scope, explicit boundaries and reliable data can often be executed autonomously. A robust agent architecture therefore does not simply decide whether an agent has autonomy; it determines, action by action, how much autonomy is justified.

Agents with the right guardrails

Do you know which actions your AI agent is allowed to take on its own?

We design the permissions, approval steps and logging so an agent only does what's justified — nothing more. Curious what that looks like for your setup?