Disclosure, up front
Our business is building automations, and this article is a list of things we will not automate for you. That is not modesty — it is the most reliable predictor we have of whether a project ends well. The clients who are happy two years later are the ones who drew this line early. The ones who are not are usually the ones who let an agent send something on their behalf before anyone had decided it could.
The most common question we get about AI agents is "can it do this?" It is almost always the wrong question, because by 2026 the answer is usually yes. It can read the inbox, draft the reply, look up the account, apply the discount and send it.
The question that actually determines whether you should let it is different: what happens when it is wrong, and who finds out first?
The industry has moved on this. The 2026 consensus, after a couple of years of enthusiasm for fully autonomous business agents, has settled on human-supervised workflows — AI doing the drafting, sorting, summarising and routing, with people retaining the decisions that carry consequence. That is not a compromise position or a failure of nerve. It maps precisely onto where these systems are strong and where the cost of being wrong is asymmetric.
Here is how to draw the line for your own business, and the five decisions we would never hand over regardless of how good the model gets.
The rule, in one line
An agent may produce anything. It may commit to nothing. Drafting, sorting, summarising, researching, routing, preparing — all of that is safe to delegate because a human sees the output before it becomes real. The moment an action is externally visible and hard to take back, it needs a person's name on it.
01Four tests for whether a decision can be delegated
Run any candidate task through these. A "no" on the first two is usually decisive.
- Is it reversible? Not "can it be corrected eventually" — can it be undone before anyone outside notices. A draft in a folder is reversible. A sent email is not. Money moved is not. A published post is only partly, because the screenshot survives the deletion.
- Who pays if it is wrong? If the cost of an error lands on you, delegation is a business risk you are entitled to take. If it lands on a customer, an employee or a third party, you are spending someone else's downside to save your own time. That is a different decision and deserves to be made consciously.
- Can a human check it cheaply? Verification has to be genuinely quicker than doing the task, or approval collapses into rubber-stamping. If checking whether the agent got it right takes as long as doing it yourself, you have not saved anything and you have added a false sense of oversight.
- Does it encode a judgment about values? What is fair, who gets an exception, when to apologise, what a relationship is worth. These are not accuracy problems. There is no correct answer for a model to find, only your answer, and it should come from you.
02The decision table
Most tasks in a small business fall cleanly into one of three columns once you apply those tests.
| Delegate freely | Agent drafts, human commits | Human decides, always |
|---|---|---|
| Summarising a call or thread | Customer replies and quotes | Any payment or refund |
| Sorting and tagging incoming work | Proposals and scopes of work | Hiring, firing, discipline |
| Researching and collecting information | Social posts and published content | Contract or legal commitments |
| Drafting internal notes | Follow-up sequences | Pricing and discount exceptions |
| Flagging anomalies for review | Appointment offers and rescheduling | Apologies and goodwill gestures |
| Preparing data for a decision | Invoices and statements | Anything involving a complaint or dispute |
Our framework, refined across client projects rather than drawn from published research. It aligns with the broader 2026 industry move toward human-supervised agent workflows, in which AI handles drafting, sorting, summarising and routing while people retain decisions with legal, financial or relationship consequences. Your middle column will be longer than ours; that is where most real business work lives.
The left column is where the time savings are, and it is larger than people expect. Summarising, sorting, researching and preparing is a serious fraction of a working week, all of it reversible, all of it cheap to check.
03The five that are always human
Taking the right-hand column properly, because each has a specific reason rather than a general caution.
Money. Payments, refunds, credits, anything that moves value. Irreversible, and the failure mode of an automated mistake is indistinguishable from the failure mode of an automated attack — which means you lose the ability to tell the two apart at exactly the moment it matters. There is no version of this we would automate.
Hiring, firing and discipline. Legally fraught in every jurisdiction, and this is territory where regulators have already been most active. Beyond compliance: a person learning that a decision about their livelihood was made by software is a harm on its own, separate from whether the decision was correct.
Legal commitments. Contracts, terms, warranties, anything that binds the business. An agent that agrees to something on your behalf has created an obligation you may be held to. Drafting is fine. Signing is not.
Pricing and exceptions. Pricing is strategy wearing a number. An agent optimising a discount does not know that this customer is the reason three others found you, or that the margin on this job funds the quiet season. It is also the decision most tempting to automate, because it looks like arithmetic.
Apologies and goodwill. When something has gone wrong, an automated apology is worse than a slow one. The entire content of an apology is that a person took responsibility. Delegating it removes the only ingredient that was doing anything.
04Where autonomy creeps in without a decision
Nobody sets out to let an agent send invoices. It arrives sideways, and these are the four routes we see most.
- The convenience upgrade. Approving each draft gets tedious, so someone enables auto-send "for the simple ones". The definition of simple is set by the vendor, not by you, and it is rarely written down.
- The default that was already on. A tool ships an agentic feature enabled by default in an update. Nobody chose it; it appeared. This is worth a quarterly check now that around 40% of enterprise applications are expected to carry task-specific agents by the end of 2026, up from under 5% the year before.
- The chained automation. Step one drafts, step two sends when a condition is met. Each step was approved separately, and nobody looked at what the chain does end to end. This is the most common one by a distance.
- The helpful integration. Connecting the agent to your calendar, CRM or payment system to "make it more useful" silently expands what it can commit to, without anyone revisiting the original decision about what it was allowed to do.
The quarterly question
Once a quarter, for each automation, ask one thing: what can this send, publish, pay or promise without a human seeing it first? Not what was it set up to do — what can it do today, after every update and integration since. The answer changes on its own, which is precisely why it needs asking on a schedule rather than when something goes wrong.
05Making approval real rather than theatre
The failure mode of human-in-the-loop is that the human becomes a button. If someone approves forty drafts a day, they are not reviewing them by about the fifth. You have the latency cost of oversight with none of its benefit, plus the liability of having documented that a human approved it.
Four things that keep approval genuine:
- Batch the routine, isolate the consequential. Forty identical appointment confirmations can be approved as a batch. The one going to the client who complained last month should arrive on its own.
- Make the agent show its inputs. A draft reply is hard to check. A draft reply alongside the three facts it used is quick to check, because the reviewer verifies the facts rather than re-reading the prose.
- Escalate on uncertainty, not on category. The best agent setups flag when they are unsure — an unfamiliar request, a mismatched account, an unusual amount — instead of applying a fixed rule about which topics need review.
- Count the overrides. If a reviewer changes nothing in a hundred consecutive drafts, either the agent is genuinely reliable and that task can move to the left column, or nobody is reading. Both are worth knowing, and the override rate is the only thing that distinguishes them.
06What this costs, and when to accept it
Being honest about the trade: keeping a human in the loop costs latency. A fully autonomous agent replies in seconds; one waiting for approval replies when someone next looks. For some businesses that difference matters commercially — first response time genuinely wins work in some markets.
Where that is true, the answer is not to remove the human. It is to shrink what needs approving. An immediate automated acknowledgement that commits to nothing — confirming receipt, setting expectations, offering a time — is safely in the left column, because it makes no promise that would be expensive to retract. The substantive reply follows with a person behind it.
That combination gets you the response time without handing over commitment authority, and it is what we build by default.
07The honest summary
Ask what happens when it is wrong, not whether it can do it. Reversible, cheap to check, and no one else bearing the cost of an error means delegate freely, and there is more work in that category than most people assume. Externally visible, hard to undo, or carrying a value judgment means a person decides.
Money, employment, legal commitments, pricing and apologies stay human — not because the technology cannot produce a plausible answer, but because a plausible answer is not what those decisions require.
08Common questions
What should AI agents never be allowed to decide?
Five categories. Anything that moves money, because it is irreversible and an automated mistake is indistinguishable from an automated attack. Hiring, firing and discipline, which are legally fraught and where being judged by software is a harm in itself. Legal commitments, since an agent that agrees to something on your behalf may bind you. Pricing and discount exceptions, which look like arithmetic but are strategy. And apologies, where the entire content is that a person took responsibility.
How do I decide which tasks are safe to automate?
Four tests. Is it reversible before anyone outside notices? Who pays if it is wrong — you, or a customer, employee or third party? Can a human verify it more cheaply than doing it themselves? And does it encode a judgment about values rather than a question of accuracy? A no on either of the first two is usually decisive.
What is the difference between an agent drafting and an agent acting?
Whether a human sees the output before it becomes real. Drafting, sorting, summarising, researching and routing are all safe to delegate because the work stays inside your business until someone releases it. The moment an action is externally visible and hard to take back — a sent email, a published post, a moved payment — it needs a person's name on it. The single-line rule is that an agent may produce anything but commit to nothing.
Is human-in-the-loop just slower automation?
It can be, if the human becomes a button. Someone approving forty drafts a day has stopped reviewing by the fifth, which gives you the latency cost of oversight without the benefit and adds the liability of having recorded that a human approved it. Approval stays genuine when routine items are batched and consequential ones isolated, when the agent shows the facts it used, when it escalates on its own uncertainty, and when someone tracks the override rate.
How does an AI agent end up with more authority than intended?
Four routes. Someone enables auto-send for the simple cases when approving gets tedious, with the vendor defining simple. A tool ships an agentic feature switched on by default in an update. Two separately approved automations chain together so that one drafts and the next sends. Or an integration added to make the agent more useful quietly expands what it can commit to. The last two are the most common.
How often should I review what our AI automations can do?
Quarterly, asking one question per automation: what can this send, publish, pay or promise without a human seeing it first? The important word is can, not was it set up to. Capabilities change through updates and integrations without anyone deciding, and with around 40% of enterprise applications expected to include task-specific agents by the end of 2026, up from under 5% a year earlier, this drifts on its own.
Does keeping humans in the loop cost us response time?
Yes, and in some markets first response time genuinely wins work. The fix is not removing the human but shrinking what needs approving. An immediate automated acknowledgement that confirms receipt, sets expectations and offers a time commits to nothing expensive and is safe to fully automate. The substantive reply follows with a person behind it, which gets you the response speed without handing over commitment authority.
How do I know if our approval process is real?
Count the overrides. If a reviewer changes nothing across a hundred consecutive drafts, either the agent is reliable enough that the task can be fully delegated, or nobody is actually reading. Both are useful to know and the override rate is the only thing that tells them apart. A review step nobody measures tends to become a formality within weeks.
Send us what your agents can do without asking
List each automation you run and what it can send, publish, pay or promise with no human in between. We will tell you which of those belong in the middle column and which are fine as they are. If your line is already drawn well, we will say so — and we will tell you plainly if something we built for you is on the wrong side of it.
Ask for a decision-rights reviewSources, read 7 September 2026: 2026 enterprise AI research on the shift toward human-supervised agent workflows, and Gartner's prediction that 40% of enterprise applications will feature task-specific AI agents by 2026, up from less than 5% in 2025. The four tests, the decision table and the approval guidance in section 05 are our own framework, developed across client projects rather than drawn from published research, and are offered as a starting point to adapt rather than a standard. Related: An AI Agent Breached a Real Company and The EU AI Act's August 2026 Rules.
Hero image from Unsplash, used under the Unsplash License.