TL;DR: The question of whether to let AI decide is usually framed as a question about accuracy, and framing it that way guarantees a bad answer. Accuracy tells you the chance of being wrong. It tells you nothing about the cost of being wrong, which is what actually determines whether delegation is sane. Two bodies of law have already worked through the harder version of this problem. GDPR Article 22 gives individuals the right not to be subject to a decision based solely on automated processing where it produces legal or similarly significant effects, along with a right to obtain human intervention and to contest the outcome. The EU AI Act, Regulation (EU) 2024/1689, goes further: Article 14(4) lists five specific capabilities a human overseer must retain, Article 14(5) requires two separate humans to confirm certain biometric identifications, and Annex III names eight domains where the law simply refuses to trust unsupervised automation. Below: those five capabilities converted into five plain questions, an original five-factor scoring test with a worked table across eleven ordinary decisions, and verified survey data on where public trust actually sits, including a Bentley-Gallup poll of 3,270 US adults fielded 4-11 May 2026 that found Americans still judge AI to perform worse than people on every task studied.
Ask most people where the line is on letting AI decide and you get an answer about capability. Let it decide the small things. Do not let it decide the big things. Wait until it gets better.
That answer sounds reasonable and is close to useless, because it does not tell you what counts as small, it does not say what “better” would look like, and it collapses two completely different risks into one word.
There is a more precise way to think about it, and it did not come from a productivity book. It came from regulators who had to write down, in enforceable language, exactly when a machine may and may not be the one that decides.
Start with what the law already worked out
This is not an appeal to authority. It is an appeal to the fact that someone has already spent years on the hard version of this question, with lawyers arguing both sides, and the output is public.
GDPR Article 22 establishes that a data subject has the right not to be subject to a decision based solely on automated processing, including profiling, which produces legal effects concerning them or similarly significantly affects them. Where such processing is permitted, the controller must implement safeguards including, at minimum, the right to obtain human intervention, to express a point of view, and to contest the decision. Regulatory guidance treats automatic refusal of an online credit application and e-recruiting practices without human intervention as examples of the effects in scope.
Read the mechanism rather than the legalese. The trigger is not that the system might be inaccurate. The trigger is the combination of two things: the decision is made solely by the machine, and the consequence for the person is significant. Accuracy is not mentioned. A highly accurate system that refuses your loan with no human in the loop is still the thing the article is about.
The EU AI Act, Regulation (EU) 2024/1689, takes the same logic and gets specific. Article 14 requires high-risk systems to be designed so they can be effectively overseen by natural persons. Paragraph 4 then lists what the overseer must actually be able to do. It is, unexpectedly, one of the better decision checklists ever written, and it translates directly to personal use.
Table 1. The five oversight capabilities of AI Act Article 14(4), translated into personal decision questions (CEOtudent editorial framework, built on the verified text of Regulation (EU) 2024/1689)
| AI Act Article 14(4) requires the overseer to be able to | The personal question it becomes |
|---|---|
| (a) properly understand the relevant capacities and limitations of the system and monitor its operation, including detecting anomalies | Do I actually know what this tool is bad at, or only what it is good at? |
| (b) remain aware of the possible tendency of automatically relying or over-relying on the output (automation bias) | If I found myself agreeing every time, would I notice? |
| (c) correctly interpret the system’s output, taking into account available interpretation tools | Can I read the output well enough to tell a good one from a confident wrong one? |
| (d) decide, in any particular situation, not to use the system or to disregard, override or reverse the output | Do I have a live alternative if I reject it, or am I committed the moment I ask? |
| (e) intervene in the operation or interrupt the system through a stop button or similar procedure | Can I stop this before it takes effect on anyone but me? |
Paragraph 5 adds one more rule worth borrowing personally. For remote biometric identification, no action or decision may be taken on the basis of the system’s identification unless it has been separately verified and confirmed by at least two natural persons with the necessary competence, training and authority. The law’s answer to the highest-stakes category is not a better algorithm. It is a second human.
Annex III lists the domains the Act designates as high-risk in the first place: biometrics; critical infrastructure; education and vocational training; employment and worker management; access to essential private and public services including creditworthiness and insurance pricing; law enforcement; migration, asylum and border control; and the administration of justice and democratic processes. Systems in these categories placed on the market after 2 August 2026 fall under the regime.
Look at what those eight have in common. They are not the hardest problems. Some are computationally trivial. They are the decisions where the person affected cannot easily undo the outcome, cannot readily verify the reasoning, and did not choose to be in the process. That is the pattern, and it is the pattern you can use.
The variable everyone gets wrong is not accuracy
Here is the case for taking the reversibility framing seriously rather than the accuracy framing.
In 2025, METR ran a randomised controlled trial with 16 experienced open-source developers working 246 real tasks in their own repositories, with tasks randomly assigned to allow or disallow AI tools. Before the study, participants forecast that AI would speed them up by 24%. After doing the work, they estimated it had sped them up by 20%. Measured, they were 19% slower. METR now labels the result historical, noting it does not necessarily reflect current tools or workflows, and that caveat should be taken seriously.
The number that survives the caveat is not the 19%. It is the 39-point gap between what expert practitioners measured and what those same expert practitioners believed, in their own domain, about work they had just finished.
That gap is the whole problem with accuracy-based delegation rules. A rule that says “delegate when the AI is reliable enough” requires you to know how reliable it was. The people best positioned to know were wrong by 39 points about themselves. The AI Act names this failure mode directly in Article 14(4)(b), which requires overseers to remain aware of the tendency to over-rely on automated output. Regulators built the checklist around it because they did not expect self-assessment to work either.
So stop asking how likely it is to be wrong. Ask what happens when it is.
The Delegation Boundary Test
Five factors. Score each 0, 1 or 2. Total from 0 to 10. It takes about ninety seconds once you have done it twice.
- Reversibility. If this turns out wrong, how cheaply can you undo it? 2 = fully reversible at near-zero cost. 1 = reversible with real cost or delay. 0 = irreversible.
- Verifiability. Before acting, can you check whether the output is correct? 2 = you can verify it directly and quickly. 1 = you can spot-check parts. 0 = you would have to trust it.
- Stakes and who holds them. Who absorbs the damage? 2 = only you, and it is small. 1 = only you, and it is significant. 0 = someone else bears it, or it is large.
- Your expertise. Would you catch a confident error in this domain? 2 = reliably. 1 = sometimes. 0 = you would not know.
- Repeatability. Does this decision recur often enough that a standing rule beats a fresh judgement each time? 2 = daily or weekly. 1 = occasionally. 0 = once, or nearly.
Bands:
- 8-10, Automate. Let the tool decide and act. Spot-check periodically rather than case by case.
- 5-7, AI drafts, you decide. The tool produces the recommendation. You make the call, and you are the one accountable for it.
- 2-4, AI informs only. Use it to gather, structure and stress-test. Do not let it produce the answer, because you will anchor on the answer.
- 0-1, Never delegate. Human decision, and for the bottom of this band, a second human as well.
Table 2. The Delegation Boundary Test applied to eleven ordinary decisions (CEOtudent editorial framework; Rev = reversibility, Ver = verifiability, Sta = stakes, Exp = your expertise, Rep = repeatability)
| Decision | Rev | Ver | Sta | Exp | Rep | Total | Band |
|---|---|---|---|---|---|---|---|
| Sorting your inbox into reply-today, later and ignore | 2 | 2 | 2 | 2 | 2 | 10 | Automate |
| Reordering the sections of a report for clarity | 2 | 2 | 2 | 2 | 2 | 10 | Automate |
| Whether to publish a finished post today or hold it | 2 | 2 | 1 | 2 | 2 | 9 | Automate |
| Which of three software tools to buy for a five-person team | 1 | 2 | 1 | 2 | 0 | 6 | AI drafts, you decide |
| What to quote a new client for a familiar piece of work | 1 | 1 | 1 | 2 | 1 | 6 | AI drafts, you decide |
| First-pass sort of 40 job applications by relevance | 2 | 1 | 0 | 1 | 1 | 5 | AI drafts, you decide |
| Which of last year’s expenses are tax-deductible | 1 | 1 | 1 | 0 | 1 | 4 | AI informs only |
| Whether to accept a job offer | 0 | 1 | 1 | 1 | 0 | 3 | AI informs only |
| Whether to sign a three-year office lease | 0 | 1 | 0 | 1 | 0 | 2 | AI informs only |
| Whether to end a business partnership | 0 | 0 | 0 | 1 | 0 | 1 | Never delegate |
| Choosing a treatment for a diagnosed condition | 0 | 0 | 0 | 0 | 0 | 0 | Never delegate |
Three things fall out of that table that are worth stating.
The top band is boring on purpose. Everything scoring 8 or above is small, reversible, verifiable and frequent. That is not a limitation of current models. It is the definition of the category. When people say AI should handle the small stuff, this is the structure underneath the intuition, and having the structure lets you find cases the intuition misses.
The job-applications row is the uncomfortable one. It scores 5, which puts it in “AI drafts, you decide”, and many organisations currently run it at “automate”. The factor that drags it down is stakes: the person absorbing a bad sort is the candidate, not you. This is precisely why Annex III of the AI Act designates recruitment and candidate evaluation as high-risk, and why GDPR guidance names e-recruiting without human intervention as an example of significant effect. When your delegation moves the cost onto someone who did not choose to be in the process, the score falls even if the task feels routine.
Low expertise pushes decisions down, not up. This is the counter-intuitive one, and it reverses how most people actually behave. The instinct is to delegate hardest where you know least, because that is where you feel least capable. But Factor 4 exists because delegation requires you to be able to catch a confident error, and in a domain you do not understand, you cannot. Not knowing the answer is a reason to use AI to learn, and a reason not to let it decide. The tax row scores 0 on expertise for exactly this reason.
What the public actually trusts, and where that diverges
Individual judgement is not formed in a vacuum, so it is worth knowing where collective intuition currently sits.
Table 3. Verified US survey data on AI use and perceived performance (Pew Research Center, February 2026; Bentley-Gallup, May 2026)
| Measure | Verified figure | Source and field dates |
|---|---|---|
| US adults who use AI chatbots | 49%, up from 33% in 2024 | Pew Research Center, 5,119 US adults, 17-23 February 2026 |
| Use chatbots daily | 24%, including 12% several times daily | Pew Research Center, same survey |
| Use them for searching for information | 42% | Pew Research Center, same survey |
| Employed adults using them for tasks at work | 38% | Pew Research Center, same survey |
| Use them for medical advice | 20% | Pew Research Center, same survey |
| Use them for emotional support or advice | 10% | Pew Research Center, same survey |
| Use them for companionship | 4% | Pew Research Center, same survey |
| Believe AI will make their personal information less secure | 71% | Pew Research Center, same survey |
| Say AI is advancing too quickly | 63% | Pew Research Center, same survey |
| Tasks on which Americans judge AI to perform worse than people | All tasks studied, in both 2023 and 2026 | Bentley University-Gallup Business in Society, 3,270 US adults, 4-11 May 2026, margin of error plus or minus 2.4 points |
| Largest year-over-year declines in “AI performs worse” | Driving a car, down 7 points; giving financial advice, down 6; providing medical advice, down 5 | Bentley-Gallup, same survey |
| Views on AI providing hiring recommendations | Essentially unchanged | Bentley-Gallup, same survey |
| Say AI assists students with homework better than people | 29%, up from 26% | Bentley-Gallup, same survey |
Set those two datasets side by side and something specific emerges.
Usage is rising steeply and confidence is not. Half of American adults use these tools, but on every task the Bentley-Gallup survey measured, more people still judge AI to perform worse than people than better. Meanwhile 20% report using chatbots for medical advice, a domain sitting near the bottom of the scoring test on every factor at once.
That combination is not hypocrisy. It is what happens when there is no shared vocabulary for the boundary. People use the tool because it is useful and available, and separately report distrust because distrust is the only word on offer. What is missing is the middle: “I used it to prepare for the appointment” and “I let it decide the treatment” are radically different acts that the survey question, and most everyday conversation, cannot distinguish.
The movement in the Gallup numbers is worth reading carefully too. The three tasks where perceived performance improved most, driving, financial advice and medical advice, are all high-consequence and low-reversibility. Perceived capability is climbing fastest precisely where the delegation test scores lowest. Capability and delegability are moving in opposite directions, and confusing them is the error the entire framework is designed to prevent.
Note what did not move: views on AI providing hiring recommendations were essentially unchanged. Public intuition has stayed put on the one item in the list where the person bearing the risk is not the person making the choice.
The rule that matters more than the score
If you retain one thing, make it this. Score reversibility first, and if it is 0, stop scoring.
An irreversible decision cannot be redeemed by a good process, an accurate tool, or a confident feeling. GDPR Article 22 is built on this: the safeguards it demands are the right to obtain human intervention, to express a point of view, and to contest the decision. All three are undo mechanisms. The law’s remedy for automated decision-making is not better automation. It is the preserved ability to reverse.
Your personal version of that safeguard is the same. Before delegating anything, ask what your undo looks like. If the honest answer is that there is not one, the decision belongs to you regardless of what any tool would have recommended.
The CEO and student read
The CEO half is delegation discipline, and it is the same discipline that applies to people. A competent executive does not delegate by seniority of the task. They delegate by whether the outcome can be inspected and corrected, and they keep the decisions where the correction window is closed. The five-factor test is that instinct made explicit, with the useful property that it produces the same answer whether the delegate is a person or a model.
The student half is harder and more important. Factor 4 says your own expertise determines what you may delegate, which means the boundary is not fixed. It moves as you learn. Every domain where you build real judgement is a domain where you can safely hand over more, because you gain the ability to catch the confident error. That is the argument against the comfortable position of outsourcing everything you find difficult: it permanently freezes the boundary where it is.
This piece is about which decisions to hand over. The adjacent question of whether a given output is right is a different skill, covered in trust calibration and when to trust AI recommendations and, at the level of assessing output quality directly, in the evaluation skill and judging AI output. For the wider architecture of deciding well, the personal decision stack sets out the layers this test sits inside, how to think in bets covers the probabilistic half that the reversibility factor deliberately sidesteps, and the judgment economy makes the case for why this capability is appreciating rather than depreciating.
Frequently asked questions
Does the test change as models improve?
Only through Factor 2, verifiability, and Factor 4, your expertise. Reversibility, stakes and repeatability are properties of the decision, not the tool. A decision that cannot be undone stays undelegatable no matter how good the model gets, which is why the framework does not need revising every model release.
What if I score a decision at 7 and it goes badly?
That is the band working as designed. In the 5-7 band the tool recommends and you decide, so accountability sits with you. If you find yourself accepting every recommendation without changing anything, you are not in that band any more, you are in the automate band without having decided to be. Article 14(4)(b) exists for this exact drift.
Is using AI for medical or legal questions always in the bottom band?
Using it to understand your situation, prepare questions and check terminology is a different act from letting it decide, and only the second is a bottom-band delegation. The distinction is whether the output determines the action or informs a decision a qualified human makes.
Does this apply to agents that act autonomously rather than chat?
More sharply, because agents collapse the Article 14(4)(e) capability, the ability to intervene before an effect lands. A system that acts as soon as it concludes has removed the stop button, which is why the AI Act treats it as a design requirement rather than a user habit. Score anything an agent will execute unsupervised at least one band lower than the same decision made in a chat window.
Why does high stakes for someone else score lower than high stakes for me?
Because you can consent to a risk you bear and cannot consent on someone else’s behalf. This is the principle behind Annex III, where nearly every listed domain involves decisions imposed on a person who did not choose to be in the process.
Is there ever a case for delegating an irreversible decision?
Only where a human being cannot do it at all, such as speeds beyond human reaction time, and those cases are engineered systems with their own safety regimes rather than personal choices. For an individual deciding in a chat window, no.
Sources
Regulation (EU) 2024/1689, the Artificial Intelligence Act, Official Journal version of 13 June 2024, Article 14 and Annex III.
General Data Protection Regulation, Regulation (EU) 2016/679, Article 22.
Information Commissioner’s Office, guidance on rights related to automated decision making including profiling.
METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, published 10 July 2025.
Pew Research Center, Americans and AI 2026: chatbots, smart devices and views on impact, published 17 June 2026, survey of 5,119 US adults fielded 17-23 February 2026.
Bentley University and Gallup, 2026 Bentley Business in Society survey, probability-based Gallup Panel web study of 3,270 US adults fielded 4-11 May 2026.
This content was compiled with the support of AI following in-depth research, then written and prepared for publication by the CEOtudent editorial team.
This post is also available in:














