TL;DR. Delegating work to an AI agent delegates the task, never the accountability. In February 2024 a British Columbia tribunal rejected an airline’s argument that its chatbot was a separate entity responsible for its own actions, calling the submission “remarkable” and holding the company responsible for all the information on its website. That principle is simple. Practice inside organisations is not. The KPMG and University of Melbourne 2025 global study of 48,340 people found that 66% of employees had relied on AI output without evaluating its accuracy, 56% had made mistakes in their work because of AI, and 57% had hidden their use of AI or presented AI-generated content as their own. Only two in five said there was a policy guiding the use of generative AI tools at work. In a separate survey reported by Stanford’s AI Index, 14% of respondents said their organisation had dedicated AI governance roles. The rules already describe what a working responsibility map looks like: the EU AI Act requires deployers of high-risk systems to assign human oversight to people with the competence, training and authority to exercise it, and the NIST AI Risk Management Framework asks organisations to document who is responsible for what in human-AI configurations. This piece joins the observed behaviour to those requirements and turns the gap into a remap managers can apply to any team using agents.
The argument that should never be made again
The case is small in money and large in principle. A customer asked an airline’s website chatbot about bereavement fares, was given incorrect information about claiming a refund after travel, and relied on it. When the customer sought the difference, the airline argued in effect that the chatbot was a separate legal entity responsible for its own actions.
The Civil Resolution Tribunal of British Columbia, in Moffatt v. Air Canada, 2024 BCCRT 149, called this “a remarkable submission”. The tribunal member wrote that while a chatbot has an interactive component, it is still part of the company’s website, and that the company is responsible for all the information on its website whether it comes from a static page or a chatbot. The tribunal found the airline had not taken reasonable care to ensure its chatbot was accurate, and ordered a total of $812.02, made up of $650.88 in damages plus pre-judgment interest and tribunal fees.
It is one decision from a small-claims tribunal in one Canadian province, and it binds no one elsewhere. Its value is that it states, in ordinary language, the principle every manager using agents needs to internalise: the tool that produced the output is not the party that answers for it.
The problem is that most teams have not rebuilt their working arrangements around that principle. The responsibility map they use was drawn for a world in which every task was performed by a person who could be asked why.
Why the old responsibility map breaks
Most teams run on some version of a responsibility matrix, formal or not: someone does the work, someone is accountable for the result, someone is consulted, someone is informed. The design assumes that the person doing the work can also explain it, notice when it is going wrong, and stop.
An AI agent breaks that assumption in three places at once.
- The doer cannot be answerable. An agent can perform a task, but it cannot be held to account, disciplined, or asked to take ownership of a consequence. The “Responsible” cell is filled by something that cannot carry responsibility.
- Review turns into approval. When the output looks fluent, checking it tends to decay into signing it off. The research on automation bias, which we covered in Human-in-the-Loop by Design, describes exactly this tendency to over-rely on automated output.
- The work becomes invisible. If people use agents without saying so, the manager’s map shows a person doing a task that a system actually did. Nobody can manage a risk they cannot see.
The data on each of these is now available.
The accountability gap inside organisations: verified data
Verified data: how people actually use AI at work, and how organisations are governing it
| Measure | Value | Source |
|---|---|---|
| Sample | 48,340 people in 47 countries; 32,352 employees answered the questions about AI use at work | KPMG and University of Melbourne, 2025 |
| Employees who intentionally use AI at work regularly | 58% | KPMG and University of Melbourne, 2025 |
| Employees who have relied on AI output without evaluating the information it provided | 66% | KPMG and University of Melbourne, 2025 |
| Employees who have made mistakes in their work due to AI | 56% | KPMG and University of Melbourne, 2025 |
| Employees who have hidden their use of AI or presented AI-generated content as their own | 57% | KPMG and University of Melbourne, 2025 |
| Employees who say there is a policy guiding the use of generative AI tools | Two in five | KPMG and University of Melbourne, 2025 |
| Organisations with dedicated AI governance roles | 14% of respondents | McKinsey survey reported in Stanford HAI AI Index 2025 |
| AI-related incidents reported to the AI Incident Database, 2024 | 233, up 56.4% on 2023 | Stanford HAI AI Index 2025 |
Reading notes. “Regular” use in the KPMG study includes people who use AI every few months, not only daily users. The 66% and 56% figures count anyone who reported the behaviour at least rarely; the share doing it habitually is smaller. The 57% is a combined measure across two separate questions. The incident count relies on media reports and is likely to understate the true number. These qualifications make the figures more precise, not less worrying: the behaviours are widespread, even if not universal.
Put the rows together and the gap is clear. Most employees are already using AI, a majority have used its output unchecked, a majority have made mistakes because of it, a majority have at some point kept that use out of sight, and most work inside organisations that have not told them the rules or named who owns the risk.
What the frameworks already require
Managers do not need to invent a responsibility structure from scratch. Two widely referenced frameworks describe one.
The EU AI Act. Regulation (EU) 2024/1689 places specific obligations on deployers of high-risk AI systems in Article 26. Deployers must assign human oversight to natural persons who have the necessary competence, training and authority, as well as the necessary support (Article 26(2)). They must monitor the system’s operation on the basis of its instructions for use, and suspend use and inform the provider and authorities if they have reason to consider it presents a risk (Article 26(5)). They must keep the logs the system generates, to the extent those logs are under their control, for at least six months unless other law provides otherwise (Article 26(6)). Employers must inform workers’ representatives and affected workers before putting a high-risk system into use at the workplace (Article 26(7)).
Two dates matter. The amending Regulation (EU) 2026/1744, part of the Commission’s digital omnibus, entered into force on 27 July 2026 and moved the application date of these high-risk obligations: they now apply from 2 December 2027 for systems classified as high-risk under Annex III, which includes employment uses, and from 2 August 2028 for systems covered by Annex I product legislation. The same amendment rewrote the AI literacy duty in Article 4 so that providers and deployers must take measures to support the AI literacy of their staff, without being required to guarantee any specific level for any individual.
Most office agents that summarise, draft and research are not high-risk systems under the Act, so Article 26 will not legally bind most readers. It is still the most carefully drafted description of a working oversight role available, and it is a sensible design benchmark for any team.
The NIST AI Risk Management Framework 1.0. Its GOVERN function is explicit about roles. GOVERN 2.1 asks that roles, responsibilities and lines of communication for managing AI risk be documented and clear to individuals and teams. GOVERN 2.3 asks that executive leadership take responsibility for decisions about AI risks. GOVERN 3.2 asks for policies and procedures that define and differentiate roles and responsibilities for human-AI configurations and oversight of AI systems.
Neither framework says “the AI is responsible”. Both say the same thing the tribunal said: a named human structure must own the outcome.
Joining the behaviour to the requirement
CEOtudent synthesis: where observed behaviour breaks the responsibility structure
| Observed behaviour | Share of employees | Responsibility element it breaks | Reference requirement | Remap fix |
|---|---|---|---|---|
| Relying on AI output without evaluating it | 66% | No named verifier with authority to reject | EU AI Act Art. 26(2); NIST GOVERN 2.1 | Name a Verifier for each class of agent output, separate from the person who briefed the agent |
| Making mistakes in work due to AI | 56% | No monitoring loop, no escalation owner | EU AI Act Art. 26(5) | Keep an error log and name who decides to pause an agent workflow |
| Hiding AI use or presenting AI output as own work | 57% | Delegation is invisible, so risk is unmanaged | EU AI Act Art. 26(7) (workplace transparency) | Make disclosure of agent use a team norm, not a confession |
| Working without a policy guiding generative AI use | About three in five do not report one | Roles and limits undocumented | NIST GOVERN 2.1 and 3.2 | Publish a one-page responsibility map per agent workflow |
| Dedicated governance roles in place | 14% of respondents | No executive owner of AI risk | NIST GOVERN 2.3 | Assign one accountable owner per workflow, at a level with authority to stop it |
| Agent actions not recorded | Not measured in these studies | No evidence trail when something goes wrong | EU AI Act Art. 26(6) logs of at least six months | Keep run logs for agent workflows that touch customers, money or people decisions |
Behaviour shares from KPMG and University of Melbourne, 2025, except the governance-roles row, which reports survey respondents in the McKinsey data cited by Stanford HAI. “About three in five” is derived by CEOtudent from the report’s finding that only two in five say a policy guides use. The mapping of each behaviour to a requirement and a fix is CEOtudent’s editorial synthesis; the frameworks do not reference these survey findings.
The Responsibility Remap
The remap replaces the single “Responsible” cell that an agent now fills with a set of human roles that together keep the outcome owned. In a small team one person may hold several roles. The rule is that each role has a name.
CEOtudent editorial framework: the Responsibility Remap for agent work
| Stage | Old map (person does the task) | Remapped role | What the role holder must be able to do | Cannot be delegated to the agent because |
|---|---|---|---|---|
| Brief | The doer interprets the request | Delegator | Specify the goal, constraints, sources and what “done” means | The agent executes the brief it gets; a vague brief is a human failure |
| Execute | The doer performs the work | Agent (tool) | Perform the task within the brief | Execution is the only stage the agent can own |
| Verify | The doer self-checks | Verifier | Check the output against the source, not against its fluency; reject it | Independent judgement is the control; the agent cannot verify itself |
| Release | The manager approves | Owner | Accept the consequence of the output going out | Accountability attaches to a person or organisation, never to the tool |
| Monitor | Problems surface through the doer | Monitor | Watch error patterns across runs, keep the log | Patterns only show up across many outputs |
| Stop | The doer stops when confused | Stop authority | Pause or switch off the workflow without asking permission | An agent does not reliably recognise when it should stop |
| Explain | The doer explains what happened | Owner, using the log | Reconstruct what the agent did and why the output was released | “The agent did it” is the answer the tribunal rejected |
This framework is CEOtudent’s editorial synthesis, drawing on the oversight roles in EU AI Act Articles 14 and 26 and NIST AI RMF GOVERN 2 and 3. It is a design pattern, not a legal compliance checklist.
Three design rules make the remap work:
- The Verifier and the Delegator should not be the same person for consequential outputs. The person who wrote the brief sees what they expected to see. For low-stakes work one person can hold both roles; for anything that reaches a customer, a contract or a decision about a person, separate them.
- The Owner must have the authority to stop the workflow. The EU AI Act’s wording is precise: competence, training and authority. An owner who cannot pause an agent is not an owner.
- Disclosure is designed in, not policed. If the 57% figure tells managers anything, it is that people hide AI use when they are unsure it is allowed. Teams that make agent use visible by default see their real risk; teams that do not see a map of work that no longer exists.
For the upstream skills this remap depends on, see How to Delegate to an AI Agent for writing the brief, The Manager-of-AI Playbook for evaluating output, and Should You Let AI Make the Decision? for which decisions should never be delegated at all.
Three failure patterns to watch for
“The agent did it.” A mistake is traced back and the explanation stops at the tool. This is the tribunal’s case in miniature. If your post-incident review ends with a system name instead of a role name, the remap is missing an Owner.
The rubber stamp. A Verifier exists on paper but approves nearly everything, because the output reads well and the queue is long. A useful signal is the rejection rate: a verification step that almost never rejects anything is either reviewing unusually reliable work or not really reviewing.
The orphaned chain. One agent’s output becomes another agent’s input, and no human owns the end-to-end result. Each step looks supervised; the whole is not. Assign ownership to the outcome, not to the individual steps.
A 30-day remap for a team
- Week 1: inventory. List every recurring task where an agent or AI tool now produces output that leaves the team. Ask people directly, and make clear that disclosure carries no penalty.
- Week 2: name the roles. For each workflow, write down the Delegator, Verifier, Owner and Stop authority. Where the same person holds all four on consequential work, split at least the Verifier.
- Week 3: set the verification standard. For each output class, define what the Verifier checks against, and start logging rejections and errors.
- Week 4: run a drill. Pick one workflow and walk through a hypothetical error end to end: who notices, who pauses it, who explains it, which log shows what happened. Fix whatever step had no name.
The result fits on one page per workflow. That page is also the document GOVERN 2.1 of the NIST framework describes, and the kind of oversight assignment Article 26(2) of the EU AI Act requires of high-risk deployers.
The CEO and the student
The CEO view. A chief executive cannot delegate accountability to a subordinate, and even less to software. What a CEO can do is design a structure in which every consequential output has a named owner who is competent, trained and empowered to stop it. Agents make that design job more important, not less, because they multiply the number of outputs a team produces without multiplying the number of people who can answer for them.
The student view. The verification role is a skill, and like any skill it decays when unused. Teams that keep checking agent output against sources keep the judgement that makes checking possible. Teams that stop checking lose it, and discover the loss only when an error has already gone out. Staying the student means treating every rejected output as information about where the agent, the brief or the reviewer fell short.
Frequently asked questions
If an AI agent makes a mistake, who is responsible?
The organisation and the people who deployed and released the output, not the tool. The Moffatt v. Air Canada decision rejected the idea that a chatbot is responsible for its own statements. It is a single tribunal decision rather than a binding precedent across jurisdictions, and this is not legal advice, but its reasoning matches the logic of the EU AI Act and the NIST framework.
Does the EU AI Act apply to the agents my team uses?
Article 26 applies to deployers of high-risk AI systems, such as systems listed in Annex III, which includes certain employment uses. Most agents used to draft, summarise or research are not high-risk. Following the digital omnibus amendment in force since 27 July 2026, the Annex III high-risk obligations apply from 2 December 2027. The AI literacy duty in Article 4 has applied since 2 February 2025 and was amended to require measures that support AI literacy.
Should employees disclose when they use AI for their work?
From an accountability standpoint, yes. Invisible delegation means the manager’s view of who did what is wrong, and a risk that cannot be seen cannot be managed. The KPMG study found 57% of employees had hidden AI use or presented AI output as their own, which usually signals unclear rules rather than bad faith.
Does every AI output need a second reviewer?
No. Match the review to the consequence. Internal drafts can be checked by the person who briefed the agent. Output that reaches customers, affects money, or informs decisions about people should have a Verifier who is not the Delegator.
What is the minimum viable version of this?
One page per agent workflow naming the Owner, the Verifier and the person who can stop it, plus a simple log of errors. That covers the core of what both frameworks ask for.
How long should agent logs be kept?
For high-risk systems under the EU AI Act, deployers must keep automatically generated logs under their control for at least six months unless other law provides otherwise. For ordinary workflows, keep logs long enough to reconstruct any output that could still be challenged.
Sources
Regulation (EU) 2024/1689 of the European Parliament and of the Council (Artificial Intelligence Act), Articles 4, 14, 26 and 113, Official Journal of the European Union, 12 July 2024, and the consolidated text of 27 July 2026.
Regulation (EU) 2026/1744 amending Regulation (EU) 2024/1689 (digital omnibus on AI), Official Journal of the European Union, 24 July 2026.
National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, January 2023, GOVERN categories 2 and 3.
Nicole Gillespie, Steven Lockey, Tabi Ward, Alexandria Macdade and Gerard Hassed, Trust, attitudes and use of artificial intelligence: A global study 2025, The University of Melbourne and KPMG, 2025.
Moffatt v. Air Canada, 2024 BCCRT 149, Civil Resolution Tribunal of British Columbia, decision of 14 February 2024.
Stanford Institute for Human-Centered Artificial Intelligence, Artificial Intelligence Index Report 2025, Chapter 3, Responsible AI.
This content was compiled with the support of AI following in-depth research, then written and prepared for publication by the CEOtudent editorial team.
This post is also available in:















