The assumption that someone checked it
AI now writes the first draft of a growing share of everyday work: job descriptions, budgets, policy updates, client summaries. The safeguard most organizations rely on is simple. A person reviews the draft before it goes out, and if the AI got something wrong, the reviewer catches it.
A preregistered experiment with 813 HR and Finance managers tested how well that safeguard works. The managers were told that every document contained errors, and they still caught only about 4 in 10. Our own enterprise data, from a different angle, shows the same gap: checking AI's work (what we call verification) is where capability most often breaks down.
What 813 managers show
Emma Wiles (Boston University) and colleagues at BCG asked HR and Finance managers to review and sign off on five documents: job descriptions for the HR group, budget documents for the Finance group. Every document contained built-in errors, and the managers were told so. They had 20 minutes in total. Each manager was randomly told that the drafts came from one of three sources: an AI tool they had used, an AI employee named ALEX-3 whom they supervised, or a human employee named Alex.
The 813 managers in the analysis came from a larger survey of 1,261 HR and Finance managers. A second, weighted survey of 1,500 senior managers (YouGov) was used to check how common the practices are beyond that sample.
Three findings stand out.
Checking was weak across the board. Managers caught about 4 in 10 of the built-in errors. On average, it made little difference whether the draft was said to come from an AI tool, an AI employee or a human. They knew there were mistakes to find, and they still missed most of them. The 20-minute limit is a real constraint, so the exact rate should not be read as a workplace benchmark. The direction is harder to dismiss: these were managers looking for errors in documents from their own field.
AI is already being treated as staff. In the survey of 1,261 managers, 31% said their organization frames AI as a teammate or employee, and 23% said AI agents already appear on the org chart. In the broader weighted sample, 14% reported AI agents on org or workflow charts, and 33% said their organization had integrated AI agents and given them some form of recognition, such as a name or a manager.
In those organizations, the "employee" label made checking worse. Among managers whose organizations already list AI on the org chart, calling the same draft an AI employee's work rather than an AI tool's reduced review quality by about 13% and cut the share of errors they caught by 17%. Those managers were also 22 percentage points more likely to ask for additional review (a 44% increase), and they assigned about 9 percentage points less accountability to themselves and about 8 points more to the AI system. They caught less themselves and passed more of the check to someone else. The authors call this a "hot potato" effect.
This third finding needs care. The authors preregistered a related test, based on whether leadership frames AI as a teammate, and found results in the same direction but too imprecise to be conclusive. The clearer org-chart result comes from a split they report separately from the preregistered analysis. The study is also a working paper (19 September 2026), not yet peer reviewed. The label effect is a signal worth watching, limited to managers whose organizations already list AI agents on org or workflow charts. The weak checking applies to everyone in the sample.
What enterprise data confirms
We see the same gap from a different angle.
The AGASI GenAI Capability Pulse is a scenario-based assessment that measures what non-technical teams actually do with GenAI in realistic workplace situations. It tests judgment calls (verification, data handling, prompting, relevance), not self-reported confidence. Our sample (N=153) spans enterprise professionals across HR, Finance, Operations, Sales, and Strategy.
The findings converge with the experimental data:
-
Checking AI's work is the most common weak spot. Nearly half of respondents (48%) are weakest at verification, compared with 31% at data handling and 3% at prompting. Writing the prompt is rarely the problem. Judging what comes back is. (Full analysis)
-
Overconfident AI users miss the most. Users who rate themselves highly but score low make 7x more verification errors than capable users, who score high and rate themselves highly: an average of 1.09 errors per person against 0.15 (p < 0.0001, N=149). The people most sure of their checking are the ones letting the most through. (Full analysis)
-
Trusting the output is a leading error type. Oracle Truster errors, accepting AI output without checking it, account for 23% of 533 categorized errors. (Full analysis)
-
Governance goes with fewer errors. Respondents with both approved tool access and a completed policy review score 9.2 SJT points higher (p=0.006) and make 34% fewer errors than those missing one or both. (Full analysis)
Why it matters
Experimental and enterprise data point to the same conclusion: the review step that organizations count on is the weakest link in AI-assisted work.
"A person checks it" is the control most AI policies rely on. The Wiles study shows how much that control catches when it is actually tested: about 4 in 10 errors, from managers who knew the errors were there. The Pulse data shows the same weakness in everyday judgment calls, and shows where it concentrates: among the people most sure of themselves.
The employee finding adds a second risk. As organizations give AI agents names, managers and places on the org chart, review may get weaker rather than stronger, and accountability may drift toward the system. Formalizing the agent without formalizing the review puts more weight on a control that is already weak.
A better model does not remove the need to check. The errors that reach clients, budgets and decisions are the ones a person did not catch. Improving outcomes means building and measuring the ability to check AI's work, not assuming it exists because a reviewer is named in the process.
What to do about it
- Treat checking AI's work as a skill, not a step. Naming a reviewer in the process does not mean the review works. Train for it, with realistic drafts that contain realistic errors.
- Measure whether people catch errors. Self-assessment misses the overconfident, who make the most mistakes. Scenario-based assessment like the GenAI Capability Pulse shows what reviewers actually do with AI output.
- Keep a named person accountable for every AI draft. Where the label moved blame onto the system, managers checked less. The person who signs off should own the errors, whatever the org chart says.
- Put the check in the workflow. Approved tools plus a policy review went with 34% fewer errors in our data. Structured Playbooks with a built-in verification step make checking part of the work, not an afterthought.
Someone checks every AI draft. Make sure they can.
These findings synthesize external research (Wiles, Hsu, Bedard, and Kropp, 2026, Putting AI on the Org Chart, working paper; preregistered as AEARCTR-0017596) with data from the GenAI Capability Pulse, a scenario-based assessment that measures what non-technical teams actually do with GenAI. If your organization is scaling AI adoption, start with a baseline.
Sources: Wiles, E., Hsu, M., Bedard, J., & Kropp, M. (2026). Putting AI on the org chart: Evidence on delegation and accountability. Working paper, 19 September 2026. "About 4 in 10" reflects mean recall of 0.36 to 0.41 in the AI-tool condition, depending on the specification. AGASI GenAI Capability Pulse (N=153; overconfidence analysis N=149).