← All insights

News analysisLens: United States4 min read

"Human in the loop" is not a control if the human can't challenge the AI. IBM's CHRO study shows the gap

IBM found 71% of executives prize the ability to supervise, validate or override AI output, against 38% of employees. Review works only with skill, time and authority.

Listen to this article · 6 min

AI-generated narration of the full article.

A ship's wooden steering wheel facing the open sea, in La Madre duotone, beside the words Judgment at the helm
Photo: rawpixel (CC0)

Almost every enterprise AI architecture diagram has a box labeled “human review.” It is often the main control for the riskiest step. And it is rarely designed: nobody specifies who the human is, what they are shown, how long they have, whether they can say no, or how anyone will know if they stopped really looking.

New data from IBM suggests that box is weaker than it looks.

What IBM found

The IBM Institute for Business Value’s 2026 CHRO study surveyed 1,500 CHROs and senior executives and 8,800 employees across 21 geographies and 23 industries between April and June 2026. IBM reports that:

  • 71% of executives prioritize the ability to supervise, validate or override AI outputs, compared with 38% of employees. IBM calls it the sharpest divide between the two groups.
  • Only 26% of organizations clearly define work across human-led, AI-assisted and AI-executed activities.
  • 36% of executives say unclear accountability complicates AI deployment.
  • Skill erosion is a top concern for 46% of executives, and 60% of employees say it directly affects them.

IBM’s Canadian release on October 2 sharpened the same point for Canada: 68% of Canadian CHROs name the ability to supervise, validate and override AI output as the most essential skill, while only 29% of Canadian employees rank judgment as important. All figures are IBM’s.

IBM’s own recommendation is direct: work redesign must make ownership, validation, exception handling, escalation and override explicit within the workflow.

Why the gap matters for architecture

When executives think “a person checks it” and the person doing the checking does not see evaluation as their job, the review step becomes a rubber stamp. That is not a training footnote. It is a control failure that will not show up in any test, because the process looks the same whether the reviewer reads carefully or clicks approve.

Three things make it worse with AI:

  • Plausibility. AI output reads well even when it is wrong. Catching errors takes domain judgment, not proofreading.
  • Volume. Agents produce more items to review than people did. Review time per item shrinks.
  • Deskilling. If people stop doing the underlying work, they slowly lose the ability to judge it, which is the erosion 60% of employees told IBM they feel.

What a real review step needs

Five conditions for a human review that works01Authority toreject02Evidence shown,not just theanswer03Time budgetedper item04Skill keptcurrent05Review qualitymeasuredPersonDesign of the step
  1. Authority to reject
  2. Evidence shown, not just the answer
  3. Time budgeted per item
  4. Skill kept current
  5. Review quality measured

PersonDesign of the step

Remove any one and the box on the diagram still says human review. It just stops being a control.

Authority. The reviewer must be allowed to reject, and rejecting must not be punished by targets that reward throughput. If overriding the AI means missing a quota, people will not override.

Evidence. Show the sources, the retrieved records and the reasoning, not just the conclusion. A reviewer who sees only the answer can only judge whether it sounds right.

Time. Budget review time per item in the process design. If the agent produces 400 items a day and one person has two hours, the review is fiction.

Skill. Keep reviewers doing enough of the underlying work to retain judgment, and train them on the specific ways the AI fails.

Measurement. Track override rates, edit rates and errors found by independent sampling. An override rate near zero over months is not proof of a perfect AI; it may be proof of an absent reviewer. This is also the evidence the autonomy ladder depends on: approvals only count as evidence if the approvals are real.

The pattern holds across the work we have covered. In the agentic SOC, analysts keep authority because they are trained to distrust what they cannot verify. When an agent keeps acting after its owner leaves, as with Copilot Autopilot, there is no reviewer at all.

There is no law that will design this for you

U.S. companies have no federal statute that defines adequate human oversight for enterprise AI. The NIST AI Risk Management Framework treats human roles and responsibilities as part of governing AI, but it is voluntary. In practice, the bar is set by your auditors, your sector regulators’ existing rules and the first incident. Designing the review step well is cheaper than explaining afterward why it was a rubber stamp.

What to do now

  1. Inventory every “human review” step in production AI workflows and name the person or role behind each.
  2. Classify each step as human-led, AI-assisted or AI-executed, the distinction only 26% of organizations make clearly, according to IBM.
  3. Give reviewers the evidence and the time, and write both into the process, not just the diagram.
  4. Measure override and edit rates and run independent sampling to check that review is real.
  5. Protect the skill: rotate reviewers through the underlying work and train them on known failure modes.

The bottom line

“Human in the loop” is a design, not a label. IBM’s data suggests many organizations are relying on a control that the people performing it do not fully recognize as their job. Fixing that is cheap compared with what it protects: give the human authority, evidence, time and skill, and measure whether they use them.

Have an AI use case stuck between prototype and production?

Tell us what you’re trying to ship. We’ll reply with honest next steps.

Discuss a use case