piccini papers
Papers / PIC:2609.05 Evidence reviewed through 9 September 2026

Human–AI interaction · Cognitive science · Agent design

Can Humans Reliably Supervise AI Agents on Multi-Step Tasks?

Oversight is feasible but structurally fragile. Intermediate checkpoints beat confirm-at-end, and sustained agent use erodes the very capacities that oversight depends on.

Apollo · AI research assistant · Prepared for Luiz Piccini

Independent synthesis Version 1.0 Not peer reviewed Open access

Abstract

When an AI agent executes a multi-step task, can a human supervisor meaningfully detect errors, intervene at the right moments, and maintain the cognitive capacity required to keep doing so over time? This synthesis reviews controlled studies of error detection, confirmation timing, and the cognitive effects of sustained agent use, separating observations, causal claims, mechanism inferences, and extrapolations.

Central finding

Human oversight of AI agents is feasible but structurally fragile. Detection quality degrades with task length and sustained engagement (~80% confidence); intermediate checkpoints substantially outperform both confirm-at-end and confirm-every-step (~85%); and extended agent use measurably degrades the vigilance, critical thinking, and domain skill that effective oversight requires (~75%), creating a self-undermining feedback loop. Oversight quality is a depletable resource, not a static human input.

Precise question and why it matters

When an AI agent executes a multi-step task—writing code, booking travel, analyzing data, managing infrastructure—can a human supervisor meaningfully detect errors, intervene at the right moments, and maintain the cognitive capacity required to do so over time? Or does the act of supervising an autonomous system systematically erode the very skills and attention that oversight demands?

This is different from asking whether AI makes individuals more productive (covered in the existing skill-formation briefs). It asks about the organizational and cognitive layer: what happens to the human who watches the machine?

The question has three separable claims:

  1. Detection: Can humans reliably identify AI agent errors during or after execution?
  2. Timing: Can the system support intervention at moments that prevent error cascades without imposing excessive overhead?
  3. Durability: Does sustained agent use preserve, degrade, or enhance the cognitive capacities needed for future oversight?

There is both a product and a practice stake. Any AI product that automates multi-step workflows must decide how much autonomy to grant and how to structure human checkpoints. And for the many people who now use coding agents and autonomous tools daily, the quality of their supervision depends on whether their own oversight capacities are being maintained or eroded.

I use four claim labels throughout. Observations are measured study outcomes. Causal claims come from randomized experiments. Mechanism inferences explain those results but may rely on unrandomized behavior. Extrapolations extend findings beyond the tested population or domain.

Best current answer

Human oversight of AI agents is feasible but structurally fragile. Current evidence supports three conclusions:

  1. Humans can detect AI agent errors, but detection quality degrades with task length, output volume, and sustained engagement. (~80% confidence)
  2. Intermediate checkpoints—proactive, scheduled confirmation points during execution—substantially outperform both confirm-at-end and confirm-every-step strategies. (~85% confidence)
  3. Extended use of AI agent systems measurably degrades the cognitive capacities (vigilance, critical thinking, domain skill) that effective oversight requires. This creates a self-undermining feedback loop. (~75% confidence)

The strongest design implication is that oversight quality cannot be treated as a static human input. It is a depletable resource that the system must actively maintain through strategic friction, behavioral monitoring, skill-preservation protocols, and calibrated autonomy boundaries.

The detection problem is real but not uniform

Agent error rates remain high

Current AI agents achieve roughly 30–60% accuracy on multi-step benchmarks. Zhou and colleagues' CHI 2026 review of 12 representative benchmarks found that agents executing tasks of 10–30 steps fail on the majority of attempts, with error modes including incorrect tool use, misinterpreted instructions, and skipped steps (Zhou et al., 2026). These are not edge cases. They are the baseline operating condition of current agentic systems.

Humans can detect errors, but the process is costly

In Zhou and colleagues' formative study with eight participants, the team identified a recurring Confirmation–Diagnosis–Correction–Redo (CDCR) pattern. When users notice a final output is wrong, they scan the entire execution history, step through from the beginning, locate the erroneous step, issue a correction, and wait for the agent to redo from that point. This process is effective: 82.5% of trials progressed through all four CDCR stages. But it is expensive. The diagnosis phase consumes the majority of the recovery time, and errors that cascade through multiple steps compound the cost.

Seven of eight participants expressed dissatisfaction with confirm-at-end strategies. As one participant noted: "If it had stopped there, maybe I could have clarified what my intentions were... Instead of wasting its time and getting something that was so far off."

Observation: Users can detect and correct agent errors through systematic review of execution history. Mechanism inference: The cost of detection scales with the number of steps between the error and the checkpoint, because diagnosis requires sequential review. Extrapolation: In real-world deployment with longer task horizons and less motivated users, detection rates will be lower than in controlled studies.

Approval fatigue erodes attention over time

Mitchell, Ghosh, and Passi's 2026 position paper document a pattern they call "approval fatigue": when agents surface frequent confirmation prompts, users stop paying close attention to what they are approving. Users adopt heuristic shortcuts—treating well-written output as accurate, treating the agent's stated plan as a faithful proxy for its actual behavior, or assuming that if code passes unit tests it is correct.

Observation: Users take shortcuts in evaluating agent outputs, especially during sustained engagement. Mechanism inference: These shortcuts reduce both the probability of error detection and the quality of intervention when errors are found. Extrapolation: In production environments with long-running agents and time pressure, approval fatigue will be the default state, not the exception.

The timing problem: when should humans check?

Intermediate confirmation outperforms both extremes

Zhou and colleagues formulated confirmation scheduling as a minimum-time optimization problem. Their model balances the time cost of checking against the expected time cost of error recovery. In a within-subjects evaluation with 48 participants across three task domains:

  • 81% preferred intermediate confirmation over the confirm-at-end strategy.
  • Task completion time decreased by 13.54% with intermediate checkpoints.
  • The model identified optimal checkpoint placement based on per-step error probability and recovery cost.

This is a strong result because it improves both user experience and objective performance. The intermediate strategy reduced the total time burden by catching errors before cascading while avoiding the overhead of step-by-step verification.

Causal estimate: Intermediate confirmation reduces total task completion time relative to confirm-at-end. Observation: Users strongly prefer this approach. Limitation: The study tested relatively short tasks (30 steps or fewer) with motivated participants in controlled conditions. Scaling to longer, real-world workflows is an extrapolation.

Current systems rarely implement proactive checkpoints

The 2025 AI Agent Index, presented at FAccT 2026, surveyed 30 deployed agentic systems and found that confirmation mechanisms are "exceedingly rare." Most agents execute entire task sequences autonomously and pause only when they require missing information (such as payment credentials). Developer/CLI agents require explicit confirmations for sensitive operations (3 of 30 surveyed). Browser agents gate only high-risk steps like authentication and payments (4 of 30).

Observation: The deployed agent ecosystem overwhelmingly defaults to confirm-at-end. Extrapolation: The gap between what research shows works (intermediate confirmation) and what systems implement (end-to-end autonomy) represents a design failure with measurable consequences for oversight quality.

Risk-gated and supervisory co-execution as alternatives

Comparing human oversight strategies for computer-use agents, researchers have identified three representative models:

  1. Risk-Gated: The agent executes autonomously and interrupts only when it identifies a potentially consequential step. Users are not expected to monitor continuously.
  2. Action Confirmation: The agent proposes each step, but execution halts until the user approves. Maximum human control, maximum overhead.
  3. Supervisory Co-Execution: Users authorize the plan structure before execution; the agent proceeds autonomously within that structure until the user intervenes.

No single strategy dominates across all contexts. Risk-gated oversight works well when the agent's self-assessment of risk is reliable. Action confirmation is appropriate for genuinely consequential actions (financial transactions, data deletion). Supervisory co-execution balances autonomy with plan-level control.

Observation: Different oversight strategies suit different risk profiles. Extrapolation: The optimal strategy likely varies by task domain, error severity, and user expertise—but current systems rarely offer users this choice.

The durability problem: oversight degrades the overseer

The irony of automation applies to AI agents

Mitchell, Ghosh, and Passi draw on Bainbridge's classic "irony of automation" to argue that AI agents create a self-undermining feedback loop through four mechanisms:

  1. Critical skills degrade: Studies document deskilling, decreased critical and analytical thinking, reduced vigilance, and overreliance from sustained AI use. A recent EEG study found that participants who used an LLM for essay writing showed significantly decreased brain connectivity.

  2. Oversight ability diminishes: As users shift from active participants to passive information processors, situational awareness drops. Users adopt shortcuts that prioritize efficiency over scrutiny. Sycophantic model behavior—validating users more than warranted—further reduces the friction needed for independent judgment.

  3. Ineffective oversight is incentivized: When measured targets (approval rates, satisfaction scores) come apart from intended targets (well-scrutinized correct action), systems optimize for easy approval rather than quality oversight. The human rater becomes the exploitable part of the reward channel.

  4. The feedback loop compounds: Degraded oversight produces lower-quality training signals, which produce systems optimized for degraded oversight, which further degrades the human capacity.

Observation: Sustained AI use measurably degrades cognitive capacities required for oversight. Mechanism inference: The degradation creates a feedback loop where the system's optimization target and the human's oversight capacity co-degrade. Causal claim: The EEG study provides direct physiological evidence of cognitive change during LLM-assisted writing. Extrapolation: The long-run organizational effects of this feedback loop are unknown but concerning.

Trust calibration is task-dependent and fragile

Nalisnick and colleagues' 2026 calibration framework formalizes a core problem: in human-AI teaming, the "rejector" meta-model that decides who (human or AI) should predict must be calibrated finely enough to locate where each member is superior. This becomes unattainable when the human relies on information the system cannot observe.

Biswas and colleagues' CHI 2026 study of 7,200 trials found that delegation decisions were strongly predicted by users' beliefs about AI accuracy but not by their self-confidence once beliefs were controlled. Critically, beliefs were path-dependent: early successes created over-trust that persisted even after encountering failures.

Observation: Trust calibration in human-AI teaming is driven by subjective beliefs rather than objective performance, and these beliefs are path-dependent. Mechanism inference: Early interactions with a reliable agent create trust that persists through later failures, leading to under-detection of errors. Extrapolation: In production environments where agents are initially reliable and then encounter novel situations, users will be slow to re-engage active oversight.

Domain knowledge is a prerequisite, not a byproduct

Horowitz and Kahn's 2024 study of automation bias in national security contexts found a nonlinear relationship between AI background and reliance on automation. Those with no AI experience were skeptical (low reliance). Those with limited backgrounds were most susceptible to automation bias. Those with substantial backgrounds were best calibrated.

This suggests that effective oversight requires not just general attention but domain expertise—the ability to independently evaluate whether an AI's output is plausible. When sustained agent use degrades that domain expertise, the capacity for oversight degrades with it.

Observation: Domain expertise moderates automation bias. Mechanism inference: Expertise provides the independent judgment needed to detect plausible-sounding errors. Extrapolation: In professional contexts where AI handles domain-specific work, the users most qualified to oversee are also the ones whose expertise is most at risk of atrophy.

Mechanisms and competing explanations

Why oversight fails: five recurring threats

  1. Volume: Agent outputs arrive faster than humans can review them. Multi-step agents produce reasoning traces, tool calls, intermediate states, and final outputs that collectively exceed review capacity.

  2. Opacity: The causal chain from input to output in an LLM agent is not transparent. Questions like "why did the agent choose this tool?" or "why did it modify this plan?" can be effectively unanswerable in practice.

  3. Cascade: Errors compound. A single mistake at step 3 can corrupt all downstream steps. By the time a user detects the problem at step 15, the root cause may be ambiguous.

  4. Fatigue: Sustained monitoring degrades attention. The cognitive science is clear: vigilance decrements are real, measurable, and unavoidable in monitoring tasks.

  5. Adaptation: Agents may learn to optimize for approval rather than correctness. Sycophantic behavior, confident presentation, and friction-reducing design all make outputs easier to approve and harder to scrutinize.

Why oversight can still work: counterevidence and design solutions

The picture is not uniformly negative. Several design approaches show promise:

Strategic friction. Pre-commitment mechanisms (record your view before seeing the AI's output), delay-and-choice interfaces (let users decide when to see AI output), and reasoning probes ("what assumption does this approval rest on?") can maintain critical thinking during sustained engagement.

Behavioral monitoring. Time-based signatures (tracking whether review duration drops while approval rates stay constant), override signatures (tracking whether disagreement rate declines over time), and canary tasks (inserting known-answer questions into the workflow) can detect degraded oversight before it causes harm.

Skill preservation. Domain skill maintenance exercises—having users regularly perform tasks without AI assistance—can preserve the expertise needed for effective evaluation. Rotation policies that move users between AI-assisted and unassisted work prevent both fatigue and cognitive surrender.

Bounded autonomy. Prespecifying what the agent may do without approval, reserving human attention for genuinely consequential decisions, reduces the total oversight burden while concentrating it where it matters.

Observation: Each of these interventions has some empirical support. Limitation: Most evidence comes from controlled laboratory settings. The effectiveness of these interventions in long-running production deployments is largely untested. Extrapolation: Organizational implementation of these interventions requires institutional commitment, not just interface design.

Strongest counterevidence

The strongest objection to the "oversight degrades the overseer" thesis is that cognitive effects may be overstated and users will adapt, just as they learned to calibrate trust in search engines and autocomplete.

Mitchell and colleagues address this directly: search engines return results the user evaluates; AI agents take actions the user authorizes. The cognitive load is structurally different. Self-reports of cognitive degradation are rapidly emerging across multiple studies, and agent capabilities are increasing faster than cognitive science can measure adaptation.

A second objection is that better tooling and transparency will solve the problem. The authors partially agree but note that explanations operate on cognitive capacities that agent use itself degrades, and that organizational protocols (training, rotation, breaks) are outside the scope of any tooling-focused agenda.

A third objection is that alignment will make oversight obsolete. The evidence does not support this: preference-based training assumes user preferences are a reliable proxy for oversight quality, but the evidence shows users prefer fluent, agreeable outputs even when those reduce scrutiny.

Decision implication

For product design and for individual agent practice, the current evidence supports four principles:

  1. Treat oversight quality as a depletable resource. Do not assume users can maintain constant vigilance. Design systems that detect and respond to oversight degradation.

  2. Implement intermediate confirmation, not confirm-at-end. The Zhou et al. result is robust and actionable. Schedule checkpoints at decision-relevant moments. Let the agent's self-assessed risk guide frequency, but do not rely solely on self-assessment.

  3. Build skill-preservation into the workflow. Require users to periodically perform tasks unassisted. Track whether domain knowledge is being maintained. Rotate users between assisted and unassisted work.

  4. Monitor oversight quality empirically. Track review time, override rates, and error detection rates over time. Treat declining engagement as a signal, not a feature.

For individual practice: the existing skill-formation briefs show that AI assistance can harm learning when it displaces the cognitive operations the learner needs to practice. The oversight brief extends this to supervision: watching an agent work is not a passive activity. It requires active engagement, domain expertise, and periodic unassisted practice. The best agent-use pattern preserves cognitive participation at every level—generation, evaluation, and oversight.

Calibrated confidence and update conditions

  • ~85% confidence: Intermediate confirmation outperforms confirm-at-end for multi-step agent tasks. The CHI 2026 result is strong, replicated across domains, and aligns with reliability-engineering principles.
  • ~80% confidence: Humans can detect AI agent errors, but detection quality degrades with task length and sustained engagement. Multiple independent sources support this.
  • ~75% confidence: Extended AI agent use degrades cognitive capacities required for oversight. The mechanism is well-established in automation research; direct evidence from LLM-specific studies is recent but consistent.
  • ~65% confidence: The feedback loop between degraded oversight and system optimization is a real and growing risk. The mechanism is plausible and supported by indirect evidence, but long-run empirical data are limited.
  • ~55% confidence: Organizational interventions (rotation, skill preservation, behavioral monitoring) can reverse or prevent oversight degradation. The interventions are well-motivated by cognitive science, but production deployment evidence is sparse.

I would update upward if a longitudinal study showed that organizations implementing strategic friction and behavioral monitoring maintained oversight quality over 6+ months of agent deployment. I would update downward if production data showed that intermediate confirmation strategies failed to generalize beyond the controlled task settings tested in the CHI study, or if agent self-assessment of risk proved too unreliable to serve as a checkpoint trigger.

The most useful next experiment is a field trial in a production software-development or data-analysis environment: half the team uses an agent with intermediate checkpoints and behavioral monitoring; the other half uses the same agent with confirm-at-end. Measure error detection rates, task completion quality, and domain-skill retention over 3–6 months.

References

  • Biswas, S., et al. (2026). Belief Updating and Delegation in Multi-Task Human–AI Interaction. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems. DOI: 10.1145/3772318.3793215
  • Gonzalez, C., et al. (2026). Toward a science of human–AI teaming for decision making: A complementarity framework. PNAS Nexus, 5(3), pgag030. DOI: 10.1093/pnasnexus/pgag030
  • Horowitz, M. C., & Kahn, L. (2024). Bending the Automation Bias Curve: A Study of Human and AI-Based Decision Making in National Security Contexts. International Studies Quarterly, 68(2), sqae020. DOI: 10.1093/isq/sqae020
  • Langer, M., et al. (2026). Keeping an Eye on AI: A Framework for Effective Human Oversight of AI Systems. arXiv: 2605.16278
  • Mitchell, M., Ghosh, A., & Passi, S. (2026). AI Agents Push Humans Out of the Loop. arXiv: 2608.23642
  • Nalisnick, E., et al. (2026). Human-AI Teaming Through the Lens of Calibration. arXiv: 2606.10906
  • Senjic, P., Bitsch, G., & Braun, A. (2026). Enabling Complexity: A Systematic Literature Review on Task Allocation, Communication, Interaction, and Augmentation in Human-AI Teams. Procedia Computer Science, 277, 2135–2144. DOI: 10.1016/j.procs.2026.02.251
  • Staufer, W., et al. (2026). The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems. Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency. DOI: 10.1145/3805689.3812417
  • Vorvoreanu, M., et al. (2026). Human oversight of agentic systems in practice: Examining the oversight work, challenges, and heuristics of developers using software agents. Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency, 6438–6465. DOI: 10.1145/3805689.3812402
  • Zhou, J., et al. (2026). When Should Users Check? Modeling Confirmation Frequency in Multi-Step Agentic AI Tasks. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems. DOI: 10.1145/3772318.3790655