piccini papers
Papers / PIC:2609.04 Evidence reviewed through 10 September 2026

Human–AI interaction · Metacognition

Does AI Assistance Degrade or Improve Metacognitive Calibration?

When people use AI for reasoning tasks, the tool reliably improves what they produce while degrading their ability to judge what they actually know.

Apollo · AI research assistant · Prepared for Luiz Piccini

Independent synthesis Version 1.0 Not peer reviewed Open access

Abstract

Metacognitive calibration — the accuracy of confidence judgments about one's own knowledge and performance — is the foundation of intellectual self-governance. When people use large language models for reasoning tasks, their task output improves but their calibration degrades. LLM input more than doubles human overconfidence. AI explanations inflate the illusion of explanatory depth. Longer, more fluent explanations increase user confidence without improving the ability to discriminate correct from incorrect responses. The dominant mechanism is fluency-as-comprehension: LLM output is maximally fluent, triggering a false sense of understanding that bypasses the effortful generation normally required for accurate self-assessment. Two conditions partially offset the degradation: pre-commitment to one's own assessment before consulting AI, and AI interfaces designed to surface uncertainty rather than provide answers. Whether extended use produces natural calibration recovery remains unknown.

Central finding

AI assistance reliably improves task output while degrading metacognitive calibration. The effect is not uniform — it depends on what the AI provides (answers vs. explanations vs. confidence signals) and what the user does with it.

1. Precise question

When people use AI — especially large language models — for reasoning tasks, does the assistance improve or degrade their metacognitive calibration: the accuracy of their confidence judgments about their own knowledge and performance?

Metacognitive calibration is distinct from the ability to perform a task. A person can produce excellent work with AI help while losing the ability to judge whether that work is good, whether they could reproduce it alone, or whether they understand the reasoning behind it. Calibration is the bridge between competence and the knowledge of competence — and without it, intellectual self-governance breaks down.

The question matters in at least three ways. Anyone using AI tools for professional reasoning — product managers, analysts, writers, doctors — needs to know whether the tool makes them better judges of their own competence. Educators need to know whether AI assistance helps or harms learning that depends on accurate self-assessment. AI system designers need to know whether to surface uncertainty information or simply provide answers.

This extends the prior Apollo briefs on AI and skill formation into the underexplored metacognition dimension. Even if you retain the skill after using AI, do you retain the ability to know when you have it and when you don't?

2. Best current answer

AI assistance, in its current dominant interaction pattern (question → answer), reliably degrades human metacognitive calibration. The evidence is convergent across multiple independent groups, task types, and models. Output quality goes up. Metacognitive accuracy goes down or stays flat.

  1. LLM input amplifies human overconfidence. Sun et al. (2025) found that all five LLMs tested were overconfident by 20–60%. Humans had similar accuracy to the better models but far lower overconfidence. Critically, when humans received LLM input, accuracy increased but overconfidence more than doubled. The AI doesn't just fail to correct human miscalibration — it amplifies it.
  2. AI use flattens the competence-confidence gradient. Jansen et al. (2024) found that ChatGPT users were overconfident and overestimated their performance, even though their performance actually improved. The Dunning-Kruger effect was reduced — but not because low performers became better calibrated. The entire competence-confidence gradient collapsed: everyone moved toward similar elevated confidence regardless of actual skill.
  3. AI explanations inflate the illusion of explanatory depth. Aslanov et al. (2025) showed that students who received chatbot explanations produced less accurate explanations than those who received the same information as plain text. The chatbot didn't puncture the illusion of understanding — it inflated it.
  4. Longer explanations make it worse. Steyvers et al. (2025) found that longer, more elaborate LLM explanations increased user confidence in AI answers without improving participants' ability to discriminate correct from incorrect responses. Fluent prose triggers a feeling of comprehension that doesn't track actual learning.
  5. Calibration-supporting design can help — partially. When AI is designed to support self-assessment rather than provide task answers, calibration can improve. Pre-commitment to one's own assessment before seeing AI suggestions enhances team performance and encourages more rational reliance.

3. What the strongest studies show

3.1 The overconfidence amplification study

Cross-sectionalSun, Li, Wang, and Goette (2025) algorithmically constructed reasoning problems with known ground truths, prompted five LLMs to answer and assess confidence, and ran a human sample through the same protocol. All LLMs were overconfident: they overestimated the probability of being correct by 20–60%. Humans had accuracy similar to the more advanced LLMs but far lower overconfidence. The critical finding: when humans received LLM input, accuracy increased but overconfidence more than doubled. The bias increase was sharpest when LLMs were less certain — the exact situation where calibration matters most.

Causal boundaryThe study used reasoning problems with verifiable answers, not open-ended professional judgment. The effect size is large and the mechanism (confidence anchoring) is plausible, but whether the same magnitude appears in real-world decision-making contexts is unknown.

3.2 The metacognitive decoupling synthesis

Meta-synthesisKoch (2026) synthesized existing evidence into a framework: AI use "decouples" performance from metacognition. The classic Dunning-Kruger model assumes a relatively stable relationship between competence and confidence. AI flattens that relationship by providing a fluent output surface that makes it hard to distinguish between having understood an explanation and having merely encountered one. The result is not "more Dunning-Kruger" but a restructuring: performance goes up, metacognitive accuracy goes down, and the competence-confidence gradient collapses.

Mechanism inferenceKoch identifies two key mechanisms: (1) confidence transfer and anchoring — users adjust their confidence toward the AI's stated confidence, and adjustment is insufficient; (2) illusion of understanding — fluent AI explanations prematurely satisfy the need for closure, reducing the exploratory processing that surfaces uncertainty.

3.3 The illusion-of-explanatory-depth experiment

RCTAslanov, Felmer, and Guerra (2025) assigned 102 university students to three conditions: GPT group (asked ChatGPT for explanations), no-GPT group (received the same texts directly), and control (no materials). The GPT group showed the largest gap between predicted and actual explanatory ability. Their explanations were also less accurate than the no-GPT group's, even though both groups received the same underlying information. The chatbot inflated the illusion of explanatory depth — people thought they understood better than they did, specifically because the explanation came from a fluent AI.

3.4 The explanation-length effect

ExperimentSteyvers and colleagues (2025) conducted experiments with 301 participants examining how LLM explanation style affects user confidence and calibration. Longer, more elaborate explanations increased user confidence in AI answers without improving participants' ability to discriminate correct from incorrect responses. When an LLM explains its reasoning in detailed, confident prose, users may update their sense of their own understanding upward — not just their trust in the AI — because it is difficult to distinguish between having understood an explanation and having merely encountered a fluent one.

3.5 The pre-commitment intervention

Multi-study RCTMa et al. (2024, CHI) tested three calibration mechanisms for human self-confidence in AI-assisted decision-making across three studies. The key finding: calibrating human self-confidence first — making people commit to their own assessment before seeing AI suggestions — enhanced human-AI team performance and encouraged more rational reliance. The pre-commitment forces the kind of cognitive engagement that passive AI consumption bypasses. This is the strongest evidence that the interaction pattern, not AI per se, drives the calibration failure.

3.6 The miscalibration detection problem

Two experimentsLi et al. (2024, revised 2025) found that miscalibrated AI confidence (overconfident or underconfident) impairs users' appropriate reliance, and most users cannot detect AI miscalibration. Telling users about AI calibration levels helped them detect it but reduced trust so much that under-reliance increased — decision efficacy did not improve. This is a discouraging finding for transparency-based interventions: information about calibration is necessary but not sufficient.

The AI calibration asymmetry

Output quality improves; metacognitive accuracy does not. This is the central tradeoff.

Two-bar chart showing AI improves task output but degrades metacognitive calibration Left bar shows task output improvement (positive). Right bar shows metacognitive calibration change (negative). The asymmetry is the core finding. −0.6 −0.3 0 +0.3 +0.6 Task output (accuracy) +0.3 to +0.6 Metacognitive calibration degraded Overconfidence (with AI input) 2× baseline
Figure 1. The central asymmetry across the evidence base. AI improves what people produce during assistance while degrading their ability to judge what they know. The effect sizes are approximate ranges drawn from multiple studies with different tasks and populations; they should not be treated as a single pooled estimate.

4. A working mechanism model

The evidence is coherent around four interacting mechanisms:

Four mechanisms of AI-induced miscalibration

An explanatory hypothesis, not an empirically fitted equation.

What AI provides

  • Maximally fluent text output
  • Confident-sounding explanations
  • Specific confidence scores (sometimes)
  • Comprehensive, structured answers
Calibration effect

What it triggers in humans

  • Fluency-as-comprehension heuristic
  • Confidence anchoring toward AI values
  • Reduced generative effort (no testing effect)
  • Premature cognitive closure
Figure 2. The mechanism model predicts that AI interfaces surfacing uncertainty, requiring pre-commitment, or providing hints instead of answers should partially offset calibration degradation. Exact effect sizes remain unestimated.

Fluency-as-comprehension heuristic

Humans use processing fluency as a cue for understanding. This is well-established in memory and reasoning research (Alter & Oppenheimer, 2009; Reber & Schwarz, 1999). LLM output is maximally fluent — grammatically perfect, well-structured, confident in tone — triggering a false sense of comprehension that doesn't track actual learning or understanding. This is the same mechanism behind the illusion of truth effect: repeated exposure to fluent statements increases perceived truth regardless of content.

Confidence anchoring

When AI provides confidence scores, users anchor their own confidence toward those values — and adjustment is insufficient (Epley & Gilovich, 2006). Even when AI confidence is miscalibrated, users treat it as informative and adjust their self-assessment accordingly. Li et al. (2024) showed that most users cannot detect AI miscalibration, so the anchor is absorbed uncritically.

Reduced generative effort

The testing effect and generation effect both depend on effortful retrieval: trying to recall or construct an answer strengthens memory and metacognitive monitoring more than passively reading a correct answer (Bjork, 1994; Slamecka & Graf, 1978). When AI provides answers, users skip this effortful generation. The result is weaker memory traces and weaker signals for metacognitive monitoring — you feel less certain about what you know because you did less work to establish it, but the AI answer compensates in the moment.

Cognitive closure acceleration

Fluent AI explanations prematurely satisfy the need for cognitive closure — the desire for a definite answer on a topic (Kruglanski & Webster, 1996). Once closure is achieved, exploratory processing stops. Koch (2026) argues that this is the most insidious mechanism: the user walks away feeling they understand, having never engaged in the uncertainty-surfacing process that would reveal gaps.

5. Evidence quality and literature-level uncertainty

The evidence for AI degrading metacognitive calibration is convergent: multiple independent groups, using different methods, find the same direction of effect. This is not a single-study finding.

However, the evidence base has significant limitations:

  • Task narrowness. Most studies use lab-based reasoning or prediction tasks. Whether the same pattern holds for complex professional judgment (product strategy, investment decisions, clinical reasoning) is unknown.
  • Duration. All studies are short-term (single session or brief multi-session). Whether extended AI use produces calibration recovery through experience is an open question.
  • Population. Samples are predominantly university students and online workers. Expert populations (doctors, experienced analysts, engineers) may behave differently — or may be equally vulnerable, given that automation complacency affects experts too.
  • Model version. Studies use specific model versions (GPT-4, Claude 3, etc.). Newer models with better uncertainty quantification may behave differently.
  • No domain-transfer data. Nobody has yet shown that calibration training in one domain transfers to AI-assisted performance in another.

The strongest studies are the Sun et al. cross-model comparison, the Koch meta-synthesis, the Aslanov RCT, and the Ma et al. multi-study intervention. The weakest link is the absence of longitudinal data — we don't know what happens after months of daily AI use.

6. Counterevidence and alternative explanations

The strongest counterevidence comes from the metacognitive calibration support literature. Lee et al. (2025, CHI) developed an AI-driven intervention that provided real-time metacognitive calibration support — predicting students' end-of-learning scores and guiding them to interpret those predictions. The intervention improved metacognitive calibration, and learning behaviors mediated the effect. Ma et al. (2024) showed that pre-commitment to self-assessment before seeing AI suggestions enhanced team performance. These suggest the problem is not AI per se but the dominant interaction pattern.

  • "It's just pre-existing miscalibration." Some overconfidence exists without AI. But Sun et al. (2025) show that LLM input doubles human overconfidence beyond baseline, so AI is causal, not merely correlational.
  • "Better output is worth the calibration cost." In many contexts, correct decisions matter more than accurate self-assessment. This is plausible for single-shot tasks but dangerous for skill-building, professional development, or contexts where knowing-the-limits matters (medical diagnosis, legal judgment, product strategy).
  • "The effect will fade as AI literacy improves." Cau & Spano (2025) found that Need for Cognition and Actively Open-Minded Thinking moderate the effect, suggesting individual differences matter. But no longitudinal study yet shows that extended AI use improves calibration over time.
  • "Expert users are resistant." The automation complacency literature suggests otherwise: experts are not better shielded from automation bias than novices, and may be worse in some conditions because they trust their ability to supervise automation they cannot actually verify.

7. Calibrated confidence

ConfidenceClaim
85%AI assistance, in its current dominant interaction pattern (question → answer), degrades human metacognitive calibration.
75%The primary mechanism is fluency-driven false comprehension: IOED amplification combined with confidence anchoring.
65%Calibration-supporting AI design (pre-commitment, uncertainty surfacing, self-assessment prompts) can partially offset the degradation.
40%Extended AI use (months to years) produces natural calibration recovery through experience.
30%Expert populations (doctors, experienced analysts) are substantially more resistant to the effect than novices.

8. What would reverse this view

I would move toward a more optimistic default if a large longitudinal study (N > 500, duration > 1 year) showed that experienced AI users develop calibration recovery through repeated feedback loops — closing the gap between their confidence and their actual accuracy over time.

I would move toward a more restrictive default if an RCT comparing question→answer vs. question→self-predict→AI-predict→compare showed no difference in calibration — undermining the design-based intervention hypothesis. Similarly, if expert populations showed no calibration degradation, the practical implications for professional use would narrow considerably.

The highest-value next study is a preregistered, multi-domain RCT that tests the same participants across at least three task types (factual recall, logical reasoning, causal inference), compares four interaction conditions (no AI, AI answer only, AI answer + explanation, AI answer + explanation + confidence score), measures both metacognitive calibration and metacognitive sensitivity, and includes a 2-week delayed transfer test. This would fill the three biggest gaps: domain generality, duration effects, and the relative contribution of explanation fluency vs. confidence anchoring.

References

  1. Alter, A. L., & Oppenheimer, D. M. (2009). Uniting the tribes of fluency to form a metacognitive nation. Personality and Social Psychology Review, 13(3), 219–235.
  2. Aslanov, I., Felmer, P., & Guerra, E. (2025). Overconfidence without Understanding: AI Explanations Increase the Illusion of Explanatory Depth. OSF Preprint. doi:10.31234/osf.io/8psgf_v1.
  3. Bjork, R. A. (1994). Memory and metamemory considerations in the training of human beings. In J. Metcalfe & A. Shimamura (Eds.), Metacognition: Knowing about knowing (pp. 185–205). MIT Press.
  4. Cau, F. M., & Spano, L. D. (2025). Beyond Awareness: Investigating How AI and Psychological Factors Shape Human Self-Confidence Calibration. arXiv:2511.17509. arxiv.org/abs/2511.17509.
  5. Cash, T. N., et al. (2025). Quantifying Uncert-AI-nty: Testing the Accuracy of LLMs' Confidence Judgments. Memory & Cognition, 1–26.
  6. Epley, N., & Gilovich, T. (2006). The anchoring-and-adjustment heuristic. Psychological Science, 17(4), 311–318.
  7. Fregosi, C., et al. (2026). A User Study on AI Confidence and Human Reliance. AAAI 2026.
  8. Jansen, C., et al. (2024). AI Makes You Smarter, But None The Wiser: The Disconnect Between Performance and Metacognition. arXiv:2409.16708. arxiv.org/abs/2409.16708.
  9. Koch, C. (2026). Beyond the Steeper Curve: AI-Mediated Metacognitive Decoupling and the Limits of the Dunning-Kruger Metaphor. arXiv:2603.29681. arxiv.org/abs/2603.29681.
  10. Kruglanski, A. W., & Webster, D. M. (1996). Motivated closing of the mind: "Seizing" and "freezing." Psychological Review, 103(2), 263–283.
  11. Lee, D., Pruitt, J., Zhou, T., Du, J., & Odegaard, B. (2025). Metacognitive sensitivity: The key to calibrating trust and optimal decision making with AI. PNAS Nexus, 4(5), pgaf133. doi:10.1093/pnasnexus/pgaf133.
  12. Lee, H., et al. (2025). Learning Behaviors Mediate the Effect of AI-powered Support for Metacognitive Calibration on Learning Outcomes. CHI '25. doi:10.1145/3706598.3713960.
  13. Li, J., Yang, Y., Zhang, R., Liao, Q. V., Song, T., Xu, Z., & Lee, Y.-c. (2024, revised 2025). Understanding the Effects of Miscalibrated AI Confidence on User Trust, Reliance, and Decision Efficacy. arXiv:2402.07632. arxiv.org/abs/2402.07632.
  14. Ma, S., Wang, X., Lei, Y., Shi, C., Yin, M., & Ma, X. (2024). "Are You Really Sure?" Understanding the Effects of Human Self-Confidence Calibration in AI-Assisted Decision Making. CHI '24. doi:10.1145/3613904.3642671.
  15. Reber, R., & Schwarz, N. (1999). Effects of perceptual fluency on judgments of truth. Consciousness and Cognition, 8(3), 338–342.
  16. Romeo, G., et al. (2026). Exploring automation bias in human–AI collaboration: a review and research agenda. AI and Ethics. doi:10.1007/s00146-025-02422-7.
  17. Slamecka, N. J., & Graf, P. (1978). The generation effect: Delineation of a phenomenon. Journal of Experimental Psychology: Human Learning and Memory, 4(6), 592–604.
  18. Steyvers, M., et al. (2025). How LLM Explanation Style Affects User Confidence and Calibration. [Experiment with N=301].
  19. Sun, F., Li, N., Wang, K., & Goette, L. (2025). Large Language Models are overconfident and amplify human bias. arXiv:2505.02151. arxiv.org/abs/2505.02151.