piccini papers
Papers / PIC:2609.02 Evidence reviewed through 3 September 2026

Human–AI interaction · Skill formation

When Does LLM Assistance Build Skill Rather Than Borrow It?

Direct answers can improve assisted work while leaving less capability behind. Scaffolding changes the treatment—and sometimes the result.

Apollo · AI research assistant · Prepared for Luiz Piccini

Independent synthesis Version 1.0 Not peer reviewed Open access

Abstract

For novices learning a cognitive skill, direct-answer LLM access can improve work completed with AI while reducing performance on a subsequent task without it. This performance–learning gap is causal in several short randomized studies. It is not a general verdict against AI assistance: tightly designed tutors that elicit reasoning, sequence hints, and retain human guidance have produced real short-term learning gains. The interaction policy is part of the treatment. Evidence on durable retention and professional transfer remains sparse, so products that claim to build capability should measure delayed, unaided performance rather than infer learning from assisted output.

Central finding

Assisted task completion and capability after assistance are different outcomes. Direct answers can separate them; good scaffolding can bring them back together.

1. Precise question

For a novice acquiring a cognitive skill, does access to a large language model improve immediate task performance at the cost of later unaided capability, and can a scaffolded interaction design avoid that tradeoff?

The question matters whenever a system is meant to coach, teach, or preserve a human’s ability to supervise automated work. A tool can look excellent on completion rate, speed, satisfaction, and assisted accuracy while weakening the skill needed when the tool is absent or wrong. The same distinction applies personally to writing, analysis, coding, and strategic reasoning.

ScopeThe causal evidence is concentrated in students, online participants, and one narrow programming task. Applying it to adult professional work is an extrapolation, not an observed result.

2. Best current answer

Direct-answer LLM access can create a performance–learning gap. In several randomized, short-duration studies, it improved work completed with AI but reduced performance on a subsequent task without AI. Carefully constrained tutors that elicit reasoning, sequence hints, and keep a teacher or tutor in the loop have instead produced genuine short-term learning gains. “LLM use” is therefore too coarse a treatment label: the interaction policy helps determine the outcome.

The evidence is much weaker about durable capability. Most experiments last minutes or weeks, use immediate post-tests, and cover one narrow domain. Longer studies and large educational panels are consistent with a cost from easy, on-demand outsourcing, but several are working papers or observational designs. The defensible conclusion is conditional, not anti-AI: use completion metrics when only output matters; use delayed, unaided holdouts when the human must retain or supervise the skill.

3. Four outcomes that should not be collapsed

  1. Assisted performance: quality, speed, or completion while AI is available.
  2. Immediate unaided performance: a post-test minutes or days later without AI.
  3. Delayed retention: what remains after weeks or months.
  4. Transfer and supervision: whether the learner can solve a novel problem or detect a plausible AI error.

An intervention may improve the first and harm the second. It may improve the second without supplying evidence about the third. Almost no current study provides a strong test of the fourth. The relevant product metric depends on which capability the user actually needs.

4. The strongest evidence for a performance–learning gap

4.1 School mathematics: answers help today and hurt on the no-AI exam

Randomized observationBastani and colleagues assigned nearly 1,000 Turkish secondary-school students to four mathematics sessions with no AI, a general GPT-4 interface, or a safeguarded “GPT Tutor” using teacher-designed hints and instructions not to reveal solutions. Relative to no AI, the general interface improved assisted practice performance by 48%, and the tutor by 127%. When access was removed for the exam, students assigned to the general interface performed 17% worse than controls. The tutor eliminated that penalty but did not create a statistically clear positive exam effect. [1]

Causal claimThe assignment establishes that access to the general interface caused lower immediate unaided performance in this setting. Observed solution copying and students’ failure to anticipate the penalty are suggestive mechanisms, not independently randomized causes.

One experiment, two phases, two different conclusions

Performance relative to the no-AI group in the Bastani et al. field experiment.

Assigned conditionAssisted practiceSubsequent no-AI exam
GPT Base48% higher performance17% lower performance
GPT Tutor127% higher performanceNo statistically clear difference
Figure 1. The practice and exam percentages refer to different phases and should not be combined into one effect or read as a ranking across studies. The comparison isolates assigned access to each system in one school context over four sessions; it does not estimate durable retention.

4.2 Programming: less knowledge without a clear time gain

Randomized observationShen and Tamkin assigned 52 experienced Python programmers who were new to the Trio library to complete coding tasks with GPT-4o or without AI, followed by a 27-point no-AI quiz. The AI group scored 4.15 points lower—about 17%, with Cohen’s d = 0.74—and did not finish significantly faster. [9]

Screen recordings associated delegation, reliance, and AI-led debugging with lower quiz performance; conceptual questions and generation followed by active comprehension were associated with preserved learning. The randomized score difference is causal for this task. The workflow comparison is exploratory because participants chose how to use the tool.

4.3 Mathematics and reading: performance falls when assistance disappears

Randomized observationAcross online experiments with 1,222 participants, Liu and colleagues tested what happened after practice with or without AI when access was removed. In the cleaner mathematics replication, the AI group solved 71% of subsequent problems versus 77% for controls (d = −0.19). In a reading task, the rates were 76% and 89% (d = −0.42); AI users also skipped more items. [6]

Assignment to AI was randomized, so the post-removal group differences are causal. Choice of direct answers versus hints was not randomized, so the association between direct-answer use and larger loss is only suggestive. These were brief online tasks with immediate, somewhat harder follow-up phases—not delayed retention studies.

5. Evidence that scaffolded assistance can build skill

5.1 A tightly engineered physics tutor

Randomized observationKestin and colleagues ran a crossover trial with 194 Harvard undergraduates in introductory physics. Students received either an established in-class active-learning lesson or an at-home AI tutor for each of two topics. The AI condition produced a 0.63-standard-deviation higher immediate learning gain in the preregistered regression. Students reported greater engagement and spent a median of 49 minutes rather than the scheduled 60-minute class. [5]

This was not a generic chatbot. Instructors supplied learning goals and worked solutions; the platform sequenced the material, required engagement, provided targeted feedback, and used a carefully developed system prompt. The result shows that an LLM inside a designed instructional system can outperform one comparison lesson on an immediate post-test. There was no delayed or transfer test.

5.2 Teacher-guided use in Nigeria

Randomized observationA World Bank trial assigned 1,328 first-year secondary-school volunteers in Edo State, Nigeria, to a six-week after-school program using Microsoft Copilot or to no program. Teachers led twice-weekly sessions with a curated curriculum and prompts encouraging reasoning and hallucination checks. The program raised an overall assessment composite by 0.31 standard deviations and English performance by roughly 0.23 standard deviations. [3]

Causal boundaryThe intervention bundled approximately 13 hours of extra instruction, teacher facilitation, digital access, curriculum, and AI. Because the control group did not receive an equally intensive non-AI program, the design estimates the package’s effect—not the model’s independent contribution. Attrition and the absence of longer follow-up further constrain the inference.

5.3 AI that coaches a human tutor

Randomized observationIn Tutor CoPilot, 783 tutors serving 1,013 K–12 students were randomized to receive real-time AI suggestions or not. Students whose tutors had access were four percentage points more likely to master the session topic on an exit ticket; the increase was nine points for initially lower-rated tutors. The human tutor chose or edited suggestions, and the system promoted guiding questions and explanations. [11]

This supports an alternative architecture: AI can augment the person providing instruction rather than answer for the learner. The outcome was same-session mastery, not delayed independent capability.

6. What the longer-run evidence adds

Randomized observationA 12-week study of more than 200 chess-club students compared system-timed AI advice with a version that also allowed help on demand. Both groups improved, but the system-regulated group improved by 64% versus 30% in the self-regulated condition. The authors’ mediation model attributes more than half of the gap to reduced productive struggle and about 15% to lower engagement. This working paper has no no-AI arm, and the mediation model does not prove that either mechanism caused the gap. [7]

Observational evidenceA 30-month panel of 26,811 Chinese secondary students associated staggered generative-AI adoption with higher homework scores and shorter homework time, followed by lower closed-book and entrance-exam performance. Losses were concentrated in patterns consistent with outsourcing. A ten-year panel of 3.2 million ALEKS interactions found a post-ChatGPT decline in time spent on text problems relative to graphical tasks and a decline on randomly assigned proctored retention items. [10] [8]

These panels have scale and duration, but neither directly randomizes or perfectly observes AI use. Adoption timing, differences between task types, and other cohort changes remain alternative explanations. They are strong signals, not clean causal effect sizes.

CounterpointIn a laboratory logic-puzzle study, every randomized group learned whether AI was unavailable, cheap to query, or costly to query. Final accuracy did not differ, although cheap access produced more requests and slower subsequent unaided solutions. The “AI” was a simulated perfect oracle revealing part of the solution rather than an LLM. The result shows that offloading need not erase learning, while leaving depth and efficiency open. [12]

7. Plausible mechanisms—and their evidential status

What the interaction preserves, and what it replaces

A mechanism map derived from the studies, not an empirically fitted equation.

Capability-building inputs

  • Useful explanation and feedback
  • Adaptive sequencing and practice
  • Reduced extraneous cognitive load
  • Motivation and time on task
Net learning return

Capability-displacing inputs

  • Replaced retrieval and generation
  • Replaced error diagnosis
  • Uncritical acceptance of errors
  • Premature escape from struggle
Figure 2. This explanatory synthesis does not assign weights or imply that the two sides are directly measurable on one scale. The optimal amount of struggle remains unestimated and depends on the target skill.

7.1 Cognitive offloading can reduce encoding

In three randomized pattern-copy experiments, increasing the cost of looking back reduced offloading and slowed immediate work but improved memory. Telling participants that memory would be tested largely counteracted the loss even under forced offloading. Offloading therefore does not mechanically erase learning: goals and attention change what is encoded. [4]

7.2 Answers can remove productive struggle and retrieval

Generating a step, retrieving a principle, and debugging an error are costly operations, but they also produce diagnostic feedback and strengthen memory. An answer-first interface can remove those operations. The base-versus-tutor contrast in the Turkish experiment is consistent with this mechanism. The chess mediation results and programmer workflows add support, but neither cleanly randomizes productive struggle itself.

7.3 Fluent output can miscalibrate confidence

When a correct-looking answer appears inside a workflow, assisted success can feel like personal competence. The stronger inference is that assisted performance is an unreliable proxy for independent capability. Evidence that a specific metacognitive illusion caused the later loss remains thinner because subjective calibration is measured inconsistently.

7.4 Good tutoring can reduce unproductive load

Not all difficulty is useful. Novices can waste working memory decoding instructions, searching blindly, or rehearsing errors. A tutor that diagnoses the misconception, selects the next step, and gives a contingent hint can preserve retrieval while removing irrelevant load. Positive tutoring trials fit this explanation, but their bundled designs cannot isolate each component.

7.5 Motivation and persistence can move in either direction

Fast answers may reduce persistence when help disappears; timely feedback may instead sustain engagement. Current studies show both patterns. Motivation is more plausibly a moderator of design effects than a one-way property of AI.

8. Counterevidence and competing interpretations

Aggregate reviews often report benefits. A 2026 meta-analysis of 35 experimental and quasi-experimental studies estimated a pooled effect of g = 0.67 for ChatGPT-supported learning. Its studies differed in comparison activities, duration, outcomes, and whether tests were genuinely unaided. [13]

A stricter preregistered review illustrates why the average is unstable. Boolzen and colleagues retained 49 studies with externally assessed cognitive outcomes and comparison groups. A conventional random-effects model was positive, but heterogeneity was extreme (I² = 96.3%) and the prediction interval ran from a large negative to a large positive effect. A robust Bayesian model accounting for publication bias estimated a mean near zero (g = 0.076 ± 0.254). Whether AI augmented or substituted for a comparable cognitive activity explained part of the variation; no tested factor explained enough to support a general effect. [2]

  • External cognition may be the correct target. A professional need not memorize everything a reliable tool will always supply. Reduced unaided skill becomes a product failure only when access can fail, the model can be wrong, or the person must transfer, supervise, or choose among outputs.
  • Workflows are usually self-selected. People who request complete answers may already be rushed, less motivated, or less skilled. Usage clusters and mediation models should not be promoted to causal mechanism estimates.
  • Publication and novelty effects can run both ways. Dramatic harm attracts attention; enthusiastic deployments may preferentially report gains. Short custom tests, researcher-built interfaces, and selective outcomes constrain both literatures.
  • Assessment choice is normative as well as empirical. A closed-book test may undervalue specification and integration skills in an AI-rich workplace. Assisted production may miss the judgment required to detect rare, costly failures. Both outcomes can matter.

9. Decision payoff: use explicit modes

Output mode. If a task is disposable and only the artifact matters, delegate aggressively. Assisted accuracy, time, and cost are appropriate metrics. There is no reason to preserve every intermediate skill.

Capability mode. If the user must retain the skill, have the human generate before the model supplies: attempt; request a hint or critique; explain the principle in one’s own words; solve a varied case; then perform an unaided retrieval or transfer check. This sequence is an evidence-informed extrapolation, not a fully tested optimum.

Supervision mode. If AI will continue doing the work but a human remains accountable, train the minimum capability needed to specify, audit, and recover. Evaluation should include seeded errors and novel cases because a person can appear productive with AI while being unable to detect a confident failure.

Product extrapolationA credible test of a learning or coaching product should randomize the assistance policy while keeping content and time as comparable as possible. It should measure assisted performance, an immediate no-AI test, delayed retention, transfer to a structurally new task, error detection, and interaction traces distinguishing attempts, hints, explanations, and answer requests.

The north-star distinction is capability after assistance, not engagement during assistance. A product may legitimately choose output mode, but it should not market assisted fluency as learning without holdout evidence.

10. Calibrated confidence and important limits

ConfidenceClaim
85%Direct-answer assistance can cause a short-run gap between assisted and immediately unaided performance in some novice tasks.
80%Well-designed scaffolded tutoring can produce genuine short-term learning gains in some contexts.
60%Unrestricted, on-demand LLM use causes a material loss of durable, transferable skill on average.
45%The magnitude of the observed effect generalizes to professional knowledge work.

The main limitation is external validity. Effects probably vary with expertise, stakes, model accuracy, prior knowledge, time pressure, and whether learners expect continued access. Positive trials often bundle the model with curriculum, interface constraints, extra time, and human guidance. The strongest longer-run negative evidence has not yet been replicated through randomized access with delayed transfer tests.

Evidence boundaryThe 85% estimate concerns the existence of a short-run failure mode, not its universal prevalence or average magnitude. The 60% and 45% estimates are extrapolations from a sparse evidence base and should move more than the first two as new long-horizon trials arrive.

11. Reversal conditions and the next experiment

I would raise confidence that unrestricted AI causes durable skill loss above 80% if several preregistered, multi-week randomized trials found worse delayed retention and transfer despite equal practice time, equal curriculum, low attrition, and direct logs showing that answer delegation mediates the effect. I would lower it below 30% if adequately powered replications found no delayed disadvantage—or a benefit—after matching time and content, especially in adult professional tasks, while showing that answer-first users could still detect errors and transfer principles.

I would raise confidence in scaffolding above 90% if factorial trials separately manipulated answer availability, attempt-before-help, hint sequencing, self-explanation, and retrieval checks, then reproduced durable gains. I would lower it below 50% if the apparent benefit vanished on delayed or novel tasks, or if matched human-guided controls showed that the model contributed nothing beyond extra time and structure.

The highest-value next study for a learning product is a multi-week randomized comparison of four policies on the same curriculum: no AI, answer on demand, hint first with a mandatory attempt, and a human coach with AI support. Delayed unaided transfer should be the primary endpoint; assisted completion should be secondary. Until then, the rational default is conditional: borrow capability freely when only output matters, but require generation, explanation, and an unaided check when the human must own the skill.

References

  1. Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). “Generative AI without guardrails can harm learning: Evidence from high school mathematics.” Proceedings of the National Academy of Sciences, 122(26), e2422633122. doi:10.1073/pnas.2422633122. Data and materials.
  2. Boolzen, C., Kuhn, J., Flegr, S., Rott, E.-M., Stausberg, N., Küchemann, S., et al. (2026). “Evidence of impact and interpretational limits of generative AI in STEM education: a systematic review and meta-analysis on cognitive learning outcomes.” Artificial Intelligence Review. doi:10.1007/s10462-026-11665-9.
  3. De Simone, M. E., Tiberti, F. H., Barron Rodriguez, M. R., Manolio, F. A., Mosuro, W., & Dikoru, E. J. (2025). “From Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in Nigeria.” World Bank Policy Research Working Paper 11125. doi:10.1596/1813-9450-11125. Reproducibility package.
  4. Grinschgl, S., Papenmeier, F., & Meyerhoff, H. S. (2021). “Consequences of cognitive offloading: Boosting performance but diminishing memory.” Quarterly Journal of Experimental Psychology, 74(9), 1477–1496. doi:10.1177/17470218211008060.
  5. Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). “AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting.” Scientific Reports, 15, 17458. doi:10.1038/s41598-025-97652-6. Study data.
  6. Liu, G., Christian, B., Dumbalska, T., Bakker, M. A., & Dubey, R. (2026). “AI Assistance Reduces Persistence and Hurts Independent Performance.” arXiv preprint. arXiv:2604.04721.
  7. Poulidis, S., Bastani, H., & Bastani, O. (2026). “Self-Regulated AI Use Hinders Long-Term Learning.” Wharton School working paper. doi:10.2139/ssrn.5604932.
  8. Rismanchian, S., Uzun, H., Matayoshi, J., Cosyn, E., & Kurd-Misto, E. (2026). “Faster Completion, Less Learning: Generative AI Reduced Study Time on Math Problems and the Knowledge They Build.” arXiv preprint. arXiv:2605.21629.
  9. Shen, J. H., & Tamkin, A. (2026). “How AI Impacts Skill Formation.” arXiv preprint. arXiv:2601.20245. Preregistration.
  10. Strömberg, D., Lei, V., & Wu, Y. (2026). “The Generative AI Learning Penalty: Evidence from Chinese Secondary Education.” CEPR Discussion Paper 21577. Official record.
  11. Wang, R. E., Ribeiro, A. T., Robinson, C. D., Loeb, S., & Demszky, D. (2025). “Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise.” EdWorkingPaper. doi:10.26300/81nh-8262.
  12. Wu, S., Belem, C. G., Fu, S., Steyvers, M., & Smyth, P. (2026). “How AI Assistance Affects Human Skill Development: A Study of Learning with Logic Puzzles.” arXiv preprint. arXiv:2608.23543.
  13. Wu, X., Zhu, P., Zhang, J., Yin, M., & Wang, Y. (2026). “ChatGPT’s impact on student learning outcomes: a meta-analysis of 35 experimental studies.” Humanities and Social Sciences Communications, 13, 684. doi:10.1057/s41599-026-07019-z.