Human–AI interaction · Learning science
Does AI Assistance Build Skill or Merely Improve Performance?
When help disappears, the capability left behind depends on which cognitive work the assistant preserved—and which it replaced.
Abstract
Generative AI has no stable, context-independent effect on unaided learning. It reliably improves many assisted outputs. After removal, however, randomized studies find both retained gains and performance-learning reversals. Direct solution provision can displace the reasoning a novice needs to acquire; structured tutoring can preserve that work and sometimes improve learning. The best current policy is therefore purpose-dependent: automate expendable output, but preserve attempts, retrieval, diagnosis, and evaluation when the skill must be owned. Long-run workplace deskilling, effects in young children, and far transfer remain largely unmeasured.
AI can improve the work produced during assistance while weakening, preserving, or improving the capability that remains after assistance.
1. Precise question
When a novice uses a generative-AI assistant while acquiring a new cognitive skill, does the assistant improve later unaided performance—and which features of the tool, task, and use pattern determine whether the effect is positive or negative?
This is deliberately narrower than “Does AI improve education?” and different from “Does AI make people more productive?” An assistant can raise the quality of an essay, solution, or program produced today while leaving the user less able to produce or evaluate the next one alone. Conversely, it can give timely explanations and feedback that accelerate learning. The outcome of interest is the capability remaining after assistance is removed, preferably measured after a delay on a genuinely new task.
The question matters in at least three settings. Extensive agent use creates a recurring choice between delegating a result and acquiring the judgment needed to supervise future results. Any learning-oriented product decision should distinguish assisted completion from retained capability. The same distinction will eventually matter when deciding how children should use AI. These applications are extrapolations: the studies do not directly estimate effects for professional agent users, product teams, or young children.
2. Best current answer
Generative AI has no stable, context-independent effect on unaided learning. It reliably makes many assisted tasks easier and often improves the output produced with it. What happens after removal depends substantially on what cognitive work it displaced or induced.
- Direct solution provision can create a performance-learning tradeoff. In randomized experiments, people produced better work or progressed more easily with AI but performed worse when assistance disappeared. This has appeared in high-school mathematics, short online mathematics tasks, and professional programming instruction.
- Structured tutoring can avoid that harm and sometimes outperform conventional instruction. Systems that sequence material, require engagement, offer hints, and give targeted feedback have produced positive immediate learning effects. Most evidence is short-term and often tests a whole instructional package rather than AI alone.
- Unrestricted access is not inevitably harmful. A 2026 classroom experiment found that ordinary AI access improved immediate and one-week unaided knowledge. That contradicts a simple “answers always cause deskilling” story.
- Variation matters more than the literature-wide average. A careful 2026 STEM meta-analysis found extremely heterogeneous raw effects and a near-zero pooled estimate after explicitly modeling publication bias. A narrower synthesis of randomized student experiments found a small positive mean. Neither supports a universal sign.
My synthesis is that AI helps learning when it supplies feedback, explanation, adaptive sequencing, or relief from extraneous work while preserving the learner’s generation, retrieval, diagnosis, and evaluation. It harms learning when it supplies the very representation or decision rule the learner needs to construct. This mechanism is plausible and fits the experiments, but no study has estimated a complete, general law connecting particular AI interactions to long-term skill formation.
Selected standardized effects on unaided outcomes
Positive values favor the AI-inclusive condition. Points are not a meta-analysis.
3. What the strongest studies show
3.1 High-school mathematics: the cleanest performance-learning reversal
Randomized observationBastani and colleagues assigned Turkish high-school classrooms to ordinary instruction, a general GPT-4 interface (“GPT Base”), or a constrained tutor designed with teachers (“GPT Tutor”). During practice, performance rose by 48% with the base assistant and 127% with the tutor. On a subsequent exam without AI, the base-assistant group scored 17% lower than control—about −0.19 standard deviations—while the tutor group was statistically indistinguishable from control at about −0.01 SD.
Causal claimIn that setting, access to the unconstrained assistant caused worse immediate unaided performance even while improving assisted practice. The guardrailed tutor removed most of the damage but did not produce an unaided advantage. Incorrect model responses contributed, yet the result is not reducible to hallucination: students could use mostly helpful answers as a substitute for mathematical work.
LimitationThe study covered four 90-minute sessions, the assessment followed immediately, and the outcome was school mathematics in one national context. It establishes a real failure mode, not a universal effect.
3.2 Undergraduate unfamiliar topics: unrestricted AI produced retained gains
Randomized observationContractor and Reyes assigned undergraduates to learn blockchain, carbon capture, and CRISPR with or without generative AI. After a 35-minute learning-and-writing session, the AI group scored 6.7 percentage points higher on an unaided knowledge test (0.27 SD). Roughly a week later, the advantage remained 5.1 points, about three-quarters of the immediate effect. Students with AI did not spend more total time; they shifted time away from producing text and toward reading and search, and reported greater enjoyment.
CounterevidenceThis contradicts the claim that ordinary AI access inherently impairs learning. The intervention improved a delayed unaided outcome, not merely assisted output.
The paper classified interaction logs into “automation” and “augmentation.” Automation users produced stronger assisted essays but showed almost no later unaided advantage; augmentation users appeared to retain more. This comparison is observational within the randomized treatment arm. People chose their style, so motivation, prior ability, or strategy may explain the difference. It suggests a mechanism but does not establish one.
3.3 Structured tutors: promising, but mostly package effects
Randomized observationKestin and colleagues ran a crossover trial with 194 students in introductory Harvard physics. A carefully engineered GPT-4 tutor produced higher immediate post-test scores than an active-learning class, with a regression estimate of 0.63 SD, while students reported greater engagement and spent less time. The tutor sequenced problems, used curated material, elicited active engagement, managed cognitive load, and scaffolded explanations.
Causal boundaryThe full tutoring package improved immediate learning relative to the classroom comparison. The study does not isolate which component caused the gain, includes only two lessons and no delayed test, and bundles medium with context because the AI condition occurred at home.
A World Bank randomized evaluation in nine Nigerian public secondary schools found a 0.31 SD overall gain from a six-week after-school program combining GPT-4 access, teacher supervision, English instruction, and digital-skills practice. This is encouraging evidence for a deployable program, not an estimate of AI alone: treatment added twelve 90-minute sessions, devices, structured activities, and human supervision. Attrition was substantial and differed between groups.
3.4 Brief solution access can reduce persistence
Randomized observationLiu and colleagues conducted three online experiments with 1,222 participants. In two mathematics experiments, participants could view accurate GPT-5-generated help before assistance was unexpectedly removed. In the first, the AI group solved fewer unaided problems than control (0.57 versus 0.73; d = −0.42) and skipped more. A better-controlled replication found a smaller but significant solution deficit (0.71 versus 0.77; d = −0.19); its skipping difference was not significant.
These experiments support a narrow causal claim: brief access to ready-made AI solutions can reduce persistence and immediate independent performance on similar problems. They do not establish lasting deskilling. Exposure lasted roughly 10–15 minutes, and direct-answer versus hint-seeking behavior was self-selected.
3.5 Programming: a small but relevant warning
Randomized observationShen and Tamkin assigned 52 experienced developers learning the unfamiliar Python Trio library to work with GPT-4o or without it. Developers with AI scored 50% on an immediate unaided quiz versus 67% for controls, a 17-point gap (d = 0.74 in the control-favoring direction), without a significant average time saving.
This preprint is directly relevant to AI-mediated professional skill acquisition, but its small sample, one library, immediate test, and preprint status warrant caution. It does not show that AI reduces expert software productivity. It shows that assisted exposure may fail to build the local knowledge needed to reason without the assistant.
3.6 Timing of help: an adjacent 12-week result
Poulidis, Bastani, and Bastani randomized more than 200 chess-club students using a custom AI-assisted training platform. One group received system-timed tips; another could request help whenever it wanted. Both improved, but the on-demand group improved by about 30% while the system-regulated group improved by about 64%, and on-demand users completed 24% fewer games as help requests rose.
This supports the idea that help timing can causally affect longer-run learning and engagement. It is adjacent rather than decisive evidence for generative AI: chess training is specialized, there was no no-AI control, and mediation estimates assigning losses to reduced “productive struggle” are model-dependent.
4. A working mechanism model
The evidence is coherent if learning return is treated as a balance rather than a property of the model:
A balance-sheet model of AI-assisted learning
An explanatory hypothesis, not an empirically fitted equation.
Capability-building inputs
- Useful explanation and feedback
- Adaptive sequencing and practice
- Reduced extraneous load
- Motivation and time on task
Capability-displacing inputs
- Replaced retrieval and generation
- Replaced error diagnosis
- Uncritical acceptance of errors
- Premature escape from struggle
The model generates testable predictions. Asking AI to critique an attempted solution should preserve more learning than requesting a finished solution before attempting. Hints delivered after initial effort should outperform unlimited answers when the target is independent problem solving. Learners with enough prior knowledge to check outputs may gain more than novices unable to recognize plausible errors.
A tool can also increase real-world performance in an always-assisted environment while reducing resilience to outages, rare errors, or novel edge cases. Only the first-order contrast—solution substitution versus preserved cognitive work—has repeated experimental support. Claims about precise expertise thresholds, ideal struggle, or long-term dependence remain speculative.
Why findings differ
The target of learning. Producing a polished essay, remembering factual content, solving equations, and constructing a software mental model require different internal operations. AI may free time for reading in one task and replace core reasoning in another.
The comparator. “AI versus no AI” can mean a chatbot versus a textbook for equal time, or an intensive supervised after-school program versus business as usual. The latter estimates a package, not a model effect.
The interface and pedagogy. A generic answer box, a tutor constrained to hints, and a sequenced curriculum powered by an LLM are different interventions. Calling all three “generative AI” hides the most decision-relevant variation.
The outcome horizon. Most studies test immediately; one tests at a week; the chess study spans 12 weeks. Almost none establish retention over months, far transfer, or the ability to detect a confident AI error.
User behavior. Random access does not randomize how people use it. Log-based findings that “automation users” learn less are vulnerable to selection. The right experiments must randomize interaction policy itself.
5. Evidence quality and literature-level uncertainty
The causal core is stronger than much commentary suggests: randomized studies measure post-assistance, unaided outcomes in real classrooms and online settings. The adverse findings are not merely correlations between heavy AI use and weak students. The positive results are not just self-reports of convenience.
But the evidence base remains immature. It overrepresents short interventions, immediate tests, higher education, text interfaces, and domains with easily scored answers. Many studies confound AI with extra instructional time, teacher support, or redesigned curricula. Usage-mode analyses are frequently post-treatment. Model capabilities and defaults also change faster than educational trials can run.
Boolzen and colleagues’ preregistered systematic review is the strongest warning against reading too much into a pooled average. It identified 85 eligible STEM studies and meta-analyzed 59 effects from 49 studies. Raw effects were positive but extraordinarily heterogeneous (I² = 96.3%), with a prediction interval from g = −1.52 to +3.20 and funnel asymmetry. A robust Bayesian model accounting for publication bias estimated a pooled effect near zero (0.076 ± 0.254). Cognitive activity relative to the comparison and whether outcomes tested knowledge or skill explained some—but only about 12–13%—of the variance.
Contractor and Reyes assembled a narrower set of 13 randomized student experiments with unaided assessments and estimated a small positive random-effects mean of 0.18 SD across 26 estimates. This was a literature comparison inside a working paper, not a standalone preregistered systematic review; multiple estimates per study and selective availability remain concerns.
The defensible conclusion is not “zero effect.” It is that a grand mean conceals interventions ranging from harmful substitution to effective tutoring and is highly sensitive to study selection and bias adjustment.
6. Counterevidence and alternative explanations
The strongest counterevidence to a harm-focused account is Contractor and Reyes: ordinary access improved a delayed unaided test. The strongest counterevidence to an enthusiasm-focused account is Bastani and colleagues: enormous assisted gains coexisted with an unaided loss, while a tutor wrapper neutralized rather than reversed that loss. Liu and Shen show similar short-term reversals in different populations. Kestin and the Nigerian program demonstrate that well-designed AI-inclusive instruction can beat relevant alternatives.
- Bad answers rather than offloading. Errors plausibly explain some harm. Yet Liu supplied accurate solutions and performance still fell, so accuracy cannot be the whole story.
- Surprise removal. Learners expecting continued access may rationally invest less in memorization. That is part of the effect of an always-available tool, but it means the outcome measures resilience rather than necessarily performance in the intended assisted environment.
- Novelty and motivation. Positive results may partly reflect novelty, attention, or enjoyment. Those can be real mechanisms but may decay with routine use.
- Assessment mismatch. A closed-book test may undervalue specification and integration skills needed in an AI-saturated workplace. An assisted production metric may miss the judgment needed to catch rare, costly failures. Both should be measured.
7. Decision payoff
The practical decision is not “use AI or abstain.” It is to name the purpose of each interaction.
- If the goal is output and the component is genuinely expendable, automate it. There is little reason to practice boilerplate the user does not need to reproduce or audit.
- If the goal is learning, preserve the target cognitive operation. Attempt first; ask for a hint, counterexample, critique, or explanation; then reconstruct the answer without the assistant. Do not let the model perform the exact inference being trained.
- If future judgment is required, close the loop. After assisted work, explain the mechanism in your own words, solve a nearby case unaided, and revisit it after a delay. Include a deliberately flawed output when error detection is mission-critical.
ExtrapolationFor product teams, a product claiming to teach should measure delayed unaided transfer, not only completion, satisfaction, or assisted correctness. A sensible default is progressive help: request an attempt, diagnose the error, give the smallest useful hint, and reveal a solution only after another attempt. This should be tested against a direct-answer condition rather than treated as doctrine.
ExtrapolationFor an individual agentic workflow, full delegation is compatible with learning only when the delegated skill is not one the user needs to own. If they must maintain the code, challenge a strategic model, or detect a dangerous exception, the relevant metric is whether they can explain and stress-test the result after the agent is gone.
8. Calibrated confidence
| Confidence | Claim |
|---|---|
| 95% | Assisted task performance and retained unaided capability are distinct outcomes; measuring only the first cannot establish learning. |
| 85% | The sign and size of AI’s effect on unaided learning depend materially on task, interface, pedagogy, comparator, and user behavior. |
| 75% | Ready access to direct solutions can reduce immediate independent performance when novices learn an unfamiliar reasoning procedure. |
| 65% | Scaffolds preserving attempts, retrieval, diagnosis, and feedback can avoid much of that harm and sometimes improve short-term learning. |
| 55% | Across well-designed current interventions, the average short-term effect is probably small-to-moderately positive, but not stable enough to drive policy. |
| 25% | Existing evidence can quantify long-run workplace deskilling, effects in young children, or outcomes after years of routine agent use. |
9. What would reverse this view
I would move toward a broadly positive default if several preregistered, multisite studies found that ordinary AI access improves six- to twelve-month retention, far transfer, and error detection across ages and domains under equal instructional time—especially if direct-answer interfaces performed as well as scaffolded tutors.
I would move toward a broadly restrictive default if similarly long studies found persistent losses under realistic use, including on assisted future tasks where users must supervise rare errors, and if losses remained after controlling for model accuracy, expectations of continued access, and prior ability.
The highest-value next experiment would randomize the interaction policy, not just access: no AI versus direct answer versus hint-after-attempt versus critique-after-solution, with identical content and time, process logs, immediate and delayed tests, novel transfer problems, and adversarial error-detection trials. Until then, progressive assistance plus explicit unaided checks is the most evidence-aligned reversible policy.
References
- Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). “Generative AI without guardrails can harm learning: Evidence from high school mathematics.” Proceedings of the National Academy of Sciences, 122(26), e2422633122. doi:10.1073/pnas.2422633122. Data and code.
- Contractor, Z., & Reyes, G. (2026). “Experimental Evidence on the Learning Impact of Generative AI.” IZA Discussion Paper 18792. Official record and PDF; arXiv:2607.08849.
- Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). “AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting.” Scientific Reports, 15, 17458. doi:10.1038/s41598-025-97652-6.
- De Simone, M. E., Tiberti, F. H., Barron Rodriguez, M. R., Manolio, F. A., Mosuro, W., & Dikoru, E. J. (2025). “From Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in Nigeria.” World Bank Policy Research Working Paper 11125. doi:10.1596/1813-9450-11125; official record.
- Liu, G., Christian, B., Dumbalska, T., Bakker, M. A., & Dubey, R. (2026). “AI Assistance Reduces Persistence and Hurts Independent Performance.” Conference on Language Modeling (COLM 2026). arXiv:2604.04721.
- Shen, J. H., & Tamkin, A. (2026). “How AI Impacts Skill Formation.” arXiv preprint. arXiv:2601.20245.
- Poulidis, S., Bastani, H., & Bastani, O. (2026). “Self-Regulated AI Use Hinders Long-Term Learning.” SSRN working paper. doi:10.2139/ssrn.5604932.
- Boolzen, C., Kuhn, J., Flegr, S., Rott, E.-M., Stausberg, N., Küchemann, S., et al. (2026). “Evidence of impact and interpretational limits of generative AI in STEM education: a systematic review and meta-analysis on cognitive learning outcomes.” Artificial Intelligence Review. doi:10.1007/s10462-026-11665-9.