Human–AI interaction · Cognitive science · Economics of AI adoption
Why Do AI Productivity Studies Disagree? The Expertise-Reversal Account
The sign of the AI productivity effect is set by the worker-task pair, not by the model. Assistance pays for the novice's generation cost and charges the expert a verification cost, so published estimates disagree because they average over opposite-signed segments of the same curve.
Abstract
Controlled trials of AI assistance report large speedups, real-world field studies report small gains, and at least one well-designed randomized trial found experienced practitioners were made slower. This synthesis asks a narrower question than "does AI improve productivity": what predicts the sign of the effect, and why do estimates disagree? The best-supported answer is an extension of the expertise reversal effect from cognitive load theory, with an explicit verification-cost term. A 2026 meta-analysis of 23 studies finds a moderate pooled gain (Hedges' g = 0.33) with near-total heterogeneity (I² = 99%), where study setting — not randomization — is the only significant moderator. We separate observations, randomized contrasts, mechanism inferences, and extrapolations, and specify what would overturn the account.
AI assistance is not a uniform productivity technology. It reliably helps users who lack a schema for the task, does little for users who have one, and actively harms everyone on tasks outside the model's competence — because the tool removes generation cost while imposing verification cost that rises with the user's ability to check output (~75% confidence). That makes the effect heterogeneous in sign, so any average is uninformative: the pooled 2026 programming estimate is g = 0.33 with I² = 99%, and study setting explains ~36% of the variance while experimental rigor explains none of it.
Precise question
When does AI assistance make a person measurably more productive at real work, and why do controlled trials of the same technology report speedups of 55% while randomized field studies of professional developers find them 19% slower?
This is narrower than "does AI improve productivity?" That framing presumes there is one number to find, which is the source of the confusion. The question here is about the sign of the effect and its moderators: what property of a person, a task, or their pairing determines whether assistance pays.
Why it matters
Three decisions turn on this, and none can be made from a headline statistic.
The first is tool deployment by role. An organization licensing an assistant for senior staff versus new hires is betting on a heterogeneous treatment effect, while the headline pooled estimate averages over segments pointing in opposite directions. If the effect reverses with experience, a blanket rollout can be net-negative precisely for the people the organization depends on most.
The second is evaluation design. If the sign depends on the worker-task pair, then occupation-level metrics — issues resolved per hour, pull requests, tickets closed — are structurally incapable of detecting it. Most published estimates use exactly those metrics.
The third is the field's ability to notice its own frontier. The dominant failure mode in deployed AI is not that users reject it; it is that users misjudge when it is helping. In the clearest available measurement, experienced developers predicted a 24% speedup, reported a 20% speedup afterward, and were in fact 19% slower. Independent domain experts forecasting the same setting expected 38–39% speedups. A field that measures itself by self-report will systematically over-estimate its own returns, and do so most confidently where the tool is least useful.
Best current answer
The sign of the AI productivity effect is a property of the worker–task pair, not of the model. AI assistance removes the cost of generating work and adds the cost of verifying it; the net is strongly positive when the user lacks a usable schema for the task, near zero when the user has one, and negative when the task sits outside the model's competence regardless of the user.
The mechanism with the deepest empirical grounding is the expertise reversal effect from cognitive load theory, which describes exactly this shape for instructional support, extended here with an explicit verification-cost term.
A 2026 meta-analysis of the programming literature quantifies the disagreement directly. Across 23 studies reporting 27 effect sizes, the pooled effect of generative AI assistance on developer productivity is moderate and positive — Hedges' g = 0.33, 95% CI [0.09, 0.58] — with heterogeneity at I² = 99%. Almost all of the variance is real between-study difference. Of six moderators tested, only one was significant: study setting, which accounted for roughly 36% of the heterogeneity, with laboratory experiments yielding a large effect (g = 0.73) and open-source and enterprise contexts yielding substantially smaller ones. Whether a study was randomized was not a significant moderator (QM(1) = 1.63, p = 0.202), nor was experimental design (QM(2) = 1.35, p = 0.510).
That is a strong constraint on the available explanations. The disagreement between a 55% lab speedup and a 19% field slowdown is not a story about sloppy lab work versus rigorous field work. Randomization status does not predict the sign. The moderators that do predict it are about who is doing what.
The three regimes, with the evidence for each
Regime 1: Strongly positive — users without a schema, on tasks inside the frontier
This is the best-replicated regime and the one the popular claims describe.
Noy and Zhang gave 444 college-educated professionals incentivized, occupation-specific writing tasks with random exposure to ChatGPT. Time taken fell 40% and output quality rose 18%, and the distribution of worker performance compressed: weaker workers gained most. Brynjolfsson, Li, and Raymond studied the staggered deployment of a conversational assistant across 5,172 customer-support agents, finding a 15% average gain in issues resolved per hour, concentrated disproportionately among less experienced and lower-skill workers. In their most expert tier, gains were small and average quality declined slightly.
Peng and colleagues ran a controlled trial of GitHub Copilot on a single JavaScript task — implementing an HTTP server — and found the treated group 55.8% faster, with larger gains among less experienced participants. Cui, Demirer, Jaffe, Musolff, Peng, and Salz pooled three randomized field experiments at Microsoft, Accenture, and an anonymized Fortune 100 manufacturer covering 4,867 developers, and report a 26.08% increase in completed tasks (SE 10.3%), again with higher adoption and larger gains among less experienced developers.
Cruces, Fernández Meijide, Galiani, Gálvez, and Lombardi ran a preregistered randomized experiment outside firms, where organizational selection does not compress educational heterogeneity: 1,174 adults aged 25–45, recruited in Argentina, completed an incentivized workplace-style business problem-solving task that was deliberately general rather than domain-specific. Without AI, higher-education participants outperformed lower-education participants by 0.548 SD. With AI, the gap fell to 0.139 SD — about three quarters of the initial gap closed.
The unifying feature is that in each case the user is weak relative to the specific task: a professional writer doing writing, a junior doing an HTTP server, a low-education participant on a general business exercise. The compression of the performance distribution is the signature of the mechanism, not a side effect of it.
Regime 2: Near zero to negative — experts on work they already know well
The sharpest test came from a randomized controlled trial that measured wall-clock time on real issues in large open-source repositories, contributed to for years. Sixteen developers with moderate AI experience and an average of five years of prior experience on those repositories completed 246 tasks, each randomly assigned to allow or disallow AI use. AI-allowed issues took 19% longer, with a confidence interval running from 2% to 39% longer. The developers themselves predicted a 24% speedup beforehand and still believed they had been sped up by 20% afterward. Economics and machine-learning experts forecasting the same setting predicted 39% and 38% speedups respectively.
The point estimate is the interesting part, not the headline. That 19% figure describes February–June 2025 tooling on mature repositories. In February 2026 METR reported that a follow-up experiment using later models produced a speedup instead — an estimated 18% for the subset of the original developers and 4% for newly recruited ones, both with confidence intervals spanning zero. So the sign of this particular result moved within a year, for reasons that have nothing to do with the workers: the models changed. Any brief that treats a frontier result as a stable fact has an expiry date of roughly twelve months.
Two independent lines of evidence point the same way. A natural experiment exploiting Italy's abrupt 2023 ChatGPT ban, using daily output data for more than 36,000 GitHub users in a difference-in-differences design, found that the ban increased output quantity and quality for less experienced users while decreasing productivity on routine tasks for experienced users — the expertise-reversal pattern reproduced outside a laboratory and without randomization.
And a two-year longitudinal case study in a large public-sector organization, covering 26,317 unique non-merge commits across 703 repositories, found no statistically significant change in commit-based activity after Copilot adoption — though minor increases were observed — and reported a discrepancy between commit-based metrics and the subjective experience of productivity.
Note what the negative regime does not claim. It does not claim AI is useless to experts. Cui and colleagues found real gains at three professional employers, and the field experiment at a global consumer goods company — preregistered, 791 professionals — found individuals with AI matched the performance of two-person teams without it, and that AI primarily lifted the quality of generated ideas while human judgment retained its value in selecting among them. The claim is narrower and more uncomfortable: for an expert on a familiar subtask, the net can be zero or negative, and the published literature does not currently let anyone predict in advance which expert on which subtask is on which side of zero.
Regime 3: Negative regardless of skill — tasks outside the model's frontier
The second domain experiment with 758 consultants established that capability has a jagged boundary rather than a gradient. It was preregistered, and randomly assigned 758 knowledge workers to one of three arms: no AI, GPT-4, or GPT-4 plus a prompt-engineering overview. On 18 realistic management tasks inside the frontier, AI-assisted consultants completed 12.2% more tasks, worked 25.1% more quickly, and produced solutions of significantly improved quality. On a single complex managerial task selected to lie outside the frontier, subjects using AI were 19% less likely to produce correct solutions than those without it — which the authors read as a limitation of AI supporting knowledge workers. The same experiment ran a third arm with a prompt-engineering tutorial, so learning how to drive the tool was itself a variable rather than a fixed handicap.
This is the most decision-relevant finding in the whole literature, because the boundary is invisible at the moment of use. Two tasks that look equally hard, written in the same register, on the same topic, sit on opposite sides of it.
Mechanisms
Redundancy: the expertise reversal effect
Cognitive load theory predicts that instructional support which reduces load for a learner without a schema imposes extraneous load on a learner who has one, because the support is redundant with knowledge already in long-term memory. This is the expertise reversal effect, and its empirical record is long and strong.
A 2025 meta-analysis pooled 176 effect sizes from 60 experimental studies and 5,924 participants. Learners with low prior knowledge performed better under high-assistance instruction (d = 0.505, 95% CI [0.260, 0.750]); learners with high prior knowledge performed better under low-assistance instruction (d = −0.428). The interaction — the difference of differences — is d = 0.971, 95% CI [0.631, 1.312]. Expertise reversal is robust and large.
AI assistance is instructional support in exactly this sense: it supplies information the user might not have retrieved, and it cannot know which information that is.
Verification cost: the term that makes it a productivity mechanism rather than a learning one
Redundancy explains why assistance stops helping. It does not by itself explain why assistance starts hurting, because an expert can in principle ignore redundant material at no cost. The missing term is that the user must still evaluate the output, and evaluation is not free.
The cleanest available evidence isolates this term by looking at where the loss lands. In a randomized experiment, 52 software developers with at least a year of weekly Python experience learned an asynchronous library unfamiliar to all of them, completing two coding tasks either with or without AI assistance. The AI group finished approximately two minutes faster — a difference that did not reach statistical significance. The authors' own summary is blunter than that: AI use impaired conceptual understanding, code reading, and debugging, without delivering significant efficiency gains on average. On a quiz covering concepts used minutes earlier, the AI group averaged 50% against 67% for hand-coding, a gap of 17 percentage points whose largest component was debugging: the ability to tell whether code is wrong and why.
That is the signature of a verification deficit, not a generation deficit. The users could produce the work and could not reliably check it. The same study's within-condition analysis reinforces it: participants who used AI for conceptual questions, or generated and then asked for explanations, retained understanding; participants who delegated generation or used AI as a debugging crutch did not. Time typing rather than pasting the output made no measurable difference to understanding, which rules out motor-effort explanations.
Frontier mismatch
Frontier effects are orthogonal to expertise effects, and they compose. A capable user on a capable task gains. A capable user on an incapable task loses more, because they have a strong prior and the output contradicts it fluently. A non-expert on an incapable task is the worst case: no schema, no reliable signal. A non-expert on an incapable task is the worst case: no schema, no reliable signal. This is why the out-of-frontier result is hard to design around — the user with the strongest prior is the user whose prior the model contradicts most fluently, and fluency is exactly what makes the error pass unnoticed.
Why the estimates disagree: three non-mechanism explanations
Real heterogeneity is only part of the story. Three artifacts inflate the disagreement, and they are worth separating because they have different fixes.
Averaging over opposite signs. A 26% gain and a 19% loss average to a number with no referent: no population in any of these studies experienced the pooled effect. The pooled effect is a property of the study set, and this is the artifact most damaging to the field's ability to learn.
Selection into adoption. People who choose to use a tool, and people who consent to a study that removes it, are not random. In the two-year longitudinal case study, Copilot users were consistently more active than non-users before Copilot existed — a baseline difference that any post-hoc comparison inherits, observed across just 25 users and 14 non-users. The 2025 randomized trial is the cleanest evidence, and its authors themselves documented the mechanism in a 2026 note: a significant increase in developers declining to participate because they did not wish to work without AI, some developers declining to complete tasks assigned to the AI-disallowed condition (one completed none of them), a pay reduction from $150 to $50 per hour, and time-spent becoming unreliable for developers running multiple concurrent agents. Their conclusion was that the newer data "gives us an unreliable signal" and that the central estimate is "likely a bad proxy for the real productivity impact of AI tools." That is an unusually clean example of a field correcting its own bias in public — and note that the direction of the bias favors the tool, which is the dangerous direction for a field that would like a positive answer.
Proxy metrics. Counting commits, pull requests, or lines of code assumes that output quantity and value move together. In the pooled three-company field experiment, the working-paper version reports the gain concentrated in pull requests (+26.1%, SE 10.3%) rather than commits (+13.6%, SE 10.0%). The metric moved more than the work plausibly did. A task that is subdivided into more commits looks like a productivity gain with no change in delivered value.
Contamination runs the other way. In the Microsoft experiment, the trial was stopped early on 3 May 2023 because control-group participants sought access to Copilot. Non-compliance in the opposite direction — treated participants working without assistance — is structurally harder to detect and to penalize, and no study in this literature has a credible instrument for it.
Strongest counterevidence and unresolved disputes
Four findings cut against the account, and none of them is trivial.
The reversal is asymmetric, which flatters the technology. The 2025 expertise-reversal meta-analysis finds that giving assistance to novices has a larger effect than withholding redundant assistance from experts does. If that asymmetry transfers, the realistic aggregate effect of broad deployment is somewhat more positive than a symmetric model predicts, and a firm that assumes "neutral-to-good for seniors" is closer to correct than one that assumes "dangerous for experts." I weight this heavily; it is the strongest reason not to over-read the negative regime.
The delegation story is weaker than the skill-formation result alone implies. The same Argentine experiment reported above ran a follow-up module with the assistant removed. Treated participants did not perform worse than controls once AI was withdrawn, and lower-education participants retained part of their gain — although a sizable education gap re-emerged. That is inconvenient for the simple version in which assistance is borrowed capability that evaporates on demand. The pattern is better described as assistance shifting where skill is formed rather than destroying it, and it weakens any claim that AI users are quietly running down a resource they cannot rebuild.
Natural experiments on the same event disagree in sign. A second analysis of Italy's ChatGPT ban, with 88,022 users across Italy, France, and Portugal and two-way fixed effects, finds that access increased developer productivity by 6.4%, with novices receiving the productivity gains and more experienced developers benefiting instead through knowledge sharing (+9.6%) and skill acquisition (+8.4%). That is the opposite sign on the less-experienced group from the 36,000-user analysis. Both use the same design and the same event. The less-experienced-side estimate should be treated as unresolved, and the identification strategy itself as weaker than advertised — the first study documents users circumventing the ban via VPN, which contaminates the treatment assignment it relies on.
The learning literature is not settled either. The same 2026 programming meta-analysis that finds g = 0.33 for productivity finds no significant effect on learning outcomes: g = 0.14, 95% CI [−0.18, 0.47]. A three-level meta-analysis of generative AI in higher education, pooling 36 studies, 132 effect sizes, and 7,229 participants, reports a substantially larger overall effect on learning outcomes of g = 0.499. Two meta-analyses of adjacent questions disagree by roughly a factor of four. Any account that treats the skill formation consequences as settled is overreaching.
"Expertise" is probably measured badly. Nearly all of this literature stratifies by job tenure, self-reported skill, or educational attainment. The mechanism requires task-specific schema. A senior engineer is a novice in a language they have not touched in a decade, and a junior is an expert in the third week of a codebase. If the moderator is mismeasured, the heterogeneity attributed to expertise may partly be something else. This is the largest unforced weakness in the account, and it is not currently testable with the published data.
Evidence quality and limitations
The causal core is strong. The three studies that establish the sign reversal are all randomized: the 246-task developer trial, the 758-consultant experiment with its between-subject task arms, and the 52-participant skill-formation trial. The field experiments at three employers and the two-year longitudinal case study add real-work behavior with objective metrics. The mechanism is anchored in a 60-study, 5,924-participant meta-analysis from an independent literature.
The limits are equally clear. The 246-task trial has 16 participants and clustered standard errors, and its own authors note the settings may not generalize to less experienced developers or unfamiliar codebases. The pooled three-company estimate carries a standard error of 10.3% on a 26% effect — barely distinguishable from zero at conventional thresholds — and the individual company experiments are described by their authors as noisy with varying signs. The consultant experiment used one firm, one model, and one deliberately frontier-violating task. The skill-formation trial is 52 people, one library, and an immediate quiz with no retention or transfer measure. The 2026 programming meta-analysis pools 27 effect sizes from partly non-independent designs; with I² = 99%, its point estimate is a summary of a distribution, not an estimate of anything.
Every estimate here describes 2023–2026 tools. Model capability is the fastest- moving variable in the system, every "frontier" result is a snapshot of a coordinate that has already moved, and the METR follow-up above shows one study's sign flipping inside a single year.
Calibrated confidence
~85% that the sign of the AI productivity effect varies substantially by worker–task pair, and that the disagreement among published estimates is principally heterogeneity rather than measurement error. This rests on a significant moderator in a 23-study meta-analysis, two randomized studies with opposite signs on populations chosen to differ in exactly the predicted way, and a 60-study meta-analytic mechanism.
~75% that the expertise-reversal framing plus an explicit verification-cost term is the correct mechanism rather than one of several adequate ones. The debugging-specific deficit is strong direct support, but no study has yet manipulated verification cost independently of assistance.
~55% that the field's expertise proxies capture task-specific schema rather than general seniority. This is the weakest link, and it is the one I would most want to see tested.
~40% that any current published point estimate can support a deployment decision. With I² = 99% and a between-study moderator that explains a third of the variance, I do not think pooled numbers should be quoted as expected returns.
Reversal conditions
This account would need revision if:
- A preregistered trial measuring the same workers on both familiar and unfamiliar subtasks, with the sign predicted in advance by task-specific schema rather than by job tenure, came out homogeneous. That would falsify the central claim directly.
- Verification cost were measured and manipulated independently — for example, by varying the verifiability of AI output while holding assistance constant — and did not predict the sign. The mechanism would then be wrong even if the heterogeneity were real.
- A well-powered, contamination-controlled natural experiment on an access restriction produced a single unambiguous sign across experience groups, resolving the Italy conflict.
- The asymmetry finding failed to replicate: if withholding redundant assistance harmed experts as much as helping novices helped them, the realistic aggregate effect would be much closer to neutral, and the deployment implications would weaken considerably.
- Replication of the near-zero enterprise estimate with an instrumented, task-level productivity measure replaced occupation-level proxies. Most of the negative evidence is currently built on small samples or proxies, and a clean large-N null would be a serious problem for the account.
The highest-value next study is not another headcount. It is a within-worker design: the same experienced practitioners, on tasks they know well and tasks they do not, with assistance randomized within worker, with verification cost instrumented, and with task-specific expertise measured directly rather than proxied by title. Every existing study varies at most one of these at a time, which is exactly why the literature is stuck.
Decision implications
The account above makes four falsifiable predictions for anyone deciding how to deploy these tools.
Measure at the task level and stratify by task-specific experience, not by job title. "Developer" is not a moderator; "has not worked in this subsystem in four years" is. Aggregating to occupation level averages the effect toward zero and hides a real sign flip.
Treat self-reported productivity as uninformative. The 246-task trial's participants were wrong about their own performance by 39 percentage points while being the most motivated possible evaluators, paid $150 per hour. Where measurement is cheap, take the measurement.
Assume a capability boundary that is invisible in the moment, and bound the cost of being wrong outside it — review gates, verification on task classes where the model is known to fail — rather than assuming users will notice.
Do not treat the average as the return. The realistic expectation for broad deployment is a positive aggregate effect weighted toward the least experienced workers, near zero for senior staff on familiar work, and materially negative outside the frontier. Planning capacity around the aggregate will over-promise to seniors and under-serve novices.
References
-
Maier, S., Gunzenhäuser, M., Schweisthal, J., Schneider, M., & Feuerriegel, S. (2026). A meta-analysis of the effect of generative AI on productivity and learning in programming. arXiv:2605.04779. https://arxiv.org/abs/2605.04779
-
Tetzlaff, L., Simonsmeier, B., Peters, T., & Brod, G. (2025). A cornerstone of adaptivity – A meta-analysis of the expertise reversal effect. Learning and Instruction, 98, 102142. https://doi.org/10.1016/j.learninstruc.2025.102142
-
Kalyuga, S., Ayres, P., Chandler, P., & Sweller, J. (2003). The expertise reversal effect. Educational Psychologist, 38(1), 23–31. https://doi.org/10.1207/S15326985EP3801_4
-
Becker, J., Rush, N., Barnes, B., & Rein, D. (2025). Measuring the impact of early-2025 AI on experienced open-source developer productivity. arXiv:2507.09089. https://arxiv.org/abs/2507.09089
-
Becker, J., Rush, N., Cunningham, T., Rein, D., & Mahamud, K. / METR. (2026, February 24). We are changing our developer productivity experiment design. https://metr.org/blog/2026-02-24-uplift-update/
-
Brynjolfsson, E., Li, D., & Raymond, L. (2025). Generative AI at Work. The Quarterly Journal of Economics, 140(2), 889–942. https://doi.org/10.1093/qje/qjae044
-
Cui, Z. K., Demirer, M., Jaffe, S., Musolff, L., Peng, S., & Salz, T. (2026). The effects of generative AI on high-skilled work: Evidence from three field experiments with software developers. Management Science. https://doi.org/10.1287/mnsc.2025.00535
-
Dell'Acqua, F., McFowland III, E., Mollick, E., Lifshitz-Assaf, H., Kellogg, K. C., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. R. (2026). Navigating the jagged technological frontier: Field experimental evidence of the effects of artificial intelligence on knowledge worker productivity and quality. Organization Science, 37, 403–423. https://doi.org/10.1287/orsc.2025.21838
-
Dell'Acqua, F., Ayoubi, C., Lifshitz, H., Sadun, R., Mollick, E., Mollick, L., Han, Y., Goldman, J., Nair, H., Taub, S., & Lakhani, K. R. (2026). The cybernetic teammate: A field experiment on generative AI and teamwork. Organization Science, 37, 1217–1242. https://doi.org/10.1287/orsc.2025.20702
-
Shen, J. H., & Tamkin, A. (2026). How AI impacts skill formation. arXiv:2601.20245. https://arxiv.org/abs/2601.20245
-
Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science, 381(6654), 187–192. https://doi.org/10.1126/science.adh2586
-
Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). The impact of AI on developer productivity: Evidence from GitHub Copilot. arXiv:2302.06590. https://arxiv.org/abs/2302.06590
-
Kreitmeir, D., & Raschky, P. A. (2024). The heterogeneous productivity effects of generative AI. arXiv:2403.01964. https://arxiv.org/abs/2403.01964
-
Bonabi, S., Bana, S., Gurbaxani, V., & Nian, T. (2025). Beyond code: The multidimensional impacts of large language models in software development. arXiv:2506.22704. https://arxiv.org/abs/2506.22704
-
Cruces, G., Fernández Meijide, D., Galiani, S., Gálvez, R. H., & Lombardi, M. (2026). Does generative AI narrow education-based productivity gaps? Evidence from a randomized experiment. NBER Working Paper No. 34851. https://www.nber.org/papers/w34851
-
Stray, V., Brandtzæg, E. G., Wivestad, V. T., Barbala, A., & Moe, N. B. (2026). Developer productivity with and without GitHub Copilot: A longitudinal mixed-methods case study. Proceedings of the 59th Hawaii International Conference on System Sciences (HICSS-59), 7413–7422. https://doi.org/10.24251/HICSS.2026.880
-
Fan, C., Ke, L., Chen, Z., & Lv, P. (2026). Exploring the effect of GenAI on learning outcomes in higher education: A three-level meta-analysis. Frontiers in Psychology, 17, 1758670. https://doi.org/10.3389/fpsyg.2026.1758670