Files
BusinessNotes/wiki/concepts/ai-productivity-evidence.md
EugeneTes 60176d2fdc all
2026-07-30 11:15:52 +02:00

5.9 KiB
Raw Blame History

AI Productivity — Measured vs. Claimed

#concept #ai #evidence

Summary

The empirical counterweight to the vault's dominant premise. Where the vault's promotional sources (twelve at this page's creation; nineteen of twenty as of 2026-07-29) assert AI has zeroed the cost of commodity software work, the rigorous evidence (RCTs, a top-5-journal field study, large surveys) shows the real effect is modest (≈1419%), context-dependent, tilted toward NOVICES not experts, and systematically overstated by self-report. It qualifies rather than demolishes the thesis — no study finds AI worthless — but it directly inverts two load-bearing claims and is the vault's first source with evidentiary standing rather than sales incentive. Anchored entirely by 2026-07-18-ai-productivity-adversarial-evidence; single-source, so Status: tentative, but its inputs are the highest-provenance in the vault.

Current Understanding

Two direct inversions of the popular thesis:

  1. "AI zeroed the cost of skilled dev work" → the strongest developer-specific RCT (METR, 2025) found 16 experienced open-source maintainers were 19% slower completing real tasks on their own repos with early-2025 AI. The cost didn't go to zero; for experts on familiar code it went up.
  2. "Seniors up, juniors irrelevant" (seniority-and-ai) → the best skill-distribution study (Brynjolfsson, Li & Raymond, QJE 2025) found the least-experienced gain most (+34% vs. ~0 and small quality declines for experts). A GitHub Copilot RCT independently found less-experienced developers benefit more. AI's productivity dividend accrues to novices, not seniors.

The perception trap (why the hype is self-sustaining). METR's developers forecast a 24% speedup, still estimated +20% after finishing, and were measured 19% — a ~39-point gap between felt and real. This matters because "it feels faster / I shipped more" is the exact evidence most promotional AI-productivity claims are built on. Self-report is not measurement.

Individual speed ≠ shipped software. DORA 2024 found AI adoption correlated with reduced delivery stability (and, that year, throughput — reversed in DORA 2025). Stack Overflow 2025 (~49k devs): only ~3% "highly trust" AI accuracy while ~46% distrust it (up from 31%), even as adoption hit 84%; ~45% say debugging AI code is time-consuming. The artifact isn't finished when it's generated.

The honest reading. Held to the same standard as the promotional sources, this evidence does not prove AI is worthless — average effects are positive (1415%). It proves the effect is smaller, unevenly distributed, and easier to overstate than "cost → zero" implies. The vault's strategic advice can survive a modest, novice-tilted AI; it cannot survive on the premise that AI is a full substitute for a developer.

⚠ Time-scope. All figures are 20232025, tied to specific tool generations. METR labels its result "historical"; the Copilot RCT used 2022 Copilot; DORA reversed a 2024 finding in 2025. This is not the live 2026 state of the art — it is a corrective against extrapolating hype, not a prediction.

Evidence

  • Whole concept (all six findings, effect sizes, designs, caveats, and the two refuted claims) — 2026-07-18-ai-productivity-adversarial-evidence
  • Primary roots cited there: METR RCT (arXiv:2507.09089); Brynjolfsson/Li/Raymond QJE 140(2) 2025 (NBER w31161) — customer-support agents, not devs (scope caveat); GitHub Copilot RCT (arXiv:2302.06590); DORA 2024; Stack Overflow 2025 Developer Survey

Contradictions / Uncertainty

  • Scope mismatch on the strongest skill result. The +34%-novice finding is customer-support agents, not developers — directional counter-evidence, not same-population proof about coding. The developer-specific studies (METR, Copilot) point the same way, which is what keeps it credible.
  • Speed, not labor-market value. Every finding measures task speed/productivity; the "juniors irrelevant" thesis is about wages, hiring, and headcount, which no source here measures. So this inverts the productivity claim without settling the employment claim — and the "experts see small quality declines / juniors lack judgment to catch AI errors" mechanism actually lends partial support to seniority-and-ai's risk argument.
  • Small samples / significance. METR n=16 (significant across 246 tasks, but a narrow population); the Copilot RCT's key experience coefficient is not significant at 0.05.
  • Time-sensitivity is the biggest limit (above). Treat every number as a dated snapshot.
  • This source has its own incentives too: GitHub/Microsoft authored the pro-Copilot RCT (COI); Stack Overflow has a mild interest in AI skepticism; DORA/Google is independent. Held to the vault's usual standard.

Next Questions

  • Does METR's 19% persist, vanish, or reverse on mid-2026 agentic tools? A properly powered follow-up RCT is the key missing piece — and would decide how much of this page survives.
  • Does "novices gain most" hold on real, complex codebases, or invert where juniors can't catch AI's errors?
  • Given this, should ai-market-shift's "$200/mo substitute" and dmitry-rodenko's "hourly dev is dead" carry a stronger Status: contested?
  • What would a pro-thesis rigorous source look like — is there RCT-grade evidence that AI does zero the cost for some real dev segment (greenfield, juniors, specific stacks)?