Files
BusinessNotes/wiki/concepts/ai-productivity-evidence.md
EugeneTes 60176d2fdc all
2026-07-30 11:15:52 +02:00

51 lines
5.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# AI Productivity — Measured vs. Claimed
#concept #ai #evidence
## Summary
The empirical counterweight to the vault's dominant premise. Where the vault's promotional sources (twelve at this page's creation; nineteen of twenty as of 2026-07-29) *assert* AI has zeroed the cost of commodity software work, the rigorous evidence (RCTs, a top-5-journal field study, large surveys) shows the real effect is **modest (≈1419%), context-dependent, tilted toward NOVICES not experts, and systematically overstated by self-report.** It **qualifies rather than demolishes** the thesis — no study finds AI worthless — but it directly inverts two load-bearing claims and is the vault's first source with evidentiary standing rather than sales incentive. Anchored entirely by [[2026-07-18-ai-productivity-adversarial-evidence]]; single-source, so `Status: tentative`, but its inputs are the highest-provenance in the vault.
## Current Understanding
**Two direct inversions of the popular thesis:**
1. **"AI zeroed the cost of skilled dev work"** → the strongest *developer-specific* RCT (METR, 2025) found 16 experienced open-source maintainers were **19% *slower*** completing real tasks on their own repos with early-2025 AI. The cost didn't go to zero; for experts on familiar code it went *up*.
2. **"Seniors up, juniors irrelevant"** ([[seniority-and-ai]]) → the best skill-distribution study (Brynjolfsson, Li & Raymond, *QJE* 2025) found the **least-experienced gain most** (+34% vs. ~0 and small quality *declines* for experts). A GitHub Copilot RCT independently found less-experienced developers benefit more. AI's productivity dividend accrues to novices, not seniors.
**The perception trap (why the hype is self-sustaining).** METR's developers forecast a 24% speedup, *still* estimated +20% after finishing, and were measured 19% — a ~39-point gap between felt and real. This matters because "it feels faster / I shipped more" is the exact evidence most promotional AI-productivity claims are built on. Self-report is not measurement.
**Individual speed ≠ shipped software.** DORA 2024 found AI adoption correlated with *reduced* delivery stability (and, that year, throughput — reversed in DORA 2025). Stack Overflow 2025 (~49k devs): only ~3% "highly trust" AI accuracy while ~46% distrust it (up from 31%), even as adoption hit 84%; ~45% say debugging AI code is time-consuming. The artifact isn't finished when it's generated.
**The honest reading.** Held to the same standard as the promotional sources, this evidence does **not** prove AI is worthless — average effects are positive (1415%). It proves the effect is *smaller, unevenly distributed, and easier to overstate* than "cost → zero" implies. The vault's strategic advice can survive a modest, novice-tilted AI; it cannot survive on the premise that AI is a full substitute for a developer.
**⚠ Time-scope.** All figures are 20232025, tied to specific tool generations. METR labels its result "historical"; the Copilot RCT used 2022 Copilot; DORA reversed a 2024 finding in 2025. This is **not** the live 2026 state of the art — it is a corrective against extrapolating hype, not a prediction.
## Evidence
- Whole concept (all six findings, effect sizes, designs, caveats, and the two refuted claims) — [[2026-07-18-ai-productivity-adversarial-evidence]]
- Primary roots cited there: METR RCT (arXiv:2507.09089); Brynjolfsson/Li/Raymond *QJE* 140(2) 2025 (NBER w31161) — **customer-support agents, not devs** (scope caveat); GitHub Copilot RCT (arXiv:2302.06590); DORA 2024; Stack Overflow 2025 Developer Survey
## Related Pages
- [[ai-market-shift]] — the claim this most directly tests (is a $200/mo AI a real *substitute*?)
- [[seniority-and-ai]] — the claim this most directly *inverts* (novices gain most)
- [[future-of-engineering-work]] — "coding cost → 0 / teams collapse" qualified by modest, uneven effects
- [[2026-07-06-sebastian-interview-ai-and-software-engineering]] — the vault's strongest statement of the thesis this counterbalances
- [[overview]]
## Contradictions / Uncertainty
- **Scope mismatch on the strongest skill result.** The +34%-novice finding is customer-support agents, not developers — directional counter-evidence, not same-population proof about coding. The developer-specific studies (METR, Copilot) point the same way, which is what keeps it credible.
- **Speed, not labor-market value.** Every finding measures task speed/productivity; the "juniors irrelevant" thesis is about wages, hiring, and headcount, which no source here measures. So this inverts the *productivity* claim without settling the *employment* claim — and the "experts see small quality declines / juniors lack judgment to catch AI errors" mechanism actually lends partial support to [[seniority-and-ai]]'s risk argument.
- **Small samples / significance.** METR n=16 (significant across 246 tasks, but a narrow population); the Copilot RCT's key experience coefficient is not significant at 0.05.
- **Time-sensitivity is the biggest limit** (above). Treat every number as a dated snapshot.
- **This source has its own incentives too:** GitHub/Microsoft authored the pro-Copilot RCT (COI); Stack Overflow has a mild interest in AI skepticism; DORA/Google is independent. Held to the vault's usual standard.
## Next Questions
- Does METR's 19% persist, vanish, or reverse on **mid-2026 agentic tools**? A properly powered follow-up RCT is the key missing piece — and would decide how much of this page survives.
- Does "novices gain most" hold on **real, complex** codebases, or invert where juniors can't catch AI's errors?
- Given this, should [[ai-market-shift]]'s "$200/mo substitute" and [[dmitry-rodenko]]'s "hourly dev is dead" carry a stronger `Status: contested`?
- What would a *pro-thesis* rigorous source look like — is there RCT-grade evidence that AI **does** zero the cost for some real dev segment (greenfield, juniors, specific stacks)?