ingest: Stanford SWEPR widening-gap study and AI-in-SDLC adoption pitfalls
Add two new sources with summaries, new concepts (developer-as-agent-manager, review-is-the-new-bottleneck), new entities (SWEPR, Nikolai Sheiko), and a query on the Stanford source; update related concept pages, overview, index, and log.
This commit is contained in:
38
log.md
38
log.md
@@ -230,3 +230,41 @@ Entry format:
|
||||
- Uncertainty flagged: the pendulum case is second-hand tweets with fuzzy company identification (tentative); the checklist is prescriptive, not observed; the author's own company is currently building a Datadog replacement — if it ships and survives, he becomes his own counterexample. Open question with no evidence either way in the corpus: does the maintenance objection survive *agents* doing the maintenance ([[agentic-loops]], [[async-by-default]])?
|
||||
- Effect on the last lint's watch item: [[emacsification-of-software]] and [[explosion-of-internal-software]] were merge candidates "if neither gains second-source support by the next lint" — both now have second-source engagement (as a bounding counterpoint), which argues for keeping them separate with [[maintenance-is-the-real-cost]] as the shared boundary page.
|
||||
- Next: the standing recommendations are unchanged (a `taste` concept page; extend [[theo-konstantin-allie]] to four lenses; the skills falsification test). New candidate question for the HR audience: which of their candidate tools (candidate knowledge base, transcribe→summarize) fall on the safe side of the build-vs-buy checklist — directly webinar-relevant if Q&A raises "should we build or buy?"
|
||||
|
||||
## 2026-07-30 — query (Stanford "widening gap" chart traced to source)
|
||||
- Intent: query
|
||||
- Input: user saved `raw/assets/G6g3O60bkAE05ZW.png` (X/Twitter screenshot of a Stanford slide, "Teams that master AI are accelerating their productivity gains, widening the gap with laggards") and asked to find the original research.
|
||||
- Pages created: [[2026-07-30-stanford-widening-gap-source]] (query — the slide is from Stanford's Software Engineering Productivity Research group (SWEPR), Yegor Denisov-Blanch; 600+ companies / 120k+ engineers since 2022; the chart is a 46-vs-46-team difference-in-differences analysis showing the AI-adopter productivity gap growing 4.8% → 19% (4×) from April 2023 to July 2025; slide matches his Sept 2025 AI Conference deck "Will AI Replace Software Engineers?"; primary links recorded on the page).
|
||||
- Pages updated: [[2026-07-14-gap-between-ai-users-irreversible]] (external-corroboration pointer under Connections — Allie's title claim gains its first measured, non-practitioner support), `index.md` (Queries).
|
||||
- Notes: the DiD result itself is talk/deck-published, not peer-reviewed — marked tentative; only the measurement methodology has a peer-reviewed paper (arXiv 2409.15152). Nuance recorded: same study finds ~15–20% average gains with AI *decreasing* net productivity in complex legacy codebases (rework, +91% PR review time) — honest-caveat material aligning with [[maintenance-is-the-real-cost]] and [[make-more-cheap-code]]. No concept pages changed (citation policy: the talk is not yet an ingested source).
|
||||
- Next: decide whether to ingest the talk/deck as a proper `raw/sources/` doc (would let concept pages and the webinar theses cite it as `wiki/sources/*` evidence); consider adding the "Stanford measured it: 4× in two years" line to the webinar's stakes beat; watch for a peer-reviewed version of the DiD analysis.
|
||||
|
||||
## 2026-07-30 — ingest (Stanford SWEPR — AI and the widening productivity gap)
|
||||
- Intent: ingest
|
||||
- Input: user — "Create the source. I'm giving you permission to write a new file in the sources folder." **Authorized raw write** (exception to rule 1, new file only): created `raw/sources/Stanford SWEPR - AI and the widening productivity gap.md`, a research *dossier* compiled from the saved slide screenshot (`raw/assets/G6g3O60bkAE05ZW.png`) plus public coverage — honestly marked as not-a-transcript, with per-claim provenance (slide-read vs secondary coverage). 12th source; the corpus's first quantitative outside study.
|
||||
- Pages created: [[2026-07-30-stanford-swepr-widening-gap]] (source), [[swepr]] (entity — Stanford Software Engineering Productivity Research group, Yegor Denisov-Blanch; 15th entity).
|
||||
- Pages updated: [[levels-of-ai-usage]] (evidence: measured team-level twin of the mastery gap, with the non-engineer-audience caveat), [[context-as-scarce-resource]] (evidence: gains collapse toward 10M LOC via context-window limits — first outside quantitative support for context-as-constraint), [[make-more-cheap-code]] (evidence: +91% PR review time / 2.6× rework = the cost moving downstream, measured; brownfield-negative caveat), [[2026-07-14-gap-between-ai-users-irreversible]] (corroboration pointer upgraded to cite the source page), [[2026-07-28-webinar-theses]] (T7 note: thesis upgraded from prediction to measurement, with a citable stage line), [[2026-07-30-stanford-widening-gap-source]] (ingest follow-up marked done), [[overview]] (11→12 sources; stakes-claim measurement added to the agree section; navigation), `index.md`.
|
||||
- Notes: the source *corroborates* rather than contradicts — no new contradictions logged, but three caveats recorded on the source page and carried into every citing page: the DiD analysis is talk-published, not peer-reviewed (tentative); the rework/review numbers are secondary coverage; the study measures engineering teams, not the webinar's non-engineer audience. Deliberately did **not** claim the study validates the *skills* rung specifically — Stanford asserts "quality of usage" as the differentiator without decomposing it, so it cannot arbitrate contradiction #1 (skills vs no-skills). The webinar theses' T7 gains the corpus's best number: gap 4.8% → 19% (4×), April 2023 → July 2025.
|
||||
- Next: watch for a peer-reviewed version of the DiD analysis (status upgrade); candidate webinar edit — the one-line "Stanford measured it: the gap quadrupled in two years" in the stakes beat (user's call, script is an authored deliverable); standing recommendations unchanged (a `taste` concept page; extend [[theo-konstantin-allie]] to four lenses; the skills falsification test).
|
||||
|
||||
## 2026-07-30 — maintenance (authorized raw edit: Stanford beat in the script's closing arc)
|
||||
- Intent: maintenance
|
||||
- Input: user — "good idea to add it to the webinar script. Do it."
|
||||
- **Authorized raw edit** (exception to rule 1, on direct instruction; same pattern as the 2026-07-14/07-28 script edits): `raw/notes/Webinar script.md` — inserted a Stanford beat into the closing arc, between "You don't buy it. You build it — one small tool at a time." and "We started this journey…". The beat: Stanford tracked 46 AI teams vs 46 matched non-AI teams for 2+ years → the teams that *learned* it pulled away from the ones that just *had* it → spread under 5% (spring 2023) → 19% (summer 2025) → "the gap quadrupled in two years" → callback to the script's own reveal: "everyone had the same models the whole time. The difference was never the model. It was who built something around it."
|
||||
- Placement rationale: the closing arc is where the script's "same model, different harness" argument lands, and the Stanford curve is that exact argument as data — it also answers "why start now" right before the final chat-box→OS callback. Phrasing kept factually careful: the widening spread is *among AI-using teams* (masters vs laggards), so the beat says "pulled away from the ones that just had it," not "AI users vs non-users."
|
||||
- Stage-safety/honesty note added inline (`_note:`): source pointer to [[2026-07-30-stanford-swepr-widening-gap]] plus the three Q&A caveats (talk-published not peer-reviewed; software teams not office workers; "quality of usage" asserted but not decomposed).
|
||||
- Pages changed: [[2026-07-28-webinar-theses]] (T7 note updated — dramatized in the script as of today, promoted from Q&A material to an on-stage beat). No other wiki pages changed; `index.md` unchanged (no catalog change).
|
||||
- Note on scope: the Stanford research itself was already fully ingested earlier today ([[2026-07-30-stanford-swepr-widening-gap]] + [[swepr]] — see the previous ingest entry); this operation only carries the number into the deliverable.
|
||||
- Next: unchanged from the ingest entry (peer-review watch; `taste` concept page; four-lens comparison; skills falsification test).
|
||||
|
||||
## 2026-07-30 — ingest (Грабли во внедрении ИИ в SDLC / Nikolai Sheiko)
|
||||
- Intent: ingest
|
||||
- Input: `raw/sources/Грабли во внедрении ИИ в SDLC.md` — viewer's conclusions from Nikolai Sheiko's 45:59 Russian YouTube talk ("why the AI is there but the results aren't"). 13th source; `raw/sources/` fully ingested again.
|
||||
- Pages created: [[2026-07-30-rakes-in-ai-sdlc-adoption]] (source), [[nikolai-sheiko]] (entity — 16th), [[review-is-the-new-bottleneck]] and [[developer-as-agent-manager]] (concepts — 26th and 27th).
|
||||
- Core of the source: models are already good enough — **people, companies and metrics throttle the gains by an order of magnitude**. The SDLC collapsed into days/hours but not around the two human "red squares" (reviewer, planner); the fix is review *with* the agent plus the one metric that resists gaming — **completed tasks without rework** (never LoC/commits/PRs; his outsourcer case: more PRs, +1% net). Error #0: the AI-developer is an IO-bound **manager of an agent-employee**, not a CPU-bound coder ("sit watching Claude Code work = bad employee") — and not everyone can or should switch. "Companies no longer need custom AI development — install Claude Code/Codex, configure, attach connectors, mind security." **Agentic Evolution**: walk the agent through hard tasks → "remember this and write the manual for the next one" → verify via a **context-free subagent** solving the task from the skill alone. Compaction curse on big codebases → best practices exist *for the agent* (locality, interfaces, AST search over grep); embeddings/RAG over code rejected flatly.
|
||||
- Why it matters to this vault: (1) **second practitioner vote for the skills layer**, narrowing the evidential asymmetry Thorsten's dissent enjoyed on [[skills-as-memory]] — and his verification protocol is the corpus's first described run of anything like the proposed skills falsification test (with-skill half only; no without-skill control, so the test question stands). (2) **Independent citation of the Stanford chart** ([[swepr]]) as his stakes slide, plus an anecdotal mirror of its +91%-review-time finding — recorded on [[2026-07-30-stanford-swepr-widening-gap]]. (3) Names the org-level bottleneck the corpus had only as an open question on [[async-by-default]] (parallel diffs pile up on a human) — now a page: [[review-is-the-new-bottleneck]]. (4) His anti-RAG-for-code stance supports the audience-driven reading of the skills-vs-RAG contradiction on [[context-as-scarce-resource]] (anti-RAG votes are about code/procedures; the pro-RAG vote is about business data).
|
||||
- Pages updated: [[solve-first-then-skillify]] (Agentic Evolution + 5-step verification protocol; next-question partially answered), [[skills-as-memory]] (evidence + asymmetry softened + falsification-test note), [[context-as-scarce-resource]] (compaction curse; codebase-stores-context converging with Thorsten from the opposite direction; AST search; RAG contradiction note), [[async-by-default]] (IO-bound-manager evidence; related links to both new concepts), [[enterprise-ai-reality]] (no-custom-AI-dev quote as the managed-harness market seconded; external-configurator vs teacher/curator anti-pattern; Cursor metered billing → team economizes, the metered-vs-subscription split observed organizationally; tokens-dearer-then-cheaper matching Eugene's prediction), [[make-more-cheap-code]] (related link to the review-bottleneck page), [[2026-07-30-stanford-swepr-widening-gap]] (independent-citation pointer), [[overview]] (12→13 sources; new "adoption side" bullet; agree/diverge updates; skills-dissent paragraph rebalanced), `index.md`.
|
||||
- New contradiction logged (on [[developer-as-agent-manager]]): Sheiko's "don't force everyone" vs the irreversible/compounding gap (Allie, Stanford) — opting out is legitimate *and* costly; no source reconciles the two. Also tentative: all client cases are anonymous self-reported anecdotes; "embeddings over code don't work" has no mechanism given; the users-vs-Agentic-Operations split is prediction, not observation.
|
||||
- Uncertainty about the speaker himself: affiliation unknown; "you don't need custom AI development" is also a consultant's pitch — flagged on [[nikolai-sheiko]].
|
||||
- Webinar relevance noted but not applied (script untouched): the Intelligence-vs-Judgment framing is close kin to the planned "what stays yours" closing beat, and the talk independently strengthens the case for the standing `taste` concept-page recommendation (Judgment = taste-built-over-years or domain expertise — a second source alongside Thorsten's).
|
||||
- Next: standing recommendations unchanged (a `taste` concept page — now with two sources backing it; extend [[theo-konstantin-allie]] to four lenses; the skills falsification test — half-run by this source, control still missing). New candidate question: what does review-with-the-agent look like concretely (no transcript of it done well exists in the corpus).
|
||||
|
||||
Reference in New Issue
Block a user