194 lines
9.8 KiB
Markdown
194 lines
9.8 KiB
Markdown
# Designing the Webinar's Verification Beat
|
||
|
||
#query
|
||
|
||
## Question
|
||
|
||
"What are your suggestions for the verification beat? What can we add?" — following [[2026-07-28-webinar-theses]], which flagged the absence of any verification moment as the current script's biggest gap. (2026-07-28)
|
||
|
||
## Answer — the headline
|
||
|
||
**Make it one beat that closes both gaps, not two.** The refresh listed two holes in the script: no verification beat, and no "what stays yours" beat. They are the same beat. Verification is *precisely* where the human's remaining job lives ([[product-ownership]]), so a single moment can answer "can I trust it?" and "then what's left for me?" at once. In a 30-minute talk that matters — two gaps, one insertion.
|
||
|
||
The frame that makes it land for non-engineers, from [[2026-07-24-non-engineer-throwaway-verification]]:
|
||
|
||
> **You don't verify by reading everything. You verify by making a second, cheap, disposable agent try to break the first one.**
|
||
> Generated *checks*, not generated *content*.
|
||
|
||
And the honest bottom line that becomes the "what stays yours" half:
|
||
|
||
> A checker catches **drift**. It cannot catch a **wrong rule**. Mechanical correctness is delegable; judgment is not.
|
||
|
||
## Placement — where the anxiety actually peaks
|
||
|
||
The script's emotional arc is capability rising as the human recedes: *I do everything* → *I click a button*. The audience's unease peaks at one exact line in the **Process** station:
|
||
|
||
> "I take my hands off the keyboard." … "Nobody is typing. It just... runs."
|
||
|
||
That is the moment to answer it — not earlier (they don't feel it yet) and not in Q&A (too late). Recommended split:
|
||
|
||
| Where | What | Cost |
|
||
|---|---|---|
|
||
| **Skills station** | Introduce the checker skill — the second species of skill | ~45–60 sec |
|
||
| **Process station** | Cash it in: the unattended process checks itself | ~20 sec (reuses an existing reveal) |
|
||
| **Closing arc** | The judgment half — what a checker *can't* do | ~30 sec, also fills gap #2 |
|
||
|
||
---
|
||
|
||
## Option 1 — Minimal: one line in an artifact you already show *(≈15 seconds)*
|
||
|
||
The Process station already reveals the prompt the shell wrote for its worker ("the AI wrote... a prompt. For another AI."). Add one visible line to that prompt:
|
||
|
||
```
|
||
After moving the cube, re-read the shelf state and confirm it matches the rule.
|
||
If it doesn't, report the mismatch instead of reporting success.
|
||
```
|
||
|
||
Then one spoken line:
|
||
|
||
> Look at the last instruction it gave itself.
|
||
>
|
||
> "Check your own work — and if it's wrong, *say so* instead of saying done."
|
||
>
|
||
> It didn't just delegate the job. It delegated the *checking*.
|
||
|
||
**Why this is the cheapest possible win:** the reveal already happens, the artifact is already on screen, and you add zero demo steps. If time is tight, do only this.
|
||
|
||
---
|
||
|
||
## Option 2 — Recommended: the checker skill *(≈60–90 seconds)*
|
||
|
||
Slots into the **Skills** station, immediately after `"save what we just did as a skill"`. The apparatus needed already exists: cube, shelves, weather widget, Override slider.
|
||
|
||
**Draft copy, in the script's voice:**
|
||
|
||
> So now I've got a skill that does the job.
|
||
>
|
||
> But here's the question you're all actually asking.
|
||
>
|
||
> If I'm not watching... how do I know it did it *right*?
|
||
>
|
||
> Let's give it a second skill.
|
||
>
|
||
> "save a skill that checks the first one — read the rule, read the weather, look at where the cube actually is, and tell me if they disagree"
|
||
>
|
||
> {AI writes `check-weather-based-movement/SKILL.md`}
|
||
>
|
||
> Two skills now. One does the job.
|
||
>
|
||
> One does nothing *but* look for the job being done wrong.
|
||
>
|
||
> Now watch — I'm going to break it on purpose.
|
||
>
|
||
> {drag the cube to the wrong shelf by hand}
|
||
>
|
||
> "run the check"
|
||
>
|
||
> {AI reports: rule says top shelf, cube is on bottom — mismatch}
|
||
>
|
||
> It caught it.
|
||
>
|
||
> And notice what that check cost me.
|
||
>
|
||
> One sentence. No code. It's a folder with a note in it — same as the first one.
|
||
>
|
||
> That's the trick nobody tells you about working with AI:
|
||
>
|
||
> you don't check the work by reading all of it.
|
||
>
|
||
> You check it by asking for something *cheap* whose only job is to find the mistake.
|
||
|
||
**Why this specific demo works:** breaking it by hand is visible, instant, and unfakeable to a live audience — they see the cube in the wrong place *before* the agent says so. It also introduces the second species of skill (producers and checkers), which is a genuine corpus finding from [[2026-07-24-non-engineer-throwaway-verification]] and costs no new level.
|
||
|
||
**Audience translation to say right after** — the demo is a cube, the takeaway must not be:
|
||
|
||
> Same move, your work:
|
||
> "Read this job description as if you were a candidate who'd be put off by it — what did you see?"
|
||
> "Read this shortlist and argue *against* my top pick."
|
||
>
|
||
> That's not asking it to do the work. That's asking it to attack the work.
|
||
|
||
---
|
||
|
||
## Option 3 — The ambiguity moment: you already wrote the perfect example *(≈30 seconds, standalone)*
|
||
|
||
The Process station's goal line is:
|
||
|
||
> "keep the cube on the **right shelf**: below 20 — top, above 20 — bottom. continuously."
|
||
|
||
**"The right shelf"** is genuinely ambiguous — *correct* shelf, or the shelf on the *right*? The colon disambiguates it, so the script is safe as written. Which means you can deliberately show the unsafe version first:
|
||
|
||
> Before I give it the real goal — watch this.
|
||
>
|
||
> {type only: "keep the cube on the right shelf"}
|
||
>
|
||
> {agent moves the cube to the right-hand shelf}
|
||
>
|
||
> That's not what I meant.
|
||
>
|
||
> I meant the *correct* shelf. It heard the shelf on the *right*.
|
||
>
|
||
> And here's the part that costs you: it didn't ask. It didn't hesitate. It just confidently did the wrong thing.
|
||
>
|
||
> {now type the full goal with the rule spelled out}
|
||
>
|
||
> Every gap you leave, it fills. And it fills it *silently*.
|
||
|
||
This is the cheapest possible dramatization of [[leave-less-room-for-imagination]] — Eugene's sharpest claim, currently thesis T12 with no demo — and it doubles as verification motivation (*this* is what a checker catches). It costs one extra typed line and one cube movement.
|
||
|
||
It also inverts cleanly, which is the durable insight from [[2026-07-24-non-engineer-throwaway-verification]]: **a fresh agent's misreading is a free ambiguity detector.** Before sending a brief to a human, hand it to a zero-context agent and ask what it thinks you meant.
|
||
|
||
---
|
||
|
||
## The closing half — what a checker *can't* do *(fills gap #2)*
|
||
|
||
The script's closing arc is currently all harness ("the model never changed… that's the harness"). Add the human half immediately before "You don't buy it. You build it":
|
||
|
||
> One last thing — because I don't want to oversell this.
|
||
>
|
||
> That checker I wrote? It was written by the same AI it's checking.
|
||
>
|
||
> It'll catch the cube on the wrong shelf. Every time.
|
||
>
|
||
> What it will *never* catch... is Marcus's rule being wrong in the first place.
|
||
>
|
||
> If twenty degrees was the wrong number, both agents agree, confidently, forever.
|
||
>
|
||
> So here's the split, and it's the honest one:
|
||
>
|
||
> the machine checks whether the thing was done right.
|
||
>
|
||
> You check whether it was the right thing.
|
||
>
|
||
> That part doesn't get automated. That part is why you're still in the room.
|
||
|
||
**Why this is worth the 30 seconds:** it is the strongest available answer to "will this replace me," it is honest rather than reassuring, and it converts the talk's ending from *capability* to *the audience's own value* — which is what an inspire talk should land on.
|
||
|
||
## Recommended combination
|
||
|
||
If you add **one** thing: Option 1 (15 sec, free).
|
||
If you add **one minute**: Option 2 + the closing half.
|
||
**Best value for ~2 minutes total:** Option 3 at Process → Option 2 at Skills → closing half. Option 3 creates the fear, Option 2 resolves it, the closing bounds the resolution honestly.
|
||
|
||
Sequencing note: Option 3 sits *later* in the script than Option 2. If you use both, move the ambiguity moment earlier — into the Skills station just before the checker — so the problem precedes its solution.
|
||
|
||
## Evidence trail
|
||
|
||
- [[2026-07-24-non-engineer-throwaway-verification]] — the non-engineer analog: generated checks not content; checker skills as the second species; fresh-agent misread tests; the tier-D-stays-human caveat
|
||
- [[make-more-cheap-code]] — the engineer form (100:1 slop-to-ship), "have AI review before humans do," reading costs attention
|
||
- [[async-by-default]] — "you're async anyway, ask for proof," and the logged limit: *proof produced by the thing being checked is evidence, not verification*
|
||
- [[leave-less-room-for-imagination]] — drift's damage is what you don't notice; the source of Option 3
|
||
- [[product-ownership]] — verification as the human's remaining craft; taste as the meta-skill
|
||
- [[2026-07-28-webinar-theses]] — T6a, and the two gaps this design closes
|
||
- Script state: `raw/notes/Webinar script.md` (Process and Skills stations, closing arc)
|
||
|
||
## Open questions / honest caveats
|
||
|
||
- **A checker written by the agent, checking the agent, is not independent.** It catches mechanical drift, not shared misunderstanding. The closing half says this out loud rather than hiding it — but if a technical audience member pushes, the real answer is that independence comes from *the human choosing the rule*, not from a second model.
|
||
- **Nina's finding is a standing counterweight:** transcript beat summary in her workflow ([[2026-07-14-nina-interview]]). A checker that reports "looks fine" is a summary. Don't let the beat imply reading is now optional — Theo's tier discipline is that some things still get read line by line.
|
||
- Untested: none of this has been run in front of a non-engineer audience. The cube demo may make verification feel mechanical in a way that doesn't transfer to judgment work — which is exactly why the audience-translation lines after Option 2 are load-bearing rather than optional.
|
||
|
||
## Changed existing pages?
|
||
|
||
No concept or entity pages changed — this is design synthesis on top of existing pages. `index.md` and `log.md` updated; [[2026-07-28-webinar-theses]] links here from T6a.
|