The 17 drafts that read 0/17 against the poisoned tree were never 17 codegen
failures. Re-gated on a clean tree: 9 banked, then re-gated TOGETHER per binary
(a draft that passes alone can still collide with a sibling in the same TU —
which is exactly what `morph_lerp` and `struct V8` did) — 9 of 9 held.
ov_SC03_118 func_80183D68 · ov_SC06_008 func_8017F8CC func_80183BF4
ov_SC06_011 func_8017EDA4 · ov_SC07_006 func_8017FDF8
ov_SC06_018 func_80184370 func_80184944 func_8018A974 func_8018EDB8
Wave 4 final: 40 targets -> 35 agent-MATCH -> 27 BANKED (68%).
One stray comment in one draft had masked eight real matches.
Wave 4 (wf_05895a19-121, 75 agents, 7.9M tok): 40 targets -> 35 agent-MATCH,
0 refuted, 5 NEAR, 0 FAIL, 0 drafts lost. 18 banked so far on the whole-binary
gate across 9 binaries; the other 17 are re-gating on a clean tree (see below).
PRIOR-NOTES SEEDING HELD AT SCALE: 10 of 12 seeded targets confirmed (wave 3
was 7 of 9). func_8017C294 — the x16 family, the largest single item on the
board — is now NEAR at **2 ins** (18 -> 11 -> 2 across three seeded attempts).
THE 4th COMMENT-BLINDNESS DEFECT OF THE SESSION, and the first with blast
radius. A crack agent annotated a decl in its draft:
extern void func_801842DC(s32 a0); /* TU:4023 INCLUDE_ASM (no decl) */
`corpus._INCLUDE_ASM_CAND` only skips lines that BEGIN with a comment marker,
so it read `INCLUDE_ASM (` out of the trailing PROSE, found no quoted path, and
refused the whole binary's stub oracle — correctly, by its own R32 contract.
That then failed gate_stage for every LATER binary in the run, because they all
walk the corpus: 12 binaries banked, then 5 were blocked by one comment in a
13th. Fixed the same way as the other three today: decide candidacy on
cdecl._mask'ed text, PARSE FROM THE ORIGINAL (the mask blanks string content
and would erase the asm path). Verified on 4 binaries incl. main (2,002 stubs).
AND THE DAMAGE IT LEFT: gate_stage raised out of the CorpusError BEFORE its
revert, stranding failed drafts spliced in src/. The 17 solo re-gates that
followed all read 0/17 — they were building a POISONED TREE, not judging their
own drafts. Residue reverted here; the 17 re-gate clean next.
The pattern is now sharp enough to state: any scanner that greps C source for a
token must mask comments and strings FIRST — and agent-authored drafts make it
far likelier, because their prose mentions the exact tokens our tools hunt for.
The §163z catalogue was 34 crack-agent note-sets claiming 190 distinct laws.
One independent skeptic per function, each required to read the full notes,
grep the whole cookbook, classify, and GRADE THE EVIDENCE:
NEW 20 · SHARPENS 62 · COVERED 80 · UNSOUND 28
byte-probed 114 · single-instance 51 · asserted 25
57% of what the crack agents flagged as novel was already in the cookbook or
does not survive scrutiny. That ratio is the lesson: a crack agent is the right
instrument for FINDING a lever and the wrong one for judging its novelty — it
has just spent hours in one function and has not read the other 497 sections.
Never bank a wave's flags directly.
Banked as §164-01..82, each carrying its verdict, what it sharpens, and its
evidence grade (70 byte-probed, 12 single-instance). Several skeptics CORRECTED
the mechanism the crack agent proposed while confirming its effect — e.g. the
"fold distributes the constant out of an index" claim, where the skeptic traced
the real site to expand_expr's MULT_EXPR EXPAND_SUM case (expr.c:5359-5375)
after showing pointer_int_sum's distributive law cannot fire on that tree.
§164z records the 28 REFUTED claims with the reason, so no future wave spends
tokens rediscovering them.
cookbook_index.py: 501 sections.
The waves flagged ~40 candidate laws. Five were byte-probed, generalizable and
actionable enough to bank; they were deduped by hand against the file (no
skeptic-agent pass this time, so each says what it sharpens and why that
section is insufficient):
- §163a decl-conflict severity is SCOPE-DEPENDENT — hard error if either decl
is at file scope, warning only if BOTH are at block scope. §8d proves the
phenomenon on D_801812A4 but never states the rule or its LEVER half: block
scope is a deliberate conflict SOLVENT, so a struct-typed draft can be banked
into a scalar-typed TU by moving the typedef AND the extern into the block.
- §163b the switch-index parameter-WIDTH oracle: sll/sra straddling the minval
subtract is a 2-insn signature of a short parameter. Read the extension, not
just the bound.
- §163c case_values_threshold is 5 — an empty `case k:` glued to default can be
the only thing that emits a table at all; jtbl[k]==default label is the tell.
- §163d cse deletes a reg-reg copy by rewriting the PREVIOUS insn's SET_DEST
(cse.c:7440-7477). This is §162j's symptom in a DIFFERENT PASS and needs a
different lever; the residual it explains had been declared "unsteerable, 30
variants all >=17" and fell to source-shape edits alone, no pins.
- §163e the frame is a PSEUDO-NUMBER oracle (reload1.c:658 alter_reg in NUMBER
order); dead-local slot order is not declaration order, and a BLKmode local
is 8-aligned while a scalar s32 is not. Sharpens §162i, which gets the pad's
SIZE right and its PLACEMENT wrong.
§163z catalogues the ~35 unvetted claims by function so a future harvest can go
straight to them, explicitly marked "one agent's reconstruction until
re-measured" (R14).
cookbook_index.py: 497 sections.
Wave 3 added 84 (27 cracks + 21 propagated families). Fleet 94.8% instr /
89.2% distinct / 96.41% fn-count. R22 clean-fleet run 6x this session, 213/213
every time. 96 commits.
Bank rate measured three times: 67% -> 79% -> 69%. The dip is the cost curve
(wave 3's tier was 29 Opus-band / 14 jr vs wave 2's 8 / 5, median reach x6 ->
x3-4), not a regression.
Two levers proved out and belong in every future wave: the hardened harness
contract (0 drafts lost vs 21) and prior-notes seeding (7 of 9 previously
failed targets converted, incl. both long-standing NEARs and all three wave-2
gate misses). NEAR is a resumable state, not a write-off.
Wave 3 (wf_2680d8ff-539, 74 agents, 8.08M tok): 39 targets -> 35 agent-MATCH,
0 refuted, 4 NEAR, 0 FAIL -> 27 BANKED on the whole-binary gate (69%).
THE HARDENED HARNESS HELD: 0 drafts missing on disk (wave 2 lost 21 of 26 to a
shared output dir). Per-agent dirs + "never touch anything outside your own
directory" + a verifier that re-runs sha1sum LAST.
PRIOR-NOTES SEEDING IS THE SESSION'S BEST LEVER: 7 of 9 seeded targets
converted, including all three wave-2 whole-binary-gate misses and both big
NEARs — func_80189540 (551 ins, was NEAR +2) and func_8017C3BC (407 ins, was
NEAR 17). func_8017C294 (the x16 family) went 18 -> 11 ins: narrowing, not a
wall.
func_80189540 also required the one host edit its agent byte-probed:
src/ov_SC04_018/ov_SC04_018_jr_80188E1C.c:3093
extern s32 func_80189540(s32 a0, s16 a1) -> (s16 a0, s16 a1)
That TU has no call site, so the edit is inert; the OTHER TUs' (s32,s16) decls
are deliberately left alone (real call sites, and an s16 prototype there would
force caller-side truncation and could de-match banked callers).
Its agent also recovered a better draft that already existed at
.run/backlog_drafts/func_80189540.c — a 551-ins MATCH that had been DE-matched
to 549 by "fixing" the definition's s16 first parameter to s32, the exact
inverse of that draft's own warning. Restoring s16 recovered the match.
The wave-2 draft declared `extern void func_8012CAE4(void *a0);` at block
scope while the host TU already declares it twice at file scope (K&R at :2790,
`s32 a0` at :2849), so the gate reported PLUMBING: conflicting types.
Dropping the decl is wrong — match_one compiles the draft STANDALONE and then
the symbol is undeclared (gcc-2.7.2 prints that with no `error:` prefix, so it
reads as CC1 FAIL). The fix is to AGREE with the TU's visible decl and cast the
ARGUMENT (`(s32)a0` — same bits in $a0, codegen unchanged).
MATCH (99 ins) standalone, then BANKED on the whole-binary gate.
Today a 28-agent wave lost 21 adversarially-verified drafts to a shared output
directory, and I wrote them off before Drew asked whether the workflow results
could just be analysed. They were all recoverable, for zero agent tokens.
Encodes the method that worked 21/21, including the two shortcuts that do NOT:
- taking each Write's content recovers only single-write drafts (8/21 — agents
refine);
- taking an Edit's new_string as a file yields a FRAGMENT, not a file.
So it replays the mutation history per (agent, file_path), snapshots after
every mutation, emits newest-first, and also scans Bash heredocs (the 21st
draft never used Write/Edit at all). --gate runs match_one newest-first and
keeps the first MATCH.
Self-test on wf_d804f25a-f6f: 3/3 including the heredoc case.
Session close state. Three parts: stage 0b (91, zero decompilation), wave 1
(26), wave 2 (116). Fleet 94.4% -> 94.7% instr, 88.3% -> 88.9% distinct,
13,345 -> 13,112 stubs. R22 clean-fleet run 5x, 213/213 every time.
The campaign now has a MEASURED rate, twice: 67% (wave 1, all-Opus) then 79%
(wave 2, 20 of 28 Sonnet) of cracks survive the whole-binary gate. The Sonnet
band beating the all-Opus wave is the session's most useful economic finding
and sets wave 3's routing.
Resume order changed on evidence, twice over:
- harden the wave harness FIRST (per-agent dirs, sha1-last verifier, and a
tools/recover_drafts.py built from the transcript-replay method that
recovered 21/21 today);
- then wave 3, sized on 79%, not on the reach-15 prior.
Error ledger grew to 6. The two that matter: I wrote off 21 verified cracks as
lost when the run transcripts held every one of them, and my first two
recovery passes both failed by reading a single tool record instead of
replaying the file's mutation history.