Wave 3 added 84 (27 cracks + 21 propagated families). Fleet 94.8% instr /
89.2% distinct / 96.41% fn-count. R22 clean-fleet run 6x this session, 213/213
every time. 96 commits.
Bank rate measured three times: 67% -> 79% -> 69%. The dip is the cost curve
(wave 3's tier was 29 Opus-band / 14 jr vs wave 2's 8 / 5, median reach x6 ->
x3-4), not a regression.
Two levers proved out and belong in every future wave: the hardened harness
contract (0 drafts lost vs 21) and prior-notes seeding (7 of 9 previously
failed targets converted, incl. both long-standing NEARs and all three wave-2
gate misses). NEAR is a resumable state, not a write-off.
Wave 3 (wf_2680d8ff-539, 74 agents, 8.08M tok): 39 targets -> 35 agent-MATCH,
0 refuted, 4 NEAR, 0 FAIL -> 27 BANKED on the whole-binary gate (69%).
THE HARDENED HARNESS HELD: 0 drafts missing on disk (wave 2 lost 21 of 26 to a
shared output dir). Per-agent dirs + "never touch anything outside your own
directory" + a verifier that re-runs sha1sum LAST.
PRIOR-NOTES SEEDING IS THE SESSION'S BEST LEVER: 7 of 9 seeded targets
converted, including all three wave-2 whole-binary-gate misses and both big
NEARs — func_80189540 (551 ins, was NEAR +2) and func_8017C3BC (407 ins, was
NEAR 17). func_8017C294 (the x16 family) went 18 -> 11 ins: narrowing, not a
wall.
func_80189540 also required the one host edit its agent byte-probed:
src/ov_SC04_018/ov_SC04_018_jr_80188E1C.c:3093
extern s32 func_80189540(s32 a0, s16 a1) -> (s16 a0, s16 a1)
That TU has no call site, so the edit is inert; the OTHER TUs' (s32,s16) decls
are deliberately left alone (real call sites, and an s16 prototype there would
force caller-side truncation and could de-match banked callers).
Its agent also recovered a better draft that already existed at
.run/backlog_drafts/func_80189540.c — a 551-ins MATCH that had been DE-matched
to 549 by "fixing" the definition's s16 first parameter to s32, the exact
inverse of that draft's own warning. Restoring s16 recovered the match.
The wave-2 draft declared `extern void func_8012CAE4(void *a0);` at block
scope while the host TU already declares it twice at file scope (K&R at :2790,
`s32 a0` at :2849), so the gate reported PLUMBING: conflicting types.
Dropping the decl is wrong — match_one compiles the draft STANDALONE and then
the symbol is undeclared (gcc-2.7.2 prints that with no `error:` prefix, so it
reads as CC1 FAIL). The fix is to AGREE with the TU's visible decl and cast the
ARGUMENT (`(s32)a0` — same bits in $a0, codegen unchanged).
MATCH (99 ins) standalone, then BANKED on the whole-binary gate.
Today a 28-agent wave lost 21 adversarially-verified drafts to a shared output
directory, and I wrote them off before Drew asked whether the workflow results
could just be analysed. They were all recoverable, for zero agent tokens.
Encodes the method that worked 21/21, including the two shortcuts that do NOT:
- taking each Write's content recovers only single-write drafts (8/21 — agents
refine);
- taking an Edit's new_string as a file yields a FRAGMENT, not a file.
So it replays the mutation history per (agent, file_path), snapshots after
every mutation, emits newest-first, and also scans Bash heredocs (the 21st
draft never used Write/Edit at all). --gate runs match_one newest-first and
keeps the first MATCH.
Self-test on wf_d804f25a-f6f: 3/3 including the heredoc case.
Session close state. Three parts: stage 0b (91, zero decompilation), wave 1
(26), wave 2 (116). Fleet 94.4% -> 94.7% instr, 88.3% -> 88.9% distinct,
13,345 -> 13,112 stubs. R22 clean-fleet run 5x, 213/213 every time.
The campaign now has a MEASURED rate, twice: 67% (wave 1, all-Opus) then 79%
(wave 2, 20 of 28 Sonnet) of cracks survive the whole-binary gate. The Sonnet
band beating the all-Opus wave is the session's most useful economic finding
and sets wave 3's routing.
Resume order changed on evidence, twice over:
- harden the wave harness FIRST (per-agent dirs, sha1-last verifier, and a
tools/recover_drafts.py built from the transcript-replay method that
recovered 21/21 today);
- then wave 3, sized on 79%, not on the reach-15 prior.
Error ledger grew to 6. The two that matter: I wrote off 21 verified cracks as
lost when the run transcripts held every one of them, and my first two
recovery passes both failed by reading a single tool record instead of
replaying the file's mutation history.
Drew's question ("can you just analyze the workflow results to get those
drafts back?") was right, and my write-off was wrong. The drafts were never
lost: every agent's tool calls are recorded in the run transcripts, content
included. Recovery is deterministic and cost ~0 agent tokens.
Method (all three passes were needed):
1. Write records -> 20 of 21 had one. Naive extraction gated only 8/21,
because agents REFINE with Edit and I was treating each edit's new_string
as a whole file.
2. Replay Write-then-Edit in order per (agent, file_path), snapshotting after
every mutation, then gate every snapshot newest-first -> 20/21 MATCH.
3. The last one (func_8017DF40) never used Write/Edit — it wrote via a shell
heredoc. Extracted the heredoc bodies from the bash tool calls -> MATCH.
Whole-binary gate on the 21: 17 BANKED, 4 NEAR.
md_SC03_073 func_801EFC94 x8 <- a MODULE exemplar, through the path fixed
earlier today (commit:1626)
ov_SC03_014 func_8017C154 x7 func_8017C6A0 x7 func_8017D1E0 x7
ov_SC03_118 func_80187180 x8 func_80187B80 x8
ov_SC06_018 func_80185C2C func_8018DE60 func_801850F4 func_801857CC
func_801880C4 func_80185DD8 (all x6)
ov_MAIN_012 func_8017DD28 x5 · ov_SC02_026 func_801814D8 x6
ov_SC03_093 func_8018171C x5 · ov_SC03_107 func_8017EB70 x5
ov_SC06_008 func_80180000 x8
NEAR at the binary gate: func_8017DF40, func_8017EEEC, func_80187960,
func_8018A974 — the §52b population (per-function MATCH, binary gate refuses).
Wave 2 now stands at 22 of 28 banked (79%) vs wave 1's 8 of 12 (67%), with
20 of 28 drafted by Sonnet.
Wave 2 (wf_d804f25a-f6f, 54 agents, 6.76M tok, 66 min): 28 targets ->
26 agent-MATCH (93%, vs wave 1's 75%), 0 refuted, 2 NEAR, 0 FAIL.
Then only 5 of the 26 drafts still existed on disk. The verifiers were NOT
lying — their evidence quotes exact instruction counts matching each target
(101/187/108/77 ins), so the files existed when they ran. Later crack agents
DELETED them while tidying the SHARED .run/wave2/ directory; one agent's own
notes say "scratch dir removed afterwards so .run/wave2/ contains only draft
.c files". 28 agents, one output dir, and a prompt line ("drafts to .run/wave2
ONLY") that invited cleanup.
Zero-token recovery: re-gated every surviving .c under .run/wave2/** whose
text DEFINES the target function -> 2 of 21 recovered.
Banked (5/5 through the whole-binary gate — every draft that survived passed):
ov_SC01_005 func_8017ED5C x5 func_8017DEFC x5
ov_SC03_093 func_801818F0 x5
ov_SC04_018 func_80188B84 x5
ov_SC06_018 func_80183F50 x6
The 21 lost cracks keep their full agent notes in .run/jr48/wave2_lost.json —
idioms, integration surface, family maps. They are re-runnable from those
notes at a fraction of a cold crack.
`jtbl_family_bank` fed every module jr member to `jtbl_carve`, which died with
`jtbl_… not found in the raw data asm`; `harvest_verify` turned that into
CARVE-REFUSED and never built. So the verdict named the TOOL, and 12 slots in
the wave-1 propagation read as a carve bug. Probing one member to the byte
level shows it is a LAYOUT the carve model does not cover:
A module binds `.rodata` at 0x0 to the SAME subseg as its code (§154-A), so
the object's rodata order IS the C file's include chain — INCLUDE_RODATA
pieces, then each INCLUDE_ASM'd function's MIGRATED table, in address order.
That reproduces the island exactly while the function is a stub. Matching it
PRUNES its .s, its table leaves the chain, and cc1 re-emits it at the END of
the object's .rodata: build 43,768 vs 43,760 bytes, first diff at 0x144
inside the island's own pointer table.
`JTBL_PADS` does not reach it either — `jtbl_rodata_pads` refuses the object
outright ("unexpected rodata content .include ... D_801EF468.s"): the carve
model covers jump tables, not an island of mixed included data.
- `migrated_tables()` detects the layout by EVIDENCE (table absent from the
data asm, present as a dlabel in the function's own .s), refuses loud with
the measurement and the design that would work (isolate the jr function into
its own subseg so its .rodata is a separate OBJECT, then ld_interleave — the
§8 machinery re-aimed at a LEADING island instead of a data tail), and
refuses a mixed carve set rather than half-carving (R32).
- Regression-checked both ways: overlay stubs classify [], modules classify
migrated.
SIZED (R37): 70 module binaries, 42 with this layout; 1,345 open module
member-slots in sibling families, of which only 44 are jr. The island work is
worth 44 slots — it is NOT the module lane's main gate.
R22 clean-fleet: 213 passed / 0 failed of 213.
`_body_open_brace` ran BOTH its scans on unmasked text. A crack agent's draft
opens with a header comment that names the function and quotes C at it:
/* func_801EE8E0 (ov_MAIN_012 / jr_801789AC) — 188 ins, byte-exact vs …
* 3) The `do { } while (0)` around the loop-1 call is a REGISTER-ALLOCATION
so `sig` matched the COMMENT's first line and `find('{')` found the comment's
`do {`. Every carried `extern` was spliced into the comment — silently
commented out — and the gate reported `'D_8011511A' undeclared`.
The sweep classified that CC1-FAIL, so it read as a property of the SIBLING
(all 4 members failed identically) when it was a property of the EXEMPLAR'S
PROSE. It had nothing to do with module binaries, which is where I had filed
it. Every richly-commented agent draft is a carrier; the trigger is any brace
inside the header comment — so this would have grown with the campaign.
- both scans now run on `cdecl._mask`ed text and index the original by the
masked offsets (§134 / R33: one masking oracle);
- refuse outright if the length invariant is broken, rather than mis-place a
declaration into live code (R32).
Measured: family func_8017CBC8 -> its 4 md_ siblings went 0/4 -> 4/4 banked.
Stage 0b closed (91, zero decompilation) + Stage-1 wave 1 (8 cracks -> 26
instances). Fleet 94.6% instr / 88.8% distinct / 96.36% fn-count.
Resume order changed on measured evidence: FIX THE md_ MODULE LANE FIRST.
16 of the wave's 42 member slots were unreachable for tooling reasons, not
matching reasons — 12 on a carve that assumes raw data lives in
<binary>/data/*.data.s (modules do not), 4 on an uncarried extern
(`D_8011511A' undeclared). Both are named with verbatim errors; probe one of
each before pricing (R37). Precedent: 0b's three repairs banked 91 for ~0
agent tokens; the wave spent 3.36M for 26.
Also recorded: the frontier re-derivation (1,955 zero-crack families /
330,622 templatable ins), the tier-ordering correction (ins-per-crack is flat
across x5-x8, so rank by templatable weight, not by tier), and the §162
harvest with its two in-place cookbook corrections.