- wave C: 35 cards (11 tell + 24 weak) -> 35/35 standalone (re-verified independently, R14)
-> 32 banked / 3 near, 91% gate, 0 symbol failures (Law 4 prevention worked)
- weak lane proven for the first time: 24/24 on haiku; 890 candidates remain
- reach measured: 32 exemplars, 8 with sharers, x2 each => ~1.25x effective (the x134
era ended in P25/29/30) -> throughput, not leverage, is now the lever
- tools/build_wave.py (pool=adapt|weak, corpus-derived open-stub filter, R35 gate guard)
- 3 self-inflicted instrument defects found+fixed+NC'd (P9, recorded not buried):
pgrep self-match via shell=True; corpus.stubs() is addr->Stub not names (nearly
declared both card pools spent); a wave fired on hand-typed placeholder cards (stopped)
- STRATEGIC: card lanes are ~0.23% of open ins/wave; the Atlas's head-crack bucket is
1,276 groups / 186k ins with high-reach groups up to 265 instances -> retarget waves
at atlas groups next
- tools/warmstart.py: --from-banked walks a banked exemplar's family's open
members, builds remapped proven-body drafts (symbol_map + build_draft), and
STREAM-classifies member-vs-seed with zero compiles; enqueues ONLY permuter-
shaped work (bucket==permuter or LENGTH-DRIFT |delta|<=2) as backlog records;
--lenmiss ingests T8's 49-route. Armed live: 59 enqueued, 120 refused by the
stream filter (the anti-92%-wasted-CPU discipline)
- grinder patch NOT needed: candidates() deliberately keeps unclassified
records ('unknown is not a reason to skip'), so pre-filtered enqueues flow
as-is — documented in the feeder docstring (YAGNI honored)
- family_cousins --weak-cards: 954 seeded-crack cards from the never-consumed
0.70-0.85 band, ins-ranked, §168 laws embedded, model-routed haiku 804 /
v3 43 / sonnet 86 / opus 21 (cheap tiers dominate), 0 unresolved .s
- match_one --emit-streams (additive; stdout-identity NC'd): word streams for the
len lane
- family_align.addr_true_rel: reloc-vs-constant range discriminator — FULL
conservative set kept for pair semantics (NC-1 157/157 regression), address-
true subset for indel eligibility only (a constant li-cluster must not read as
reloc-in-indel); synthetic probes green both directions
- tools/len_tells.py: aligned classification + §172b tell tagging (EXTPAIR/
SELECT/NOP) on target-side indels; detectors imported from atlas_features
(R33); cookbook text embedded in cards
- tools/lenmiss_route.py: pool-parallel (A8) — 587 audit LEN rows re-verified
live + routed in 24s: redraft 345 (frac>0.35, APPEND-ONLY backlog
reclassification — near-miss metrics stop lying) / permuter-length 49 (grinder
fuel) / cards 192 incl 14 tell-tagged (the audit's own detectors had emitted
ZERO) / mechanical 0 — an HONEST NULL: stored drafts rarely get constants
wrong; LEN drift is shape, family_align's value here is classifier/detector
- R32 accounting 587/587
- tools/family_align.py (NEW module — classify_member's return contract untouched,
the remap_hseq silent-pass trap avoided by design): SequenceMatcher alignment
over FC.tok streams; li-cluster reconstructor (lui/lui+addiu/lui+ori/li-from-$0
chains, split-cluster absorb for the rs-changed addiu partner); verdicts
LEN-LI/LEN-NOP/LEN-JTBL/LEN-STRUCT/STRUCT-ALIGNED/PURE/IMM; aligned imm engine
mirroring imm_map_tier1 (ordinal deliberately out in v1)
- NC-1 verdict-equivalence 157/157 banked pairs — the NC caught two real gaps:
R-type non-shift sa diffs are STRUCT; registers tested BEFORE the reloc skip
(a reloc-slot word with a different register is STRUCT). NC-2 parity 21/21
- R37 PROBE REFUTED the planned mechanical driver before it was built: 0/26
LI-ONLY cards classify mechanically (regfields x19) — cousins are 0.85-similar
DIFFERENT functions; §168 law 1 re-derived by measurement; no driver written
- family_align re-scoped: its consumer is T8's LEN+N near-miss pile (draft vs
its OWN target = same function); reloc-vs-constant range discriminator parked
for T8. decision-log entry (R31)
- tools/plumbing_groups.py: derives the honest still-open pool from the classified
ledgers (R38) — '1,217 PLUMBING' collapsed to 237 (SELF 109 / CALLEE 48 / OTHER
48 / DATA 32)
- recover_integration: PER-GROUP ISOLATION (git-checkout binary TUs between groups
— one TU-stage edit was poisoning every other group's whole-binary gate with a
phantom shared error; per-group banked_from_source capture) + new stages
'macro-externs' (§121 draft-tier, via family_sweep.macro_def_sig_map, R33) and
'tu-scope' (§103 STU binary-tier, the sweep-only lever)
- the probe (ov_SC03_107): raw 0/14 -> root-caused (poisoning + stale seed
symbols; rtu_match MATCHes them — blind to reloc names, R34) -> symfix-first
-> 9/14 BANKED (64%)
- sweep finding (Law 3): the no-draft majority (ov_SC02_037 44/44, most of
ov_MAIN_012) had verdicts from transient sweep remaps never persisted — family-
lane fuel, not recovery fuel; the stored-draft class is consumed
- cookbook §173 (symfix-first / per-group isolation / verdicts-without-drafts);
index 518 green; R22 clean fleet 213/213; phase total 17 banked @ 0 agent tokens
- decision-log: the P31 re-charter entry (organize-before-grind; R37/R38/R39
ratified at gate-1) per R31
- harvest_verify.py: import guard — a bare import now RAISES loud instead of
running a full gate (CLI unchanged, verified both directions)
- sig-resident: bootstrap boundary artifacts fixed (fused +0 data word with
func_800CEDFC; func_800D33E0 dropped past a glued tail) -> ELF-seeded per the
S45 pattern, exactly 145 fns; true denominator confirmed 145 (progress was
right); audit-corpus 0 PHANTOM + 0 TRUNCATED; all three oracles agree
- family maps regenerated at HEAD commit:2161: 11,025 open non-main members
reconciles EXACTLY with 12,059 - main 1,034 (102 stale phantoms cleared);
adapt cards 704, aprop cards 204 (full emission)
- main fuel-gap finding: 2,001/2,002 main stubs already have cached Ghidra-C
(only func_80049600 missing) — the roadmap '0/2,096' note was stale
- tools-health OK (dedup 2,063/0; C1 254,521/254,521; audit-digest green)
- PhaseEnd_Phase30.md written; CURRENT_PHASE.md archived to phase-ends/logs/Phase30.md (R19)
- 25 sessions (S26-S50), 947 commits, stubs 28,296 -> 12,059, dedup 2,061/0
- milestone met on both clauses: >=95% instr AND every remaining stub on a named ledger
- v1.28.0 -> v1.29.0
The §172a/§172b tells + repaired instruments swept over all 892 open near-misses:
- 33/95 stored drafts re-verified MATCH and banked through the whole-binary gate
(aprop_symfix caught 40/108 carrying stale seed symbols before gating — §171 at scale)
- 19/20 hand/mech fixes banked: four pure lhu<->lh s16 flips; the lhu+sltiu->lh+slti
shared-global quadruplet (D_80126B5E/B66/CB0, D_80126CB0 are s16 FLEET-WIDE); one xor-eq
rewrite; 11 per-location literal swaps (mask/threshold constants from sibling binaries)
- 1 refusal (func_8017EE78) stays as redraft fuel
Stubs 12,111 -> 12,059. Fleet 95.3% instr / 90.0% distinct / 96.68% fn-count.
Veins mapped for next waves: ~400 LEN+N drafts, 13 ambiguous-symbol, 7 multi-literal.
Audit ledger: .run/c294/audit_results.json (classifier derives from match_one's own sig).
Final S50 state: 307 instances banked, stubs 12,468 -> 12,161, fleet 95.3% instr / 90.0%
distinct / 96.65% fn-count. R22 clean rebuild 4x, check-all 213/213 every time.
- tools/aprop_autodraft.py + tools/draft_prechecks.py: seed body + symbol_map + a MINIMAL
synthesized preamble. The seed's decl layer never travels — that layer is family_sweep's
dominant failure (331 of 458 S49 verdicts). 256 banked at zero agent tokens, against the
~20M the same work would have cost as a wave.
- Macro seeds (567 of 1196 members, all 3737 de-macroize) take the DEFINITION only; the block
stays the decl source. Pasting it whole measured 28% vs inline's 68% — func_8016AB6C's macro
is 1,891 lines of which 108 are the function.
- IMM is a second engine, not a wall: T2a's imm_map_tier1 resolves a per-location LITERAL like
symbol_map resolves a per-location SYMBOL. 131 of 275 IMM members resolve.
- draft_prechecks negative-controlled against ALL 205 banked drafts: zero false positives,
catches 39 of 67 known failures. That control found two bugs in the checks themselves —
C89 `f()` declares UNSPECIFIED parameters (not zero), and a member's own definition read as
a call to itself. Conservative by design: a pre-check that discards good drafts is worse
than one that lets a few builds fail.
- The A-prop pool is now priced exactly: PURE 437/37,376 ins, IMM 275/8,849, STRUCT 238/4,259.
- Cookbook §171a; SETUP rows; CURRENT_PHASE S50 FINAL checkpoint.
- The blocker was carried as "one missing file-scope extern gates 83 PURE members". Both
halves were wrong (R14): corpus.stubs says 4 open members, and D_801ED98C is a DEFINED
const Blk8 whose rodata lives inside the member's own nonmatchings .s — replacing the stub
deletes the data with it. gather_externs can carry an extern DECL, never a DEFINITION,
which is why it reported "no file-scope decl" for a symbol md_SC05_023 defines on line 114.
- Fix: paste typedef + const definition + body per sibling (data bytes verified identical
across md_SC05_024/025/028/029). 4/4 banked.
- aprop_symfix: new `local-only` class — draft-DEFINED identifiers that merely carry a
vram-looking suffix (Blk8_…, S8_…, L_call_…) are not stale symbols. Measured: that is every
non-clean case in the whole wave-7a/7b stored-draft residue, which holds ZERO stale-symbol
recoveries (a clean negative result — the defect was A-prop-specific).
- cookbook index regenerated (tools-health fails closed on a stale index — it caught §171).
- R22 clean rebuild: check-all 213 passed, 0 failed of 213. Stubs 12,445 -> 12,441.
- REFUTES §170's open hypothesis (batched cards concentrate members into one TU ⇒ §169
collision): 5-draft groups banked 5/5; 11 of 35 unbanked drafts were already one-per-TU;
and the two "concentrated" groups banked 12/12 and 10/10 once the real defect was fixed.
- The cause: a per-location data symbol carried out of the seed body unrebased. match_one
compares instruction ENCODINGS and is blind to a relocation's target NAME, so it scores
MATCH standalone and dies at link in the host TU. 24 of 24 concentrated failures, all 1:1
rewritable at one constant vram delta (0x4128).
- tools/aprop_symfix.py: audit + --fix, emits a gate_lane-shaped slate; deterministic and
build-free, so it runs BEFORE the gate. The R34 second oracle for the class match_one
cannot see.
- family_cousins.py --aprop-cards: members now carry sym_map, the explicit {seed -> member}
renames, read from the seed's C BODY (a matched seed has no .s of its own) vs the member's
.s. Two case-mismatch defects fixed while wiring it (sig lowercase vs splat uppercase).
- 23/24 banked. Stubs 12,468 -> 12,445. Fleet 95.2% instr / 89.9% distinct / 96.57% fn.
R22 clean rebuild: check-all 213 passed, 0 failed of 213. dedup 2,043/0.
- A-prop's true conversion is 87% (79/91); the 320 batched members are unblocked.
- Cookbook §171 + §170 struck in place; SETUP row; decision-log (R31).
- checkpoint block refreshed for a fresh session (R30/checkpoint-before-pause): the TU-spread
test is the named FIRST action, the regen chain and gate contract are spelled out, and the
>=16 head's three open items are listed with their evidence.
- promote gate_lane.py into tools/ (it was scratchpad-only): accepts slate OR confirmed shapes,
explicit draft paths, R32 coverage assertion (refuses to report 0 as a result), dirty-tree
abort, no outer timeout, per-function propagation after.
- session: 179 commits, 177 instances banked, 5 R22 clean-fleet gates all 213/213.
- NEW family_cousins.py --aprop-cards + tools/wave/aprop_wave.js: lane A (1,700 open fns /
76,419 ins) had NO card type — cousin diffs are empty for h_seq-identical members, so the card
is a positional WORD diff vs the matched sibling, grouped BY FAMILY (one agent, N drafts).
Head cards: 13 families / 433 members, median TWO differing words each.
- calibration 9 batches / 108 members: 98 agent-MATCH (91%, best of any wave) -> 56 BANKED (57%),
~80k tok/banked fn vs 157k (cousin card) vs 400k+ (crack wave). R22 213/213 BYTE-IDENTICAL.
- HONEST GAP (R14): 91% agent -> 57% gate is the worst conversion measured; 14 groups banked 0.
Hypothesis TESTABLE not proven — family batching concentrates members per destination TU, the
§169 collision. Re-gate unbanked ONE PER TU before scaling the remaining 320.
- >=16 head diagnosed: 3 of 4 blockers are plumbing — the --band substantial default hid 5 of 13
families from every prior sweep; one missing file-scope extern (D_801ED98C) gates 56 PURE
members; dedup_extend is macro-only. Only func_8017C294 is a genuine crack.
- fleet 96.56% fn / 95.2% instr / 89.9% distinct; stubs 12,535 -> 12,468; dedup 2,043/0.
- cookbook §170.
- thresholds relaxed to <=6 blocks/<=16 tokens UNION edit-fraction <=0.20: cards 518 -> 721,
MIXED 310 -> 50 skeletons; the 753-ins func_8017BEBC (0.987 sim) became reachable.
- 59 cards -> 48 agent-MATCH (81%) -> 44 BANKED (92% MATCH->bank, 75% end-to-end), 6.9M tok.
- FINDING (the actionable one): 7b's bank rate crushed 7a's because it SPREAD 48 drafts over 35
destination TUs; 7a's failures were per-TU declaration collisions between sibling drafts.
Cookbook §169 updated with the spread law.
- R22 213/213 BYTE-IDENTICAL from clean; fleet 96.55% fn / 95.2% instr / 89.9% distinct;
stubs 12,584 -> 12,535; dedup 2,035/0.
- incidents 3 & 4 recorded: an agent wrote a TRACKED header (guard caught it, prose is not
enforcement); my own gate_lane filtered on the wrong key and printed 'gating 0 drafts' as a
result (R32 silent skip) — fixed with a coverage assertion that refuses to report 0.
- pilot 30 cards -> 25 agent-MATCH (0 refuted) -> 16 banked; 29 instances banked tonight
(89 incl. propagation); 2.7M tokens haiku-tier ~= 30k/banked instance vs a crack wave's ~75k.
- R22 213/213 BYTE-IDENTICAL from clean; fleet 96.53% fn / 95.1% instr / 89.8% distinct;
stubs 12,613 -> 12,584; dedup 2,029/0.
- R14 CORRECTION: a banked cousin usually does NOT propagate (2 of 8; cousins are byte-variant).
The card 'reach' column is cousin fuel, not dedup copies — priced wrong in my earlier framing.
- FINDING: the 9 gate failures are per-TU INTEGRATION (standalone-MATCH, host-TU-rejected),
clustered 5+2 in two binaries — the reconcile-ladder class, not codegen.
- TWO INCIDENTS (mine): an outer timeout tighter than gate_stage's own scaled timeout killed a
healthy 5-bank group mid-write AND orphaned its dedup_propagate child, which kept rewriting
src/ through a git checkout. Killed, inspected, reverted; the same 5 drafts banked 5/5 untimed.
Law: never wrap a self-timing tool in a tighter cap; kill process GROUPS, not pids.
- cookbook §169 (the lane + the three laws + the threshold sizing table).
- FINDING (Drew's smell, byte-verified): the '4,513 unique singletons' picture is substantially
an h_seq exact-hash artifact — 86/120 near-pairs in the 0.85-0.99 band differ by PURE
insertion/deletion (li-expansion tell in 25). Specimen: ov_SC06_010:0x8017bebc (753 ins,
'singleton') is 0.987-similar to a MATCHED fn in the same binary.
- NEW tools/family_cousins.py: distinct open skeletons -> shingle index -> >=0.85 union-find ->
matched-seed attachment -> .run/family_cousins.json + docs/family-cousins.md. R32 BOTH ways
(independent stub recount fails loud on a stale map — negative-control-proven; partition
assert). Reproduced the probe within +-1%; totals EXACT (11,627 inst / 584,448 ins).
- Unit table: A-prop 197u/68,729ins · seeded 418u/50,422 · cousin-multi 1,552u/249,799 ·
cold 3,240u/215,498 — the genuinely-unique tail is 37% of the remainder, not 90%.
Main's 'structurally barren' HOLDS at the similarity tier (94% mass <0.70).
- --targets wave slate: .run/wave7_targets.json = 40 targets / 33,304 unit ins (+33% vs
family-ranked), 9 resolved seed C paths, size-routed 2 haiku/20 sonnet/18 opus.
- LAWS (§168): a cousin is a SEEDED CRACK never a remap; rank waves by UNIT weight; discount
short-fn similarity. Byte-gate stays the sole arbiter (G3/P9).
- docs/family-hseq.md: this session's frontier regen (post-S48 propagations) rides along.
- cookbook §168 + SETUP inventory row (R16/R21/R30); CURRENT_PHASE S49 entry.
Fresh-session safe. 197 commits; six waves (bank rates 67/79/69/68/73/60%);
cookbook 508 sections with all six idiom harvests banked (§162-§167).
Resume order: T5/phase close is a LIVE option (the >=95% instr milestone is met
and holding); otherwise wave 7 from tools/wave/crack_wave.js after a regen.
§167's saturation signal is the strategic finding: COVERED+UNSOUND went
57% -> 64% -> 76% across three harvests, five new laws from 197 claims. Stop
mining waves for idioms; spend the tokens on cracks.
Carries the session's central lesson: four times a tool asserted a conclusion it
never reached, and I repeated the pattern once myself. A confident wrong label
costs more than a missing one.
Wave 6 added 110 (24 cracks + 23/24 families propagated). Fleet 95.1% instr /
89.7% distinct / 96.51% fn-count. R22 clean-fleet run 9x this session, 213/213
every time. Bank rate across six waves: 67/79/69/68/73/60%.
Records §166a (the destination-TU oracle) and the four-instance pattern it
completes: a tool asserting a conclusion it never reached. A confident wrong
label costs more than a missing one.
Wave 5 added 152 (29 cracks + 28/29 families propagated) — the session's
largest. Fleet 95.0% instr / 89.6% distinct / 96.48% fn-count; stubs
13,345 -> 12,771. R22 clean-fleet run 8x this session, 213/213 every time.
P30's milestone is '>=95% instr fleet, or every remaining overlay stub on a
named ledger'. THE FIRST HALF IS NOW MET — T5 (phase close) is a live option.
Bank rate across five waves: 67% -> 79% -> 69% -> 68% -> 73%.
Wave 4 added 105 (27 cracks + 27/27 families propagated). Fleet 94.9% instr /
89.4% distinct / 96.44% fn-count. R22 clean-fleet run 7x this session, 213/213
every time.
Bank rate now measured four times: 67% -> 79% -> 69% -> 68%. Prior-notes
seeding 10/12 (was 7/9). func_8017C294 — the x16 family, largest item on the
board — is NEAR at 2 ins after three seeded attempts (18 -> 11 -> 2).
Also records the 4th comment-blindness defect and its blast radius (one draft
comment refused a binary's stub oracle, failed 5 later binaries, and left
drafts spliced in src/ so 17 re-gates read a poisoned tree as 0/17), and that
the wave harness now lives in tools/wave/ with its contracts written down.
Wave 3 added 84 (27 cracks + 21 propagated families). Fleet 94.8% instr /
89.2% distinct / 96.41% fn-count. R22 clean-fleet run 6x this session, 213/213
every time. 96 commits.
Bank rate measured three times: 67% -> 79% -> 69%. The dip is the cost curve
(wave 3's tier was 29 Opus-band / 14 jr vs wave 2's 8 / 5, median reach x6 ->
x3-4), not a regression.
Two levers proved out and belong in every future wave: the hardened harness
contract (0 drafts lost vs 21) and prior-notes seeding (7 of 9 previously
failed targets converted, incl. both long-standing NEARs and all three wave-2
gate misses). NEAR is a resumable state, not a write-off.
Session close state. Three parts: stage 0b (91, zero decompilation), wave 1
(26), wave 2 (116). Fleet 94.4% -> 94.7% instr, 88.3% -> 88.9% distinct,
13,345 -> 13,112 stubs. R22 clean-fleet run 5x, 213/213 every time.
The campaign now has a MEASURED rate, twice: 67% (wave 1, all-Opus) then 79%
(wave 2, 20 of 28 Sonnet) of cracks survive the whole-binary gate. The Sonnet
band beating the all-Opus wave is the session's most useful economic finding
and sets wave 3's routing.
Resume order changed on evidence, twice over:
- harden the wave harness FIRST (per-agent dirs, sha1-last verifier, and a
tools/recover_drafts.py built from the transcript-replay method that
recovered 21/21 today);
- then wave 3, sized on 79%, not on the reach-15 prior.
Error ledger grew to 6. The two that matter: I wrote off 21 verified cracks as
lost when the run transcripts held every one of them, and my first two
recovery passes both failed by reading a single tool record instead of
replaying the file's mutation history.
Stage 0b closed (91, zero decompilation) + Stage-1 wave 1 (8 cracks -> 26
instances). Fleet 94.6% instr / 88.8% distinct / 96.36% fn-count.
Resume order changed on measured evidence: FIX THE md_ MODULE LANE FIRST.
16 of the wave's 42 member slots were unreachable for tooling reasons, not
matching reasons — 12 on a carve that assumes raw data lives in
<binary>/data/*.data.s (modules do not), 4 on an uncarried extern
(`D_8011511A' undeclared). Both are named with verbatim errors; probe one of
each before pricing (R37). Precedent: 0b's three repairs banked 91 for ~0
agent tokens; the wave spent 3.36M for 26.
Also recorded: the frontier re-derivation (1,955 zero-crack families /
330,622 templatable ins), the tier-ordering correction (ins-per-crack is flat
across x5-x8, so rank by templatable weight, not by tier), and the §162
harvest with its two in-place cookbook corrections.
2,197 instances banked, derived from the STUB ORACLE (15,542 -> 13,345), not from summing
per-batch reports (my running total said ~2,307 — summing drifts, the oracle does not).
35 commits, R22 213/213 after every batch, fleet 94.4% instr / 88.3% distinct / 96.33% fn-count.
THE STRATEGIC FINDING: the sweep residue started the session ~5:1 PLUMBING:DIFF and ends at DIFF
170 of 575 (30%), larger than the next four classes combined. The declaration-axis vein is spent;
from here the mover is volume with multipliers, not more plumbing. The Fable frontier analysis
predicted this and the residue confirmed it rather than my framing.
Landed: ~1,900 functions from plumbing at ~0 agent tokens (8 declaration axes, the symbol-KIND fix
at 205, cdFileLocTable, memcpy, the cpp-derived type map, the alias-drop fix) plus 163 from the
reach-15 agent wave at a 15x effective multiplier (reach-ordering, not the 3.55 mean, produced
that). Seven instruments repaired, cookbook to 469 sections.
Error ledger: 8, all caught. The pattern in most of them is that I proposed a mechanism before
reading evidence that was already written down — four were in files I had open.
Resume at Stage 0b (JTBL_PADS, 122 slots at ~100% measured conversion), then Stage 1's
reach-ordered sibling campaign. Stage 0a is done enough; what remains of it is bounded, not
compounding. tools-health was launched at close — confirm it before banking.
Independent frontier analysis re-derived every headline number from family_hseq.json (all reproduce
exactly) and corrected four claims, one of them mine from this session.
THE SEQUENCE: Stage 0 tooling (0a the DATA SIDE of the template engine — the 207 undefined-refs,
54/56 parse errors, 116 data-bundled .s and the F2 collisions are ONE mechanism, and unlike the
JTBL_PADS fix it COMPOUNDS across ~7,500 future member banks; 0b JTBL_PADS; 0c sig-main, recommended
yes, needs Drew; 0d a free-CPU permuter probe with a kill rule). Then Stage 1, the reach-ordered
sibling campaign to ~97.5%. Then Stage 2 singletons. Special projects LAST.
CORRECTIONS: (1) my G2 main finding was inflated — 968 of sig.main's 2,002 rows are SDK-region
stubs that fall to LINKED conversion, so the real main templating pool is ~123 families / ~303 fns
/ ~9k ins, not 207/748/11,537. (2) docs/family-hseq.md stamps its own scope as '212 OVERLAYS only
(no main, no resident)' while the map contains resident and 70 md_* modules — the §159 coverage law
violated by the file that documents coverage. (3) the close-1-4 backlog is 103, not 157, and its
reach column is TOTAL sharers not LIVE. (4) 'zero-crack' means opposite things in the roadmap and
the current map — a 30x mis-scope risk.
FRAMING: my cost-per-crack thesis holds for the singleton half only. True singleton pool is
~5,200-5,600 cracks / ~345k ins (45%); the other 55% rides on ~2,100 exemplar cracks where ORDER and
LEAK-RATE decide the calendar. Do not cross-price the two economies: 5:1 plumbing:DIFF is a property
of the residue queue, while W1's fresh wave converted 81%.
Also recorded: the 15-step reach-ordered sibling loop, with the three steps whose omission destroys
the multiplier called out (regen before targeting, regen after cracking, --band all).
W1b — the 3 targets whose agents died on API rate limiting, retried with cookbook §160 in the
prompt: func_801EFBF4 (reach 12), func_801EFDC8 (12), func_8018CC40 (10, jr). 3/3 confirmed by an
independent verifier, all banked, R22 clean-fleet 213 passed / 0 failed of 213.
func_8018CC40 failed the first gate with `too many arguments to function func_80178970` — which its
own crack agent had PREDICTED in its report, naming the §17a-1 remedy. Dropped the draft's
empty-paren externs and cast 6 call sites instead; banked. Read the agent's integration notes
before diagnosing a gate failure — it has already seen the TU.
Cookbook §161a-c (index 469 sections):
§161a case 0: break; is LOAD-BEARING when a jump table is indexed from zero. The natural
case 1..5 makes gcc-2.7.2 pick minval=1, emit `addiu $v1,-1`, and shift every table index —
58 of 77 mismatched on a byte-perfect body. Tell: the table's FIRST entry points at the
function's own end address. Family-wide (10 members).
§161b aliasing a parameter into a local can force a SECOND callee-saved register (+8 frame,
+3 ins) even when uses are mutually exclusive. Suspect it before reaching for register pins.
§161c loose-prototype engine helpers: don't fight the TU's (void) decl, cast at the call site.
G2 — THE MAIN EXPERIMENT. family_hseq excludes main as "structurally barren — zero h_exact
overlap". True and irrelevant: an h_exact claim guarding an h_seq tool. There is not even a
sig-main target — main had never been signed for this pipeline. Signed it (2,002 fns, seeded from
splat boundaries via corpus.stubs rather than --bootstrap, which glues functions around jtbl
dispatch and would have corrupted the hashes under test).
Result: main is ~85% singleton work, not 100%.
internal h_seq families (>=2): 207 families / 748 fns / 11,537 ins (13.7%)
shapes shared with the fleet: 161 fns / 1,346 ins (1.6%)
genuine x1 remainder: ~71,034 ins (84.6%)
IMMEDIATELY ACTIONABLE: 44 classes / 151 main functions / 1,239 ins already have a matched exemplar
in the fleet — free propagation, invisible only because main is not in the map.
Long-term: 748 of main's 2,002 functions (37%) are templatable once one exemplar per family is
cracked, which refutes "2,002 independent cracks" as the planning assumption for the 79k-ins tail.
OPEN, deliberately not done unilaterally: adding a sig-main target and dropping main's exclusion
from family_hseq.load() changes a fleet-shared oracle every targeting tool reads. Needs Drew's call.
The reach-57 exemplar func_801EDC18 is banked and R22-green, but its family sweep returned 0/56.
54 of 56 failed 'parse error before buffer': the remapped sibling carries neither the draft's own
Blk8 typedef nor any declaration of the per-member data symbol.
Two gaps. (1) family_remap does not gather draft typedefs — the T7-S1 class, named a phase ago and
still unbuilt; cdecl.strip_provided_typedefs is NOT the culprit, it correctly keeps a typedef the
target lacks, so the loss is in unit extraction. (2) NEW: this family's data is PER-MEMBER — each
sibling's .s carries its own rodata bytes, so a symbol remap cannot produce it; the bytes must be
decoded per member and emitted as that member's definition. No existing tool does this.
Bounded: 116 of 12,583 .s files are data-bundled. The md_* module TUs cannot fall back on the
shared Blk8 (engine_types.h:497) because they include only common.h — tu_scope is 53 entries there
versus 4,190 for an overlay.
The 44 data-symbol conflicting-types failures are NOT the cdFileLocTable duplicate-typedef class.
They are genuine per-view type differences: D_80078EB4 is s16 at 2,409 sites and u16 at 1,341;
D_800AE620 is Blk20/s32/Mat32. For data the declared type drives the load (lh vs lhu), so
canonicalising would rewrite thousands of already-banked sites' codegen. The answer is one type
PER VIEW — the §37 asm-label alias the fleet already hand-writes for D_800AE620 (9 sites).
Shipped: scope_data_externs now auto-aliases a conflicting extern —
extern s16 aD80078EB4 __asm__("D_80078EB4");
keeping the draft's own type (byte-truth for that body) while the private C name makes collision
impossible and the asm label pins the emitted symbol, so codegen is unchanged. Fires only where
the TU declares that symbol with a DIFFERENT type text; same-type and already-aliased drafts are
untouched (4 controls, 2 of them negative). Applied on both scope paths — a staged draft's externs
arrive indented, so a demote-path-only fix reached 1 of 44.
R22 clean-fleet: check-all 213 passed / 0 failed of 213.
WHY ONLY 2 BANKED — the collision is DRAFT-vs-DRAFT, not draft-vs-TU. The failing draft declares
`extern s32 D_80114F24;` and ov_MAIN_012.c declares that symbol nowhere; the error lands at the
splice point. family_sweep stages every member of an (overlay, split) group into one TU before
gating, so two templated bodies with different views of one symbol collide with each other.
scope_data_fix sees one draft plus the pre-splice TU and structurally cannot see the others.
The real fix belongs in the staging loop, which knows the whole group: alias any data symbol
declared with >=2 distinct types across the drafts staged together. Not attempted here.
Three wrong inferences on this one task before reading a failing draft: scoped as the typedef
class; aliased only the demote path against the file's own comment; targeted the wrong collision.
harvest_verify.classify_fail kept only stderr lines containing the word `error`. gcc-2.7.2 emits
no `error:` prefix on hard errors, so lines like
src/…/ov_SC02_037_jr_8013B83C.c:447: multiple storage classes in declaration of `tail_…'
src/…/ov_SC06_025_jr_8012ACE0.c:2217: `tbl_D_80187044' undeclared (first use this function)
never survived the filter, `errs` held nothing but make's `Error 33` wrapper, and every hard error
was labelled CC1-FAIL(no-diagnostic) — "the compiler failed and we cannot see why". Measured cost
this session: 132 siblings of func_80132018 classified that way by one missing declaration, which
reads as a codegen wall and gets a family deprioritised. rtu_match had the same blindness repaired
at T0(b); the fix was never propagated here.
Fix: a diagnostic is a POSITION, not a vocabulary — `<file>:<line>: <text>`, plus the assembler's
`{standard input}:<line>:` (_SRC_DIAG). Context lines carry no `:<line>:` and are skipped.
Five controls pass, including the two that guard against over-fixing: PLUMBING still wins on a
declaration conflict, and a warnings-only failure still returns no-diagnostic.
Re-swept: 0 no-diagnostic remain. The 93 resolve to 35 redefinition-note, 11 D_801202A0
undeclared, 10 too-many-arguments, 4 too-few-arguments, 4 func_8001534C undeclared — every one a
cheap declaration/arity class, not a wall.
Residue now fully named (737): 207 undefined-reference across 42 symbols (a link-stage REMAP gap,
now the largest class), 138 DIFF (real divergence, 19% — the honest floor), 44 data-symbol
conflicting-types, 35 redefinition-note, 26 memcpy, 24 redeclared, 15 undeclared, 14 arity.
The one-line fix committed ahead of this run (CdFileLoc_80128C98 aliasing CdFileLoc) cleared the
largest remaining propagation-sweep class. Re-sweep: 138 member-matches banked, failures 875 -> 737,
`conflicting types for cdFileLocTable` gone entirely (136 -> 0).
Derived net (138 INCLUDE_ASM removed, 0 re-added) equals the report's 138 — they agree.
R22 clean-fleet: check-all 213 passed / 0 failed of 213.
Fleet 94.2 -> 94.3% instr / 87.9 -> 88.1% distinct / 96.15 -> 96.21% fn-count; stubs 13,780.
Residue reclassified — no symbol dominates any more: 227 PLUMBING-other, 125 DIFF (real byte
divergence, 17%), 93 CC1-FAIL(no-diagnostic), 26 memcpy, then a tail of small data-symbol
conflicts (D_80114F24 12, D_800AE620 11, D_800183E0 9, D_80126B58 6, D_80078EB4 6).
CC1-FAIL rose 77 -> 93 and that is NOT a regression: members that previously died earlier on the
cdFileLocTable conflict now reach a different compile error. Those 93 are hard gcc errors whose
text the sweep's classifier discards because it greps for `error:`, which gcc-2.7.2 never emits on
hard errors. That classifier is now the highest-value instrument fix left — three times today a
no-diagnostic verdict concealed something cheap.
Found while scoping the jr carve for the 3 newly-onboarded binaries. My scoping said "one
unplaceable construct" — it was the first of four layers. Three are fixed here; the fourth is out
of this tool's scope and leaves the carve blocked.
1. asm_label_aliases: the scan could START inside a #define. `#define gte_SetRotMatrix(r0)
__asm__ volatile ("lw $12, 0( %0 );" ...)` is textually `ident(...) __asm__(...)`, and
cdecl._mask blanks string CONTENT — deleting the `;`s that would stop the greedy [^;{}]*.
The match ran 116 lines and swallowed the real `aF8012EFB8 ... __asm__("func_8012EFB8");`,
so the alias never entered the map and addr_of returned None. jr_isolate_all then refused to
carve (R32, correctly), which presented as 112 isolate-fails that looked like a per-binary wall.
Fixed: _mask_cpp_directives() — a preprocessor directive is the other place a match must not
start. Masking comments/strings fixed the comment case and left this one.
2. _split_macro_body returned a `static inline` internal HELPER as the macro's definition, so
_proto_from_lines hoisted `extern static inline void tail_8012F274(...);` into all 41 regions:
invalid C (multiple storage classes) AND the wrong function — the exported definition sits
below the helper and lost its implied declaration. Fixed: skip static definitions
brace-balanced on the masked body. A static helper needs no hoisted declaration at all.
3. A declaration that WRAPS across continuation lines was taken as one line, so half became a
`;`-less extern and the continuation was read as the definition header, producing
`extern __asm__(""); void aF801466F0(...);` in 22 regions. Fixed: accumulate until the
statement terminates, tested on the masked text. Same wrapped-declaration blindness
family_remap._alias_decl_for records fixing at S33 — never propagated here (§134/§139).
Not fixed, and why: jtbl_rodata_pads reports "consumed 0 rodata .align(s) but 4 pad spec(s) given
— table-count drift vs the carve". The carve moves jtbl-owning functions into _jr_ regions but
leaves the pad specs on the residual gap object. jr_isolate_all's docstring states this class is
NOT isolate-fixable; it needs JTBL_PADS repointing in overlays.mk. 122 jr member-slots stay blocked.
Regression-checked: 1,948 macros parse with 0 malformed externs; alias maps unchanged on three
already-carved overlays. No build impact (splitters run offline). Carve reverted, tree clean.
scope_data_externs §8d drops the draft's decl of any symbol the TU already declares at file scope.
It keys on the SYMBOL, but a §37 asm-label ALIAS binds a DIFFERENT C identifier to that symbol:
the TU declares `D_801851BC`, it does NOT declare `tbl_D_80187044`. Dropping the alias left the
body referencing an undeclared name, which cc1 reports with no `error:` prefix — so the sweep
classified all 132 siblings as CC1-FAIL(no-diagnostic), i.e. as a codegen wall.
The bitter part: the alias exists PRECISELY BECAUSE the TU declares that symbol with a conflicting
type (a `void (*[])(void)` dispatch table vs this function's 20-byte-stride view). The drop rule
fired on exactly the declarations written to survive it. Why 1 of 2 died was fully determined:
tbl_D_80187048's symbol is not in the TU, so it demoted normally.
Fix: is_asm_alias() — an alias is demoted into the body, never dropped (the identifiers differ, so
it cannot collide with the TU's decl). Control-tested 6 ways incl. self-labels and plain externs.
Measured: func_80132018 3/135 -> 135/135; full re-sweep +16 more. Total +148 members.
R22 clean-fleet 213 passed / 0 failed of 213. tools-health OK, dedup-check 1949/0.
Fleet 96.11 -> 96.15% fn-count, 87.8 -> 87.9% distinct; stubs 14,120 -> 13,972 = -148 (2nd oracle).
CORRECTION TO MY OWN CLAIM (R14): after the probe I said the 58% aggregate was concealing a broad
problem. The re-sweep refuted it — only 16 more banks fleet-wide. The alias class really was one
family; the first read ("outlier") was right and the correction was wrong.
875 sweep failures classified: 231 PLUMBING-other, 141 DIFF (real divergence, only 16%),
136 `conflicting types for cdFileLocTable` (ONE symbol — biggest single class left),
77 CC1-FAIL(no-diagnostic), 26 memcpy, 12 D_80114F24, 11 D_800AE620, 9 D_800183E0.
STILL UNFIXED, and the most dangerous instrument left: the sweep's failure classifier greps for
`error:`, which gcc-2.7.2 never emits on hard errors. Every hard error therefore reads
CC1-FAIL(no-diagnostic). That is how a missing declaration looked like a codegen wall across 132
functions. rtu_match was fixed for this at T0(b); this classifier was not.
family_sweep --hseq --band all -j 8 over every matched-exemplar family: 553 families /
203 overlays / 1,419 banked / 1,023 failed (58%). R22 clean-fleet 213 passed / 0 failed of 213.
tools-health OK, dedup-check 1949 validated / 0 failed.
Fleet: 93.9 -> 94.2% instr / 87.2 -> 87.8% distinct / 95.72 -> 96.11% fn-count.
Second oracle (R34): INCLUDE_ASM stubs 15,542 -> 14,120 = -1,422, equal to the diff-derived net
(1,451 removed - 29 re-added = 1,422 = 1,419 sweep + 3 probe). Three independent counts agree.
B -> C -> P IS ONE CHAIN, NOT THREE WINS. 1,102 of the 1,422 landed in ov_SC02_037 (409),
ov_SC03_107 (364), ov_MAIN_012 (329) — the three newly-onboarded binaries from C, which had never
been wired into the shared-body ecosystem, so every matched exemplar was unreachable from them.
B fixed the declarations, C wired the include, P poured through the opening. A repeat sweep will
NOT pay like this; the opening was one-time.
S47 total: 1,481 functions banked with zero agent drafting, all from removing plumbing.
Two findings recorded, neither fixed (deliberate, costed):
- --band defaults to `substantial`: the first probe returned a confident {"families": 0,
"banked": 0} on a real 135-member `mid` family. Always pass --band all.
- The alias-gather defect: probe on 0x80132018 banked 3/135, all 132 failures classified
CC1-FAIL(no-diagnostic) because gcc-2.7.2 emits no `error:` prefix. Real error is
`tbl_D_80187044' undeclared` — the exemplar declares TWO §37 asm-label aliases and uses both,
family_remap carried one. T7-S1's "gather" class. Measured as an OUTLIER (aggregate 58%),
which is why the sweep ran before the fix.
Refused by design, all named: 50 jr families / 183 member-slots (§53 interlock — it printed its
own coverage and reason), 264 STRUCT, 112 unresolved immediates, 3 not-stub.
C, unblocked by B's declaration conform. 23/23/16 banked across ov_SC03_107, ov_MAIN_012,
ov_SC02_037 — the first non-zero result on this population (S46 got 0/142, then 0/129).
R22 clean-fleet: check-all 213 passed / 0 failed of 213. tools-health OK.
dedup-check 1949 validated / 0 failed, C1 coverage 249295 (= 249233 + 62, independent
confirmation of the count). Fleet 93.9% instr / 87.2% distinct / 95.72% fn-count.
THE TOOL REPORTED "BANKED 0 / 129" AND WAS WRONG. gate_stage's ladder hands the same
--verified-out path to harvest_verify on every rung, and each rung opens it for write: stage 0
banked 23 and wrote them, then a later rung that banked nothing truncated the file to 1 byte.
The in-memory list uses += and stayed correct, which is why the JSON verdict listed all 23 names
while the file said nothing. dedup_extend read the file, printed BANKED 0, and took its
`if not banked:` branch — skipping add_members_surgical, so the registry was missing 62
memberships for functions already spliced in and byte-verified.
- Registry repaired by deriving the banked set from git diff (+DEFINE_func_*), not from the
broken file. Post-check: 0 missing.
- ensure_include_revert did NOT fire (added_include False, include already present) — the
P29-S19 defect that once stripped a load-bearing include from 135 binaries stayed closed.
- gate_stage now writes verified_out once at the end from the accumulated truth.
Caught only because bank truth is derived from source (§55b), never from the gate report.
Residue (67) is consistent with the symbols B deliberately left: memcpy 17, ApplyMatrixSV 12,
gte_SetRotMatrix 4, plus 21 CC1-FAIL and 3 DIFF. Not separated: how much of the 62 is B's
conform vs the ladder's own recovery rungs.
Task B, re-scoped from evidence. The 129 dedup_extend failures are 106 conflicting-types /
21 CC1-FAIL / 4 undefined-ref / 3 DIFF — real byte divergence is 2%, and memcpy is 17 of 106,
not the story. Direction reversed too: the byte-true DEF of func_80128ED8 is what the target
.c files already declare; engine_core.h's macro-local extern was the stub-era guess.
Conformed 8 axes to byte-truth (func_8012F14C 2843, func_8012E5CC 2052, func_8012F038 2214,
func_8014C568 1816, func_80128ED8 1524, func_8012C750 406, func_8012C0EC 50, func_80144A04 25).
R22 clean-fleet: check-all 213 passed / 0 failed of 213. Zero functions banked by design.
Tooling (R33/R35) — three guards that asserted completeness over a narrowed population:
- NEW tools/macro_draft.py: a deduped fn has no definition in any .c (body lives in a DEFINE_
macro), so conform_decls had been refusing the largest class it was built for.
- conform_decls skipped engine_core.h wholesale as "a defining TU": 10 stale externs survived
while 1,514 fleet sites moved, and it still printed "axis complete". Skip now scoped to the
defining macro's span.
- Return-axis compare was literal: typedef int/s32 and a missing `extern` faked a return change.
Now compares normalized types.
- §85 consumer scan under-reported (the dangerous direction): a cast between `=` and the call
hid `s0 = (s32 *)func_80144A04(...)`. Now classified by position, validated both ways.
Corrections to my own predictions (R14): the documented scalar-narrowing hazard was benign
across 2,052 sites; the breaks were arity (6 call sites, fixed with §17a-1 fn-ptr casts) and
the consumer-guard gap. A header-only first probe broke ov_SC01_000 — §85 is literal.
Not done, named: memcpy (builtin codegen), ApplyMatrixSV (no DEF), gte_SetRotMatrix (link bug),
func_80147364 (unparseable macro), D_800AE620/D_80126CC4 (data axis). Cookbook §159.
- BANKED: 11 functions at 400-952 ins from the cascade (func_8017D898 952, func_8017CE58 733,
func_801902EC 673, func_8018C2D8 673, func_8018A8D4, func_8017C6F4, func_800CBB38,
func_800CF3A4, +3). check-all 213/213 from a clean tree. 6 near = jr/switch (§53 separate
banking step), 1 failed. The cascade agents wrote 6 new cookbook sections incl. §158.
⚠️ tools-health UNVERIFIED at commit (stale cookbook index fixed, confirming re-run
interrupted) — run it first next session. check-all is the byte oracle and it is green.
- WASTE PREVENTION (Drew: "prevent this from ever happening again, however you need to"):
* tools/validate_targets.py (NEW) — names 5 defect classes (NO-ASM / MID-BODY /
OUT-OF-RANGE / ALREADY-DONE / NO-BOUNDARY), exits non-zero.
* WIRED INTO wave_snapshot so it fails closed — every wave passes through there for its .s
files, so no path from target list to spawned agents bypasses validation. Negative-control:
a 3-target bad list is refused with the exact mid-body offset (+72 bytes of 100).
* The cascade `done()` predicate now short-circuits on SKIPPED as well as MATCH. It tested
only MATCH, so a non-existent target fell Sonnet -> Opus -> Fable and three agents each
proved the same phantom absent: ~29 invalid targets x 3 tiers = 87 of 119 agents, ~9.7M
tokens. A tier that cannot act must END the pipeline, not escalate emptiness.
* docs/accelerators.md A9, including that wave_snapshot's own R32 assertion REFUSED that list
(24 of 57 found) and was routed around — the one instrument warning that was right and ignored.
- B RE-SCOPED (S46-10) and deliberately NOT done: the extend blocker is INTRA-HEADER, not
target-side. engine_core.h declares memcpy FOUR incompatible ways across its DEFINE_ macros;
two in one TU collide. NOT a safe cleanup — the in-tree note at ov_MAIN_012.c:14333 records
that `extern memcpy` disables gcc's builtin and turns an inlined block-move into a CALL, so the
declaration CHANGES CODEGEN. Probe one macro in one binary and byte-gate before any sweep.
- C (dedup_extend over the 129) stays blocked on B. Full context for both in the checkpoint.