Files
BFM-decomp/phase-ends/CURRENT_PHASE.md
T

457 lines
84 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CURRENT PHASE — Phase 31: The Frontier Atlas & Wide-Tolerance Campaign
**Started:** 2026-08-14 · **Plan approved:** 2026-08-14 (gate 1; Drew) · **Effort doctrine:** xHigh default / Max deep (T5, T7, synthesis) / Ultracode waves (R26/R27 prompts) / Fable-tier only for new wall classes.
**Approved plan:** `/home/musashi/.claude/plans/fable-5-set-max-goofy-seahorse.md` (the full design; this file is the crash-recovery log).
**Approval also ratified R37 (probe before costing), R38 (read recorded failure verdicts first), R39 (negative-control new refusal-checks) — now binding.**
## The re-charter (one paragraph)
Instead of roadmap-v2 P31's per-function grind, Phase 31 organizes the 12,059 remaining stubs (main 1,041 · resident 14 · ov 9,832 · md 1,179) into **crack groups**: a deterministic per-function feature layer + multi-tier similarity atlas (li-normalized exact tier, seed sweep vs the 2,719 matched skeletons, calibrated warm tier for the 3,238-unit cold tail + main, kNN neighborhood graph), evidence-joined to a lever label per group; plus **widened mechanical lanes** (LEN-tolerant aligned remap with a fully-mechanical LI class, §172b EXTPAIR/SELECT detectors + routing, PLUMBING campaign, permuter cluster warm-start, weak-seed cards); then a **campaign loop to ceiling** — deterministic $0 lanes first, agents only for exemplars, velocity-ledgered, closed on measured decay. Milestone shape = P30 (campaign to ceiling; every remaining stub on a named ledger at close). Main fully included from day one.
## Task checklist
- [x] **T0 — Pivot log + freshness + hygiene** — DONE 2026-08-14. (xHigh)
- [x] **T1 — Integration quick-bank sweep** — DONE 2026-08-14 (pending final R22 log line). **8 banked, 0 agent tokens.** (xHigh)
- [x] **T2 — References** — DONE 2026-08-14. (xHigh)
- [x] **T3 — Main enablement** — DONE 2026-08-14. (xHigh)
- [x] **T4 — atlas_features.py** — DONE 2026-08-14. (xHigh)
- [x] **T5 — atlas.py** — DONE 2026-08-14. **THE ATLAS EXISTS.** (Max)
- [x] **T6 — PLUMBING campaign** — DONE 2026-08-14. **+9 banked (probe 64%); recipe + 3 laws distilled (§173); sweep proved the no-draft majority routes to family lanes.** (xHigh)
- [x] **T7 — family_align.py built + NC'd; the mechanical-cousin premise REFUTED by its probe (0/26)** — engine re-scoped to T8's LEN+N same-function pile; no driver built (correctly). (Max)
- [x] **T8 — LEN+N lane** — DONE 2026-08-14. **Pile routed 587/587; 345 wrong-drafts reclassified; 49 permuter + 192 card fuel staged; mechanical lane = honest null.** (xHigh)
- [x] **T9 — Warmstart + weak-cards** — DONE 2026-08-14. **Grinder queue armed (49 lenmiss + 10 seeded drafts, 120 refused by the stream filter); 954 weak-cards emitted, cheap-tier-dominated. No grinder patch needed.** (xHigh)
- [ ] **T10+ — Campaign loop to ceiling** (repeating sessions; velocity ledger; close on measured decay). (Ultracode waves / xHigh orchestration / Fable new-walls)
- [ ] **Tclose — PhaseEnd** (gate 2). (Max)
## Standing verification (every task)
R22 clean-fleet **213/213** after every banked batch · tools-health green · 0 NON_MATCHING (G4) · dedup-check 0 failed · R32 coverage assertions on every new scanner · R39 negative controls on every refusal check · R37 probes before pricing · R38 ledgers before experiments · commit per task (task + this log in the same commit; Drew pushes).
## Progress log
- 2026-08-14 — Phase planned and approved (3 Explore + 2 Plan agents; full design in the plan file). Task list built (harness tasks #1–#12). T0 started.
- 2026-08-14 — **T0 COMPLETE.** (1) R31 decision-log entry (the re-charter WHY + R37–R39 ratification). (2) `harvest_verify.py` import guard: a bare import now RAISES instead of running a gate (verified both directions; CLI behavior unchanged). (3) **Resident ±1 RESOLVED + FIXED**: `--bootstrap`'s linear partition had fused the +0 data word with `func_800CEDFC` (row `0x800CEDF8` nins=18) and dropped `func_800D33E0` past a glued tail — the true denominator is **145** (progress was right, the sig wrong). `sig-resident` now ELF-seeds (S45 pattern: unique 4-aligned T-symbol addrs inside the `resident_TEXT_START/END` markers → exactly 145; bootstrap fresh-clone fallback). All three oracles now agree (sig 145 · corpus matched 131 · progress byte-ident 131); `audit-corpus` 0 PHANTOM + 0 TRUNCATED. (4) Family maps regenerated at HEAD `commit:2161`: **11,025 open non-main members = 12,059 − main's 1,034 EXACT** (the stale map's 102 phantoms cleared); cousins totals now A-prop 1,125 / seeded 1,293 / cousin-multi 5,373 / cold 3,234 inst; adapt cards 704, aprop cards **204 (full emission)**. (5) **Main fuel gap is DEAD**: 2,001/2,002 main stubs have cached Ghidra-C (only `func_80049600` missing) — the roadmap's "0/2,096" note was stale. (6) `make tools-health` → OK (dedup 2,063/0; C1 254,521/254,521; audit-digest confirms the fleet digest; resident fix moved instr num+denom by the same +9).
- 2026-08-14 — **T1 COMPLETE: 8 banked for 0 agent tokens.** R38-first: partitioned the MATCH-108 pile against current stubs (75 still open) and against the S50 gate history (62 gated-and-failed with verdicts · 13 never-gated). The lanes and their measured outcomes:
- **Never-gated 13** → gate_lane: 0/13, but the verdicts decomposed to 11× `undefined reference to D_*` = the §171 stale-seed-symbol class. **Extended `aprop_symfix` with STALE-DELTA** (n:n uniform-delta rebase; R39 synthetic + snapshot NCs, zero false positives; the delta test even refused a pair my hand-check wrongly accepted) → 4 rebased, **4/4 banked** (`func_8016BCC0`, `func_8017F1C8`, `func_80186BD8`, `func_80186BF8`). Cookbook **§171-D** written in-session.
- **SELF-decl PLUMBING 7** → `recover_integration --stages demacroize --max-tier binary`: **4/7 banked** (`func_80139BE0`, `func_8014ED28`, `func_80161D88`, `func_801659DC`); 3 stay near.
- **Stored-draft re-gates** (no-verdict 7 + close=0 8 + diff 1 + 9 STALE→clean world-motion drafts): **0/23-ish banked** — the ~8% A10 stored-verdict law held again; all re-verdicted fresh.
- **Handed forward with fresh classifications**: CALLEE-decl 15 + CC1-FAIL 14 → T6 (cast-callees/tu-scope stages); UNDEF-DATA/OTHER 9 → the §171b-1 data-definition carry (T7/T8); md CARVE-REFUSED 8 → campaign side-quest ledger. immfix pile: fully consumed (0 open). fix20: 19/20 consumed in S50.
- Tool fixes landed: `gate_lane` propagate-commit tag now derives from GATE_PHASE (was hardcoded phase-30 S49).
- Rate lesson for the velocity ledger: fresh-fix lanes (STALE-DELTA 4/4, demacroize 4/7) vastly outperform blind stored re-gates (0/23) — the campaign loop's L2 ordering is confirmed by measurement.
- 2026-08-14 — **T2 COMPLETE (references).** (1) **PsyQ dev-CD extracted**: walked the on-disk Track-1 image (MODE2/2352) with the frozen `tools/bfm_extract/iso9660.py` (R33 — no new extractor; walker = iter_directory/read_extent with out-of-range extents skipped) → `tools/reference/psyq-sdk/` (gitignored): 2,374 files / 231.6 MB, **400 C sources (373 in PSX/SAMPLE/** across CD/GRAPHICS/SOUND/MODULE/CMPLR/…); only 7 out-of-track `.DA` audio skipped. (2) **Provenance find**: `GNU/SNGNUVER.TXT` = SN Systems' gcc build history (`2.7.2.SN32.3.7.0002`, 14.5.97) with per-build changelog of SN's patches vs vanilla — only `UNROLL.C` (parameterised max unroll insns) is codegen-relevant; recorded in the idiom notes as the first-look suspect if a loop-unroll residual ever defies the vanilla model. (3) **gcc-2.7.2 reference completed**: +6 files from GNU ftp (`calls.c` + `caller-save.c` — both cited by §172's producer model, previously missing — + integrate/optabs/varasm/recog), tarball sha256 `7cd8bce5…` recorded. (4) `docs/psyq-sample-idioms.md` seeded (inventory, provenance, first style conventions, the lane hook); SETUP §5.6 rows added (R21).
- 2026-08-14 — **T3 COMPLETE (main enablement).** (1) `sig_image` seed-ends extension: `--seeds` now accepts `0xADDR NINS` lines (and jsonl `nins`); a seeded nins is authoritative — the slice is exactly `[addr, addr+4·nins)`, bypassing `func_end` (whose heuristic mis-sliced 3/40 main samples). (2) `corpus.s_ins_count()` factored from audit() (R33, one counter) + `corpus.py <bin> --seed-ends` CLI. (3) **`make sig-main`**: 2,002 main stubs signed at splat-true lengths → `.run/sig.main.jsonl` (deliberately splat-SEEDED — the atlas needs the boundaries a match must hit; NOT the independent second oracle, which stays deferred per `docs/second-oracle.md`). (4) **Full word cross-check: 2,002/2,002 slices byte-faithful** (EXE bytes vs `.s` comment-column words; 0 SLICE-SUSPECT; note the `.s` word field is byte-order hex, not the LE-decoded value — first checker draft compared wrong and read 1,999 false suspects). (5) `family_remap` main special-cases: `vram_of('main')` = code-segment `vram − start` from `splat.us.exe.yaml` = 0x8000F800; `img_path('main')` from the same yaml's target_path (R33). `stream_words('main')` == `.s` words on **25/25** samples. (6) Regression: `sig-resident` re-run **byte-identical** after the shared `read_seeds` change.
- 2026-08-14 — **T4 COMPLETE (`tools/atlas_features.py`).** Per-function features memoized per distinct `h_exact` (`.run/feat_memo.json`) and fanned out 1:1 with sigs to `.run/feat.<bin>.jsonl` ×213 (main reads the splat-true `sig.main.jsonl`; registry untouched). **363,525 rows / 92,855 distinct bodies in 21 s** (vs the 3–6 min estimate). Features: nins/band · o0 prologue tell · frame/callee-saved set+order/fp · CFG skeleton (nblk/ncond/nback/ret_n via branch-target scan; jal = call, never an edge — position-independence proof in the docstring) · mid_jr/jalr · stable-call-sequence hash (fixed main+resident ranges) · reloc-kind sequence hash · 16-bucket ophist · **§172b tells as importable shared functions (extpair/dupselect/sign_mix/magic_div — T8's len_tells imports these, R33)** · **li-normalized skeleton `h_seqn`** (non-anchor lui dropped, ori→addiu class). Verifies: R32 join asserted per binary at write; determinism 0 mismatches on 200 re-derived bodies; **mid_jr cross-check vs family_hseq's independent oracle: 6,444/6,444 exemplars agree** (first verifier draft compared ZERO rows — a string-vs-int addr type mismatch, the R32 silent-no-op class caught in my own verifier; now indexed + fails loud if compared==0). o0-tell vs `corpus.is_o0` on 13,019 open fns: 41 both / **94 tell-only = candidate undiscovered -O0 fns in -O2 TUs (§116 — o0_subsplit carve fuel for the campaign)** / 14 src-only outliers.
- 2026-08-14 — **T5 COMPLETE (`tools/atlas.py` + `make atlas`) — THE FRONTIER ATLAS EXISTS.** Survey in 92 s, all assertions green: **5,139 groups cover ALL 12,058 open instances / 613,710 ins** (12,058 = progress's 12,051 stubs + 7 NON_MATCHING exactly — the +1-blob and +8-banked discrepancies were both chased and resolved: data blobs excluded, T1 banks confirmed absent; main open reconciles to classify()'s stubs+NM with a hard assert). Tiers: **calibration measured** (positives = cousin ≥0.85-raw merges, negatives = size-matched cross-unit pairs; the li-normalized metric is so discriminative that recall ~99% holds down to 0.55; rule = smallest t with neg-accept ≤0.2% ∧ recall ≥95% → **THRESH_WARM=0.70** at 99.1%/0.18% — chosen because a false warm-merge wastes an exemplar crack, a miss only routes cheaper) · T1.5 h_seqn 2 merges · **warm tier 1,019 merges** · **seed sweep: 4,702/7,247 open skeletons (65%) carry a ≥0.55 matched seed** from the 2.7k-skeleton pool · kNN graph (top-8 open + matched neighbors) · tiny-direct. **Lever table (the strategic map):** head-crack 1,283g/186.9k ins · UNKNOWN 1,964g/138.6k (honest) · **extend-tell 575g/76.7k** · redraft 46.8k · **jtbl-carve 190g/45.7k** · integration 575 inst/23.4k · seeded-crack 23.4k · len-vein 778 inst/16.8k · swaprepeat 9.2k · plumbing 161 inst/8.1k · o0-lane 6.6k · cc1 6.4k. Top group: 268 drifted single-member skeletons unified (per-location ~23-ins family, 0.72 seed). Evidence joins accounted (audit/backlog/ledgers/cards; residuals EXCLUDED-STALE until regenerated); `--targets` resolves 12/12 `.s`. `make atlas` = the full regen chain. docs/frontier-atlas.md committed.
- 2026-08-14 — **T6 IN PROGRESS (PLUMBING campaign) — probe converged after two R14 corrections.** `plumbing_groups.py` derived the honest pool: **237 still-open** PLUMBING rows (the "1,217" was ledger-vintage inflation): SELF 109 (3 concentrated binaries) · CALLEE 48 · OTHER 48 · DATA 32. Probe (ov_SC03_107 SELF-21 → 14 with drafts): first run **0/14 with a PHANTOM shared error** — root-caused to **cross-group TU-edit poisoning** (a demacroize edit in TU-A persists while TU-B's drafts gate; every whole-binary build compiles ALL TUs) → fixed `recover_integration` with per-group isolation (git-checkout binary TUs between groups, gate_lane's proven pattern; engine_core.h untouched so fleet-tier arity persists) + added the **`macro-externs` draft stage** (§121: one draft's guessed `extern int f()` vs the TU's `DEFINE_`-macro definition — lifted from family_sweep, R33) + **`tu-scope` stage** (§103 STU, the sweep-only lever). Isolated re-run exposed the TRUE class: `undefined reference to D_*` = **the §171 stale-seed-symbol class** (rtu_match MATCHes these drafts — it is blind to relocation names; the R34 two-oracle disagreement exactly as documented) → **symfix-first**: 11/11 rebased (mixed 1:1 + n:n deltas incl. the S49 Δ0x4128) → **9/14 banked (64%)** where the raw path scored 0. The standing recipe: `aprop_symfix --fix → recover_integration --stages macro-externs,demacroize,tu-scope` (per-group isolated). Full sweep over the remaining ~200 rows now running.
- 2026-08-14 — **T6 COMPLETE.** Sweep over the remaining ~200 rows: **the decisive finding is Law 3 (§173)** — the biggest groups (ov_SC02_037 44 rows, most of ov_MAIN_012's 42) have **zero stored drafts**: their PLUMBING verdicts came from transient family-sweep remaps never persisted to the backlog. A recovery lane can only recover what was stored — verdict-only rows are family-lane fuel (already atlas-labeled), not recovery fuel; the true stored-draft class was largely consumed by the probe (**9 banked, 64%**). Singleton tail with drafts: 0/~20 (fresh per-fn verdicts recorded). Cookbook **§173** (symfix-first · per-group isolation · verdicts-without-drafts) + index regenerated (518). **R22 clean fleet: 213/213** with all T6 banks. Phase total so far: **17 banked, 0 agent tokens** (stubs 12,059 → 12,042).
- 2026-08-14 — **T7 COMPLETE (honest refutation).** `tools/family_align.py` built: aligned classifier (SequenceMatcher over FC.tok; LEN-LI/LEN-NOP/LEN-JTBL/LEN-STRUCT/STRUCT-ALIGNED/PURE/IMM verdicts) + li-cluster reconstructor (with the split-cluster absorb for the rs-changed addiu partner) + the aligned imm engine. **NC-1: 157/157 verdict-equivalence** with classify_member on banked pairs — the NC itself caught two real classifier gaps (R-type non-shift sa = STRUCT; registers tested BEFORE the reloc skip — a reloc-slot word with a different register is STRUCT). NC-2 parity 21/21. **The R37 probe then killed the planned driver before it was built: 0/26 LI-ONLY cards classify mechanically** (regfields drift ×19) — cousins are 0.85-similar DIFFERENT functions; §168 law 1 ("a cousin is a seeded crack, never a remap") re-derived by measurement. The engine's true consumer is **T8's LEN+N pile** (draft vs its own target = same function). Parked for T8: the reloc-vs-constant range discriminator for lui-bearing clusters. Decision-log entry written (R31).
- 2026-08-14 — **T8 COMPLETE (LEN+N lane).** Built: `match_one --emit-streams` (additive, stdout-identity NC'd) · `family_align.addr_true_rel` (the reloc-vs-constant range discriminator — full conservative set kept for pair semantics/NC-1 [157/157 regression], address-true subset for indel eligibility; synthetic probes: lui-bearing constant cluster now LEN-LI, address-anchored indel still refuses) · `tools/len_tells.py` (per-draft aligned classification + §172b tell tagging on target-side indels, detectors imported from atlas_features R33, cookbook text embedded in cards) · `tools/lenmiss_route.py` (pool-parallel per A8: **587 audit LEN rows re-verified live and routed in 24 s**). **Routing:** redraft **345** (frac>0.35 — the draft is not the function; reclassified APPEND-ONLY in the backlog, so near-miss metrics stop lying) · permuter-length **49** (|Δ|≤2 clean drift — grinder fuel) · cards **192** incl. **14 EXTPAIR/SELECT/NOP-tagged** (the audit's own detectors had emitted zero — these are new signal) · **mechanical 0 — an honest null**: stored drafts rarely get CONSTANTS wrong (agents copy them from asm); LEN drift is shape, so the LI-cluster swap lane has no fuel in this pile and family_align's value here is the classifier/detector. R32 accounting 587/587.
- 2026-08-14 — **T9 COMPLETE.** `tools/warmstart.py` (the permuter/grinder FEEDER): `--from-banked` walks a banked exemplar's h_seq family's still-open members, builds remapped proven-body drafts (`symbol_map` + `aprop_autodraft.build_draft`, refusing on reloc-count mismatch), **stream-classifies member-vs-seed with ZERO compiles** (masked_diff-shaped dicts from ground-truth bytes → `residual_class.classify_streams`), enqueues ONLY permuter-shaped (bucket==permuter or LENGTH-DRIFT |Δ|≤2) as backlog near-records; `--lenmiss` ingests T8's 49-route. **Grinder patch NOT needed** (its candidates() deliberately keeps unclassified records — "unknown is not a reason to skip" — so pre-filtered enqueues flow as-is; documented in the feeder's docstring). Armed live: **49 + 10 enqueued, 120 refused** by the stream filter (the anti-92%-wasted-CPU discipline working). `family_cousins --weak-cards`: **954 units** (the 0.70–0.85 annotate-only band, never before consumed) as seeded-crack cards, ins-ranked, §168 laws embedded, model-routed **haiku 804 / v3 43 / sonnet 86 / opus 21** (cheap tiers dominate — the token-efficiency shape), 0 unresolved `.s`.
## SESSION S53 (2026-08-16, ultracode) — wave R + the leftover-draft harvest
- 2026-08-16 — **S53-1 PREFLIGHT + ATLAS REGEN.** Tree clean at `commit:2416`, no gate in flight. `make atlas`
regenerated at HEAD (the S52 atlas predated waves P+Q, whose banks grew main's matched seed pool 175→293):
**4,985 groups / 11,352 open instances / 579,571 ins**, warm merges 880, seeded 4,714/6,911 skeletons, all
assertions green. (Watch-for confirmed live: my own `pgrep -f atlas` waiter self-matched its shell wrapper —
the bracket trick is mandatory.)
- 2026-08-16 — **S53-2 LEFTOVER-DRAFT HARVEST (R38: read the recorded verdicts before designing anything).**
`.run/s53_scan_leftovers.py` re-verified every wave-P/Q main draft on disk with `match_one` rather than
trusting the journals (R14). Of 141 drafts: **72 are already banked** (their `.s` is gone — the honest
signal that splat stops emitting a matched function), **34 still verify MATCH and are still stubs**
(3,075 ins of finished work nobody had banked), **35 are NEAR**. Oracle cross-check: all 34 MATCH and all
35 NEAR are in `corpus.stubs('main')`; all 72 ERROR are not. Zero agent tokens.
- 2026-08-16 — **S53-3 PRE-GATE LADDER ON THE FREE SLATE (the S52 protocol, applied cold).** `reloc_identity
--batch`: **33 AGREE / 1 MISMATCH** (`func_80034C24` names `D_80078F20` where the target references
`cdReq_sectorHdrBuf+0xE0`; its `--dry-fix` rename to `D_80078F10` is unverified, so it was dropped, not
guessed). `fragment_check`: **1 FAIL** — `MoveImage` DEFINES `SYS_OBJ_8F4`, a stub that still has its own
`.s` (exactly the enclosing-function trap that cost wave Q a 3-hour bisect; caught in milliseconds).
`reconcile_slate --apply` on the pruned 32: **11 compatible, 21 refused as human decisions** (7 TYPE,
5 SIGNATURE, 3 DIFFERENT-STRUCT, 3 BROKE-MATCH, 1 DEF-SIDE-RETURN, 1 resourceIdMap TYPE, 1 alias).
**The 11 were NOT gated on their own** — §176h.C2 says banking a subset hardens the rest's conflicts
(last session: 1 of 18 parked drafts survived that), so all 32 ride one slate after repair.
- 2026-08-16 — **S53-4 WAVE R BUILT — and the main mass band is measurably SPENT.** `build_wave_atlas
--only-bins main --min-ins 60 --max-ins 200 --rank mass` returned **917 ins / 10 cards** against a 6,500
target: waves O/P/Q consumed main's mass band. Probes (R37 before costing): main widened to 30–400 yields
44 cards / 4,600 ins in 10 gate groups; **fleet-wide 60–200 yields 63 cards / 6,525 ins in exactly 2 gate
groups = 31.5 drafts per rebuild**, from a 1,546-candidate well. Wave R therefore = the fleet-wide draw
(`ov_SC06_029` ×41, `ov_SC02_005` ×22) + main's last 10 mass cards, **7,442 ins across 73 fresh cards**.
- 2026-08-16 — **S53-5 WAVE R LAUNCHED (110 agents, 3 lanes, run `wf_c070b4fa-53b`).** Lane 1 = 73 fresh
mass cracks. Lane 2 = the 21 declaration-blocked drafts (already byte-MATCH; the agent's job is to make the
DRAFT agree with the TU, never the reverse, then re-verify — a declaration change is a codegen change).
Lane 3 = 16 near-miss repairs, whose residuals were classified deterministically first: **10 of 16 are the
§177 epilogue signature** (`addiu $sp / jr $ra / nop` vs `jr $ra / addiu $sp`, all in `800c3`), 4
SCHEDULE-REORDER, 1 BRANCH-POLARITY, 1 OPCODE-MIXED. The prompt carries §177 as law 4 and §178's
"REGALLOC-PERM is the most over-diagnosed class" as law 5.
## SESSION S52 TASK LIST (2026-08-15, ultracode) — the monitorable view (no TaskCreate tool in this harness build)
- [x] **S52-1 — Preflight + wave-selector repair.** Tree clean @ `commit:2390`, no gate/grinder in flight. Fixed `build_wave_atlas.py`: (a) the already-waved set was a hardcoded `'abcdefghijkl'` wave-letter literal → now a `glob('.run/wave_*_cards.json')` derivation (R33); **NC: old 634 → new 726 taken, strict superset, +92 previously-missable cards from waves m/n**; (b) `--exclude-bins` defaulted to `main` carrying the REFUTED link-defect rationale → default now empty, help corrected to the real (gate-path) reason; (c) new `--only-bins` allow-list (main waves need it — `gate_main` rebuilds once per SLATE, so main has no per-TU gate cost).
- [x] **S52-2 — Atlas regen** (`make atlas`, $0) — the selector was stale by ~448 banks (last regen predates waves J–N).
- [x] **S52-8 — `gate_lane` crash-vs-empty fixed.** A non-zero rc or a missing JSON line is now labelled `‼ CRASH`, prints the stderr tail, records `{'CRASH':True,…}` in the results JSON, lists the never-gated groups in the summary, and exits non-zero. R39 NC both ways: crash→CRASH, honest-empty→not-CRASH.
- [~] **S52-3 — Wave O: MAIN + the UNKNOWN probe** — LAUNCHED (`wf_278b05de-bae`, 49 cards / **6,266 ins**, 2 gate groups). **The selection finding that reshaped this wave: main's agent-lane pool at ≥20 ins is nearly SPENT (21 left after 113 already waved) — main's remaining 52,714 ins are overwhelmingly UNKNOWN-lever, as are 138k fleet-wide.** So wave O is deliberately three arms: 21 main head-crack (avg 145 ins, opus) + 18 main UNKNOWN (avg 122) + 10 overlay UNKNOWN (`ov_SC04_011`, avg 103, sonnet). The UNKNOWN arms are an R37 probe of the biggest unclaimed block on the atlas — 1,239 candidates at 60–120 ins alone. Wave prompt gained §175/§176A–C (statement-order-around-a-call first; pin the interloper; **a pin cannot schedule across a call and can silently DELETE an instruction**).
- [ ] **S52-4 — Independent re-verify (R14) + `gate_main.py --apply`** (one clean rebuild per slate).
- [ ] **S52-5 — R22 clean fleet 213/213 + commit** (task + this log together).
- [ ] **S52-6 — Wave P: overlay TU-packed**, drafted concurrently with the main gate (drafting is tree-free).
- [x] **S52-7 — 11 conflict-dropped main drafts recovered** (waves J/K/L). All 11 re-verified MATCH standalone; the TU-aware `gate_main` then named **7 real TU conflicts the old check missed entirely**, all repaired by adopting the TU's declaration verbatim + casting at the use site — including a NEW variant, casting the CALLEE through a function pointer when the TU prototypes it `(void)` and your call must pass an argument (`((void (*)(s32))func_8001C9D0)(a0)`). All 11 re-verified MATCH after repair; slate now 11/11 compatible, staged for the next main rebuild. **None needed a codegen change — every one was plumbing.** Cookbook **§176d**.
- [x] **S52-10 — `tools/reloc_identity.py` (NEW): the oracle that DISAGREES with `match_one` about symbol identity** (R34). match_one masks relocations, so it cannot see a wrong callee or a wrong global (§174 law 1c). But the target `.s` comment column is the FINAL LINKED WORD, so the true address behind each masked field is recoverable arithmetically and `symbols*.txt` maps it back to a name — $0, no rebuild, and it names the fix. **Three of its own bugs were caught by its negative controls before any verdict was believed:** splat-derived `func_`/`D_` names aren't in the symbol files (first run checked ZERO relocs while reporting clean — R32); a 0x4000 nearest-symbol window mislabelled `func_8001C9D0` as `SsGetMute+0xC50`; and **MIPS o32 REL relocations keep the addend IN THE INSTRUCTION**, so reading it off the operand string fabricated a mismatch for every struct-field/array access (the `+1/+2/+3` signature on `func_801F0734` was a byte-array walk, not three symbol errors). Also refuses to answer when streams aren't index-aligned.
- [~] **S52-11 — The 88-draft "MATCH but gate-rejected" pile, triaged deterministically for $0.** Result: **68 AGREE · 12 MISMATCH · 8 COMPILE-FAIL**. The 12 split into two named, actionable classes: **uniform-delta stale seed symbols** (§171 — `func_80130D48` all 5 relocs off by exactly `0xD1EC`, `func_8016D688` all 5 by `0x65450` → `aprop_symfix` STALE-DELTA rebase) and **wrong field offsets** (`func_80185D6C` +0x10, `func_8017F060` +0x4, `func_80188720` +0xC). `--fix` then mechanically repaired **10 of the 12** (2 correctly REFUSED as **SYMBOL-COLLAPSE** — one draft `extern` standing in for two distinct globals, which a rename cannot fix). MISMATCH 12 → 2, AGREE 68 → 77.
- [x] **S52-12 — LANE KILLED BY MEASUREMENT (R37).** Re-gating 20 of the AGREE drafts (5 groups) banked **1 — 5%**, statistically identical to the project's existing **A10 stored-verdict law** (~0–8%; T1 measured 0/23 on the same kind of pile earlier this phase). **The null is the finding: symbol verification does NOT improve stored-draft re-gate conversion** — a stored draft's rejection is almost never identity, it is TU plumbing or staleness. So `reloc_identity`'s real home is as a **pre-gate check on FRESH drafts** (seconds, removes a whole failure class before the rebuild), NOT as a backlog resurrection tool. The remaining 30 groups are NOT worth 30 rebuilds; lane closed. Banked `func_80186C44` + propagation. **R38 self-note: the 0/23 prior was already in this very log — I should have started from ~8%, not from optimism.**
- [ ] **S52-8 — Tooling debt: `gate_lane` swallows `gate_stage` stderr** (reports a crash as `0/0/0`).
- [ ] **S52-9 — Bank idioms into the cookbook + refresh this checkpoint** BEFORE any pause (memory `bank-idioms-before-checkpoint`).
## 📐 THE WAVE DOCTRINE (adopted 2026-08-15 by Drew, after wave O) — 6,000+ INSTRUCTIONS PER WAVE
**A wave is sized by INSTRUCTION MASS, not by card count.** The public metric is instruction-
weighted, so a wave is worth what its instructions are worth. The card lanes (12–42-ins cousins)
carried ~1,400 ins/wave ≈ 0.011pp of fleet ⇒ ~440 waves to finish. Wave O carried **6,266 ins for
the same gate cost and the same draft rate (47/49)**. That is the shape from here on.
**The recipe** (`--target-ins` implements it; the tool now refuses to under-fill silently):
```
.venv/bin/python tools/build_wave_atlas.py .run/wave_<id>_cards.json 80 \
--target-ins 6500 --min-ins 60 --max-ins 200 --max-bins 4 \
--levers head-crack,seeded-crack,redraft,len-vein,integration,family-sweep,UNKNOWN
```
- **`--target-ins 6500`** — draw cards until the instruction budget is met (capped by `n`).
- **`--min-ins 60 --max-ins 200`** — the mass band. Draft rate barely decays with size (wave M 98%
at avg 51, wave N 92% at avg 65, **wave O 96% at avg 128**), so size is nearly free mass.
- **`--max-bins 4`** — gate cost scales with (binary, TU) GROUPS, not drafts.
- **UNKNOWN is now a first-class lane** (see below). `main` needs `--only-bins main` + `gate_main`.
**THE UNKNOWN UNLOCK (wave O's strategic result).** UNKNOWN is not a difficulty label — it is
"the atlas could not name a lever". Wave O ran 22 UNKNOWN cards as an R37 probe and they drafted
like any other lane. That moves ~138k ins into reach and re-scopes the whole endgame:
| pool (agent-draftable, incl. UNKNOWN) | fns | ins | waves @6k |
|---|--:|--:|--:|
| **mass band 60–200 ins** | 1,762 | **164,357** | **27** |
| 40–59 ins | 1,752 | 84,009 | 14 |
| <40 ins (the old card lanes) | 5,585 | 132,466 | 22 |
| >200 ins | 125 | 36,493 | 6 |
| **total agent-draftable** | **9,224** | **417,325 = 70% of all open ins** | **69** |
Non-agent levers hold the remaining ~175k ins (extend-tell, jtbl-carve, cc1, o0-lane, swaprepeat,
needs-autopsy, plumbing, frame-172) and still need their own lanes.
**Order of work: the mass band first** — 27 waves covering 164k ins, and it is where the
instruction-weighted metric moves fastest per agent spent.
**THE PRE-GATE PROTOCOL (do all five, in order — wave O proved each one earns its place):**
1. **Independently re-verify every claimed MATCH with `match_one`** (R14). Agent self-reports run
optimistic; wave O happened to agree exactly (47/49), earlier waves did not (wave C claimed
35/35 → 32 banked; the main probe claimed 6/6 → 4).
2. **`tools/reloc_identity.py --batch`** — the symbol-identity check `match_one` structurally
cannot do. Seconds, $0, and it removes a whole failure class before a rebuild is spent.
(Wave O: 46 AGREE / 0 MISMATCH — the first wave of the campaign with zero symbol errors.)
3. **`gate_main.py <slate>` DRY RUN, and iterate until `N -> N compatible, 0 dropped`.** Conflicts
surface one layer at a time; each fix reveals the next.
4. **Reconcile declarations toward the form the MATCH needs, never arbitrarily** (§176f) — then
re-verify every converted draft, because a declaration change is a codegen change.
5. **Gate.** main → `gate_main.py --apply` (one clean rebuild per slate). Overlays → `gate_lane`.
## Campaign velocity ledger (T10+)
| wave | lane | cards | standalone MATCH | banked | tokens | tok/bank | lesson banked |
|---|---|--:|--:|--:|--:|--:|---|
| **O** | **mass — main head-crack + main UNKNOWN + overlay UNKNOWN (49 cards / 6,266 ins)** | 49 | **47/49 = 96%** / **6,040 ins** (my independent re-verify agreed exactly) | **50** — 46 main in ONE clean rebuild (`143dbb89` byte-identical) + 4 overlay | ~12.8M | ~256k | **THE 6k-INS WAVE SHAPE + THE UNKNOWN UNLOCK.** 4.4× the card lanes' mass at the same gate cost; UNKNOWN drafts like any lane (⇒ 138k ins re-scoped into reach). First wave with **0 symbol errors** (`reloc_identity` pre-gate). Declaration reconciliation took the slate 5-dropped → **0 dropped**, and 3 of 4 conflicts were load-bearing CODEGEN (§176f). Gate cost 8 attempts: 2 my errors, 3 real tool defects now fixed (typedef ordering, silent bisect-on-build-failure, `short`≠`s16` over-refusal), 3 reconciliation rounds |
| A | adapt SMALL-EDIT | 24 | 18 (75%) | **12** (+1 prop) | ~1.40M | **~117k** | zero stale seed-symbols (the §171 prompt-law works); 6 gate-fails all INTEGRATION shapes (3× decl-type vs TU, 1 arity, 2 TU-context DIFF) → wave-B prompt adds match-the-TU's-existing-decl |
| grinder-1 | permuter | 10 | — | 0 | $0 | — | func_800CB4CC parked at best-1 (warmstart seed); queue enriched +59+6 records since |
| C-probe | tell (3) | 3 | 1 (33%) | **1** | ~300k | ~300k | §174 **Law 4**: the TU's decl of YOUR OWN fn constrains the def sig — standalone MATCH gated 0/1 until canonical-sig + cast-at-use; tell attribution unreliable (2/3 residuals were a different class) |
| C | tell 11 + weak 24 | 35 | **35 (100%)** | **32** (8 propagated ×2) | ~2.39M | ~75k | Law 4 in-prompt → 0 symbol fails, 91% gate; **weak lane 24/24 on haiku — the 890-card vein is live**; 3 NEAR → grinder |
| E-probe | mass/main (6) | 6 | 4 (67%) | **0** | ~474k | — | **main is agent-draftable but LINK-BLOCKED**: 1 byte-correct fn ⇒ 2-byte whole-EXE diff, a `jal` retargeted game-code→PsyQ-archive symbol. Not a matching wall. `gate_lane` main-blindness fixed en route |
| D | adapt (48) | 48 | **47 (98%)** | **45** (40 main gate + 5 late-repair, 2 gates) | ~4.48M | **~100k** | repair stage rescued 5/6 first-pass DIFFs → gate them in a SECOND slate (the first slate is built before the repair stage lands); 23 gate groups for 42 drafts = the throughput ceiling → `build_wave_atlas.py` now packs by **(binary,TU)** = the real gate-group key; 1 NEAR at **close=1 DELAY-SLOT** → grinder |
| **MAIN** | **mass/atlas main probe (4 drafts, re-gated CLEAN)** | 4 | 4 | **4 BANKED — the first main-EXE functions of the campaign** | $0 (re-used the wave-E probe drafts) | — | `func_80013228`/`func_8001CB00`/`func_800142C8` (src/800.c) + `func_8001099C` (src/boot.c, -O0). `make clean && make extract BINARY=main && make build BINARY=main` → **143dbb89 BYTE-IDENTICAL**. These are the SAME drafts the incremental gate rejected 0/4 — the drafts were right, the gate path was wrong |
| **M** | mass/atlas — overlay, LARGER band (44) | 44 | **43/44** shape-verified / 2,212 of 2,249 ins | **40 banked**, ONE gate group; R22 213/213 | ~4.41M | ~110k | **The band test.** Cards avg 51 ins (up to 112) vs the 12–42-ins cousins the night opened with, and the draft rate HELD at 98%. Since the public metric is **instruction-weighted**, this is the band that moves it — and it is reachable by haiku/sonnet, not only the frontier tier. 1 NEAR (close=9, beqz+delay-slot reorder) → grinder |
| **L** | mass/atlas — MAIN (44) | 44 | **44/44 = 100%** | **42 banked** (2 dropped: dup `SVECTOR` typedef + `u8[]`-vs-`char[]`) — main 133 → **175** | ~1.96M | ~47k | First wave carrying **law 1c**. The compile-error shortcut named `SVECTOR at src/800.c:94` + the exact draft in SECONDS where the old bisect burned 28 min. Drove the `gate_main` typedef-stripper fix |
| **K** | **mass/atlas — MAIN (44)** | 44 | **44/44 = 100%** / 674 ins | re-gating (41 compatible / 3 conflict-dropped) — **an earlier "43 banked" report of mine was WRONG, see below** | ~2.36M | — | Third perfect sweep, all haiku. **`tools/gate_main.py` (NEW) drove it**: substitute batch → `make extract BINARY=main` → `make build` → SHA, with bisection. **MY TOOL BUG, caught by R22 not by the tool:** `gate_main` said BYTE-IDENTICAL, then the clean fleet check failed `[FAIL] main`. Cause: my `typesig()` split on the symbol name and kept only the PREFIX, so **`u8 D_x` and `u8 D_x[]` compared EQUAL** — three drafts declared `D_80078D98` inconsistently (1 scalar, 2 array) and the conflict reached the build. Fixed to keep the declarator suffix; NC'd both ways (wave-K conflict now caught; wave-J answer unchanged). **Both failure modes are now on record in the tool: v1 too STRICT (compared parameter names → discarded 2 good drafts), v2 too COARSE (ignored `[]` → passed a real conflict).** |
| **J** | **mass/atlas — MAIN (40)** | 40 | **39/40 = 98%** / 917 of 940 ins | **34 BANKED** (5 dropped for in-TU decl conflicts, 1 DIFF) — main now **38 matched** | ~2.65M | ~78k | **First full wave against the main EXE, and it drafts like any overlay (98%, all haiku).** Gated by the clean-rebuild batch path: substitute → `make extract BINARY=main` → `make build BINARY=main` → **143dbb89 BYTE-IDENTICAL**. **NEW CLASS — in-TU cross-draft decl conflicts:** batching N drafts into ONE `.c` means their `extern`s must agree with EACH OTHER, not just with the file (`D_800A4ED4` s16-vs-u16; `func_8001C9D0` void/void*/s32). Resolved greedily (keep-in-order, drop incompatible): cost 5. Dropped set is recoverable next session via cast-at-use — the same lever that fixed `func_80037368` (`extern u8 D_80076251;` verbatim from `src/shared/clearTbl40.h` + `(&D_80076251)[i]`) |
| I | mass/atlas (44) | 44 | **44/44 = 100%** / 1,602 ins, symfix clean ×44 | **44 BANKED**, ONE gate group | ~4.27M | ~97k | Second perfect sweep; 41 of 44 drawn by **haiku**. Gate reported `0/0/0` twice — an unreported `gate_stage` CRASH (`corpus.CorpusError`: 16 missing `.s`, fallout from MY `make clean`), not a result. Binary then built byte-identical with all 44 substituted; R22 213/213 confirmed (batch included a fleet-shared `engine_core.h` arity fix) |
| H | mass/atlas (40) | 40 | **38 (95%)** / 1,820 of 1,902 ins | **34** of 38, ONE gate group | ~5.05M | ~149k | 2 NEARs banked as permuter fuel with strong diagnoses. **SAFETY FINDING: a `register __asm__("$2")` PIN CAN BE A CORRECTNESS BUG, not just a scheduling choice** — on `func_80182EB0` a value was written before a call and read after it; the hard-reg pin made gcc treat the pre-call store as dead across the call and **silently DROP** `addiu v0,zero,-1` (49 vs 50 ins), then read garbage. Dropping the pin + storing the constant directly before the call recovered it and closed 19/25. Pins are our most-used lever — this failure mode needs a cookbook note |
| G | mass/atlas (36) | 36 | **36/36 = 100%** (independently re-verified) / 2,554 ins | **32** of 33 gated **in ONE gate group**; fleet 95.3→**95.4%** | ~3.16M | ~99k | **Best draft rate of the campaign, on FRESH CRACKS.** 3 held by the symfix audit for inventing symbol names where the target calls **PsyQ `RotTransSV`/`RotMatrixY`** — the just-fixed non-hex handling earned its keep immediately (pre-fix, one such name crashed the whole audit). Pattern: agents reconstruct the CODE reliably and guess PsyQ SYMBOL NAMES unreliably; the deterministic audit is what catches it |
| F | **mass/atlas (60)** — the TU-packed experiment | 60 | **55 (91%)** pre-repair → **59/60** post-repair / 3,508 ins | **56** total (50 + 6 stragglers) of 60, **1 rebuild each** | ~7.16M | ~143k | **THE WAVE SHAPE FOR THE REST OF THE CAMPAIGN.** Fresh-crack lane (no proven body to edit) converts like the seeded lanes; ~3× mass/card (avg 48 ins vs 12–42); **1 rebuild for 50 banks** vs wave D's 23 rebuilds for 45. 2 genuine STALE (agents invented `S80131E00`/`Mat32_…` where the target calls **PsyQ `Square0`/`RotMatrixY`**) held back; 5 "local-only" flags are just local type names (harmless). Straggler batch pending: the repair stage rescued ~6 more AFTER the slate was built (same lesson as wave D — **build the slate after the repair stage lands**) |
- 2026-08-14 — **CAMPAIGN OPEN (T10+, Ultracode).** Wave A: 24 adapt cards → 18/24 standalone (75%) → symfix-first (0 stale — the prompt-law worked) → gate **12 banked + 1 propagation** (~117k tok/bank e2e; stubs → 12,030). 6 NEARs (3 at close ≤3) enqueued to the grinder. Wave B (48 cards, all haiku, 29 binaries) launched with the decl-matching lesson. Grinder pass 1: 0/10 but func_800CB4CC at best-1.
- 2026-08-14 (resume session) — **Checkpoint resume steps 1–2 DONE.** (1) **R22 owed proof banked: `make clean && make extract-all && make check-all` → 213 passed, 0 failed of 213** (`.run/r22_p31_resume.log` EXIT=0) — wave B's banks verified fleet-clean. (2) **`make atlas` regenerated at HEAD `commit:2234`**: 363,525 feature rows / 92,855 distinct bodies (0 newly computed — memo hit); **5,144 groups / 11,994 open instances / 612,325 ins**; warm merges 1,014; seeded 4,710/7,247 skeletons; all assertions green. Open count 12,058→11,994 reconciles with the session's 65 banks. (3) Wave C fuel verified: `.run/wave_p31c_cards.json` = 38 cards (14 tell/sonnet + 24 weak/haiku); grinder idle (heartbeat `done`); gate free. Drew's re-extraction question answered (disc/splat layers gain nothing — deterministic + continuously regenerated; sig/atlas layer regenerated by this step; Ghidra-C re-analysis = targeted-probe candidate only, R37). NOTE: this harness build has no TaskCreate tool — this checklist + log is the monitorable task view (R28 noted plainly).
- 2026-08-14 (resume session, cont.) — **WAVE C COMPLETE: 32 banked, R22 213/213.** Probe→wave cadence (§174 law 3) paid: the 3-card probe cost ~300k tok and returned Law 4, which the 35-card wave then converted at 91% gate with zero symbol failures. Independent re-verify of all 35 drafts by me (R14) before gating: 35/35 MATCH held. **Reach measured honestly: 32 exemplars, only 8 had any sharer, each ×2 → ~1.25× effective.** The ×134 era is over (P25/29/30 harvested the shared cores); fleet-% now moves ~1:1 with exemplars banked, so THROUGHPUT is the lever, not leverage.
- **THREE instrument defects found+fixed in my own new tooling this cycle** (the R32/R35/R39 class, and the reason R39 exists): (1) `build_wave.py`'s gate-guard used `pgrep -f` via `shell=True` — the wrapping `sh -c` carries the pattern in its own cmdline so it self-matched and refused forever; fixed by invoking pgrep without a shell. (2) The open-stub predicate did `fn in corpus.stubs(binary)` — but `stubs()` returns **addr→Stub**, so every card looked "already-banked" and BOTH pools reported ~0 candidates (I nearly concluded the 622-card adapt pile and 954-card weak pile were spent). Fixed to `{st.symbol for st in ...values()}` and **negative-controlled against the known population** (35 wave-C cards, 32 banked → exactly 3 still open: PASS). (3) I fired wave D once with hand-typed placeholder card names instead of the generated file — caught and stopped within a minute, relaunched from `.run/wave_p31d_args.json`. Recorded, not buried (P9).
- **Pools are NOT spent (post-fix, honest): adapt 569 candidates, weak 890 candidates.**
- **STRATEGIC (Drew's question "why are our waves so weak?" — he is right):** the card lanes I have been drawing from are the *mechanical leftovers* (12–42-ins cousins), ≈1,400 ins/wave against 612,325 open ins ≈ **0.23%/wave**. The Atlas already maps where the mass actually is: **head-crack 1,276 groups / 4,248 inst / 186,291 ins** (≈3.3 inst per group — a group head crack banks its members), **UNKNOWN 1,964 / 138,597**, extend-tell 575 / 76,730, redraft 280 / 46,769, jtbl-carve 190 / 45,673. And high-reach groups DO remain: gid `96f8b78ec8d3` = **265 instances / 6,084 ins** (integration lever, 31-ins exemplar), `75cec648ba15` = 116 inst / 3,002 ins, `152d63c1086d` = 64 inst / 1,559 ins. **Next: build the wave generator off `.run/atlas.json` groups (top instruction-mass first, routed by the group's measured lever) instead of the leftover card piles.**
- 2026-08-14 (resume session, cont. 2) — **THE TARGETING CORRECTION (Drew's "why are our waves so weak?" — he was right).** Measured the arithmetic: **635,744 ins open (4.7pp of fleet) across 11,994 fns**; the adapt/weak card piles I had been drawing are the 12–42-ins best-seeded tail ⇒ ~1,440 ins/wave = **0.011pp/wave ≈ 440 waves to finish**. The mass is elsewhere: cousin-multi 294k ins (avg 44), **cold 183k (avg 82, no seed)**, main-only 38k (avg 87). TWO fixes:
1. **`tools/build_wave_atlas.py` (NEW)** — wave selection off `.run/atlas.json`, on two measured principles: **(a) gate cost scales with (binary,TU) GROUPS, not drafts** — each group is a whole-binary rebuild, and wave C was 35 drafts over 27 groups = 1.3 drafts/rebuild (~50 min of gate for 32 banks); the atlas selector concentrates a wave into few binaries (**96 drafts → 1 group**, ~70× the gate efficiency); **(b) mass beats count** for the instr-weighted metric. 7,430 draftable candidates available in the agent-lever bands.
2. **`gate_lane` was STRUCTURALLY BLIND TO `main`** (R36/R33): it located a stub's home .c by `glob('src/<binary>/*.c')`, but main's sources live at `src/*.c` → every main draft grouped under `src=None`. Latent because **main has never been wave-gated** (main = 79,510 weighted ins at 0.5%, the largest coherent mass left). Fixed to derive from `corpus.stubs()[..].path`; NC'd 3 ways (still-open wave-C drafts 3/3 agree · overlay sample 96/96 agree · main now resolves `None`→`src/800.c`).
- **MAIN PROBE (6 cards, R37 — never spend 96 agents on an unproven path): 4/6 standalone MATCH (67%)**, including the `-O0` boot-module fn `func_8001099C` (the main-specific trap: `asm/nonmatchings/boot` → `src/boot.c` is **-O0** per §6, so match_one needs `--o0`; flagged in the mass-lane prompt). **Main is agent-draftable.** Gate result pending.
- **RULE VIOLATION CAUGHT (P9, recorded):** a wave-E agent wrote its body directly into `src/boot.c` instead of its draft dir — caught by `git status`, reverted, gate re-run clean. A dirty tree ABORTS the shared gate for the whole wave, so the wave prompt's HARD RULES were hardened (explicit "never write into the tree, not even to test; no state-changing git commands").
- Wave D (48 adapt cards) drafted 47/48 concurrently — drafting and gating overlap safely (drafts land in `.run/`, the gate writes `src/`).
- 2026-08-14 (resume session, cont. 3) — **MAIN'S BLOCKER DIAGNOSED — it is LINK-LEVEL, not matching.** Main probe: 4/6 standalone MATCH (my independent re-verify; the workflow's own repair stage claimed 6/6 — the gate and my check disagree with the agents, R14), then **gate 0/4 banked, 4 near**. Reproduced one (`func_80013228`, a clean 27-ins `Square0` wrapper) through `harvest_verify` and byte-diffed the built EXE against the original: **exactly 2 bytes differ in 413,696**, and NOT inside the drafted function — a `jal` at **vaddr 0x80060E74** retargeted from **`func_80061FA8`** (game code, `build/src/800c2.o`) to **`firstfile`/`firstfile2`** (**PsyQ libapi object `build/psyq/apicard/A66.o`**, symbols.us.txt line: `firstfile2 = 0x80062248`). So adding ONE byte-correct C function to `src/800.c` perturbs **symbol resolution between game code and the LINKED PsyQ library objects** — the C is right; the link binds a call to a different definition. This is main-specific (main is the only binary with `psyq_integrate` archive objects, per the `ifeq ($(BINARY),main)` blocks) and explains why main has sat at 0.5%: it is an INTEGRATION wall, not a matching wall. **Main is NOT ready for bulk waves; it needs a named link-resolution investigation lane** (candidate leads: duplicate `.NON_MATCHING` symbol definitions visible in the map at both `Square0` and `func_80061FA8`; archive-member selection order when a new undefined ref appears in `src/800.o`). Recorded here + decision-log (R31). The 4 main drafts are preserved in `.run/wave_p31e/main/` as fuel for that lane.
- 2026-08-14 — **Campaign returned to the lanes that bank.** Wave D (48 adapt cards) drafted; independent re-verify **42/48 MATCH (87.5%)**; 6 DIFF handed to the near/grinder path. Gating now.
- 2026-08-15 — **Wave D closed: 45/48 banked. R22 clean fleet 213/213. Stubs 12,059 → 11,876** (183 banked this phase; ~118 tonight). Committed `commit:2324`. **Wave F LAUNCHED — the first TU-packed wave**: `build_wave_atlas.py` now packs by **(binary, home .c)** — the real `gate_lane` group key — giving **60 drafts / 3,508 ins in ONE gate group** (vs wave D's 42 drafts / 23 groups). Target `ov_SC02_011` `jr_8017AE2C`, avg 48 ins (vs the card lanes' 12–42), 70 head-crack + 17 redraft + 4 seeded + 2 integration + 2 len-vein + 1 family-sweep; 32 sonnet / 28 haiku. **This is the throughput experiment**: same gate cost as ~2 wave-D groups for 60 functions of larger mass. 6,981 candidates remain in the atlas's agent-draftable levers (`main` excluded).
- 2026-08-15 — **WAVE F RESULT: the TU-packing thesis is CONFIRMED, and the fresh-crack lane works.** 60 atlas `mass` cards (no proven seed to edit — agents decompile from the .s) → **55/60 = 91% standalone**, **3,203 of 3,508 instructions**, in **ONE gate group**. Wave D needed 23 whole-binary rebuilds for 42 drafts; wave F needs 1 for 60. Per-card mass is ~3× the card lanes'. **This is the wave shape for the rest of the campaign** — `build_wave_atlas.py --max-bins 1..6`, 60–96 cards.
- **NEW TOOL DEFECT (logged, not yet fixed): `aprop_symfix` CRASHES on non-hex symbol names.** `deltas[int(new[-8:],16) - int(old[-8:],16)]` assumes every symbol is `func_XXXXXXXX`/`D_XXXXXXXX`; a draft calling **PsyQ `Square0`** raises `ValueError: invalid literal for int() with base 16: 'Square0'` and takes the whole audit down. This will recur on every PsyQ-calling draft. **Fix wanted:** treat a stale pair whose names are not hex-suffixed as a direct 1:1 rename (no delta), and never let one unparseable pair abort the batch (R32: a crash is not a coverage answer). Workaround used: gate the 53 clean drafts, hold the 2 genuine STALE for a hand rename.
- **Grinder ran concurrently with drafting** (free CPU, no tree writes during permute) on the close≤6 seeds: 3 ILS cycles each on the md_MAIN_* band, **0 banked, all parked at best-1** — plus it re-surfaced the Phase-22 split-file blindness (`func_800CB270: no .s under md_MAIN_027 — skip`). Stopped via the STOP sentinel to hand the tree to the wave-F gate (single-writer discipline; the gate is chained to start on grinder release).
- 2026-08-15 — **Wave G: 36/36 drafted (100%), 32 banked, fleet 95.3→95.4%.** Two TU-packed waves now confirm the shape. **`aprop_symfix` non-hex fix landed** (`commit:2330`): a curated PsyQ name (`Square0`) made `int(name[-8:],16)` raise and abort a 55-draft audit *after* the renames had already succeeded; now counted as a `named-1:1` rename and the batch survives. Verified on the literal incident values (old aborts, new completes + still buckets the hex pair by its `0x484c` delta); honest caveat — the live slates could not re-trigger the path because those drafts are banked, so the changed expression was exercised directly.
- **Standing pattern worth keeping in the prompt:** agents reconstruct CODE reliably (91–100% standalone) but guess **PsyQ symbol NAMES** unreliably (wave F: `S80131E00`→`Square0`, `Mat32_…`→`RotMatrixY`; wave G: `SRM_…`/`STM_…`→`RotTransSV`, `Blk20_…`→`RotMatrixY`). The deterministic symfix audit — not the model, and not `match_one` (which masks relocations and is BLIND to a wrong callee name) — is what catches this every time. Consider adding "if the target's .s calls a PsyQ symbol, use that exact name" to the wave LAWS.
- 2026-08-15 — **`aprop_symfix` FALSE-POSITIVE class found (R39), and it cost real banks — my error, not the tool's alone.** I withheld 3 wave-G drafts from the gate because symfix reported them `STALE`/`AMBIGUOUS`. On inspection **all 3 already used the correct PsyQ names** (`RotMatrixY`, `RotTransSV`); what symfix flagged as "draft-only symbols" were **local identifiers** — a typedef (`Mtx8_8017DE10_8017E710`), inline-asm macro names (`SRM_80186334`/`STM_80186334`), and local struct typedefs (`Blk20_…`/`Vec32_…`). Gated unchanged: **3/3 banked** (wave G → 35/36). Two defects behind it: (a) the classifier counts local typedef/macro names as symbol references; (b) its draft-symbol extraction misses some `extern` declaration forms, so a name the draft *does* declare (`RotMatrixY`, line 8) still shows as `asm-only`.
- **OPERATING RULE (adopt now): symfix flags are ADVISORY; the whole-binary gate is the arbiter.** Never withhold a standalone-MATCH draft from the gate on a symfix flag alone — gate it and let the bytes decide. R39's own wording applies to *me* here: a refusal check that silently discards good work is worse than one that lets a few failures through. (Law 1b in the wave prompt remains correct as *prevention* — the wave-F cases were genuinely wrong names.)
## Blockers
- 🔴 **CORRECTION (2026-08-15, R14/R35/P9 — I got this WRONG earlier tonight and reported it with confidence).** The "main is blocked on a LINK-RESOLUTION defect" conclusion below is **REFUTED**. It is a **stale-artifact / missing-extract** condition in the GATE'S BUILD PATH — the R22 corollary this project already documented at Phase 20 — not a linker bug and not anything to do with the drafts.
- **The control that killed it:** with `src/` fully reverted and **NO draft at all**, `make build BINARY=main` still produced the "broken" `c4546248` and the identical 2-byte diff. A defect that reproduces with zero drafts is not caused by drafts.
- **The fix, byte-proven both ways:** `make extract BINARY=main && make build BINARY=main` → **`143dbb89…` BYTE-IDENTICAL**, reproduced twice. `make build` alone → `c4546248` + the 2-byte `jal` diff, deterministically.
- **Why main and not overlays:** main's `extract` runs the EXE-only `psyq_integrate` + `ld_interleave` steps (the `ifeq ($(BINARY),main)` Makefile blocks) which **rewrite the linker script**. `gate_stage`/`harvest_verify` build **without** re-extracting, so main gates against a stale `.ld` — exactly the "a reverted config needs a re-extract, not just a rebuild" corollary. Overlays have no such step, so they are unaffected (and every overlay bank tonight is R22-verified from a genuinely clean tree).
- **What this means: main is very likely NOT blocked at all.** Its 79,510 weighted ins @ 0.5% are gated behind a TOOLING gap in the main gate path, not a compiler or linker wall. **Next step (do this first):** teach the main path to re-extract (or re-run `psyq_integrate`) before the gate build — then re-gate the 4 preserved drafts in `.run/wave_p31e/main/` (2 of which reference no PsyQ symbol at all). The earlier "archive-member selection" evidence (the broken map gaining ~20 PsyQ symbols + `firstfile` at `0x80062248`) is a **downstream symptom of the stale `.ld`**, not the cause.
- **The false-lead ledger, kept honestly:** I built a 3-hypothesis theory on a measurement whose instrument I had not controlled, and only the null-draft control exposed it. Same lesson as the `corpus.stubs()` misread and the symfix withholding earlier tonight — three instrument errors in one session, all mine.
- *(superseded, kept for the trail)* ~~main is blocked on a LINK-RESOLUTION defect~~ — see the 2-byte `jal` retarget above. Needs its own lane before any main wave is worth running. Overlay/md lanes are unaffected and continue to bank.
- **NARROWED 2026-08-15 (a specific, testable lead).** The call site is `asm/nonmatchings/libmcrd1/func_80060D9C.s` +0xD8, and it references the callee **BY NAME**: `jal func_80061FA8`. Original encodes `0x80061FA8`; our build with one extra C function encodes `0x80062248`. **The delta is exactly `0x2A0`, which the map shows is precisely the `.text` SIZE of `build/src/800c2.o`** — the object whose `.text` *starts* at `0x80061FA8` (map line 3712: `.text 0x80061fa8 0x2a0 build/src/800c2.o`, with `func_80061FA8` and `func_80061FA8.NON_MATCHING` both bound there, and `func_80062144` inside it). So the name `func_80061FA8` resolved to the **END** of that object instead of its start — i.e. one object's length later, landing on the next section's first symbol (`firstfile`, `build/psyq/apicard/A66.o`). **Hypotheses to test, in order:** (1) the duplicate `func_80061FA8` / `func_80061FA8.NON_MATCHING` pair — a second definition winning under a changed link order; (2) `800c2.o` being dropped/reordered when a new undefined ref appears in `src/800.o`, so the name binds to the following object; (3) an `undefined_syms_auto.txt` / `symbols.us.txt` absolute (`firstfile2 = 0x80062248`) shadowing the object-provided symbol. Everything needed to test is in `build/us/SLUS_007.26.map` + `.run/wave_p31e/main/*.c` (4 preserved main drafts).
## 🛑 SESSION CHECKPOINT — S52 FINAL (2026-08-15/16). Phase 31 CONTINUES. NOTHING IN FLIGHT.
**Tree CLEAN at `commit:2415`. No process running. R22 verified 213/213 from a clean tree after the last bank.**
### Banked this session: 131
Wave O 46 main + 4 ov · re-gate probe 1 · wave P 32 main + 8 md · wave Q 40 main.
**main 175 → 293 matched · stubs 1,881 → 1,763.** Fleet 213/213 byte-identical, 0 NON_MATCHING.
### Measured wave economics — the numbers to plan with
| wave | carded | drafted | BANKED | yield |
|---|--:|--:|--:|--:|
| O | 6,266 ins | 96% | 5,166 | **82%** |
| P | 6,589 ins | 97% | 4,501 | **68%** |
| Q | 6,249 ins | 58% (stopped early) + repair | ~3,000 | ~48% |
**~4,800 banked ins/wave when a wave runs to completion (≈75% of carded mass)** — NOT 6,000. I
quoted the *draft* rate for most of the session and that overstated it. Still ~5× the old card
lanes; ~87 waves for the 417k-ins agent-draftable pool.
### THE FOUR RESULTS THAT OUTLIVE THE COUNT
1. **UNKNOWN is not a difficulty label** — it means the atlas could not name a lever. A 22-card R37
probe drafted it like any other lane ⇒ ~138k ins (a quarter of all open instructions)
reclassified as ordinary wave fuel. Agent-draftable pool: **9,224 fns / 417,325 ins = 70%**.
2. **Matching is solved at this scale; INTEGRATION is the entire cost.** 96–97% draft rates with
zero symbol errors, then ~14 clean rebuilds to bank them. Every failure was declaration plumbing.
3. **Reconcile BEFORE the first gate (§176h.C2).** Of 18 parked drafts still verifying MATCH, only
**1** survived the conflict check after their wave banked, versus 5 before. Post-bank recovery
banked **0**. Budget reconciliation into the wave.
4. **§177 — the epilogue return-delay slot is decided by the SAVED-REGISTER SET**, not scheduling
(`mips.c:5376 mips_epilogue_delay_slots`). Eleven functions sat 1–3 instructions from banked,
filed by every agent as an intrinsic wall. ~600 ins unblocked by forty lines of compiler source.
### Tooling built this session (all committed, all NC'd)
**NEW** `reloc_identity.py` (symbol identity — the oracle `match_one` structurally cannot be) ·
**NEW** `pregate_check.py` (validates a slate in 0.7s vs a 5-min rebuild) ·
**NEW** `reconcile_slate.py` (drives a slate to 0-dropped; auto-reverts any repair that moves a byte) ·
**NEW** `fragment_check.py` (the enclosing-function trap: a draft that SUBSUMES another symbol, or
that REDEFINES another stub's symbol in asm — the second cost a 3-hour bisect) ·
**NEW** `bisect_slate.py` (**null control FIRST**, per-step logging, true binary search — found the
culprit in 7 steps / 176s where `gate_main`'s built-in bisect ran 3 hours and named nothing) ·
`build_wave_atlas` `--target-ins`/`--only-bins`/`--rank mass` + 2 selector bugs · `gate_lane`
CRASH≠empty · **`gate_main` ×8 defects**.
### Cookbook banked: 543 → 564 sections
§176d–k (TU-seeded conflicts · symbol identity computable offline + its 5% null · declaration FORM
as a matching lever · the 6k-ins doctrine + 5-step pre-gate protocol · the batch-substitution hazard
map incl. **C2 reconcile-before-gating** · what a static pre-gate can/cannot prove · the cost of
stopping a wave + the measured repair-pass yield · two selector bugs) ·
**§177** epilogue delay slot ← saved-register set ·
**§178** six levers from the wave-P journals (the $0-add opaque copy vs `make_regs_eqv`;
return-const as a priority-1 hard-reg set; `birthing_insn_p` single-set rule; narrow-type copy
elision; the zero-offset alias hole; `MEM_IN_STRUCT_P` asymmetry) — leads with **"REGALLOC-PERM is
this project's most over-diagnosed class"** ·
**§179** eight more, harvested by 12 readers over 172 journal findings (loop-walked pointer
parameter → giv; the maspsx transcription checklist; no-epilogue functions; `gte_stflg` clobber;
struct-assignment block copy; pinning disables strength reduction; a pin creating a combine
LOG_LINK; mid-body `.global` fragment slicing).
### NEXT SESSION — in order
1. **Wave R the new way.** `build_wave_atlas --target-ins 6500 --min-ins 60 --max-ins 200
--rank mass`, then **iterate `reconcile_slate --apply` → `fragment_check` → `pregate_check` →
`gate_main` dry-run until `N -> N compatible, 0 dropped` BEFORE the first rebuild.** That
sequence is the whole difference between 68% and ~95% yield, and every tool in it now exists.
Use `bisect_slate.py`, never `gate_main --apply` without `--no-bisect`.
2. **Apply §177 to the eleven epilogue near-misses** (`800c`/`800c3`, closeness 1–3, ~600 ins).
Pure lever application, no drafting: change what is live across the call, re-verify, gate.
3. **4 immovable-TU-declaration drafts** (`func_8002D034`, `func_8001ABBC`, …) need their own pass:
edit the declaration in `src/800.c`, ONE clean rebuild, R22. They cannot ride a slate because
`gate_main` reverts `src/` before every build.
4. Grinder fuel: `func_80015F04` at closeness 2 with a fully-derived sched1/sched2 LUID model and
three seeds in `.run/p31p_15F04/`; plus wave-Q leftovers in `.run/wave_p31q/main/`.
### Watch-fors (all bit this session)
`gate_main`'s built-in bisect is near-linear and silent — use `bisect_slate.py`. · A wave stopped
mid-flight loses its in-flight tail; a repair-only pass recovers ~⅓ of it (12/39, 579 ins), but
resuming the workflow re-runs unfinished agents from scratch at full cost. · `pgrep -f` self-matches
its own shell wrapper — use the `[g]ate_main` bracket trick. · Closeness must be COUNTED, not read
off the first differing index. · A clean `pregate_check` is a licence to build, not a prediction of
success: link errors and byte mismatches are outside what any text check can see.
## 🛑 (superseded) SESSION CHECKPOINT — S52 mid-day
**Tree CLEAN at `commit:2407`. R22 verified 213/213 from a fully clean tree after the last bank. Nothing owed, nothing in flight.**
### Banked: 91 functions
46 main + 4 ov (wave O) · 1 (re-gate probe) · 32 main + 8 md (wave P). **main 175 → 253 matched · stubs 1,881 → 1,803.** Fleet 213/213 byte-identical, 0 NON_MATCHING.
### The two waves, measured honestly
| wave | cards / ins | drafted (my re-verify) | symbol errors | BANKED ins | yield |
|---|---|---|--:|--:|--:|
| O | 49 / 6,266 | 47/49 (96%) | 0 | 5,166 | 82% |
| P | 60 / 6,589 | 58/60 (97%) | 0 | 4,501 | 68% |
**The number to plan with is ~4,800 BANKED ins/wave (75% of carded mass), not 6,000** — I quoted the draft rate for most of the session and that overstated it. Still ~5× the card lanes' ~1,400. Revised projection: **~87 waves** for the 417k agent-draftable pool, not 69.
### THE THREE RESULTS THAT OUTLIVE THE COUNT
1. **UNKNOWN is not a difficulty label** — it means the atlas could not name a lever. A 22-card R37 probe drafted it like any other lane ⇒ ~138k ins (a quarter of everything open) reclassified as ordinary wave fuel. Agent-draftable pool is now **9,224 fns / 417,325 ins = 70% of all open instructions**.
2. **Matching is solved at this scale; INTEGRATION is the whole cost.** 96–97% draft rates and zero symbol errors across two waves, then ~14 clean rebuilds to bank them. Every failure was declaration plumbing — N standalone drafts having to agree with each other and with a TU none of them can see.
3. **Reconcile BEFORE the first gate (§176h.C2).** Measured: 18 parked drafts still MATCH, but only **1** survived the conflict check after their wave banked (vs 5 before). A banked draft's declarations become the TU's, so a sibling clash becomes a file clash, which is stricter. The post-bank recovery pass banked **0** — this law cost real work to learn.
### Tooling built/fixed (all committed, all NC'd)
**NEW** `tools/reloc_identity.py` (symbol identity, the oracle `match_one` structurally cannot be) · **NEW** `tools/pregate_check.py` (validates a slate in **0.7s** instead of a 5-min rebuild) · `build_wave_atlas --target-ins/--only-bins` + glob-derived taken-set · `gate_lane` CRASH≠empty · **`gate_main` ×8**: TU-seeded + per-file conflicts, definition-aware, trailing-comment-blind regexes (×2), typedef alias normalization, address-order walk, build errors surfaced instead of bisected, body+position-aware typedef handling.
### Cookbook banked
**§176d** TU-seeded conflicts + callee function-pointer cast · **§176e** symbol identity is computable offline (+ its honest 5% null) · **§176f** declaration FORM is a matching lever · **§176g** the 6k-ins doctrine + 5-step pre-gate protocol · **§176h** the batch-substitution hazard map (7 under-reporting holes, the 3 wrong typedef strategies, the spelled-name limit, **C2 reconcile-before-gating**).
### NEXT SESSION — in order
1. **Wave Q the NEW way**: build with `--target-ins 6500 --min-ins 60 --max-ins 200 --max-bins 4`, then **iterate `pregate_check` + `gate_main` dry-run to `N -> N compatible, 0 dropped` BEFORE the first rebuild.** That is the whole difference between 68% and ~95% yield.
2. The 4 **immovable-TU-declaration** drafts (`func_8002D034`, `func_8001ABBC`, …) need their own pass: edit the declarations in `src/800.c`, ONE clean rebuild, R22.
3. 5 genuine NEARs → grinder. `func_80015F04` is at **closeness 2** with a fully-derived sched1/sched2 LUID model (9/9 probes predicted) and three seed candidates in `.run/p31p_15F04/`.
4. Latent, unfixed: conflict detection compares spelled type NAMES; comparing struct **bodies** (the auto-reconciler already does this) is the real fix.
## 🛑 (superseded) checkpoint — S52 mid-session
**State at checkpoint:** tree CLEAN at `commit:2402` + wave-O bank commit. **R22 verified 213/213 from a fully clean tree** after wave O. Nothing owed.
### What S52 banked
**51 functions** — 46 main (ONE clean rebuild, `143dbb89` byte-identical) + 4 overlay (`ov_SC04_011`) + 1 from the re-gate probe. **main 175 → 221 matched, stubs 1,881 → 1,835.**
### THE TWO RESULTS THAT MATTER MORE THAN THE COUNT
1. **UNKNOWN IS NOT A DIFFICULTY LABEL.** It means "the atlas could not name a lever", and it had been routed as needing its own bespoke lane. Wave O's 22-card R37 probe drafted it like any other lane ⇒ **~138k ins (a quarter of everything open) reclassified as ordinary wave fuel.** With UNKNOWN in, **9,224 fns / 417,325 ins = 70% of all open instructions** are agent-draftable.
2. **THE 6k-INS WAVE DOCTRINE (adopted by Drew).** A wave is sized by INSTRUCTION MASS, not cards — see the doctrine section above for the recipe and the pre-gate protocol. Draft rate barely decays with size (M 98% @51 ins · N 92% @65 · **O 96% @128**), so mass is nearly free. Projection: **~69 waves**, mass band first (27 waves / 164k ins), vs ~440 under the card lanes.
### IN FLIGHT AT CHECKPOINT
**Wave P** (`wf_faa2b5e5-a80`, 60 cards / 6,589 ins, 2 gate groups: main ×51 + md_SC07_004 ×9) — hit the **weekly limit** at 18/60 drafted, then RESUMED (`w8xrn5y3i`). Cached agents replay; the 42 failures re-run. On completion: run the 5-step pre-gate protocol, then `gate_main --apply` for the main half and `gate_lane` for md_SC07_004.
- The 18 already-drafted are all MATCH-claimed; **my independent re-verify is still OWED** (the classifier rate-limited Bash mid-check). Do it before gating.
### Tooling fixed this session (all committed, all NC'd)
`reloc_identity.py` (**NEW** — the symbol-identity oracle `match_one` structurally cannot be, §176e) · `build_wave_atlas` (`--target-ins`, `--only-bins`, glob-derived taken-set, refuted main-exclusion default) · `gate_lane` (CRASH ≠ empty result) · `gate_main` ×4 (TU-seeded + per-file conflict table · typedef walk in ADDRESS order · build errors surfaced instead of silently bisected · `short`≡`s16` alias normalization).
### Cookbook banked this session
**§176d** TU-seeded conflicts + the callee function-pointer cast · **§176e** symbol identity is computable offline (+ the honest null: it does NOT rescue stored drafts, 5%) · **§176f** the declaration FORM is a matching lever · **§176g** the 6k-ins doctrine + the 5-step pre-gate protocol.
### Known-open items
1. `func_8002D034` — verified MATCH but needs `src/800.c`'s `D_800A4E74` decl changed u16→s16; `gate_main` reverts `src/` before building, so it cannot ride a slate. Recover as its own commit + verifying rebuild.
2. `func_80016224` — verified MATCH but requires `volatile D_800B9A02`, which is fatal to two other drafts. Near-fuel until the TU's form settles.
3. 6 of the 10 `ov_SC04_011` wave-O drafts failed the gate on TU plumbing (that overlay's own decl landscape).
4. The AGREE re-gate lane is CLOSED (measured 5% ≈ the A10 law). Do not reopen it.
## 🛑 (superseded) CHECKPOINT — OVERNIGHT CAMPAIGN CLOSED 2026-08-15 (morning)
**Cron `be8fb48c` (23-min overnight heartbeat) is CANCELLED.** No wave will fire on its own. Nothing is in flight at handoff.
### What this session did (waves C–N)
**~448+ banked** · stubs **12,059 → 11,549** · fleet **95.4% instr** · **213/213 byte-identical after every single batch** · 0 NON_MATCHING · **main 0 → 175 matched**.
Draft rates 91–100% across twelve waves, overwhelmingly **haiku/sonnet writing byte-exact C from raw MIPS with no reference body**. Four perfect sweeps (36/36, 44/44, 44/44, 44/44).
### The three things that actually changed the campaign
1. **MAIN IS OPEN — and was never hard.** It sat at 0.5% for the whole project, written up as the largest/hardest remaining mass. The blocker was that `gate_lane`/`gate_stage` build INCREMENTALLY while main's `make extract` runs the EXE-only `psyq_integrate`/`ld_interleave` steps that REWRITE the `.ld` → false diff (R22's own rationale). Four byte-correct drafts were rejected; I diagnosed a "linker defect" and built 3 hypotheses on it. **A null-draft control killed it** (the defect reproduced with ZERO drafts substituted). Use **`tools/gate_main.py <slate> --apply`**: substitute batch → extract → build → SHA; ONE clean rebuild verifies a WHOLE batch (40+ per rebuild). ~913 main stubs remain and they draft at 98%.
2. **GATE-GROUP PACKING is the throughput lever.** Gate cost scales with **(binary, TU) groups**, not drafts — each group is a whole-binary rebuild. Wave D: 42 drafts / 23 groups. Waves F–N: 40–56 drafts / **1 group**. `tools/build_wave_atlas.py` packs by TU and ranks by instruction mass; `--min-ins 40` targets the band the public instr-weighted metric tracks (wave M held 98% at avg 51 ins; wave N 92% at avg 65).
3. **`match_one` VERIFIES SHAPE, NOT SYMBOL IDENTITY.** It masks jal/HI16/LO16, so a draft calling the wrong function or storing to the wrong global reports a clean MATCH (wave K: `func_8002A234` had two globals swapped — 5 gate attempts). Only the whole-binary gate catches it. Now law 1c in the wave prompt.
### Tooling built/fixed this session (all committed, all NC'd)
`tools/gate_main.py` (NEW — batch clean-rebuild gate for main; in-TU decl-conflict resolution on TYPE SIGNATURES ONLY; duplicate-typedef stripping; compile-error culprit naming instead of bisection; **rm-output+returncode check after it once reported a FALSE PASS off a stale binary**) · `tools/build_wave_atlas.py` (NEW — TU-packed, mass-ranked selection; `main` excluded by default) · `tools/build_wave.py` (NEW — adapt/weak pools) · `gate_lane` home-TU resolution derived from `corpus` (was blind to main) · `aprop_symfix` survives curated PsyQ names · cookbook **§174 law 1b/1c**, **§175** (caller-saved pins can DELETE an instruction across a call).
### Idioms banked before this checkpoint (Drew's rule, 2026-08-15 — memory `bank-idioms-before-checkpoint`)
**Everything learned this session is in `docs/matching-cookbook.md`, not just in commit text.**
- **§174 Law 1b/1c, Law 4** — PsyQ symbol names; `match_one` verifies SHAPE not SYMBOL IDENTITY; the DEF-side prototype constraint.
- **§175** — a pin to a CALLER-SAVED register can silently DELETE an instruction when the value's live range crosses a `jal`.
- **§176a/b/c** (mine, process-level) — the verification-layer laws (what each check can and cannot prove); batch-gating mechanics (gate cost scales with (binary,TU) groups; batched drafts must agree with each other; compile errors name their own culprit); main cannot be gated incrementally.
- **§176 A–F** (agent-discovered, mined from all 14 wave journals by a 15-agent workflow; 26 novel of 81, each cross-checked against the existing cookbook first):
- **A — statement order around a call** is the FIRST check for any schedule/delay-slot/±1 residual. Two functions that first-pass agents filed as "irreducible tie-break / permuter fuel" went to MATCH by moving ONE statement above a call.
- **B — a small REGALLOC-PERM is usually not allocation** (narrow-symbol aliasing; pin the interloper, not the contested value).
- **C — 🔴 WALL REFUTATION, source-verified by me:** `gcc-2.7.2 sched.c:1704` tests `call_used_regs[i]` where every neighbouring line uses `regno + i`, so for a 1-word register it always tests `$zero` (call-used on MIPS) ⇒ **every hard-reg SET in a block gets a REG_DEP_ANTI on the last call**, while the pseudo arm is guarded by `reg_n_calls_crossed`. **A PIN CANNOT SCHEDULE AROUND A CALL — sometimes the fix is to UNPIN.** Refutes the universality of `sched.md` S11 step 1 and the "always try pins" reflex.
- D/E/F — CSE levers in reverse, two cc1-probed spellings, and four residual verdicts that were lying.
- The section ends with an explicit **"What is NOT banked here"** listing 7 mined items judged too thin — including two whose functions are still `INCLUDE_ASM` (so the lever is unverifiable) and one whose narrative **contradicts** the banked C. Nothing was silently dropped.
### Open work, in priority order
1. **Run main waves** — highest value, ~913 stubs, 98% draft, `gate_main` handles it.
2. **`gate_lane` swallows `gate_stage` stderr** — reports an unhandled crash as `0 banked / 0 near / 0 failed`, indistinguishable from an honest empty result (cost 2 cycles). Make it surface stderr / distinguish CRASH from NOTHING-BANKED.
3. **Recover ~10 conflict-dropped main drafts** (waves J/K/L) — verified-correct, need cast-at-use.
4. Grinder queue has fresh seeds incl. **close=1 DELAY-SLOT** (`func_80183578`) and count-exact `func_8017DAEC`.
### The methodological lesson (worth more than the count)
Every serious stall traced to **an instrument trusted without a control**, never to gcc: the main "linker defect"; `corpus.stubs()` read as names when it returns addr→Stub; 3 good drafts withheld on an ADVISORY symfix flag; `gate_lane` crash-as-zero; my conflict checker too strict then too coarse; my verifier passing without building. **Before believing a measurement, run the control that would make it fail** — a null input, a known-answer population, or an independent oracle. The counterweight: the safety architecture held every time. R22 caught the false pass, `corpus` refused to guess, and the byte-gate never accepted a wrong match.
## (superseded) mid-flight checkpoint
**Phase 31 CONTINUES.** Overnight campaign running under Drew's "waves and banking all night long" directive (Opus 5, ultracode, 23-min cron heartbeat `be8fb48c` as the loop's safety net).
**This session's arc:** resume R22 213/213 → atlas regen (5,144 groups / 11,994 open) → wave-C probe (3 tell) → wave C (35: 11 tell + 24 weak) **32 banked** → main probe (6) **0 banked, blocker diagnosed** → wave D (48 adapt) **47/48 standalone, gating now**. Session banked ≈ **59+** (18 mechanical/probe + 32 wave C + wave D in flight).
**THE TWO STRATEGIC FINDINGS (read these first on resume):**
1. **Reach is spent: ~1.25×.** 32 wave-C exemplars → only 8 had a sharer, ×2 each. The ×134 era ended in P25/29/30. Fleet-% now moves ~1:1 with functions banked ⇒ **throughput is the lever**, and the throughput bottleneck is the GATE, whose cost scales with **(binary,TU) groups, not drafts** (wave C: 35 drafts/27 groups; wave D: 42/23). `tools/build_wave_atlas.py` (NEW) concentrates a wave into few binaries (96 drafts → 1 group) and weights by instruction mass — **use it for every future wave**; the adapt/weak card piles are the 12–42-ins tail (~0.011pp fleet per 48-card wave ≈ 440 waves to finish).
2. **main is BLOCKED on a LINK defect, not on matching** (see Blockers). `--exclude-bins main` is the default in `build_wave_atlas.py`.
**🔑 THE NIGHT'S HEADLINE — MAIN IS OPEN (2026-08-15).** main (79,510 weighted ins, was 0.5%) was NOT blocked by a linker defect; that was my misdiagnosis off an uncontrolled measurement. It needs a **CLEAN REBUILD** to gate (`make clean && make extract BINARY=main && make build BINARY=main`), because main's extract runs the EXE-only `psyq_integrate`/`ld_interleave` steps that rewrite the `.ld` — exactly the trap **R22's own rationale** describes. **4 main functions are banked** (byte-identical, full clean build), drafted by the ordinary mass lane. **NEXT SESSION'S HIGHEST-VALUE TASK: give the main gate path a clean-rebuild mode** (one clean build verifies a whole BATCH — that is how 4 banked at once), then run main waves as ordinary campaign fuel. main has ~1,030 open stubs.
**📐 THE ONE METHODOLOGICAL LESSON OF THE NIGHT (worth more than the function count).** Every serious stall traced to *an instrument I trusted without controlling*, never to gcc:
- "main is link-blocked" → **refuted by a null-draft control** (the defect reproduced with ZERO drafts substituted). I had already written a 3-hypothesis linker theory on top of it.
- `corpus.stubs()` read as names when it returns **addr→Stub** → nearly reported both card pools exhausted.
- 3 good drafts withheld on an **advisory** symfix flag → all 3 banked unchanged when gated.
- `gate_lane` reporting an unhandled crash as **"0 banked / 0 near / 0 failed"** → indistinguishable from an honest empty result; cost 2 cycles.
- `gate_main.typesig()` **too strict** (parameter names) then **too coarse** (dropped `[]`) → discarded good work, then passed a real conflict that R22 caught.
**The rule that would have prevented all five: before believing a measurement, run the control that would make it fail.** A null input, a known-answer population, or an independent oracle. R35 says fix the instrument first; the sharper form is *confirm the instrument can even answer, and that it answers correctly on a case whose answer you already know.*
**⚠️ TOOLING DEBT FOUND TONIGHT (fix before the next long run):**
1. **`gate_lane` swallows `gate_stage`'s stderr** — it reported `banked 0, near 0, failed 0` twice while `gate_stage` was actually raising `corpus.CorpusError`. An unreported crash is indistinguishable from an honest empty result. It must surface stderr and distinguish CRASH from NOTHING-BANKED. (Running `gate_stage` directly gave the precise cause instantly.)
2. **`aprop_symfix` false-positives on LOCAL identifiers** (typedefs, inline-asm macro names) and its draft-symbol extraction misses some `extern` forms → it reported `STALE`/`AMBIGUOUS` for 3 drafts that were already correct. **Symfix is ADVISORY; the gate is the arbiter** — never withhold a standalone-MATCH draft on a flag alone.
3. **Never `make clean` mid-campaign** without immediately re-running `make extract-all` — it wipes every binary's `asm/` and every downstream tool then fails in confusing ways (cost 2 gate cycles tonight).
**📊 OVERNIGHT RESULT (2026-08-15, waves C–M).** ~408 banked · stubs 12,059 → **11,589** · fleet **95.4% instr** · **213/213 byte-identical after every batch** · **main 0 → 175 matched** (913 stubs left). Wave draft rates 91–100% across ten waves, mostly **haiku**, writing byte-exact C from raw MIPS with no reference body. Four perfect sweeps (36/36, 44/44, 44/44, 44/44).
**🔧 THE MAIN LANE — HOW TO RUN IT (this is the night's unlock).** main was 0.5% and written up as the hardest remaining mass; it was never hard, it was never *gated correctly*. Use **`tools/gate_main.py <slate.json> --apply`**: substitute the batch → `make extract BINARY=main` → `make build` → SHA. ONE clean rebuild verifies the WHOLE batch (42–43 banked per rebuild). It reports in-TU declaration conflicts and, on a compile error, names the culprit instead of bisecting. **Do NOT gate main through `gate_lane`/`gate_stage`** — they build incrementally and main's extract rewrites the `.ld`, producing a false diff (R22's own rationale).
- **Known next improvement (mechanical, recurring):** `gate_main` should strip DUPLICATE TYPEDEFS on substitution the way `harvest_verify` already does. `src/800.c` now carries local typedefs (e.g. `SVECTOR`) from previously banked functions, so any later draft defining its own collides and costs a draft per wave.
- The 5+3+2 conflict-dropped main drafts across waves J/K/L are **verified-correct and recoverable** with cast-at-use (adopt the other declaration verbatim, adapt at the use site).
**RESUME STEPS:**
1. `pgrep -f tools/gate_lane` — never run two gates, and never run `build_wave*.py` during one (R35 guard: `corpus.stubs()` misreports substituted drafts).
2. Gate the late-repaired wave-D drafts: `.run/wave_p31d_late_slate.json` (5 verified MATCH, rescued by the repair stage after the main slate was built).
3. R22 (`make clean && make extract-all && make check-all` → 213/213), then commit (task + this log together).
4. Next wave: `.venv/bin/python tools/build_wave_atlas.py .run/wave_p31f_cards.json 96 --max-bins 8` → convert to args (`{wavedir,cards_file,cards}`) → Workflow `scratchpad/p31_wave.js`. Lanes: `mass` (atlas), `adapt`/`weak` (`tools/build_wave.py <pool>`; adapt 569 + weak 890 candidates remain, both verified live after the predicate fix).
5. Grinder queue has 3 fresh high-value seeds incl. **`func_80183578` close=1 DELAY-SLOT** (§60a precedent: a close=1 delay-slot banked in ~6 min) and `func_8017DAEC` count-exact 113=113.
**Watch-fors (all bit tonight):** agent self-reports run OPTIMISTIC — always re-verify with `match_one` yourself, then the gate (wave C claimed 35/35→32 banked; main probe claimed 6/6→I measured 4/6→0 banked). An agent once wrote its body straight into `src/boot.c` (reverted; prompt hardened) — a dirty tree ABORTS the shared gate. `pgrep -f` self-matches its own shell wrapper (invoke without `shell=True`). `corpus.stubs()` is **addr→Stub**, not names.
## (superseded) checkpoint — 2026-08-14 pre-overnight
**Phase 31 CONTINUES (campaign-to-ceiling; do NOT close).** Session totals: **65 banked** (17 mechanical @$0 + wave A 12+1prop @~117k tok/bank + wave B **35/37 gated, 95% conversion** @~88k tok/bank — the decl-matching lesson nearly eliminated integration failures). Stubs ≈ **11,995** (from 12,059). All banks byte-gated + committed; wave banks propagated where sharers existed.
**RESUME STEPS (fresh session, after the standard load order):**
1. `git log --oneline -20` to see the wave-B bank commits; run **R22** (`make clean && make extract-all && make check-all` → expect 213/213) — it was NOT run after wave B (context ran out; the per-bank gates each verified their own binary, but the standing clean-fleet proof is owed FIRST).
2. `make atlas` (regenerates maps + atlas post-banks, ~15 min, $0).
3. **Wave C is STAGED, not launched**: `.run/wave_p31c_cards.json` (14 §172b tell-cards [sonnet] + 24 weak-seed haiku cards — measures the two untested agent lanes). Launch via the persisted workflow script `workflows/scripts/p31-adapt-wave-a-wf_2fbef223-859.js` pattern (args = the cards; NOTE the tell/weak cards have different fields than adapt cards — adapt the prompt per lane or write a v2 script). Needs `/effort ultracode` (R27).
4. Adapt pile remains ~630 SMALL-EDIT cards — the proven 73-95% lane; wave D+ = next 48 by the same selection (exclude banked; see `.run/wave_p31{a,b}_cards.json` for taken).
5. Grinder queue armed: 59 warmstart + 6 wave-A NEARs + 8 wave-B NEARs (4 at close ≤3). Relaunch: `GATE_PHASE=phase-31 .venv/bin/python tools/grinder.py --once --batch 15 …` (single-writer: never while a gate runs).
6. Ledger discipline: velocity row per wave (the table above); distill lessons per R16/R30; close the phase ONLY on measured multi-session yield decay (plan file §Leg-C).
**Watch-fors:** gate_lane aborts on dirty src/config (clean first); symfix-first before every gate (§173); the safety-classifier can rate-limit under 48-agent bursts (harmless — retry).
## (superseded) previous checkpoint (end of build arc)
**T0–T9 ALL COMPLETE AND COMMITTED** (through `commit:2181`). Phase totals: **17 banked, 0 agent tokens**; stubs 12,059 → 12,042; R22 213/213 verified twice (post-T1, post-T6). The machine: the Atlas (5,139 groups, `make atlas`), the widened lanes (symfix STALE-DELTA, recover_integration isolation + macro-externs/tu-scope, family_align + len_tells + lenmiss routing), the armed queues (grinder: 59 warmstart records; cards: 954 weak + 192 len + 704 adapt; permuter-49). NEXT = **T10+ the campaign loop**: L3 grinder running in background (launched at checkpoint time); **card/crack WAVES need Drew's `/effort ultracode` toggle first (R27)** — prompt and WAIT. Campaign cadence + close criterion: the plan file §Leg-C. If resuming fresh: read the approved plan + this log; check `.run/auto/grinder_heartbeat.json`; run `make atlas` to refresh; continue the loop.