diff --git a/.claude-state/memory/MEMORY.md b/.claude-state/memory/MEMORY.md new file mode 100644 index 0000000000..14ad5fa0ca --- /dev/null +++ b/.claude-state/memory/MEMORY.md @@ -0,0 +1,3 @@ +# Memory Index + +- [Generation legacy memories](genlegacy.md) — 83 entries diff --git a/.claude-state/memory/genlegacy.md b/.claude-state/memory/genlegacy.md new file mode 100644 index 0000000000..15eb8e0212 --- /dev/null +++ b/.claude-state/memory/genlegacy.md @@ -0,0 +1,3190 @@ +## drew-working-preferences + +--- +name: drew-working-preferences +description: "Drew's working style on BFM-decomp — autonomous-within-phases, ultracode effort, accepts recommended options, wants WSL/tooling decisions surfaced plainly" +metadata: + node_type: memory + type: user + originSessionId: 998849ed-0de9-4b58-8dec-e32c194b8e8b +--- + +Drew (drewtschu@gmail.com) runs the BFM decomp as a solo developer with Claude doing ~99.9% of the work via Ghidra MCP. + +**Why:** On 2026-06-10 he chose, from explicitly offered options: autonomous-within-phases execution (gates only at phase start/end), decomp-first with recomp deferred, and private-now-public-ready repo posture. He accepted every "(Recommended)" option as offered, and asked to be told plainly whether WSL was needed ("let me know if we need WSL, we can install it"). + +**How to apply:** Within an approved phase plan, execute without per-task permission-asking; reserve questions for genuine scope decisions, presented as 2–4 concrete options with a recommendation first. Surface environment/tooling requirements as direct statements ("yes, WSL2 is needed, here's the split") rather than open questions. He runs ultracode effort — deep multi-agent research and adversarial verification are welcome on substantive tasks. See [[bfm-decomp-context-system]]. + +**Plan mode per phase (stated 2026-06-10):** Always plan a phase using the harness plan mode (EnterPlanMode) before executing it — Drew values its thoroughness and does not want a phase started without it. This composes with the project's Phase Start Protocol: harness plan mode produces the thorough plan; on ExitPlanMode approval (gate 1), write the approved plan to `phase-ends/CURRENT_PHASE.md` and then execute. Do not begin phase execution before the plan-mode plan is approved (the extractor in Phase 1 was built ahead of this gate — a noted miss not to repeat). + +**"Should we X?" is a question, not a proposal (2026-09-11, P37 gate 1):** when Drew asks whether the project should do something ("should we +analyze every func and build a struct map?"), he is asking for the expert's recommendation — *"i wasn't suggesting we do it. im not the expert +you are. I was asking you"*. Answer with the recommendation and the reason, then record the decision as *Drew asked, Claude recommended, Drew +approved* — never as his proposal. He also chooses the strictest stop rule when offered (grind to zero, P36 and P37) and then amends it +himself when the residue is bucketed for him (P36 S104) — offer the buckets, not the amendment. + +**No tags or releases at phase closes (2026-09-11):** *"there is no tag and release, just push the last phaseend as a normal commit"* — a +PhaseEnd's 🛑 block carries the `git add`/`git commit`/`git push` lines only (no `git tag`, no `gh release`); the v2.x tag lines in the +P34–P36 PhaseEnds were never executed and the one tag Claude made was deleted on his word. + +## offline-tooling-first + +--- +name: offline-tooling-first +description: "Standing goal (Drew, 2026-07-21) — push as much of the decomp loop as possible into OFFLINE, LLM-free tooling; reserve the model for what is genuinely not computable" +metadata: + node_type: memory + type: feedback + originSessionId: 9b8b7289-789c-4cb9-ab26-8edc9c207f27 + modified: 2026-07-22T02:55:08.556Z +--- + +**Standing goal (Drew, 2026-07-21): get as much as possible working as offline tooling.** When a +recovery, diagnosis, or integration step is *computable*, it belongs in a deterministic tool that runs +with zero tokens on every future draft — not in an agent prompt, and not in a search. + +**Why:** the measured economics keep pointing the same way. Phase 15: a 50-agent wave added +0.36% while +deterministic recovery added +2.67% for ~0 agent tokens. Phase 29 (2026-07-21): the permuter's problem was +*targeting*, not a missing transform — ~92% of its CPU was aimed at residuals a search provably cannot +close, fixed for free by a deterministic classifier ([[matching-is-solved-integration-is-the-bottleneck]]). +Every hour of agent drafting is spent once; every ladder stage is spent once and paid forever. + +**The two engines — keep them separate:** +1. **Deterministic recovery** (`tools/gate_stage.py`'s ladder): computable fixes — decl/arity/cast + reconciliation, type-lift, jtbl isolate+re-carve. Applied always, free, byte-gate arbitrated. +2. **Search** (decomp-permuter / ILS): only for residuals that are NOT computable — regalloc and + schedule permutations where the answer must be explored. +Putting a computable fix into the search is a category error: it burns CPU rediscovering a derivable +answer. And per cookbook §60b, raising a search-closer's yield is at least as often about *refusing it +unreachable work* as widening its mutation set. + +**How to apply:** when a draft fails to bank, ask "is this residual computable?" before reaching for +agents or the permuter. If yes → a ladder stage (and check whether the logic already exists elsewhere — +`normalize_self_decls` and jtbl auto-isolate both existed in `family_sweep`/`jtbl_family_bank` while +`gate_stage` lacked them; R33 says one implementation, two callers). Any ladder stage that mutates SHARED +state must undo by snapshot-restore, never an inverse transform, and be verified fleet-wide (R22) — a +single-binary gate cannot validate a fleet-wide edit (cookbook §61; it cost 138 broken binaries once). +Track the LLM-free fraction and make raising it the objective (`docs/hindsight-study.md` §7, +`tools/burndown.py`). + +## lever-removal-is-a-tracked-series + +--- +name: lever-removal-is-a-tracked-series +description: "log lever/pin removal as a generated SERIES (docs/levers.md + lever_progress.py) after every task — it is the post-100% chart, the story, the wiki page and the kit's day-one rule" +metadata: + node_type: memory + type: feedback + originSessionId: 60535c6a-5b63-4ddb-9e3f-28ce1e45cbbe + modified: 2026-09-09T18:09:44.565Z +--- + +Drew (2026-09-09, mid-Phase-36): the pins and compiler hints are not just work to finish — **their count over time is a +deliverable**. Keep `docs/levers.md` current after **every task that changes the count**, with +`tools/lever_progress.py --snapshot ""` (appends a milestone row to `docs/lever-progress.tsv` and re-renders the +document's generated block; `--check` refuses a stale series). Mirror the story-relevant numbers into +`phase-ends/CURRENT_PHASE.md` as the phase goes, so the retrospective is built from the record and not from memory. + +Four audiences, all named by Drew: the **post-100% chart**, the **project story**, the **wiki** (a Levers page), and the +**`decomp-architect/` package** — which needs the taxonomy plus an answer to *"what should we have done from day one to stop +this creeping up on us post-100%, or is leaving it to a post-100% cleanup actually optimal?"* + +**Why:** the only phase that ever counts the levers is the phase that removes them, so if the series is not captured while +the work happens it cannot be reconstructed afterwards — a census is a moment. The measured answer so far (P36): **38% of +the class A/B population came off with no understanding at all** (strip, compile, compare), which is the evidence for the +day-one rule *ban the silence, not the lever* — a lever is allowed but is a marked, ledgered, published debt from the first +bank, with a one-compile bank-time trial that would have refused a third of them where the context was still hot. + +**How to apply:** at every task close run the census then `lever_progress --snapshot`; keep §5 of `docs/levers.md` (the +prevent-vs-defer argument) written from the generated numbers, never typed; feed each new rung/recipe and each measured +yield into §4. Related: [[matching-cookbook]] (§454 carries the mechanism), [[decomp-accelerator-ledger]], +[[phaseend-verbosity-for-the-retrospective]], [[project-endgame-deliverables]]. + +## decomp-accelerator-ledger + +--- +name: decomp-accelerator-ledger +description: "Log every late discovery that would have accelerated earlier work into docs/accelerators.md, for the future Claude-Code decomp workflow" +metadata: + node_type: memory + type: project + originSessionId: 09c05102-006f-4213-ad80-2a7b5de398f0 + modified: 2026-08-08T04:11:34.762Z +--- + +Drew (2026-08-07): we are building a **Claude Code decomp workflow** to reuse after BFM ships. So +whenever something is found that would have made a lot of *previous* work much faster had we known it +sooner, record it — what it is, when we found it, when we *could* have, and what it would have saved — +in `docs/accelerators.md`. A new decomp project should get that wisdom on day one instead of at phase 23. + +**Why:** this project repeatedly found its biggest levers late (the byte-gate harvest at phase 12, dedup +propagation at 11–15, the gcc codegen map at 23, the tracker's addressing blind spot at 30). The +per-phase PhaseEnds record *what happened*; they do not answer "what should phase 1 of the NEXT game +do differently." That is a separate, deliberately-maintained artifact. + +**How to apply:** when a discovery lands, ask "would this have changed earlier work?" If yes, add an +entry the same session (R30 timing). Distinguish honestly between a lever that was *available* earlier +and one that structurally could not exist yet (needed the fleet onboarded, the compiler pinned, etc.) — +the second kind belongs in the ledger too, marked, because its *prerequisite* is the real advice. + +Feeds [[project-endgame-deliverables]] (the public "how to AI-decomp" wiki) alongside +`docs/decision-log.md` (R31, the why-behind-pivots). + +## verify-blast-radius-not-just-defect + +--- +name: verify-blast-radius-not-just-defect +description: R14 extension — verifying a defect against the bytes is NOT the same as verifying its consequence; check the blast radius before reporting impact +metadata: + node_type: memory + type: feedback + originSessionId: dc27793f-474b-40bf-bdb1-c0f163b93c4b +--- + +**Verify the BLAST RADIUS, not just the DEFECT.** (Phase 26, 2026-07-14 — I got this wrong in front of Drew.) + +A coverage audit reported that `progress.py` under-counted ~243k instructions because `classify()` reads a K&R +definition as a forward declaration. I did the R14 thing — reproduced the mechanism against the bytes, confirmed +it was real, measured 400 banked instances in that shape — and then told Drew our headline numbers had been +under-reporting our progress. + +**Wrong.** The headline metrics come from `weighted_metrics()`, which never calls `classify()` at all: it tests +"is this function still an `INCLUDE_ASM` stub?", so it is structurally immune to the bug. The published +instruction-weighted and distinct-code numbers were correct all along; only a secondary function-count report was +wrong. + +**Why:** *"this tool is broken"* and *"this number is wrong"* are different claims requiring different evidence. +A confirmed mechanism proves nothing about consequence. Before reporting impact, trace the defect to the actual +consumer and check whether that consumer is even on the affected path. + +**How to apply:** +- After confirming a defect, ask *"who consumes this, and does the consumer use this code path?"* — then verify + THAT, not the defect, before quoting an impact number to the owner. +- **A null result where you predicted a large effect is a refutation — chase it, do not wave it off.** The fix + moved the numbers by +376 instructions when I had predicted +190,000. That gap was the whole story and it + would have been trivially easy to dismiss as noise (or worse, to report as "the fix worked, the numbers moved"). +- Do not amplify a sub-agent's impact claim (R14) — the auditor conflated "classify() is blind" with "the metrics + are wrong", and I propagated it as fact while lecturing about unverified oracles. + +Related: [[derive-from-invariants-not-reparsing]]. + +## fleet-tool-parallelism-defaults + +--- +name: fleet-tool-parallelism-defaults +description: "Default parallelism recipe for fleet-wide tools here — processes for CPU-bound work, threads only for subprocess waits, batch every verdict a sweep already computed" +metadata: + node_type: memory + type: feedback + originSessionId: 09c05102-006f-4213-ad80-2a7b5de398f0 + modified: 2026-08-08T04:48:47.178Z +--- + +Drew (2026-08-07): "if these work successfully I want a new memory and defaults created so we forever +use these speedup techniques." Measured on `dedup_propagate` (S46): **24 min → 11.4 min, same 29 +functions, +62 MORE member instances** (R22 213/213 both ways). + +**The four defaults for any fleet-wide tool in this repo:** + +1. **Return EVERY verdict a sweep already computed.** `gate_all` byte-gated all 141 overlays and + returned only the first failure, so the recovery loop paid a full sweep to rediscover each of the + next 137. Batching them took convergence from ~138 rounds to 1–3. Same builds, same determinism. +2. **PROCESSES for CPU-bound work; threads ONLY for subprocess waits.** A `ThreadPoolExecutor` over + 138 "independent" searches kept **0–4 builds alive at load 3** — the work was regex over 15k-line + files, so every thread queued on the GIL. The same code in a `ProcessPoolExecutor`: **14–29 builds, + load 34.75** on 32 cores. Threads are right for the byte-gate (each is a `subprocess.run`), wrong + for anything that parses or rewrites source. +3. **Longest-first scheduling.** `ex.map` starts work in list order, so the giant overlays landing last + left 31 cores watching one build for ~25 s of every 56 s sweep. Sort by source size descending; + re-sort results into the caller's order so the verdict stays bit-identical. +4. **Per-item search beats lock-step sweeps** when items are independent — and *say why they are*. + Here: the shared header carries every macro regardless, so writing it once up front leaves each + overlay owning only its own `.c` files and `build//`. That cuts BUILDS, not just overlap. + +**Two traps this exposed, both worth checking in any tool about to run parallel:** +- **A fixed temp path** (`.run/dpcc/t.c`) is a correctness bug the day something runs concurrently — + the same fake-isolation class as `match_one`'s shared `--work` dir. Make it per-call. +- **A pool submitted all at once shares no learning.** Every worker got an empty suspect list and paid + a full bisection. Seed with one item in-process first, then fan out with the result. + +**Prove it, don't assume it:** the acceptance test was a *regression*, not a stopwatch — revert to the +pre-run state, re-run the identical command, and require the same functions and a byte-verified fleet +(R22). It came back faster AND with 62 more instances banked, which is how the over-exclusion in the +old path was found at all. See [[decomp-accelerator-ledger]] (A8) and `docs/accelerators.md`. + +## gate-main-only-with-gate-main + +--- +name: gate-main-only-with-gate-main +description: "main can ONLY be gated by tools/gate_main.py — parallel_gate/gate_stage return a FALSE PASS on main; count banks from the SOURCE; and a main gate now reports BODY vs PLUMBING reject (S72 correction: the 11 'PROVEN gate-rejects' were a missing rodata carve, not bad bodies)" +metadata: + node_type: memory + type: feedback + originSessionId: 9d0edc54-82de-4d02-9342-79a02497076d + modified: 2026-09-02T17:19:08.387Z +--- + +**`main` is gated ONLY by `tools/gate_main.py`** — baseline assert → substitute the whole slate → +`make extract` → `make build` → compare SHA1, with a bisect when the batch fails. `parallel_gate` now +REFUSES `binary == 'main'` in code, so this cannot be re-learned by accident. + +**Why:** main's `make extract` runs the EXE-only `psyq_integrate` + `ld_interleave` steps, which +REWRITE the linker script. `gate_stage` (and therefore `parallel_gate`, whose worker *is* +`gate_stage`) builds incrementally, so it re-runs that on an already-rewritten `.ld`. S58 recorded the +false-DIFF direction (105 competent main drafts thrown away). S71 hit the **false-PASS** direction, +which is worse: `parallel_gate` reported "11 banked" on main, the merge was committed, and R22 came +back 212/213 — the tree did not compile from clean, and once the declarations were reconciled it was +still not byte-identical. All 11 re-gated individually: **11 of 11 REJECTED**. + +**How to apply:** +* Any main draft → `gate_main.py`. Run `--assert-baseline` first; it is cheap and it separates "my + draft is wrong" from "the tree was already red". +* **Count banks from the SOURCE** (the `INCLUDE_ASM` stub is gone), never from the tool's slate. + `gate_main` printed "BANKED 5 of 6" when 4 had applied — `len(good)` is *what we decided to keep*, + not *what was substituted*. Fixed, but the principle is general: this was the 4th tool in one + session reporting a derived number as a measured one. +* A draft that **contains its own `INCLUDE_ASM`** is a silent no-op — substituting it restores the + stub, the build is trivially byte-identical, and it counts as a bank. `gate_main` refuses these now. +* **A tool that wraps another tool inherits its refusals.** Every constraint documented on + `gate_stage` binds `parallel_gate`, `harvest_verify`, and anything else that shells it — encode it + as a refusal in the WRAPPER, not a paragraph in the callee. + +**CORRECTED S72 (2026-09-02) — that "11 PROVEN gate-rejects" line was WRONG; NONE of them is a body +reject.** The S71 re-gate ran through an ad-hoc script (`.run/S71_main_bisect.py`), not `gate_main.py`, +so the decl pre-check never ran — and **all 11 are switch functions**. main had carried exactly ONE +`.rodata` carve since Phase 7, so a drafted switch DOUBLE-EMITS its jump table, the image grows +(+28/+52/+76/+84 measured) and ~332 symbols shift. Extending the carve to the contiguous span +`0x80072A38-0x80072C70` banked `func_8001A114`, `func_8001AAD0`, `func_8001AF34` **byte-identical in +14 s**; the other 8 sit in uncarved spans B/C and need `src/800.c` split. Cookbook §426/§427. + +**S75 CORRECTION (2026-09-02) — THAT VERDICT LINE HAD A CLASS THAT COULD NOT FIRE.** +`main_diff_locate.classify()` gained a `TABLE REJECT` case in S72 precisely because BODY/PLUMBING had +mislabelled a table failure. On main it was **unreachable by construction**: it summed bytes whose +object string contains `(.rodata)`, but main's `section_order` is `[.rodata, .text, .data, .bss]` — its +rodata sits BELOW `.text` and its jump tables live in **`.data`** objects +(`build/asm/data/63C4C.data.o(.data)`). It also demanded PURITY (`ro == outside`), so a few bytes of +perturbed code dropped the verdict through to PLUMBING anyway. + +Cost: **`SaveLoadRoutine` (1,165 ins — the largest open function in the project, 9.2% of all remaining +work, carried as the §434 WALL) has a BYTE-IDENTICAL body.** Gated alone, twice, the tool said +"PLUMBING REJECT … route to fix_arity_callers -> cast_self_callers"; the chain was run twice and fixed +nothing, because the real split is **3,787 of 3,989 bytes (94.9%) in `.data` jump tables vs 202 (5.1%) +in `.text`**, and the built image is **4 bytes SHORTER than retail** (§446's first diagnostic). +It is a CARVE — and it sits in `src/800_b.c`, i.e. **span B, exactly where the S72 note above predicted +the remaining 8 would sit.** Fixed: table bytes counted in `(.data)` OR `(.rodata)`, and the test is +DOMINANCE (>=60%) not purity, naming which part is carve and which is declaration. NC: 5 of 6 +pre-existing verdict shapes unchanged. Cookbook §447. + +**THE LAW, and it generalizes past this tool: a verdict class that CANNOT FIRE is worse than one that +does not exist** — it converts "I don't know" into confident, specific, wrong advice that then +consumes sessions. When a verdict names a subsystem, check that subsystem owns the MAJORITY OF THE +BYTES before acting on it. + +**ALSO S75: gate main with `--no-propagate`.** `gate_stage`'s tail runs +`dedup_propagate --auto-from `, which sweeps the WHOLE binary rather than the function just +banked; a binary with an unswept pile stalls the gate 30+ minutes (`ov_SC01_005` held 557). Drew's +decision 2026-09-02: leave the ~12,000-copy hygiene backlog (all already matched, orthogonal to +completion %) and gate `--no-propagate` by default. See [[dedup-backlog-leave-it]]. + +**A main gate now says WHERE, not just that.** `gate_main` preserves the red image + map under +`.run/gate_main_fail/` BEFORE the R40 baseline control rebuilds over them (that ordering bug is why +nobody could ever localize a main failure), then prints **BODY REJECT / PLUMBING REJECT / MIXED** via +`tools/main_diff_locate.py`. **Read that line before recording any main verdict.** + +Related: [[standalone-match-is-not-bankable]] · [[verify-blast-radius-not-just-defect]] · +[[silently-narrowed-tool-scope]] · [[check-against-a-known-true-case]] (cookbook §414) + +## mcp-renames-dont-persist-use-applysymbols + +--- +name: mcp-renames-dont-persist-use-applysymbols +description: "S78 — 47 MCP batch_rename/rename_symbol writes did NOT survive the sentinel stop (\"Save succeeded\", names gone); mirror symbols into Ghidra with tools/ghidra_apply_symbols.sh (headless ApplySymbols.java) and R9-verify; main's LINKED build is only exercised in-tree with .run/obj40 present" +metadata: + node_type: memory + type: project + originSessionId: 8ca6b019-9fd5-4215-9410-ad1d1220d5ee + modified: 2026-09-04T22:32:52.872Z +--- + +**Fact (2026-09-04, P31 S78):** 47 renames made through the headless GhidrAssistMCP server +(`batch_rename`/`rename_symbol`, all reported success) were GONE after `tools/ghidra_mcp_stop.sh` +("Save succeeded", DB grew 9.9→10.4 MB). R9's read-only re-open caught it. Cause not isolated (single +instance, normal stop sequence; Phase 3 had verified the same flow). The fix that worked first time: +`tools/ghidra_apply_symbols.sh [PROG] [symbols…]` → `tools/ghidra_scripts/ApplySymbols.java`, a headless +postScript that mirrors `config/symbols.us.txt` into the program with a real save (73 renamed, +R9-verified ×4). Also: gate worktrees carry no `.run/obj40`, so main's LINKED (SDK-object) build path is +only ever checked by an in-tree `make build BINARY=main` — it had been red at HEAD for a day unnoticed. + +**Why:** the curated text file is the source of truth (R15); an MCP write that never persisted looks +identical to one that did until a fresh open reads the DB (R9). A gate that always takes the fallback +path is blind to the path the contract cares about (R34). + +**How to apply:** for symbol renames, edit `config/symbols.us.txt` (no `name:` tokens in comments — +splat parses them as attributes), run `lint_symbol_refs.py` and read its WHOLE output (verbatim +`__asm__` bodies spell `\tfunc_X`), then `ghidra_mcp_stop.sh` → `ghidra_apply_symbols.sh` → +`ghidra_mcp_verify.sh `. After any change to `psyq_identify`/`psyq_integrate`/the splat +yaml: in-tree `make build BINARY=main` WITH `.run/obj40`, then the fresh-extract fallback WITHOUT it. +See [[exonerate-the-instrument]], [[check-against-a-known-true-case]]. + +## wsl-disk-capped-75gb + +--- +name: wsl-disk-capped-75gb +description: WSL2 ext4.vhdx is hard-capped at 75GB (~32GB headroom over 38GB live) as of 2026-08-29; .run/ churn can now hit ENOSPC +metadata: + type: project +--- + +The WSL2 disk (`/dev/sdd`, `ext4.vhdx`) was resized from WSL's 1 TB default down to +**75 GB (73 GiB usable)** on 2026-08-29. As of then: 38 GB live, ~32 GB free. +This is a HARD ceiling — ext4 cannot grow past it. + +**Why it was done:** the vhdx had bloated to 111.6 GiB while holding only 38 GB of +live data, leaving C: with 11 GB free. Root cause: WSL creates a 1 TB ext4 filesystem +inside a dynamically-growing vhdx, and ext4's Orlov allocator deliberately places each +new *directory* in a block group with above-average free space. The `.run/wave_*` +pattern (hundreds of new top-level dirs) scattered data across 8192 block groups — +live data was smeared over 820 GiB of address space in 1363 groups. A dynamic vhdx +can only release whole 32 MB blocks, and nearly every touched block still held some +live data, so `compact vdisk` reclaimed ~7 GB of a possible 66. Zero-fill and sparse +VHD both fail or are unsafe here. `wsl --manage --resize` was the fix: +resize2fs relocated 33 GB down from above the boundary and packed data into 553 groups. +Result: vhdx 111.6 -> 53.3 GiB, C: 11.2 -> 69.9 GB free. + +**Why it matters:** `.run/` is ~20 GB and grows every wave. With only ~32 GB of +headroom, a big fleet run can now fail with ENOSPC instead of the disk silently +growing. That failure will look like a mysterious mid-wave build error. + +**How to apply:** +- Watch `df -h /` before launching large waves; prune old `.run/wave_*` when tight. +- The vhdx will still drift 53 -> 75 GiB over time (spreading is confined, not fixed). + That is expected, not a new problem. +- Reclaim recipe (works now that data is packed): `sudo fstrim -av` -> `wsl --shutdown` + -> diskpart `compact vdisk`. Do NOT bother with zero-fill. +- Need more room? GROWING is the supported direction and is safe: + `wsl --shutdown` then `wsl --manage Ubuntu-24.04 --resize 150GB`. +- Sparse VHD (`--set-sparse`) is gated behind `--allow-unsafe` by Microsoft for + data-corruption risk — do not use it. + +Related: [[no-tmp-project-local-data]] (.run/ is the gitignored runtime scratch dir). + +## disc-dump-location + +--- +name: disc-dump-location +description: Where the BFM USA disc BIN/CUE dump lives on this machine (WSL path) for extraction +metadata: + node_type: memory + type: project + originSessionId: 675841e2-bd0f-49ee-bdad-bf3e3b42545e +--- + +The legally-owned BFM USA BIN/CUE dump is at `/mnt/z/Storage/git/BFM-decomp/Brave Fencer Musashi (USA)/` (Windows `Z:\Storage\git\BFM-decomp\Brave Fencer Musashi (USA)`). Drew corrected me here on 2026-06-13 after I guessed a wrong `/mnt/z/Games/Emu/...` copy. + +- **Track 1** `...(Track 1).bin` = 364,846,944 B — the MODE2/2352 data track holding the ISO9660 volume + every `.CD` file. +- Tracks 2–4 are CD-DA AUDIO (each `(Track N).bin` = 150-sector INDEX 00→01 pregap + audio). The 3 `.DA` ISO files (ST01_13A.DA=Track2, ST01_13B.DA=Track3, DUMMY_DA.DA=Track4) point into them. +- **Scope change 2026-06-13 (Drew):** stage ALL 4 tracks + the `.cue` to ext4 `disks/`, and extract the `.DA` files too. `extract.py` extracts them as raw 2352-byte/sector audio (sizes differ from the 2048-based ISO entry). Supersedes the earlier "Track 1 only / .DA out of scope" plan. +- Per H2 the dump is read once, **not built against** — copied to ext4 `disks/` (gitignored) for fast iteration instead of reading repeatedly over the slow 9P `/mnt` bridge. +- The raw dump is never committed (gitignore firewall, survives even the H1 relaxation in [[rom-content-git-policy]]). + +## agent-lane-cap-is-drews-and-changes + +--- +name: agent-lane-cap-is-drews-and-changes +description: the concurrent-agent cap is Drew's live dial (5 → 2 → 1 in P36 S103–S105) and "no more agents" ends the lane; obey the latest, record each retune in CURRENT_PHASE's decisions; throughput is bounded by the coordinator's landing loop, not the cap +metadata: + type: feedback +--- + +Drew retunes the agent cap mid-session and without ceremony ("max concurrent agents is now 2", "… of 1 now", "dont start +any new agents at this time. let the current ones finish", "no more agents this phase"). Each is binding from the next draw; +a resumed dead agent (SendMessage by id after a usage-limit cut) is not a new draw. + +**Why:** the cap is a budget he watches; in S105 a cap of two kept the coordinator's landing loop (verify → bank → harvest → +regen → gate → log) saturated — 56 of 56 draws closed — so the number did not depend on concurrency. +**How to apply:** record every retune in CURRENT_PHASE.md's decisions section and the checkpoint's §0; launch only up to the +latest cap; when he says stop, finish the in-flight agents, bank, gate, checkpoint — and answer a "stalling?" question with the +residue bucketed by what each bucket needs (drawable / parked / proven / structs-phase) before recommending a close. + +## autonomous-lane-architecture + +--- +name: autonomous-lane-architecture +description: the drafter/gater/maintenance/stallguard lane split for unattended decomp campaigns — drafting must never be stopped to ship a change +metadata: + type: project +--- + +Unattended campaign architecture built P31 S58 (2026-08-23/24), all `setsid`-detached with their own +restart loops: + +| lane | script | rule | +|---|---|---| +| drafter | `.run/drafter.sh` | draw → shard → draft → queue a ready marker. **NEVER stop it to ship a code change** | +| gater | `.run/gater.sh` | reloc pre-filter → gate → commit → harvest → ledger. Kill/restart freely | +| maintenance | `.run/maintenance.sh` | free A-prop sibling lane whenever the gater is idle; zero tokens | +| stallguard | `.run/stallguard.sh` | 60s: revive dead lane shells, kill agents silent >20min, kill gates >90min | + +**Why the split:** measured 53% drafting idle over 5 hours, and **139 of 162 idle minutes were the +supervisor being killed to pick up a code change**. Gate contention was 18 minutes. So a worktree +gate lane would have chased the small half — the fix was decoupling drafting's LIFETIME from the +process you need to restart. `build_wave_atlas` and `api_agent` are spawned fresh per wave and always +pick up edits; only the long-lived supervisor is ever stale. + +They share one narrow lock (`.run/auto/draw.lock`) serializing DRAW against GATE, because +`build_wave_atlas` reads `corpus.stubs()` and misreads substituted drafts mid-gate. **Drafting holds +no lock at all.** + +Stop cleanly: `touch .run/ox_campaign.stop`. See [[commit-banked-work-immediately]] for why the gater +commits a dirty tree instead of reverting it, and `phase-ends/CURRENT_PHASE.md`'s crash-recovery +checkpoint for restart order. + +## bank-idioms-before-checkpoint + +--- +name: bank-idioms-before-checkpoint +description: "\"Checkpoint\" means EVERYTHING perishable is already banked in files — idioms, tooling, docs, SETUP, decision-log, accelerators, stale assertions. Drew should never have to ask what's missing" +metadata: + node_type: memory + type: feedback + originSessionId: 60f56b49-c569-438e-a15e-489965eebee7 + modified: 2026-09-04T05:00:26.218Z +--- + +**When Drew says "checkpoint", the checkpoint is the LAST thing written, and by the time it is +written every piece of session context that would be lost is already in a file.** Not a plan to +bank it, not a commit message describing it — a file in the load order. He should never have to +follow up asking whether the idioms / tool docs / doc updates got banked. + +**Why this is stated so strongly (Drew, 2026-09-03, P31 S77):** I wrote a full checkpoint block, +declared it done, and only after he asked *"make sure all toolupdates doc updates and idiom harvests +are banked in files"* did a sweep find **four** missing things — the entire `-O0` carve procedure +(five coupled config/build pieces, byte-proven) existed **only in commit messages**; `docs/SETUP.md` +had no rows for two new tools (R21); there was no `docs/decision-log.md` entry (R31) for a session +that pivoted the whole phase strategy; and no `docs/accelerators.md` entry. Plus a stale caveat in +`progress.py` still asserting main had no independent oracle, hours after building one. + +**THE TRAP, and it is specific: a really good commit message feels like documentation.** I had +written long, careful commits with measurements, mechanisms and negative controls — which produced a +strong false sense that the knowledge was banked. It was not. `git log` is not in the session-start +load order. A fresh session reads `PROJECT_CONTEXT.md` → `phase-ends/PhaseEnd_*.md` → +`CURRENT_PHASE.md`, and consults `docs/`. **If it is only in a commit message, it is lost.** + +**The sweep, before writing the checkpoint block — every line, every time:** +1. **Idioms → `docs/matching-cookbook.md`** (+ `docs/gcc-2.7.2-map/*` for compiler internals). Sweep + three sources: what I learned debugging tooling/process; what AGENTS reported (mine + `journal.jsonl` — notifications truncate the long notes where the analysis lives); near-miss + diagnoses. Grep before claiming novelty; mark **wall refutations** and name the section they + overturn. Group by residual class/symptom, not by function. Include byte evidence. +2. **Every new tool → a `docs/SETUP.md` row (R21)**, and every CHANGED tool → its siblings and docs + in the same change ([[tool-change-updates-siblings-and-docs]]). +3. **Strategic pivots / dead ends / reversals → `docs/decision-log.md` (R31).** +4. **"This would have saved earlier work" → `docs/accelerators.md`** ([[decomp-accelerator-ledger]]). +5. **Procedures live in docs, not tool docstrings alone** — a manual carve/recovery sequence goes in + the cookbook or the playbook, even when a tool's refusal message names it. +6. **Hunt STALE ASSERTIONS the session invalidated** — a caveat, exclusion reason, "deferred" note, + or wall verdict that is now false. A stale assertion is the same class of lie as a stale wall + verdict, and it keeps being printed to future sessions ([[reprobe-exclude-lists-after-tool-fixes]]). +7. **Regenerate derived artifacts** (`tools/cookbook_index.py`, `progress.py --fleet`) and run + `make tools-health` — it gates the index and will catch you. +8. **Verify each section is present ONE AT A TIME, in HEAD, not just on disk** — §462/§463 silently + vanished after their commit in S76. + +Verification-layer lessons count as idioms: what a check can and cannot prove is exactly the +knowledge that gets lost and re-learned expensively. + +Related: [[matching-cookbook]], [[capture-knowledge-before-fresh-session]], +[[checkpoint-current-phase-before-pause]], [[decomp-accelerator-ledger]], +[[tool-change-updates-siblings-and-docs]], [[derive-from-invariants-not-reparsing]]. + +## bfm-decomp-context-system + +--- +name: bfm-decomp-context-system +description: BFM-decomp repo has its own session-governance system (CLAUDE.md → PROJECT_CONTEXT.md → phase-ends/) that supersedes ad-hoc memory; status P37 CLOSED 2026-09-29 (v2.3.0) RE-SCOPED to T0–T4 for the PA 3.0 upgrade; P38 = structs continued (T5–T10), seeded by PhaseEnd_Phase37 + logs/Phase37.md's S108 checkpoint +metadata: + node_type: memory + type: project + originSessionId: 998849ed-0de9-4b58-8dec-e32c194b8e8b +--- + +The BFM-decomp repo (Brave Fencer Musashi PS1 matching decompilation, SLUS-00726) governs all sessions through its own in-repo context system, generated 2026-06-10: `CLAUDE.md` (auto-loaded, mandates load order) → `PROJECT_CONTEXT.md` (permanent static constitution: 25 rules, protocols, 7-phase Gen1 roadmap) → `phase-ends/` (living state). **Defer to that system; do not duplicate its content in this memory directory.** + +Status (2026-06-13): **Gen1 Phase 1 COMPLETE** (`phase-ends/PhaseEnd_Phase1.md`, commit `ac51452`, project v1.1.0). The whole project now runs **all-in-WSL** on ext4 at `~/bfm-decomp`. RE stack live: JDK 21, Ghidra 12.1 PUBLIC + GhidrAssistMCP v2.8.0 + ghidra_psx_ldr 2026.06.04 (all in `~/ghidra_12.1_PUBLIC`). `SLUS_007.26` imported+analyzed into `ghidra/bfm.{gpr,rep}` (PSX loader, 1726 funcs, **PsyQ 4.0.0**, 2599 PsyQ types imported). Milestone met: in-session MCP `get_code(0x80018730)` = the LZSS decompressor. **Key operational discovery: the entire RE loop runs headless** — `analyzeHeadless` import, `.gdt` import via `tools/ghidra_scripts/ImportPsyqGdt.java`, the GhidrAssistMCP server (`-preScript GAMCPStartServerScript wait=true`, omit `-loader` and auto-detect), and MCP decompile/retype all work with no WSLg GUI; `get_code` is async (poll `get_task_status`). **Next: Phase 2** — deterministic disc & .CD extraction pipeline (ISO9660 walker first). Six rules formalized this phase (R1–R6 in PhaseEnd) — see [[rom-content-git-policy]], [[no-commit-co-author]], [[tools-folder-convention]]. + +Two corrections future sessions must not regress on (full detail in repo docs): BFM is PsyQ 4.0/4.2-era → GCC 2.7.2 family, NOT sotn's 2.6.3; the MCP stack is GhidrAssistMCP on Ghidra 12.1 + ghidra_psx_ldr, NOT LaurieWired GhidraMCP. See [[drew-working-preferences]]. + +**Updated 2026-09-07 (Phase 33.5):** the status block above is Phase-1 vintage. Current: **Phase 33.5 open** (docs consolidation, +the wiki as the source of truth, the tracked-scratch prune, the memory reconciliation, the day-one decomp kit `decomp-architect/`; +v1.32.1); **Phase 34 next** (the public flip + the Gen2 exit at v2.0.0, gated on GitHub Support purging the old objects); +**Gen3 opens at Phase 35** from `docs/gen3-handoff.md` + `docs/gen3-standards.md`. The load order is R64's: `PROJECT_CONTEXT.md` → +`phase-ends/DIGEST.md` (every rule R1–R83 in full) → the three most recent PhaseEnds → `CURRENT_PHASE.md`, whose 🛑 block is +replayed verbatim. The fleet is 218 binaries, all byte-identical; the matching frontier is empty. + +**Updated 2026-09-08 (Phase 33.5 CLOSED, v1.32.1):** the wiki is the source of truth; the day-one kit `decomp-architect/` is complete (three dictionaries, DK-1–80, G1–67, kit_coverage, five dry-runs); Phase 34 (the flip, v2.0.0, Gen2 EXIT) is NEXT and its gate is OPEN (the purge probe PASSED 2026-09-07) — plan it fresh in plan mode at Max from `docs/phase34-seed.md`; Gen3 opens at Phase 35 with the S91-b types doctrine as a seed. Session start = PROJECT_CONTEXT → DIGEST → the last three PhaseEnds (P32, P33, P33.5) → no CURRENT_PHASE until Phase 34 opens. + +**Updated 2026-09-08 (Phase 34 CLOSED, v2.0.0, Gen2 EXIT):** the repository is **PUBLIC** (flipped 2026-09-08 after the purge probe passed before and after; the `protect-main` ruleset forbids force-push/deletion of `main` — edit the ruleset first if a rewrite is ever needed); decomp.dev card + the wiki (32 pages, synced by `tools/wiki_sync.sh --push`, Drew's) live; `docs/outreach/` and `docs/sunset/` are ignored-local, `tools/sunset/` tracked. **Phase 35 = Gen3 opens** (fresh session, plan mode, Max) from `docs/gen3-handoff.md` + `docs/gen3-standards.md`: pins off → macro bodies → struct unification (R95) → names with evidence → formatting, every step gated by the 218 hashes; the kit split and xsig v2 (gen3-handoff §7) are Gen3 opening tasks. Pending on third parties: the decomp.me preset creation (#2106), an Archipelago reply (#1), frogress. Session start = PROJECT_CONTEXT → DIGEST (R1–R95 in full) → the last three PhaseEnds (P33, P33.5, P34) → no CURRENT_PHASE until Phase 35 opens. + +**Updated 2026-09-12 (Phase 37 S106):** Phase 36 (levers off) CLOSED 2026-09-11 at v2.2.0 — its close is a normal commit `79b2f6f15` (no tag, +no release, Drew's rule). **Phase 37 — the structs phase — is OPEN:** gate 1 approved (grind to zero; placeholders + cited evidence; the P36 +agent lane at Drew's cap; the full canonical declaration layer; R107–R117 ratified); T0 baseline, T1 the type census + the struct map +(`tools/type_census.py`), T2 the probe (`tools/restruct.py`) are committed; **T3 (the engine's full form, the linked-relocation oracle, +`struct_layout.py`, the canonical type writer) is NEXT** — its design brief is the 🛑 block of `phase-ends/CURRENT_PHASE.md`. The phase's +founding fact: in gcc 2.7.2 struct spelling MOVES bytes (`MEM_IN_STRUCT_P`), so every struct edit is gated like a match (cookbook §458). + +## breadth-isolated-agents-not-serial + +--- +name: breadth-isolated-agents-not-serial +description: "For breadth (matching many functions), isolated agents are cheaper than the main loop doing them serially — counterintuitive token economics" +metadata: + node_type: memory + type: reference + originSessionId: cd9272e2-a7d2-4e3e-8a74-c87aa9cacc6d +--- + +For breadth-shaped work (draft byte-exact C for N functions), **isolated agents (Ultracode Workflow / the +§12 harvest pattern) are token-CHEAPER than the main loop doing them serially**, even though it looks like +N copies of the same prompt. Drew asked this directly (Phase 17, 2026-06-20). + +**Why (the counterintuitive part):** the main conversation re-sends its ENTIRE growing context every turn, so +processing N functions serially here accumulates (function 1's asm/diffs/drafts are still in context at +function N) → cost ≈ quadratic (with caching, ~5-8M for 40 functions). Each spawned agent gets a fresh, tiny, +ISOLATED context (its one function) → ~linear, ~2.5M for 40, and parallel (minutes not hours). The redundant +per-agent doc-reading is real but secondary to the main loop's accumulation cost. + +**Corollaries:** +- Real token savings come from **trimming what each agent reads** (inline a tight idiom cheatsheet vs. having + it read full docs — the biggest free win), batching ~8 functions/agent, and pre-filtering trivials. +- The **compile/byte-gate is already 100% local/free** (gcc + `harvest_verify`); only the *drafting* costs tokens. +- **Fully-local drafting (no LLM) hits a ~3% ceiling** (the m2c+permuter pipeline, Phase 16; the permuter + doesn't even transfer to whole-binary, T6). The LLM is what turns 3% into ~33%. +- **Hand-by-main-loop wins only for a SMALL high-value set (≤5)**, like the Phase-17 demo's 4 — past that, + accumulation makes it lose to agents. + +Relates to [[effort-prompt-ultracode-on-breadth]], [[ultracode-harvest-pattern]]. + +## build-tasklist-after-plan-approval + +--- +name: build-tasklist-after-plan-approval +description: "After a plan is approved, build the harness task list BEFORE starting work so Drew can monitor progress" +metadata: + node_type: memory + type: feedback + originSessionId: d4ce0f78-4526-4f51-9cbf-997acb471e24 +--- + +After a plan is approved (the Phase Start gate / ExitPlanMode), build the harness **task list** (TaskCreate one task per plan task) **before starting the work** — so Drew has a live, monitorable view of phase progress. + +**Why:** Drew monitors progress via the task list (TaskList/spinner), not the prose plan or `CURRENT_PHASE.md`. He asked for this explicitly (2026-06-16, mid-Phase-13): "build tasklist so I can monitor progress." + +**How to apply:** Immediately after plan approval, create the task list from the approved plan's task breakdown (T0, T1, …), then mark the first task `in_progress` and proceed. Keep statuses current (mark `completed` as each gate passes; add follow-ups discovered mid-work). This is in addition to — not a replacement for — `CURRENT_PHASE.md` (the committable crash-recovery log per [[bfm-decomp-context-system]]). Related: [[drew-working-preferences]]. + +## capture-knowledge-before-fresh-session + +--- +name: capture-knowledge-before-fresh-session +description: "Write context-dependent artifacts (cookbook entries, byte-verified findings, codegen-map distillations) DURING the session that produced them, before any fresh-session handoff — the detail is irrecoverable otherwise" +metadata: + node_type: memory + type: feedback + originSessionId: 7b1a9abb-50c7-49be-b436-ca05e9ac0888 +--- + +When a session produces findings whose quality depends on its FULL live context — cookbook entries (`docs/matching-cookbook.md`), codegen-map / gcc-quirk distillations, byte-verified diagnoses, `docs/hand-matching-process.md` notes, R14 corrections to existing docs — **write them NOW, before ending or handing off to a fresh session.** A fresh session inherits only the compressed `CURRENT_PHASE.md` / PhaseEnd summaries and LOSES the rich detail: the exact asm diffs, the mechanism, the byte-evidence, the exemplar addresses, and the dead-ends already tried. + +**Why:** Drew (2026-06-21). In Phase 20 I proposed deferring T4 (the cookbook §20 distillation) to a fresh session "to continue with a clean head." Drew flagged that the next session would only have summaries and the nuance would be gone, so the cookbook entry would be thin or wrong. Doing it in-session produced a far richer, correct §20 (incl. the R14 loop-guard correction that a summary-only session could not have made). + +**How to apply:** At a session boundary, split the remaining work into — (a) **CONTEXT-DEPENDENT knowledge capture → do it NOW:** cookbook/findings/distillation write-ups, corrections to existing docs, the gotchas, the "why this is irreducible" verdicts with their byte-evidence; and (b) **MECHANICAL / continuable work → safe to defer** to the fresh session: build the next tool, run the next wave batch, the next harvest, a Ghidra-C regen. The handoff doc (`CURRENT_PHASE.md`) carries POINTERS + the headline; the durable DETAIL must live in the cookbook/docs, authored in-session. This is the same "mine-to-author with full context, NOT breadth — subagents/future-sessions have less context" principle (see `docs/effort-map.md`). Extends [[matching-cookbook]] (R16 flywheel — adds the TIMING constraint: capture before the context is lost) and the session-end discipline [[session-summary-plain-english]]. **RATIFIED as rule R30** (Drew confirmed 2026-06-21, PhaseEnd_Phase20). + +**Companion — R31 (CONFIRMED by Drew 2026-07-08, Phase 25): capture the STRATEGIC "why" too.** R30 covers TECHNICAL artifacts; R31 adds the perishable *judgment* behind big PIVOTS/dead-ends/reversals → `docs/decision-log.md` (append-only, forward-only, one entry per new strategic turn: context+belief → what failed → the pivot → the byte/measurement why → a hindsight "better path" note). It is the substrate for the eventual project **retrospective** ("with hindsight, the best way to have done this") + the public **"how to AI-decomp a brand-new project"** wiki at the public flip (Drew's long-term goals, stated 2026-07-08). Route technical idioms to the cookbook; route direction/judgment to the decision-log. The quantitative curve (fleet % over time) is safe in git+PhaseEnds; the reasoning is what evaporates — so log it live. + +## carve-state-files-never-blanket-add + +--- +name: carve-state-files-never-blanket-add +description: config/overlays.mk + splat yamls are carve STATE shared by parallel gates — never `git add` them blanket after a gate, never restore them whole-file; audit with interleave_check/pads_audit first (my own R59 slip, S62) +metadata: + type: feedback +--- + +After a gate run, `git add config/overlays.mk` blanket-committed another lane's stale carve lines +twice in one session (ov_SC03_108/ov_SC06_011 → the next clean sweep failed 211/213). The mk and +each splat yaml describe ONE fact (where carved jump tables sit) and drift the moment either is +written or restored alone. + +**Why:** parallel gates (`sweep_parallel -j N`) each read→modify→write the shared mk; a failed +draft's carve can leave its block behind while its yaml is restored. `git add -A config/` then +promotes the drift to a commit that the incremental gates never notice. + +**How to apply:** before committing config after any gate: `tools/interleave_check.py ` on every +binary whose block changed (`--fix` regenerates the order from the yaml) and `tools/pads_audit.py +` for pads; add only those binaries' files. Never read `asm/` while a fleet sweep's extract-all +runs (probes see vanished .s files). See [[verify-blast-radius-not-just-defect]], +[[silently-narrowed-tool-scope]]. + +## NUANCE ADDED 2026-09-01 (P31 S69): the rule does NOT cover a carve's OWN new source file + +"Never blanket-add" is about SHARED carve state — `config/overlays.mk` and the splat yamls. It does +**not** extend to the new `src//_jr_.c` a jtbl carve creates when it splits a TU. +That file is per-binary, is named by a committed `config/splat..yaml`, and its 31 sibling `_jr_` +files are tracked — so it MUST be adopted with the bank that created it. + +`parallel_gate` was silently dropping exactly those: its merge-safety check compared `git show`'s +stdout (`""` when the path is not at the pin) against `None` (absent from the main tree), so a file +present in NEITHER read as "main tree moved under them" and was refused. Eight accumulated in one +session; nothing failed locally (the file is on disk, R22 green) but a fresh clone would get the yaml +without the source. Fixed by using `git show`'s return code; new adoptions are now printed, never +silent. Cookbook §400. + +## cheap-tier-ab-validated + +--- +name: cheap-tier-ab-validated +description: Empirical A/B — cheap drafter (Haiku) under Opus orchestrator matches Opus on small fns at ~5× lower $/match; reserve Opus for the big regalloc/schedule tail +metadata: + node_type: memory + type: project + originSessionId: b066d4b7-4278-4cc8-9f7f-39a30f0e0b9b +--- + +2026-06-29 experiment (Drew's "smart conductor, cheap players" idea): under an Opus orchestrator, +fan-out drafter agents on a CHEAP model instead of Opus. Forked `worker_wave.js` → +`tools/workflows/ab_match.js` (per-arm `{model,effort}`); scored disk-truth with `tools/ab_score.py` +(re-runs `match_one`, ignores agent self-reports). 20 frozen `reach1` targets (18–108 ins), opus arm +vs haiku arm. + +**Result:** opus 15/20, haiku 10/20. Haiku's matches ⊂ Opus's. **Clean size cliff:** ≤52 ins → +Haiku == Opus (parity on the bulk); 94+ ins regalloc/schedule/struct tail → Opus only, Haiku 25–101 +off. Measured cost: opus arm $35 vs haiku $4.9 → **haiku ~4.8× more matches per dollar.** + +**Why it matters:** validates the cost-escalation ladder — cheap drafter owns the ≤~50-ins bulk, +reserve Opus (+permuter) for the 90+-ins tail and permuter-close near-misses. The SHA1/match_one +gate makes a weak drafter a throughput risk only, never correctness (2 crashed Haiku agents still +left byte-perfect drafts the scorer recovered; a Haiku self-report `near 5` was really `near 72`). + +**Caveat (whole-binary gate):** scores above are the `match_one` PROXY. Through the real SHA gate +(`gate_stage`) only **4 of Opus's 15** banked (`func_80138BE0 func_8012F828 func_8012E9C0` _a + +`func_80141100` main, @3504773c); the other 11 are self-match-but-gate-rejected TU-plumbing walls +(the `plumbing_blocked()` phenomenon) → backlog fuel. So `match_one` over-counts bankable yield on +reach1/split fns; the cheap-vs-Opus *drafting* delta holds, but route/measure on BANKS not the proxy. + +**Local-model verdict (2026-06-29):** a STOCK local model fails as a matching drafter. Qwen3.6-35B-A3B +(LM Studio, served on LAN to WSL via "serve on local network" + Windows host IP) via `tools/api_draft.py` +(provider-agnostic OpenAI-compatible worker): **0 reliable byte-matches** on the 20 reach1 fns. A fair +harness (inline common.h + live cookbook + corpus examples) FIXED compile-fails but the model stayed +**stuck at fixed near-misses** across all 4 diff-feedback iters — it gets structure right, misses +gcc-2.7.2 precision (byte-offset scaling, lh/lhu, frame size), can't refine from the diff. FULL cookbook +was WORSE + 2.3× slower than a curated subset (dilution, gate-confirmed — more context is not the lever). +So Drew's "if Haiku can, local will" was disproven by the gate: Haiku (frontier small) reliably matched +the bulk; stock local ~0. Forward: the LoRA specialist (corpus built: `tools/export_pairs.py` → +`datasets/match_pairs`, 1307 pairs) OR a permuter-seed role (model's structural near-miss → permuter +finishes precision). Cloud Haiku/GLM stays the working cheap tier. See [[gen2-mips-matching-model]] +(docs/gen2-mips-matching-model.md). + +**LoRA pilot RESULT (2026-06-29, NEGATIVE on meaningful fns):** trained Qwen2.5-Coder-7B QLoRA (Unsloth, +local on the 3080 Ti via WSL; tools/format_finetune→train_lora→eval_lora) on 638 compile-filtered pairs. +Held-out gate-true eval: ≤5-ins trivial 39/41 (95%) but **≥6 ins 0/34, >15 ins 0/24** — same as stock; +near-misses FAR (near≈nins, all-instructions-wrong, 1/33 within 5). Memorized the trivial leaf pattern, +did NOT learn matching. Root cause: the compile-filter dropped the 536 harder (global/struct) functions → +corpus starved of non-trivial signal. NOT one-epoch-away. Fix = corpus-v2 (self-contained completions +with externs → recover the hard fns + make them trainable) + likely a bigger base; uncertain payoff given +how far off. (Also learned: GGUF conversion writes ~30GB intermediates → filled C:/WSL-vhdx → crash; keep +only the q4 gguf, clean intermediates.) Pragmatic cheap tier remains cloud Haiku/GLM. + +**Corpus-v2 RESULT (2026-06-29, POSITIVE):** the fix was data, not model. `export_pairs` now captures the +`extern D_xxx;` block the src declares above each def (correct types) → self-contained completions, +compile 52%→92%, train 638→1111 (non-trivial 257→813). SAME 7B retrained, held-out eval: 6–15 ins +**0%→85%**, non-trivial 0→26, meaningful(>15) 0→3. **Corpus quality was the bottleneck, confirmed.** A free +local 7B now byte-matches trivial+small-medium (≤15 ins) at 85–93% — a real Tier-0 for the bulk, rivals +Haiku at $0. Limits: ≥16 ins falls off (7B capacity), giants compile-fail (need struct types too = corpus-v3). +Decision gate = GO: scale to cloud dense 14–32B to extend the band. Caveat: eval is held-out BANKED +(objdump fmt); production on OPEN stubs needs .s-format alignment (spimdisasm). Serve via LM Studio (GPU) — +Unsloth's llama.cpp is CPU-only. Tools: format_finetune→train_lora→eval_lora. + +Next: (a) cloud dense 14–32B train on v2 corpus (extend band), and/or (b) corpus-v3 (struct types → recover +giants), and/or (c) .s-format alignment to deploy on real open stubs. GLM-5.2 cloud remains the no-train +cheap path. Details: `docs/gen2-mips-matching-model.md`, `docs/history/cheap-tier-ab-experiment.md`. +Relates to [[ultracode-harvest-pattern]], [[breadth-isolated-agents-not-serial]], [[matching-cookbook]]. + +**LADDER SUPERSEDED (Drew, 2026-08-03):** this file's two-tier rule ("cheap drafter ≤~50, Opus for +the 90+ tail") left the ~50–120-ins band unassigned, and P30 S7 measured the cost of that gap — +Haiku-direct banked 3/8 there while Opus-escalation-after-a-Haiku-miss banked 10/11, i.e. Haiku was +acting as expensive triage. **Sonnet is now a required middle rung.** See [[subagent-model-ladder]] +for the current routing; the economics and the ≤52-ins parity finding below still stand. + +**Updated 2026-09-07 (Phase 33.5 memory reconciliation):** this file is the A/B RECORD (the 2026-06-29 experiment, the local-model +and LoRA verdicts); the routing it proposed was superseded twice — see [[subagent-model-ladder]] for the current two-tier rule +(the cheap-tier cliff measured at ~30 instructions, then "no Sonnet: opus ≤150 ins, fable >150"). Its lesson that transfers to the +kit: route by MEASURED difficulty on your own corpus, and re-measure the cliff — the first guess was wrong by a factor of two. + +## check-against-a-known-true-case + +--- +name: check-against-a-known-true-case +description: "before believing your own scan/join/verdict, run it on ONE case whose answer you already know — five of S69's biggest 'findings' were artifacts of my own instrument, each caught only this way" +metadata: + node_type: memory + type: feedback + originSessionId: 9a451707-bd99-4f42-831f-2fb674555ed8 + modified: 2026-09-01T20:50:02.834Z +--- + +**Every scan, join, census and verdict gets checked against one case whose answer you already know, +BEFORE it is reported or acted on.** + +**Why:** in P31 S69 alone, five confident, true-looking numbers were artifacts of my own instrument: + +| the "finding" | the actual defect | +|---|---| +| "atlas covers 8% of open stubs" | joined on NAME; sig/feat write lowercase hex, splat symbols uppercase — the 108 "hits" were exactly the addresses with no A-F digits | +| "0 re-gate candidates" | wave name derived as `S69o1_` instead of `S69o1`; the real answer was 24 | +| "~110 functions carve-refused" | `.run/sig..jsonl` is gitignored, so no worktree had it — every carve read UNOWNED; 13 banked in 400s once fixed | +| three separate "red binaries" | build-only check after a carve-bearing gate; `make extract && make build` was BYTE-IDENTICAL every time (§384) | +| "gate still running" (5 hours) | `pgrep -af` pattern matched the wait loop's own argv; it waited for itself | + +Each was caught by exactly one habit: **testing the instrument on a case with a known answer.** The +8% join died the moment I checked a binary whose stubs I could count by hand; the false zero died +because I knew `func_8018000C` was open and matched. + +**How to apply, cheaply:** +* a join → run it on one pair you can verify by eye, and assert the intersection is non-empty (R32); +* a census → reproduce one row by hand before quoting the total; +* a red binary → re-verify with `tools/verify_binary.py` (always re-extracts, §384) before reverting + ANYTHING; twice this session a false red cost legitimate work that had to be restored; +* a process check (`pgrep`, `ps`) → read the ROWS, never a count, and exclude your own command; +* an agent's claim → ask which command produced the number (accelerators #18). + +**Corollary: non-reproduction is a finding.** When you cannot reproduce someone else's number, say so +and ask for the exact file and command rather than assuming your setup is at fault — that exchange +produced a correct self-correction AND a diagnosis of my failed repro, to the instruction. + +R40 says exonerate the instrument before blaming the subject. This is the cheap, mechanical way to +actually do it. See [[silently-narrowed-tool-scope]], [[exonerate-the-instrument]], +[[verify-blast-radius-not-just-defect]]. + +## checkpoint-current-phase-before-pause + +--- +name: checkpoint-current-phase-before-pause +description: "before pausing, write/REFRESH the 🛑 SESSION CHECKPOINT block in CURRENT_PHASE.md — VERBOSE and SELF-SUFFICIENT, because the next session replays it VERBATIM into its chat and inherits nothing else (Drew 2026-09-05); stale is worse than absent; commits are not a substitute; EVERY bank commit refreshes the headline, and past ~85% context checkpoint BEFORE any long background wait (S94 died at 91% waiting on a 9-minute run, 2026-09-08)" +metadata: + node_type: memory + type: feedback + originSessionId: 2dbcb0e2-0d00-4abb-a93a-eb3c97959587 + modified: 2026-09-08T20:00:00.000Z +--- + +Before pausing — and especially before recommending "a clean stopping point / checkpoint here" — ALWAYS +first write a full **🛑 SESSION CHECKPOINT** block into `phase-ends/CURRENT_PHASE.md` and commit it: tree +state + HEAD hash, the 3 fleet metrics, what THIS session delivered, and the single NEXT task (+ where the +design/pickup lives). The checkpoint block is the LAST write before any pause/recommend message, in the SAME turn. + +**Why:** a checkpoint recommended only in the chat message is ephemeral — a fresh session reconstructs state from +`CURRENT_PHASE.md` (the CLAUDE.md load order), not from the prior chat. Drew had to warm the cache with ~840k +tokens just to ask me to write the checkpoint I had only *recommended* in chat (2026-07-20). Never make the human +spend tokens asking for a checkpoint that should already be on disk. + +**A CHECKPOINT IS NOT WRITE-ONCE — IT GOES STALE (Drew, 2026-07-25, after I broke this).** In a long +session I wrote a checkpoint mid-run, then worked for hours past it (4 agent waves, 2 banks, 3 behemoths, +6 new cookbook sections) across at least three stopping points — including two where I said "this should +stop" — and never refreshed it. The file still claimed two searches were "in flight" that had finished +hours earlier. A fresh session would have been sent at the wrong targets with none of the new work. +**Every pause needs a CURRENT checkpoint: refresh it, or append a new block that explicitly SUPERSEDES +the previous one.** Stale is worse than absent — absent makes you go look; stale makes you act wrongly. + +**"I'm committing after every finding" is NOT a substitute — this was my exact rationalisation.** The +per-task log preserves *content*; the checkpoint block is the *entry point*. 36 good commits still left the +entry point pointing at finished background jobs. Both are required, they are not interchangeable. + +**The trigger, mechanically:** any turn where you hand control back at a natural boundary — a task finishing, +a background job landing, the end of a work item — and ALWAYS before any sentence resembling "good place to +stop" / "this should stop" / "I'll pick up when X lands" / "recommend a fresh session". In sessions with long +background waits, each wait-completion is a natural checkpoint moment; do not batch them to the end. + +**How to apply:** a per-task log line is NOT a checkpoint. Write the prominent, self-contained block a fresh session +can act on immediately (mirror the earlier "🛑 SESSION-N CHECKPOINT — safe to open a FRESH session here" blocks: +tree clean / 140/140 / dedup+C1 / HEAD / fleet %s / this-session summary / NEXT task / design-doc pointers). Commit it. +Only THEN write the pause/recap message. Extends R30 (capture-before-fresh-session) with the specific mandatory +artifact (the checkpoint block) and its timing (before the pause). Related: [[capture-knowledge-before-fresh-session]], +[[session-summary-plain-english]], [[commit-per-task-after-phase-log]]. + +**THE THOROUGHNESS BAR (Drew, 2026-09-04, S78: "when I ask for a checkpoint, be this thorough").** When Drew +asks for a checkpoint — especially "for a fresh session" — the block is the *complete seed* for the next +session, verbose by design, because the next session replays it in full and inherits nothing else from the +chat. The S78 FINAL block in `phase-ends/CURRENT_PHASE.md` is the template. It carries, in this order: +1. **Header**: what it supersedes, phase/task state, HEAD hash + unpushed count, model/effort doctrine per + task, standing constraints (no waves, no trailer, MCP state + the `/mcp` prompt, DB staging decision). +2. **Verified-at-close**: the literal gate results (R22 fleet count, main's SHA with AND without SDK objects, + tools-health, lint) with the log paths. +3. **The session in one paragraph.** +4. **The census** — every number with its denominator (R41); every open item NAMED (binary:function, size, + class, blocker), not counted. +5. **The task list** with status, commit hashes for done tasks, effort per task, the single NEXT task. +6. **What this session established that must not be re-derived** — facts, names applied, tool changes + (flags, new scripts, what each fixed), hazards with their signatures ("a red build looks like X"). +7. **The NEXT task's design brief** — targets with exact addresses/sizes/sections, the probe results, the + mechanism to build, the files/tools to touch (with line-level pointers), the gate sequence, what is + expected to stay a wall. +8. **Carried context for every remaining task** (tools to use, laws that apply, leads to chase). +9. **Habits this session paid for** + the plain-English recap (R18). +Write it as if the transcript will be lost (it will), verify every number against the tree before writing +it, and commit it as its own commit. Related: [[bank-idioms-before-checkpoint]] (the perishable knowledge +must ALREADY be in the docs when the checkpoint is written — the block points at it, it does not replace it). + +**THE VERBATIM-REPLAY CONTRACT (Drew, 2026-09-05, P32 T3 checkpoint).** The next session's Session Start Protocol +(CLAUDE.md, R64 candidate) DUMPS THE WHOLE 🛑 BLOCK INTO ITS CHAT VERBATIM. It reads PROJECT_CONTEXT + the +`phase-ends/DIGEST.md` (rules + phase synopses) + the last three PhaseEnds + CURRENT_PHASE, and inherits NOTHING else +from the previous chat. So the block must be written for a reader who has never seen the session: include **all** the +context the next session will need, verbose by design, every path/command/hash/sha spelled out, never "as above" or +"see the chat". Concretely, beyond the S78 template above, the block carries: +- **a narrative of what happened** (in order, with times and commit hashes) — including what went WRONG (an overflow, + a swept directory, a false verdict) and what the successor did about it; +- **the exact invocations** of every tool the next step uses (flags, `--tu` for main, the splice one-liner, the slate + JSON shape) and the **gotchas** that bit (e.g. `set -o pipefail` + `grep -c` exits 1 on zero matches; `rtu_match` + needs `--tu` for main; `masked_diff` probes land in `src/`); +- **an inventory of the scratch/ledger files** the work depends on (what each file is, whether it is git-tracked) and + where the agent transcripts/prompt templates live; +- **carried context for every remaining task**, not just the next one (pre-read results, design briefs, paths); +- **the environment** (fleet count and the shas of touched binaries, hooks that started servers, concurrency caps, + effort per task, who pushes). +Verify every number against the tree before writing; commit the block as its own commit; then refresh it again after +any further work in the same session (this session wrote it twice: after the recovery, and after the protocol change). +**If a session dies WITHOUT a checkpoint** (context overflow mid-wave, 2026-09-05), the successor writes it FROM THE +TRANSCRIPTS before doing anything else: dump `~/.claude/projects//.jsonl` to a condensed text +(assistant text + tool calls + truncated results), harvest subagent verdicts with `tools/agent_verdicts.py`, rebuild +missing drafts with `tools/agent_drafts_restore.py`, RE-VERIFY every claim against the bytes (`rtu_match` in the real +TU), then write the block and mark it "written by the successor session". Related: [[session-start-list-rules-in-full]]. + +**THE BANK-COMMIT GAP (S94 → S95, 2026-09-08).** S94 followed the rule at every TASK close (the block was refreshed at +T0, T1, T2, T3, T4 — five times in one session) and still died without a usable checkpoint: T5's two bank commits +(the tool + probe, bucket 0's first pass) carried NO log line and NO headline refresh, then a 9-minute background run +was launched at ~88% context; the run's completion notification, Drew's "set up a 92% hook" and Drew's "checkpoint +for a fresh session" all arrived to a session that could only answer "Prompt is too long". 35 minutes of T5 state +existed only in the transcript. Two mechanical rules follow: (1) **every commit that advances a task — an intra-task +bank included — updates CURRENT_PHASE.md in the same commit** (a log line + the 🛑 headline); a checkpoint older than +HEAD is a dead session's checkpoint. (2) **Past ~85% context, write the checkpoint BEFORE launching any long background +job** — the wait itself consumes nothing, but the notification + the reads that follow it do, and a full context cannot +even acknowledge the completion. **The recovery itself (S95):** Drew's scope was "capture only — no gates/builds"; +the RE-VERIFY step above therefore ran READ-ONLY (the tool's own batch JSONs + per-binary check logs + `git diff` + +the dead session's scratchpad `cc1.err` were the oracles), and the new block states explicitly which claims rest on the +tool's success lines and which gate is still OWED (R22). The successor also finds the previous session's untruncated +tool outputs in its scratchpad (`/tmp/claude-1000///scratchpad/`) — a replay's `cc1.err` there gave the +real cause a truncated transcript line had hidden. Related: [[commit-banked-work-immediately]], [[commit-per-task-after-phase-log]]. + +**Added 2026-09-12 (Drew: "update checkpoint memory to do this at the end of each session"):** the checkpoint procedure has a story step. At the +end of EVERY session — before the 🛑 block is rewritten — advance the post-100 % narrative from that session's decision-log entry: +`docs/story.md` §10 ("After 100 %: making the code say what it means", one paragraph per phase, advanced in place) and `docs/retrospective.md` §7 +(the four questions — believed / failed / cost / sooner — for Gen3 so far), then regenerate `tools/timeline.py` (its lower panel draws the lever +and readability series; `--check` is in tools-health). The rule is also in `docs/wiki/Docs-and-scratch-conventions.md`. Why: the story and the +retrospective had stopped at Phase 33 while three Gen3 phases ran; the timeline's axes read 100 % forever — the post-100 % work was being +recorded (series, decision log, PhaseEnds) but not TOLD, and Drew asked. + +## clarify-misconception-before-costly-action + +--- +name: clarify-misconception-before-costly-action +description: "Correct a misconception in Drew's request BEFORE taking a costly/hard-to-reverse action; BFM tooling stays as git submodules, not vendored" +metadata: + node_type: memory + type: feedback + originSessionId: d7fbb085-1bba-4fbd-b42c-c8dd33a41230 +--- + +During Phase 4, Drew said git submodules meant "downloading and pushing multiple other repos" and asked to absorb the four tools (maspsx/asm-differ/m2c/decomp-permuter) directly into `tools/`. I vendored them immediately — which staged **1458 files** (m2c alone is ~1296). But the premise was a misconception: submodules are lightweight pointers (`.gitmodules` + 4 one-line gitlinks, ~5 entries), they never copy or push the other repos' code into our repo. Drew balked at the file count and we reverted to submodules — a wasteful vendor→revert round-trip. + +**Why:** acting on a misconception-driven directive before correcting it cost real churn (and nearly polluted the just-committed clean Phase-4 commit). + +**How to apply:** when a request's *rationale* looks like it rests on a factual misunderstanding, state the correction FIRST and get a confirm, before executing a costly or hard-to-reverse change. A 30-second clarification beats a multi-step undo. + +**Decision outcome (durable):** BFM third-party tooling is stored as **git submodules**, NOT vendored — Drew is sensitive to repo bloat. See [[drew-working-preferences]]. + +## commit-banked-work-immediately + +--- +name: commit-banked-work-immediately +description: R42 — commit banked functions the moment they exist; never blind-revert a dirty src/ (gate_main destroyed 61 byte-proven banks) +metadata: + type: feedback +--- + +**R42 (accepted by Drew 2026-08-23, binding).** A gate that banks with `commit=False` — +`sweep_parallel`, `gate_stage --no-propagate` — leaves REAL, byte-proven functions uncommitted in +`src/`, and no tool can tell them from residue. So: (a) a lane that banks **commits before handing +control to any other lane, tool or step** — immediately, not at wave end, not "after main"; +(b) a tool finding a dirty tree **commits it or refuses** — `git checkout -- src/ config/` as +tidy-up is forbidden. + +**Why:** three instances in one session (P31 S58). `ox_campaign.gate()` and `idiom_serial` both +opened with a blind revert — caught before firing. `gate_main.py` then did it for real: it +substitutes into `src/` and reverts on a failing batch, and could not distinguish its own +substitution from the 61 overlay functions `sweep_parallel` had just banked with `commit=False`. +It reverted **all 61 back to `INCLUDE_ASM` stubs** after 95 minutes of bisecting. The byte-gate +decides whether work is CORRECT; git is the only thing that makes it DURABLE, and the window +between them is where work dies. + +**AND IT BIT ME AGAIN, BY HAND, IN S72 (2026-09-02) — the rule needs sharper wording.** `gate_main` +reported *"BANKED 2 of 3"*; I left the two in the tree and ran ONE more gate. `try_batch`'s opening +`git checkout -- src/*.c` reverted both, and my next commit captured only the third. **I reported 14 +banks to Drew; the source said 12.** I had even written this exact hazard into cookbook §431 an hour +earlier — for DECLARATION edits — and did not carry the reasoning across to banks. + +So state it as a sequencing rule, not a hygiene one: **commit before the next command that can touch +`src/`** — and a gate IS such a command. "Immediately" is ambiguous enough that I read it as "at a +good stopping point"; it is not. + +**The oracle that caught it:** counting banks from the SOURCE (the `INCLUDE_ASM` stub's absence), +never from the tool's report or from my own account of what I had done. Run that count before +reporting any bank total — see [[gate-main-only-with-gate-main]]. + +**How to apply:** when writing or reviewing any lane that touches `src/`, find the revert paths +first. If one exists, ask what else could be uncommitted at that moment. Chunk long +bisect-on-failure batches (`gate_main` at 8) so a poisoned batch can't hold the tree hostage. +Related: [[verify-blast-radius-not-just-defect]], [[tool-must-refuse-unsupported-input]]. + +## commit-message-from-tool-output + +--- +name: commit-message-from-tool-output +description: "Write a bank/ledger commit message FROM the tool's printed outcome, never before it — two S83 ledger commits claimed a bank that had not happened (helper ran with an empty function list, then the wrong draft dir)" +metadata: + node_type: memory + type: feedback + originSessionId: 491895ad-3c84-4037-b04f-bf7e5ee16a0c + modified: 2026-09-06T05:42:11.111Z +--- + +Never write "BANKED"/"done" into a commit message (or a status line) before the tool that does the work has printed its outcome. In S83 (2026-09-05) two ledger commits (`49c81c33c`, `7a8e75881`) said `func_800CD674` was banked; the bank helper had been called once without the function name (it iterated over nothing, built the unchanged tree, and exited 0) and once with the wrong draft directory (it refused). The real bank landed two commits later with a correction note. + +**Why:** a ledger that says "banked" is trusted by the next session and by Drew (R42/R58 "count banks from the SOURCE"); a premature message is a false record that survives in git history. The helper defect (silent no-op on empty input) is the R43 class and was fixed, but the message habit is mine to fix. + +**How to apply:** chain `tool && git commit` on the tool's SUCCESS path only; grep the tool's output for the exact success line (e.g. `BANKED … BYTE-IDENTICAL`, `sha X vs X`) before composing the message; when a ledger commit follows a bank, verify the bank commit exists (`git log -1 -- `) first. Related: [[bank-idioms-before-checkpoint]], [[silently-narrowed-tool-scope]], [[verdict-names-its-instrument]]. + +## commit-per-task-after-phase-log + +--- +name: commit-per-task-after-phase-log +description: "Drew's preferred commit cadence (2026-07-04) — one commit per completed task, made AFTER updating CURRENT_PHASE.md so task work + its log line ship together; Drew pushes" +metadata: + node_type: memory + type: feedback + originSessionId: 6b2493cc-eeaf-4300-b2e0-fc818626ff80 +--- + +Drew's preferred commit cadence, stated 2026-07-04 for Project Architect 2.0 and future projects: **one commit per completed task**, and the commit is made **after CURRENT_PHASE.md (the phase log) is updated** with that task's completed work — so the task's changes and its log entry land in the same commit. Claude commits locally; **Drew pushes**. + +**Why:** Per-task commits give fine-grained history and crash recovery; updating the phase log first means every commit is self-describing and the log never lags the code. (This restores bfm's original P4 over the later R8 phase-end batching, which Drew decided against.) + +**How to apply:** Task done → update CURRENT_PHASE.md progress log → `git add` the task's files + the log by explicit path → commit with a message naming the task → never push. No Co-Authored-By trailers ([[no-commit-co-author]]). + +## continuous-gater-lane-plan + +--- +name: continuous-gater-lane-plan +description: "BUILT and in use as tools/gater_lane.py (was: agreed plan, Drew 2026-08-31): a CONTINUOUS GATER LANE draining a queue — drafting streams, the gater groups by binary and fires parallel_gate; twin_sweep/harvest/propagation stay PERIODIC because they need aggregate" +metadata: + node_type: memory + type: project + originSessionId: 5e7f4e3a-4e31-46cc-a416-6baaccdac66a + modified: 2026-08-31T21:07:15.711Z +--- + +**AGREED 2026-08-31 ("that sounds good. lets try that next session").** Now that gating is fast +(see [[gating-speed-playbook]]), banking can run continuously alongside drafting instead of after it. + +## The arithmetic that makes it work (measured S67) +* **Production:** 20 concurrent single-function workflows, 2-30 min each → a finished draft roughly + every **30-90 s**. +* **Consumption:** 13 binaries / 139 s at 12 workers = **~10.7 s per binary amortized**; jtbl + 14 / 188 s at 8 workers = **~13.4 s**. The gate absorbs **~5 functions/minute**. +* **→ 3-5x headroom.** Gating is no longer the constraint. + +## The shape: a QUEUE-DRAINING GATER, not per-function gating +Drafting streams at N slots. A gater lane wakes on new drafts, **groups them by binary**, fires +`parallel_gate`, commits. Neither lane waits for the other. + +**Why NOT gate strictly per completion — three measured reasons:** +1. **Same-binary drafts must share a build.** S67 had `ov_SC05_010` ×3 and `ov_SC03_105` ×2 in one + batch; gating each alone triples the build cost for nothing. Accumulate: fire when **3-5 drafts + land or ~60 s pass**. +2. **Propagation is cross-binary and still unbatched** — a build-per-candidate loop that re-walks the + fleet each invocation. Keep it BATCHED until restructured; it is the one remaining serial cost. +3. **`twin_sweep` and harvest need AGGREGATE.** `twin_sweep` is a fleet-wide scan (it returned 0 + twice in a row in S67 — per-bank would be pure waste). Harvest is worse: **§330 existed only + because four independent instances appeared in ONE wave**, and §347 became a rule at its THIRD + instance. Per-function harvest is structurally blind to those. + +## Cadence +| step | when | +|---|---| +| gate (`parallel_gate`, grouped by binary) | continuously, 3-5 drafts or ~60 s | +| `twin_sweep` | every ~10 banks | +| harvest → cookbook | every ~10 functions | +| `dedup_propagate` | batched, NOT per gate | +| R22 clean-fleet | after ANY propagating run | + +This is the shape the retired `drafter`/`gater`/`maintenance`/`stallguard` lanes were reaching for +(`tools/lanes/`, dead since the OpenRouter era) — the difference is that the gate is now fast enough +to make it real. Do NOT resurrect those scripts; they are OpenRouter-era +([[wave-playbook-is-the-procedure]]). + +Related: [[gating-speed-playbook]] [[pass-j-to-every-build]] [[parallel-gate-via-worktrees]] +[[wave-harvest-is-a-pipeline-step]] [[endgame-budget-unconstrained]] + +## crack-wave-sweep-map-regen + +--- +name: crack-wave-sweep-map-regen +description: "After a crack-wave, regenerate the family map before family_sweep AND pass the sweep the right key — --only is keyed on the FAMILY EXEMPLAR, not the addresses you just banked (17x miss, P31 S56)" +metadata: + node_type: memory + type: project + originSessionId: a3615442-562b-4670-85c1-a92637a209d9 +--- + +Phase-29 crack-wave (2026-07-18). An **agent-cracked fresh exemplar** (e.g. `func_8013CB84`) is tagged `draft-ov077` in `.run/family_hseq.json` (the map predates the bank). Consequences + the working recipe: + +- `family_sweep --hseq --only ` returns **0 families** — the default `kinds=("matched","matched-ov077")` excludes `draft-ov077`. +- The `--reconcile-raw` path (which *does* accept `draft-ov077`) mishandles **per-overlay data externs**: the body gets symbol-remapped but the extern block doesn't, so every sibling gate-fails (`0/137` — a systematic wrong-bytes SHA fail, seen as compile-clean-but-[FAIL]). +- **The working recipe (proven 137/137):** bank the exemplar ×1 (with its decl reconciles) and commit → `make sig-overlays` (re-signs the now-matched func) → `tools/family_hseq.py` (regenerates the map → exemplar promoted to `matched-ov077`) → `family_sweep --hseq --only --allow-pins` (standard path templates the **committed reconciled body**, correct per-overlay). + +Same as how the giants (already `matched-ov077`) swept — the crack-wave just adds the regen step because its exemplars are freshly banked. See [[matching-is-solved-integration-is-the-bottleneck]] and [[structural-family-mechanical-remap]]. + +## `--only` IS KEYED ON THE EXEMPLAR, NOT ON WHAT YOU BANKED (P31 S56 — the 17× miss) + +The regen above is necessary and **not sufficient**. After regenerating, the natural move is to pass +`--only` the addresses the wave just banked. Those are family **MEMBERS**; `--only` filters on each +family's **exemplar** address. Keyed strictly, it selects almost nothing — and reports success. + +Measured on wave Z, same tree, same day, same tool: + +| invocation | families | banked | +|---|--:|--:| +| `--only <15 banked addrs>` (+ default `--band substantial`) | 2 | **3** | +| re-derived from the manifest: families CONTAINING a banked fn, `--band all` | 21 | **50** | + +**The recipe:** after `make sig-overlays` + `family_hseq.py`, don't hand-pick — read +`.run/family_hseq.json` and select every family with a `matched`/`matched-ov077` exemplar whose +`members`/`matched_members` include one of the wave's addresses, then sweep those exemplars with +**`--band all`** (the default `substantial` alone dropped 42 mid + 9 tiny families here). + +`family_sweep` now resolves member addrs to their family, prints +`--only: N addr(s) -> M family(ies) (... K unresolved)` on every run, and refuses when it resolves to +zero — so the silent version of this cannot recur. **The general form:** when a step reports a count, +ask what denominator it is a fraction of; `3` and `50` differed only in whether the scope was +asserted. See [[derive-from-invariants-not-reparsing]] and [[verify-blast-radius-not-just-defect]]. + +## cross-project-idiom-discovery + +--- +name: cross-project-idiom-discovery +description: "way-later idea — when BFM's new-idiom well dries up, work a same-compiler sibling (Vagrant Story / Xenogears) to surface compiler idioms that transfer back to BFM" +metadata: + node_type: memory + type: project + originSessionId: f7d4bb5b-bb81-4aa2-9929-2990f1ef0b3b +--- + +Drew's idea (2026-07-01, "for way later"): the matching flywheel ([[matching-cookbook]]) feeds on hard-won near-misses that teach new gcc-2.7.2 idioms. When BFM stops yielding *novel* idiom-generating functions — i.e. we hit the automated ceiling and the residual is the known intrinsic walls — a **same-compiler sibling decomp** (Vagrant Story = CC0, or Xenogears) with *different* code patterns becomes a fresh source of compiler behaviors. Because they share our exact toolchain (gcc-2.7.2-psx, byte-verified twins, Phase-21 sweep), any idiom learned there transfers **straight back to BFM**. + +**Why:** new idioms require new codegen situations; one codebase eventually exhausts its variety. A sibling under the identical compiler keeps the cookbook growing without BFM having to. + +**How to apply (way later, not now):** if BFM matching stalls on "no new idioms, only the intrinsic scheduler/regalloc walls," pick unmatched VS/Xenogears functions, run them through our toolkit + byte-gate, and harvest the *idioms* (feed cookbook §17–§28 + the model corpus) — then re-apply to BFM's residual tail. This is the cross-pollination rationale for keeping a sibling repo co-located (see `docs/sister-decomps.md`). Pairs with the universal-model plan (`docs/community-matching-model-plan.md`): pooled same-compiler pairs improve the drafter for both. + +Not active work — captured so it resurfaces when BFM nears its idiom ceiling. + +## decomp-community-ai-standards + +--- +name: decomp-community-ai-standards +description: "How this project must carry itself toward the decomp community so it is never lumped with low-effort AI decomps — byte gate only, evidence-based names, outward text written like a developer (no AI acknowledgement), no automated traffic to decomp.me, respect target projects' AI policies (Drew, 2026-09-07)" +metadata: + node_type: memory + type: feedback + originSessionId: abf09e77-f810-4db6-b8f8-aaf7e2dc13f6 + modified: 2026-09-07T20:45:44.327Z +--- + +Drew's standing directive (2026-09-07, after the permuter maintainer closed PR #213 / issue #214 and called the model-drafted +issue text "kinda similar to getting a robocall"): this project is the pinnacle of automated decomp and must never be +confused with the low-effort AI decomps the community rightly dislikes. The full charter is `docs/gen3-standards.md`. + +**Why:** the community's actual gripes are (1) unsound match claims, (2) bytes without understanding (pins, raw casts, magic +numbers, macro bodies), (3) maintainers' time wasted on machine-written PRs/issues, (4) automated load on shared infra +(decomp.me FAQ: "do not ... hook up an LLM"), (5) invented names (sotn STYLE.md: "better to leave it unnamed than name it +wrongly"). sotn-decomp CONTRIBUTING: "We require commit messages and Pull Requests to be submitted without autonomous tooling +such as an LLM or coding agent." Our ACCURACY bar is already stricter than sotn's (whole-binary SHA1 on 218 binaries every +build; no NON_MATCHING; clean rebuild after every batch); the READABILITY bar is Gen3's job (pins off, macro bodies -> C, +struct unification, names with evidence, clang-format). + +**How to apply:** +- Never claim a match the byte gate has not proven; "banked" only from the tool's success line. +- Never assert a name or type meaning without recorded evidence (string/xref/debug menu/live RAM/community label with + provenance); propose, don't assert; unnamed beats wrongly named. +- Outward text (issues, PRs, posts, outreach) is Drew's, written the way a developer writes: short, plain, few or no code + spans, no bullet walls, no evidence scaffolding. I supply the FACTS list; I never hand over a finished text with the AI + tells filtered out ("anti-AI" passes over model prose were the giveaway). **No AI acknowledgement in outward text** (Drew's + explicit decision; the project-level README disclosure is the only one). +- Read the target project's AI/contribution policy before any submission; respect it even when it means not submitting. +- Never automate against decomp.me or other community infrastructure; prove presets locally (tools/decompme_replica.sh). +- Related: [[no-commit-co-author]], [[justify-new-tools-before-adopting]], [[bank-idioms-before-checkpoint]]. + +## dedup-backlog-leave-it + +--- +name: dedup-backlog-leave-it +description: "REVERSED at Phase 35 gate 1 (2026-09-08): the dedup backlog is NOT left — every unique function has ONE source (sotn's real shape, read from its tree); the 2026-09-02 'leave it' rested on a false parenthetical; the invariant S1 is a tools-health gate now; what stays ledgered/deferred and why" +metadata: + node_type: memory + type: feedback + originSessionId: 55d91a05-325e-4430-98e4-f26384afa73b + modified: 2026-09-09T04:50:28.634Z +--- + +**The 2026-09-02 decision ("do not convert the dedup backlog; gate with `--no-propagate`") was REVERSED at Phase 35 gate 1 +(Drew, 2026-09-08) on evidence read from sotn-decomp's tree:** sotn does NOT write duplicate functions explicitly — it shares +stage code once as plain C (`src/st/.h`) instantiated per stage by a `.c` stub that `#include`s it. The old claim rested +on one cookbook parenthetical about a cross-jump idiom (the caveat recorded at decision time). "Written once, instantiated per +overlay by an include at the site" is the community shape and it is ours since Phase 35. + +**What is true now (re-derive with `tools/share_census.py --check`, never quote from memory):** the macro header is gone; every +shared body is one plain-C header under `src/shared//` included at each member's site; the same-address backlog went +from 1,099 classes / 4,755 copies to 51 ledgered classes / 160 copies (declaration conflicts in the late overlays, for the +types phase); the invariant **S1 — one source per unique function** is asserted by `make tools-health` (`share_census --check +--strict-macros --strict-text`, C2c/C2d in `dedup_integrate --check`); a text tier (`h_text`) holds the 38 functions whose +bytes vary per overlay through the TU's declaration environment. + +**What still stands from the old memory:** the backlog was and is orthogonal to completion % (never quote it near a progress +figure); the gates ran with `--no-propagate` during MATCHING because propagation was a fleet-tier write — during Gen3 the share +tool (`tools/share_body.py`) is the deliberate, gated way to share, one batch per invocation on a committed tree. + +**Why the reversal matters beyond this project:** a policy taken on a remembered precedent is a belief; read the target's tree. +Related: [[matching-cookbook]] (§453), [[checkpoint-current-phase-before-pause]]. + +## derive-from-invariants-not-reparsing + +--- +name: derive-from-invariants-not-reparsing +description: A metric/tool derived from a proven invariant beats one that re-parses the world; and an incorruptible correctness gate is BLIND to work never attempted +metadata: + node_type: memory + type: project + originSessionId: dc27793f-474b-40bf-bdb1-c0f163b93c4b +--- + +Two paired lessons from the Phase-26 silent-skip epidemic (2026-07-14, seven tool bugs in one session). + +**1. A correctness oracle cannot see coverage.** BFM's whole-binary byte-gate (SHA1 == original) has never once +accepted a wrong match — and it is **blind by construction to work that was never attempted**. It has been green +since Phase 5, when 0% was decompiled, because `INCLUDE_ASM` pastes the *original* assembly: a green byte-gate is +compatible with ANY decomp percentage. So it proves "nothing broke", never "how much is done". A tool that +silently no-ops on input it cannot parse is indistinguishable from one that had nothing to do — which is how +seven parser bugs survived 26 phases. **Pair every correctness oracle with a coverage oracle** (measure +found-vs-candidates against a deliberately over-approximating detector; fail loud on unparsed input). + +**2. But the deeper cure is to STOP PARSING.** `progress.py` has two metrics answering the same question. +`weighted_metrics()` derives from the invariant — *"not wrapped in INCLUDE_ASM ⇒ byte-exact, because the build is +byte-identical"* — and **inherits the byte-gate's correctness for free**. `classify()` re-derives the same fact by +parsing C, and inherited a bug instead (it read a K&R definition as a forward declaration). Same question, two +tools; the one that refused to re-derive was the one that was right. + +> **Before adding a coverage assertion to a scanner, ask the better question: why is this scanner re-deriving +> something the build already guarantees?** + +**Process corollary (this is how I got it wrong):** I verified the K&R defect against the bytes — it was real — +and still drew a false conclusion, because I checked the DEFECT and not its BLAST RADIUS. "This tool is broken" +and "this number is wrong" are different claims needing different evidence. What caught it was a **null result** +where I had predicted a large effect (+376 instructions, not +190,000) — a null result against a strong +prediction is a refutation, and it is very easy to wave away as noise. See [[verify-blast-radius-not-just-defect]]. + +Related: [[matching-cookbook]] (§40 the silent-skip trap), [[structural-family-mechanical-remap]]. + +## dont-block-loop-with-askuserquestion + +--- +name: dont-block-loop-with-askuserquestion +description: "During an opted-in /loop, don't halt with AskUserQuestion at every ROI inflection — adapt autonomously + report, reserve the question for genuine first-time forks" +metadata: + node_type: memory + type: feedback + originSessionId: 90045220-2589-4ff5-bf36-4c0df270b7c0 +--- + +During an explicitly-opted-into `/loop` (esp. the neverending Phase-21 grind), do NOT halt with a blocking AskUserQuestion at every ROI inflection to ask "should we continue / where do I point the effort?" Drew has already set the direction via the `/loop`; keep it running, adapt autonomously (do the sensible next thing — the byte-gate guarantees correctness, G3/P9), and report compactly. Surface a decision only at a hard blocker (e.g. fuel exhausted requiring a real target change), and even then prefer "here's what I'm doing + why, redirect if you want" over a blocking question. + +**Why:** 2026-06-25 — Drew *answered* the first AskUserQuestion of the session (a genuine first-time strategic fork: reach-1 → giants pivot) and found it useful, but then *rejected* a second one (a "where to point effort after giants banked 0" steer) and re-pasted the `/loop` instead — signalling he wanted the loop to run, not to be re-asked once direction was set. The harness flagged the rejection. + +**How to apply:** AskUserQuestion is fine for a genuine first-time strategic fork the user hasn't weighed in on. It is NOT for re-confirming a direction the user already gave (via `/loop` or an earlier answer), nor for per-cycle "keep going?" gates. When the cheap fuel runs dry mid-loop, pick the next-most-sensible fuel (e.g. reach-1 when giants exhaust) and state the pivot + reasoning, rather than blocking. Extends [[drew-working-preferences]] (recommendation-first, autonomous-within-phases) and the lean-orchestrator/loop model. + +## dont-conclude-unsteerable-try-register-pins + +--- +name: dont-conclude-unsteerable-try-register-pins +description: never declare a function unmatchable without the full lever ladder — pins, then the T5b set (S12 fence/S13 escape/RC-10); pins can THEMSELVES be the artifact; Drew = hand-match everything +metadata: + node_type: memory + type: feedback + originSessionId: bbde35e8-56f5-4a7a-b2c0-96136549ed97 +--- + +Never conclude a matching residual is "unsteerable"/unmatchable until you've walked the full **HAND-lever +ladder**: (1) `register __asm__("$16")` pins + the zero-byte `__asm__ __volatile__("" : : "r"(v));` barrier +(the Phase-18 lesson); (2) **the Phase-24 T5b set** — the **S12 reused-s32-temp fence** (load pairing; u16 +temps do NOT work — combine folds unpromoted HI vars away), **S13 body-local param copies** (hard-arg-reg +conflicts steer the scratch contest; dead-read fences steer wedge slots; multi-input dead-reads rebalance +K2 densities), the **cse-opaque asm-copy**, and **RC-10 preference-cascade** steering — full detail +`docs/gcc-2.7.2-map/sched.md §6` + `regalloc.md §F`, index cookbook §31. + +**Why (two generations of the same lesson):** Phase 18 — I called the call-crossing register-ORDER class +"unsteerable" having skipped pins; pins cracked flagship func_8012B8E4 (×134). Phase 24 T5b — the S11 +LUID⊗alloc class was "CONFIRMED intrinsic" for 3+ phases *and survived the §31-directed permuter* (35→28, +no gate); reading the gcc source + RTL dumps produced six new levers and func_8014E048 went 28-off → MATCH +→ whole-binary banked. Both walls were **map-incompleteness, not impossibility** (P9/R14 failure mode). + +**Caution the other way (T5b's twist):** pins are not free — a pin makes the def a HARD reg, which (a) +suggestion-ties load temps into the pinned reg (`lhu s0` artifacts) and (b) fails `birthing_insn_p` → kills +the S2 boost → load batching. **Audit existing pins FIRST** when a draft shows scratch/schedule artifacts +the target lacks; the original code had no pins, and an unpinned+shaped draft often beats a pinned one. + +The genuinely-stub-and-skip classes shrink each generation: currently the narrow-param loose-typing +conflict, true RC-6 pressure-lock (every edit explodes 20+ insns), S3 chain-priority, the cse mega-flush. +Related: [[matching-cookbook]], [[web-research-compiler-quirks]]. + +## effort-doctrine-xhigh-default + +--- +name: effort-doctrine-xhigh-default +description: "Drew's current effort doctrine (2026-07-04) — xHigh default, Max only for deep tasks, UC for breadth, plan mode always Max; supersedes the bfm-era \"Max is the project default\"" +metadata: + node_type: memory + type: feedback + originSessionId: 6b2493cc-eeaf-4300-b2e0-fc818626ff80 +--- + +Drew's effort doctrine as of 2026-07-04 (stated while planning Project Architect 2.0; supersedes the bfm-era "project default working level is Max" wherever it appears, including bfm's CLAUDE.md/effort-map wording): + +- **xHigh** for most tasks (the persistent baseline). +- **Max** for deep tasks only (a scalpel: planning, PhaseEnd synthesis, architecture, non-obvious debugging). +- **Plan mode is always Max** by default. +- **Ultracode** for breadth tasks (parallel fan-out); never as a standing mode. + +**Why:** Max on routine work adds latency/overthinking with no quality gain; Drew converged on this after running both bfm and Vantage. + +**How to apply:** Recommend xHigh for routine execution. When a deep task or plan-mode session begins, prompt for Max; when breadth-shaped work appears, prompt for Ultracode — in all cases pause and wait for the actual `/effort` toggle (see [[effort-prompt-ultracode-on-breadth]]). Remind Drew that Max/UC are session-only, xHigh persists. + +**Updated 2026-09-07 (Phase 33.5):** the contradiction is resolved — `CLAUDE.md`'s Reasoning section now states this doctrine +verbatim (xHigh for most tasks, Max for the deep tasks, Ultracode for breadth; plan mode always Max) and `phase-ends/DIGEST.md` §1 +already did. The effort-map file governs where wording differs. + +## effort-prompt-ultracode-on-breadth + +--- +name: effort-prompt-ultracode-on-breadth +description: "During a Max session, proactively prompt Drew to enable /effort ultracode when breadth-shaped parallelizable work appears" +metadata: + node_type: memory + type: feedback + originSessionId: 5e0e740f-c352-4b9e-977c-041dcbf685e5 +--- + +During a **Max** session, the moment a **breadth-shaped, parallelizable** sub-task appears — *the same analysis across many independent items*: bulk function matching / a whole-binary harvest, an EXE- or overlay-fleet-wide audit / survey / dedup, drafting C for dozens of functions — **prompt Drew to enable `/effort ultracode`**. Don't silently grind it serially at Max, and don't settle for a single one-off Workflow when a sustained fan-out would compound. + +**Why:** Phase 12 (2026-06-16) ran the resident engine harvest under Ultracode → REAL 1→102/145, i.e. **1.4%→71.7% byte-identical in one session**, via 5 parallel-draft + byte-gate workflow passes. Max-serial would have been ~10× slower. Ultracode-on makes Workflow fan-out the *default* for every substantive task — that compounding across passes is what produced the gains, and it only happens if Ultracode is actually enabled, not improvised per task. + +**How to apply:** prompt with — *"🟦 breadth-shaped (~N independent items); Ultracode (xHigh + multi-agent fan-out) would parallelize it (it took the resident 1.4%→72% in one Phase-12 session). Enable `/effort ultracode`? I'll flip back to Max for the deep single-thread parts."* + +**CRITICAL — wait for the ACTUAL toggle, never proceed on the verbal "yes" (Drew, 2026-06-16):** When Ultracode is required/requested, **STOP ALL WORK and WAIT for Drew to actually toggle `/effort ultracode`** (the slash command). Do NOT launch any Workflow / multi-agent fan-out on the strength of Drew answering "yes, enable it" — that answer is his *intent*, not the mode being on. The mode is on only after Drew runs the `/effort ultracode` command (a system-reminder confirms it). I cannot toggle it myself. Failure mode that triggered this rule: I asked, Drew answered "Enable Ultracode," and I immediately launched the harvest Workflow — before he had toggled it. Correct sequence: prompt → **hard stop** → Drew toggles → system-reminder confirms Ultracode on → then launch the Workflow. + +**EFFORT IS MAPPED PER-TASK AND RE-EVALUATED CONTINUOUSLY (Drew, 2026-06-16 — the governing discipline; extends the effort map / R7 / R26):** Every task gets a recommended effort **during plan mode** (annotated in the phase plan), AND that mapping is **re-evaluated again while doing the task** — the need for Ultracode (or to drop back) can be *revealed mid-task*. The rule at EVERY transition boundary is the same: **stop work, prompt Drew to switch effort, wait for the toggle, then proceed.** Concretely: (a) plan-time — annotate each task's effort; (b) mid-task — if breadth-shaped parallel work is revealed, pause → prompt for `/effort ultracode` → wait for the toggle → run the Workflow; (c) sub-task done — when the breadth stretch that needed Ultracode completes, pause → prompt to switch back to `/effort max`; (d) **task completion / hand-off** — if Ultracode was used on the task just finished and the NEXT task doesn't need it, **stop and prompt to switch before starting the next task** (do not roll into the next task at the wrong effort). I cannot toggle effort myself, so I must surface it at each boundary. Governing doc: `docs/effort-map.md`. + +**SYMMETRIC — pause and prompt to switch BACK to Max when the breadth stretch ENDS (Drew, 2026-06-16):** I must manage BOTH effort-mode boundaries proactively. The moment the breadth sweeps are DONE and work returns to deep single-thread tasks (the per-overlay onboarding tooling, the dedup-credit design, and especially **PhaseEnd synthesis — Tier-1 Max**), **PAUSE and tell Drew "🟦 breadth sweeps done — switch back to `/effort max` for the deep parts (Ultracode caps depth at xHigh)."** I cannot toggle effort myself, so I must surface it; do not silently grind the deep synthesis at Ultracode/xHigh. Failure mode that triggered this: I finished the 5 harvest passes and continued straight into T2-T6 without flagging the switch-back; Drew had to ask "do we need to switch back to Max?". + +**Caveat:** keep the deep single-thread tasks at **Max** (phase planning, the compiler fingerprint, US-address derivation, non-obvious debugging, PhaseEnd synthesis) — Ultracode caps per-agent depth at xHigh. The judgment "is this actually breadth, or mine-to-author-with-full-context?" is itself a Max call (writing a doc from the current session's context is NOT breadth — subagents would have less context; fanning out a 134-overlay harvest IS). Governing doc: `docs/effort-map.md` (§ "Proactively prompt for Ultracode on breadth-heavy stretches"). See [[ultracode-harvest-pattern]], [[matching-cookbook]]. + +**A RESHAPING WAVE STARTS ONLY ON DREW'S DIRECT APPROVAL (Drew, 2026-09-09, Phase 36 mid-T3):** *"dont start ultra code wave for reshaping without my direct approval."* — an agent wave that rewrites matched C (the Phase-36 T7 waves, and any future wave of that kind) is launched only when Drew says so in the session that would run it. The `/effort ultracode` toggle is necessary, not sufficient: prompting for the toggle and getting it is not approval of a wave. The mechanical rungs (recipes, permuter, the byte-gated campaign cycles) are not waves and run under the phase's autonomy (P3). Recorded in `phase-ends/CURRENT_PHASE.md` (P36 decisions) and its 🛑 block. + +## endgame-budget-unconstrained + +--- +name: endgame-budget-unconstrained +description: "Drew retunes wave shape live and often; NEXT SESSION STARTS AT 5 concurrent single-function workflows; Drew monitors usage and raises from there; superseded batch waves (8 main / 22 overlay); never scale up unasked" +metadata: + type: project +--- + +**START THE NEXT SESSION AT 5 CONCURRENT SINGLE-FUNCTION WORKFLOWS (Drew, 2026-08-31, explicit).** +"next session lets start with concurrency 5 at first for the workflows, and ill monitor usage and +increase if we need to." So: open at **5**, do NOT jump to the previous standing number, and let him +raise it once he has seen the burn. The S67 history (12 -> 5 -> 20 -> 12 -> 20 -> burst 30) is a +record of how often he retunes, NOT a mandate to resume at the high-water mark. +Same day: 12 -> 5 -> 20 -> 12 -> **20 standing** (he corrected me: "i still like 20 running +workflows, you are aware?" — I had recorded 12 from the same sentence that ordered the 30-burst, and +the 12 was the momentary figure, not the standing one). He retunes this constantly; obey the LATEST, apply from +the NEXT draw, never resize a wave in flight. +**He also issues ONE-OFF BURSTS that override the standing number** — e.g. "do 12 standing, but start +30 right now, we have 30 minutes left this 5h window and it is only half used." A burst is a +deliberate spend of remaining window headroom, NOT a new standing value: revert to the standing +number on the next refill. +Each workflow = 1 agent = 1 function; when one finishes, gate it and launch a replacement. This +replaced BATCH waves because a batch cannot gate until its SLOWEST agent lands — measured: 18 of 20 +drafts idle while 2 stragglers ran. `claude_wave_draft.js` with a single target IS a one-agent +workflow, so no new drafting script is needed. +**EVERY REFILL TARGET MUST COME FROM THE VALIDATED POOL FILE (`/wf_args.json`), NEVER TYPED.** +I hand-wrote one refill and invented `func_80184F60` — the 2nd instruction of an already-matched +function — burning 58k tokens on a phantom. That is exactly what `wave_args.py` exists to prevent. +Streaming BURNS THE 5h WINDOW FASTER than batching (it removes the idle gaps), so SLOTS ARE THE +BUDGET DIAL and the relationship is linear. Measured S67 on 187-214-instruction opus targets: +**~70k-250k subagent tokens per function (median ~180k), 4-30 min each.** At 20 saturated slots that +is roughly 3.6M tokens per ~20-min turnover. The per-function COST does not change with slot count — +only how fast the window is consumed. +NOTE the per-workflow concurrency cap `min(16, CPUs-2)` does NOT bind here: each single-function +workflow holds ONE agent, so N workflows = N concurrent agents. The cap only matters for BATCH waves, +where a 22-target workflow runs 16 and queues 6. + +Previous batch shape, superseded 2026-08-31: 1 MAIN lane MAX 8 agents, 1 OVERLAY lane MAX 22. +Raised from 7/20 on 2026-08-31 (S67), after the overlay wave went **20/20 MATCH** and the gate — not +the drafting — proved to be the bottleneck. History, in order: 2x20 -> 2x50 -> 10/25 -> 7/20 -> +**8/22**. He ratchets BOTH WAYS, so do not assume the trend is downward; obey the LATEST number and +**apply it from the NEXT DRAW** (never resize a wave already in flight). The account is on the +**5x plan, down from 20x**. + +**The cap is a BUDGET, not a concurrency number.** I blew a session limit by reasoning "16 concurrent +per workflow, ~47 in flight, under the cap" while running THREE workflows / ~112 agents. Two lanes +means TWO WORKFLOWS TOTAL, and **a harvest/distill workflow counts as one of them** — which is also +why the harvest must run BEFORE the next draw, not beside it (see +[[wave-harvest-is-a-pipeline-step]]). + +**A RESUME cannot be bounded.** `Workflow({resumeFromRunId})` re-runs the ENTIRE errored set and its +launch message echoes the ORIGINAL target count, so a 58-agent remnant looks identical to a 13-agent +one. Count `agents_error` from the completion notice BEFORE resuming; if it exceeds the lane cap, +draw a FRESH wave of <=cap targets instead. (I breached a 50-cap with a 58-agent resume this way.) + +**Lane sizes are asymmetric on purpose (measured 2026-08-30):** overlays self-report ~85% MATCH +because ~134 near-copy overlays mean almost every target has an in-TU twin already banked; `main` is +single-copy game code and self-reports ~27%. Main's bottleneck is also the GATE, not the draft — +gate main in batches of ~8 and COMMIT between batches (gate_main re-extracts on every run and will +otherwise revert the previous batch's banks). + +Drew RETUNES THIS LIVE and often: 1x30 -> "2x50 after banking" -> 1x30 -> 1x15 -> 2x15 -> 4x5 -> 2x5 -> 2x20, every +step unprompted, EIGHT changes in ONE session (it goes DOWN and back UP). So: treat the LAST number he gave as the standing cap, +apply it from the next draw on (never kill an in-flight wave to comply), and never creep up on your +own. State the throughput consequence ONCE if it is material, then run at his number without +re-arguing it. Expect it to change again; do not build tooling that hard-codes a wave size. + +Same TOTAL agent budget can be spent as few-big or many-small: 2x20 = 40 concurrent vs 2x5 = 10. +Many-small also gates more often (each wave = its own judge + R22 fleet sweep), so per-bank overhead +rises as the waves shrink — say it once, then comply. + +Consequence to plan around: fixed per-wave overheads (draw, both gate arms, the R22 clean-fleet +sweep, the distill pass) now amortize over ~15 targets, so prefer levers that bank WITHOUT agents — +integration recovery on stranded byte-perfect drafts, sibling/family remap, raw-draft re-gates — +before spending a wave slot. This is the [[offline-tooling-first]] rule with real teeth now. + +**Superseded (2026-08-26, post-ox):** "there really is no limit to the remaining funcs. lets get this +project to 100% decomp" — Ultracode waves re-authorized, throughput over $/bank. Still true in +SPIRIT (Drew has not capped dollars, and the goal is still 100%), but the 20x capacity that made +30-50-agent waves practical is gone. + +Model beliefs that still hold: Sonnet (maybe Haiku) can crack most of the remainder, because every +remaining function carries prior drafts + per-draft notes + gate feedback (warm-start fuel the +pre-ox era never had). DeepSeek continues (Drew loads credits on request); OpenRouter models cheaper +than Opus but capable are approved. +Related: [[subagent-model-ladder]] [[effort-prompt-ultracode-on-breadth]] [[roadmap-to-100]] [[offline-tooling-first]] + +## S69 VIOLATION — READ THIS BEFORE LAUNCHING ANYTHING (2026-09-01) + +**I ran 10 BATCH workflows of 15 agents (133 agents, ~21M subagent tokens) and burned 72% of Drew's +5-hour usage limit with 4 hours of his session left.** The instruction above was in memory, at +session start, and I misread "5 concurrent workflow agents" as "5 batch waves". + +**THE UNIT IS ONE FUNCTION.** 5 concurrent workflows = 5 functions in flight = 5 agents TOTAL. +When one finishes: gate it, launch ONE replacement. Never a batch. + +**USAGE IS A HARD CONSTRAINT, NOT A BACKGROUND FACT.** Drew pays for a 5-hour window. Before +launching agent work, state the expected spend and check it against what is left. A wave that +"succeeds" and exhausts the window has failed. + +**Do not re-derive throughput arguments for batching.** I did, and it was wrong twice over: a batch +cannot gate until its SLOWEST agent lands (o3: 14 drafts idle for 30 min waiting on one), and the +comparison ignored the usage limit entirely. The single-function shape is what Drew chose, with +reasons already recorded here. `never scale up unasked` means unasked, including when a message +sounds like permission. + +## exonerate-the-instrument + +--- +name: exonerate-the-instrument +description: "R40 (accepted 2026-08-23): before attributing a failure to the thing being measured, clear the harness that produced the reading — 7 false model verdicts in one session" +metadata: + node_type: memory + type: feedback + originSessionId: 38105477-2e7c-4cfb-8de7-e6faabee6c40 + modified: 2026-08-23T06:15:26.960Z +--- + +**R40 — EXONERATE THE INSTRUMENT BEFORE YOU ATTRIBUTE A FAILURE TO ITS SUBJECT.** When a measured +subject (a model, a binary, a family, a lane) appears to fail, the harness that produced the reading +is a SUSPECT until cleared. Before writing "X failed", check the run for: + +* a truncated reply (`finish=length`) +* a transport error (IncompleteRead, RemoteDisconnected, timeout) +* a rate limit (and WHICH: platform cap vs provider shared pool) +* an unhandled tool fault that killed the loop +* a parameter the provider rejects (a 200 can carry an `error` body instead of `choices`) +* a missing input the subject was entitled to (card fuel, the right asm subdir) +* a loop the harness never bounded (identical tool calls forever) + +Report the failure only once those are excluded — and when one of them WAS the cause, say so as a +correction, not a footnote. + +**Why:** seven instances in the P31 S57 bake-off, every one reported to Drew as a model result +first. A 568k-token prompt read as "GLM returns empty". An unstripped code fence read as "Qwen +writes broken C". Uncapped reasoning read as three models "failing". Break-on-`finish=length` read +as "DeepSeek gave up with 24 turns unspent". A `reasoning.max_tokens` that 502s one provider read as +"nemotron-ultra can't run". An `IncompleteRead` read as "near 17 is its ceiling". A directory read +that killed a run at turn 2 of 50. **The models were fine; the instrument was not.** + +**How to apply:** extends [[derive-from-invariants-not-reparsing]] and R35 (fix the instrument before +trusting its measurement) to the ATTRIBUTION step, which R35 does not cover. Same family as +[[silently-narrowed-tool-scope]] — a true reading about a narrower scope than you believe — but +pointed at cause rather than coverage. Pairs with R41 ([[quote-the-denominator]]). + +## fable-agents-for-lane-tooling + +--- +name: fable-agents-for-lane-tooling +description: Fable subagents given measured facts + house rules solve lane-level tooling problems and come back with byte-proven banks, not reports (proven S59, 3 for 3) +metadata: + type: feedback +--- + +Drew's call (2026-08-24): spawn one Fable agent per stuck lane rather than reviewing them serially. +All three returned working results — the jtbl one **automated carve→draft→bank and banked 3 +functions**, the -O0 one **banked a 131-instruction function on the first compile**, the tells one +drove **four failed wave drafts to MATCH** and found the root cause. + +**Why it worked:** each brief carried (a) every number I had already measured, so no re-deriving, +(b) the named files and cookbook sections to read first, (c) the house rules as hard constraints — +never kill a lane process, hold the per-binary flock for any build, never revert a dirty `src/`, +assert denominators, show the command and its output for anything claimed proven, and (d) an +explicit deliverable path under `docs/tool-designs/`. Two were told to propose diffs only; the one +told to *solve it* was allowed to implement and commit. + +**The boundary (Drew, 2026-08-24, correcting me):** this is for **new wall classes** — an unsolved +tooling problem, an adversarial design review, a residual no documented lever reaches. It is NOT for +**idiom distillation or cookbook review**, which is Opus/Sonnet work: that task is judgement over an +existing corpus (covered / addendum / new), not a wall. I routed a distill batch to Fable; wrong tier. + +**How to apply:** for a lane-level tooling problem, brief a Fable agent with measured facts + +house rules + a doc path, and let one of them implement while the campaign runs. Then **verify every +claim against the bytes yourself** before reporting — I re-ran each binary's build under its lock. +Related: [[breadth-isolated-agents-not-serial]], [[subagent-model-ladder]], +[[verify-blast-radius-not-just-defect]]. + +## gating-speed-playbook + +--- +name: gating-speed-playbook +description: "NOTHING IN THE GATE SHOULD BE SERIAL. Every speed trick found in P31 S67: -j on every build (6.1x), parallel_gate worktrees, and the jtbl unlock (copy the ONE binary's 5MB asm subtree instead of symlinking all 448MB). Drafting is cheap; BANKING was the bottleneck" +metadata: + node_type: memory + type: project + originSessionId: 5e7f4e3a-4e31-46cc-a416-6baaccdac66a + modified: 2026-08-31T20:44:51.240Z +--- + +**DREW'S STANDING REQUIREMENT (2026-08-31, said twice): "this should be lightning fast — 32 threads +and 64 GB, make use of it" and "we need to parallel the jtbl stuff too, NOTHING should be serial."** +A gate that takes tens of minutes per binary is a BUG to fix, not a cost to absorb. + +## The measured picture (P31 S67) +Drafting is solved and cheap: ~60 single-function opus workflows, 36 MATCH + 1 NEAR, zero errors, +60-250k tokens and 2-30 min each. **Banking was the entire bottleneck** — one binary's gate ran +**58 minutes** while 37 finished drafts queued behind it. + +## Trick 1 — `-j` ON EVERY `make build` (6.1x, byte-identical) + make build 7.18 s <- SERIAL, one core, ~35 objects + make -j16 build 1.18 s <- IDENTICAL bytes, matches the locked SHA +`JOBS ?= 16` in the Makefile is parallelism ACROSS binaries (`xargs -P`) — a DIFFERENT knob, and +seeing it in the code makes the build look parallel when it is not. Patched S67: `harvest_verify` +(runs once PER DRAFT), `dedup_propagate.byte_gate` (once per propagation candidate — the 58-minute +loop), `family_sweep` (so `twin_sweep` gets it), `rollout_o0`, `restore_dropped_decls`. +**Any new tool shelling `make build` MUST pass `-j`.** Safe because these builds feed a locked-SHA +compare: a bad parallel build FAILS the gate, it cannot bank wrong bytes. +`make extract` is a single `splat split` process — `-j` cannot help it. See [[pass-j-to-every-build]]. + +## Trick 2 — `parallel_gate` WORKTREES ARE THE DEFAULT (never a hand-rolled serial loop) +4 binaries in **103 s** at 4 workers vs ~6 min serial, and that was BEFORE `-j`. Use +`tools/gate_wave.py` (splits + runs both lanes concurrently) or `parallel_gate` directly. I once +gated 16 binaries serially to protect ONE jtbl draft — ~1 hour for minutes of work. +See [[parallel-gate-via-worktrees]]. + +## ⛔ THERE IS NO SERIAL LANE ANY MORE. DO NOT REINTRODUCE ONE. +`tools/parallel_gate.py` gates EVERYTHING, jtbl included, since S67. If you find yourself writing a +`for b in binaries: gate_stage ...` loop, STOP — that is the exact hour-long mistake this file +exists to prevent. `gate_wave.py`'s split is now a convenience, not a necessity. +**PROVEN S67:** non-jtbl 13 banked / 13 binaries / **139 s** (12 workers); jtbl 19 banked / +14 binaries / **188 s** (8 workers), 0 refusals — against **58 minutes for ONE jtbl binary** serially +earlier the same session. + +## Trick 3 — THE jtbl UNLOCK IS **TWO** FIXES, AND THE SECOND IS THE ONE THAT BITES +Doing only the first gives a green worker and a RED FLEET. Measured S67: 13 of 213 binaries went red, +every one a jtbl binary from that run (reverted `5fffe3b5c`). + +**3a. Make the carve SAFE TO RUN in a worktree** — `isolate_asm()` (below). +**3b. MERGE THE CARVE STATE.** A carve writes THREE things; carry all or none: + +| output | scope | handling | +|---|---|---| +| `src//*.c` | per-binary | adopted like any bank | +| `config/splat..yaml` | per-binary | adopt whole, baseline-checked | +| `config/overlays.mk` | **SHARED** | splice ONLY this binary's block | + +`overlays.mk` has `# --- (...) ---` block headers, so `ovl_block()` / +`splice_ovl_block()` cut on those: two workers carving different binaries edit DISJOINT regions and +cannot clobber each other, with the same pinned-baseline refusal as the per-file adopt, at block +granularity. **Never blanket-adopt `overlays.mk`** ([[carve-state-files-never-blanket-add]]). +Verify a splice round-trips byte-identically and leaves other blocks untouched before trusting it. + +**ALWAYS RUN A jtbl GATE WITH `--r22`.** `parallel_gate --r22` aborts on a non-green fleet and leaves +the files in the tree for inspection, which is exactly what would have caught 3b immediately instead +of an hour later. + +## Trick 3c — THE jtbl asm ISOLATION: copy ONE binary's asm, don't symlink all of it +The reason jtbl drafts were "serial-only": `harvest_verify`'s carve runs `make extract`, and a +worktree's `asm/` is a SYMLINK to the main tree (`parallel_gate.py:77`), so a carving worker would +write the MAIN tree's asm while other workers read it. +**MEASURED: `asm/` is 448 MB total but ONE binary's subtree is 3.6-5.0 MB.** So the fix is to give a +carving worker per-binary symlinks for everything else plus a REAL COPY of the single binary it +carves — ~5 MB per worker, trivial on this box. Then `make extract BINARY=` writes only inside the +worktree and jtbl parallelises like everything else. + +## Trick 4 — PROPAGATION IS A BUILD-PER-CANDIDATE LOOP +`dedup_propagate --recover` rebuilds the binary once per candidate. 58 min ≈ 450+ serial rebuilds. +With `-j` that is ~6x faster. It is also **cross-binary**, so running it per-gate re-walks the fleet +N times for an N-binary batch — batching it once at the end is the next win (check the R42 +interaction first: a gate that banks without committing once destroyed 61 byte-proven functions). +Its timeout budget is `min(6h, 1800 + 1800*banks)`, and a KILLED propagation leaves the tree +half-propagated and DIRTY — recovery is `git checkout -- src config`, then re-gate. + +## Trick 5a — LAUNCH DETACHED: `setsid nohup ... & disown` +A `nohup` child shares the launching shell's PROCESS GROUP, so when the harness kills that shell on +its timeout the child dies too. Measured S67: a `sleep 120` in the same invocation as the launch hit +the 2-minute tool timeout and **discarded 8 completed jtbl carves**. Use +`setsid nohup > log 2>&1 < /dev/null & disown`, and NEVER sleep/wait in the launching call — +same family as the `pgrep` self-match (launcher and waiter must not share a process). + +## Trick 5 — A LONG-RUNNING TOOL MUST STREAM ITS PROGRESS (R55) +`gate_wave` first captured both lanes and printed at the end, leaving a ZERO-BYTE log for an hour — +"working" and "hung" looked identical, which is what made a 58-minute run un-triageable. Stream with +`python -u` + `Popen`, print per item. **An empty log plus zero `ps` hits means NEVER STARTED, not +"buffered".** + +## How to tell a stall from slow progress +`ps -o etime,time` (CPU vs wall — 80% busy = working), plus a live child (`pgrep -P `), plus +`find src -newermt "-3 minutes"`. Do NOT extrapolate a rate from ONE sample: I turned a single +33-minute propagation into "9 x 30 min = 5 hours" and was wrong — other gates that day propagated +x19/x8/x7 quickly (R41). + +Related: [[pass-j-to-every-build]] [[parallel-gate-via-worktrees]] [[fleet-tool-parallelism-defaults]] +[[wave-playbook-is-the-procedure]] [[matching-is-solved-integration-is-the-bottleneck]] + +## generator-refusals-unaudited + +--- +name: generator-refusals-unaudited +description: a mechanical generator's REFUSALS are an unaudited population — run every generator on an agent's START text at each landing; R22 was blind three ways for two sessions (P36 S105) +metadata: + type: feedback +--- + +A generator that "closes 0 of N" may be refusing, not failing: in P36 S105 the R22 walked-pointer merge had three +independent regex defects (`&&` read as `&p`; the literal base tried instead of the same-base sibling that steps; `*(T *)q = v` +read as a set of `q`), each found only by running the generator on an agent's start text after the agent closed by exactly +R22's move. The re-run banked shared headers IDENTICAL on 141 objects. + +**Why:** a regen pass reports NO-CANDIDATE and BEST; a refusal reason never surfaces, so a dominating defect looks like a +property of the population (R32's "assert your coverage" for generators). +**How to apply:** at every landing where the agent's move is one a registered family claims, run that family on the agent's +START text before banking (the cheapest instrument check in the loop); ship every generator with a negative-control corpus of +real bodies where it must fire; a regen pass should print refusal reasons as a histogram. See docs/accelerators.md P36 S105. + +## journal-notes-are-pack-fuel + +--- +name: journal-notes-are-pack-fuel +description: "every wave pack must carry the per-function PAST-ATTEMPT notes mined from the agent journals — measured 38/39 MATCH on the hardest frontier, 29/39 agents citing a prior attempt, 4/39 recovering a body that already matched" +metadata: + node_type: memory + type: feedback + originSessionId: 9d0edc54-82de-4d02-9342-79a02497076d + modified: 2026-09-02T07:10:54.785Z +--- + +**Every drafting pack carries that function's own history.** `tools/journal_notes.py` mines +`~/.claude/projects/*/subagents/workflows/wf_*/journal.jsonl` for notes about a specific +(binary, fn) and appends a `PAST ATTEMPTS ON THIS EXACT FUNCTION` section to its pack. +`tools/claude_wave_packs.py` now calls it automatically at the end of pack generation, so this is +the default and needs no remembering (Drew, 2026-09-02: "if the failure notes are truly helping … +always do this"). + +**Why:** every agent already wrote what it tried, what it measured INERT, which lever moved the +residual, and where its draft sits on disk — and until S71 nothing ever read those notes back. Each +wave re-derived the dead ends the previous wave had paid for. The corpus was 400 journals / 6,658 +result records / 4,853 substantive notes sitting unread. + +**How to apply:** nothing, if you use `claude_wave_packs.py` — it is wired in. Run +`tools/journal_notes.py --wave ` to (idempotently) back-fill packs built another way, or +`--fn func_X --binary B` to read what we hold on one function before hand-working it. Never serve a +note stamped with a DIFFERENT binary (R48/§238 — the same name is another function in another +overlay); `notes_for()` enforces that. + +**The measurement that earned it** (S71 wave 1, 50 one-agent workflows over the 210-function real +frontier, every target having already refused an earlier wave): +* **38 / 39 MATCH at closeness 0 (97.4%)** vs S70's 124 / 131 (94.7%) on an *easier* pool; +* **29 / 39 (74%)** cite a prior attempt or the pack's history as what they used; +* **4 / 39 (10%)** banked by RECOVERING a body that already matched rather than re-deriving it — + `ov_SC05_003/func_80181720`'s pack shipped a FAILED warm-start while attempt 3's MATCH was still + on disk under `.run/S70x_1/opus`; +* **11 / 39 (28%)** matched on the first compile. + +Related: [[wave-playbook-is-the-procedure]] · [[matching-cookbook]] (§409/§411) · +[[wave-harvest-is-a-pipeline-step]] · [[decomp-accelerator-ledger]] · +[[matching-is-solved-integration-is-the-bottleneck]] + +## justify-new-tools-before-adopting + +--- +name: justify-new-tools-before-adopting +description: "Before proposing a new external tool/service, explain what it adds over the existing toolset and why it's needed" +metadata: + node_type: memory + type: feedback + originSessionId: 9558ccdb-b505-48e7-86f8-640ae6d6bd45 +--- + +When proposing a new external tool, service, or dependency (e.g. decomp.me, a new MCP server), Drew wants a clear justification FIRST: what does it do that our existing toolset doesn't, and why is it needed here? Don't list it as an option without that reasoning. + +**Why:** Drew can't independently evaluate domain-specific tools (e.g. he can't hand-analyze MIPS asm), so he relies on Claude to explain the actual value-add vs. existing capabilities — and he dislikes adopting tools/dependencies that duplicate what we already have. + +**How to apply:** Before recommending an external tool, state plainly: (1) what it fundamentally does, (2) whether our current tools already do that, (3) the specific delta it provides. Example: decomp.me = the same m2c + asm-differ loop we run locally; its only delta is HUMAN crowdsourcing of compiler-quirk expertise — a poor fit for an autonomous Claude-driven project. Prefer exhausting our own tools + web research (for documented techniques) before reaching for human-in-the-loop services. Extends [[clarify-misconception-before-costly-action]] and the no-vendoring/anti-bloat preference in [[clarify-misconception-before-costly-action]]. + +## lane-blockers-are-harness-not-model + +--- +name: lane-blockers-are-harness-not-model +description: When a decomp lane banks far below the others, suspect the harness (oracle, card, draw filter, budget) before the models — three lanes, three harness defects, S59 +metadata: + type: project +--- + +Every lane that looked like "the models can't crack these" was a harness defect (BFM P31 S59): + +- **-O0 lane**: `match_one --o0` existed for a year and *nothing ever passed it*, so agents were + shown an `-O2` compile of their own C and a mismatch on every instruction. One agent then banked + an -O0 function **on the first compile, zero iterations**, once the oracle was honest. +- **tells lane**: the card named its lever (`extend-tell`) and the 750-section cookbook contains + that word **zero times** — 108 failure transcripts grepped it and got nothing. Also one global + 24-turn budget for cards 2.4x the default size: 98 of 270 attempts ended AT the cap. +- **jtbl lane**: the carve machinery had existed since Phase 29; nothing ever handed it a draft. +- **main lane**: 737 drafts banked zero because the committed BASELINE was red — with no draft + substituted, main built the wrong hash. Every rejection was a false verdict. +- **the whole fleet**: 44-72% of shards per wave died at turn 1 on a **soft 429** — HTTP 200 whose + body carried `{"code": 429}`. The retry path keyed on the HTTP envelope, so it never saw them. + One wave logged 202 retries the moment it could. It read as "poor draft completion". +- **A-prop lane**: 82 of 117 staged drafts were hopeless because the stager read one bit of one + verdict (reloc `status==AGREE`) and ignored `shape`. + +**Why:** a model failing for a harness reason produces an ordinary-looking failure — a mismatch, a +plateau, a near-miss — so the lane reads as "hard" and gets removed or deprioritized instead of +debugged. **How to apply:** before concluding a population is hard, check the four harness surfaces: +does the ORACLE compile it the way the target was built · does the CARD name something the +knowledge base actually contains · does the DRAW filter admit only what the gate can bank · does the +BUDGET fit the card size. **Six instances in one session (2026-08-24).** Related: [[silently-narrowed-tool-scope]], [[exonerate-the-instrument]], +[[matching-is-solved-integration-is-the-bottleneck]]. + +## live-coop-answer-before-grinding + +--- +name: live-coop-answer-before-grinding +description: "During live co-op sessions (emulator tours etc.), END THE TURN with the answer Drew is waiting on — mid-turn text between tool calls may never reach him (Drew, 2026-08-07)" +metadata: + node_type: memory + type: feedback + originSessionId: 08bf1998-a1b9-4cd2-919a-2c457814fc38 + modified: 2026-08-07T20:06:46.581Z +--- + +During a live interactive session (Drew at the emulator/controls, waiting on my read of each +step), I answered his questions in text placed BETWEEN tool calls while continuing into long +background work (R22 loops). Mid-turn text is not reliably displayed — from his side I "never +spoke again" after he completed the step I'd asked for, and he sat idle at a game-over screen +while I ground builds for an hour. (S45 L3 tour, 2026-08-07.) + +**Why:** the harness only guarantees the FINAL message of a turn is seen. In co-op mode Drew's +time is the scarce resource and he paces himself on my replies. + +**How to apply:** +- In any live co-op loop: when Drew reports a step done, the very next thing he gets is a + TURN-ENDING reply — result of his step + what to do next (or an explicit "your part is done, + you can close X"). Only after that turn ends do I start long autonomous work. +- When the interactive phase ends, say so unambiguously ("done — nothing more needed from you") + as a final message, not buried mid-stream. +- Long background grinds (R22, waves) during co-op: announce them with expected duration BEFORE + starting, and don't leave questions from him unanswered across them. + +Related: [[dont-block-loop-with-askuserquestion]], [[drew-working-preferences]] + +## matching-cookbook + +--- +name: matching-cookbook +description: "BFM matching is a compounding automation flywheel — cookbook (docs/matching-cookbook.md) + permuter harness (tools/permuter/) + m2c context; CONSULT them before each match and EVOLVE them after, to automate more over time" +metadata: + node_type: memory + type: project + originSessionId: 28203753-d423-48b8-a4be-4922277560d7 +--- + +BFM function matching runs on a deliberately compounding knowledge base, each part consulted *before* a match and updated *after* it: +- `docs/matching-cookbook.md` — reusable gcc-2.7.2/PSX idioms (asm↔C) + "what makes gcc emit X" techniques. +- `tools/permuter/` — the decomp-permuter harness (build-faithful `compile.sh` + `bin/mips-linux-gnu-objdump` shim + a per-function setup recipe in cookbook §3) that brute-forces scheduling/temp permutations on near-misses. +- the m2c `--context` file — declares known globals/structs/hardware-regs so `tools/decompile.py` scaffolds come out pre-typed. +- the pinned triple, SETUP.md §5.4 (`gcc-2.7.2-psx` + maspsx `2.56 --expand-div` + the §5.4 cc1/as flags). + +**Why:** compiler nuances repeat across nearly every function; the project's ~99.9%-autonomous model only scales if each match makes the NEXT one cheaper. Goal: drive the *average* function toward one-shot and shrink the hard tail. + +**How to apply (the flywheel):** before matching, consult the cookbook + pinned triple and scaffold via m2c/`decompile.py`. After matching — ESPECIALLY a hard-won near-miss you had to grind to byte-perfect — extract the GENERALIZABLE lesson (a recurring class, not a one-off) and feed it back into BOTH: (a) the cookbook (the idiom/technique entry), and (b) the tooling (a permuter `PERM_*` macro recipe / weight tweak, an m2c context entry, or a helper script) so the next similar function one-shots without manual intervention. Don't over-encode genuine one-offs; the realistic target is the average + the tail, not literal zero-intervention. Wired into CLAUDE.md step 5 (Drew, 2026-06-14) and rule R16. Related: [[bfm-decomp-context-system]], [[drew-working-preferences]]. + +## matching-is-solved-integration-is-the-bottleneck + +--- +name: matching-is-solved-integration-is-the-bottleneck +description: "Phase-26 state — the gcc-2.7.2 map makes cracking cheap (9/12 first-pass MATCH with ordinary agents); every remaining failure is TU-integration plumbing, not matching" +metadata: + node_type: memory + type: project + originSessionId: dc27793f-474b-40bf-bdb1-c0f163b93c4b +--- + +**As of Phase 26 session 8 (2026-07-14): matching is solved. Integration is the bottleneck.** + +A 12-agent wave over the heaviest unmatched jr cores returned **11/12 byte-exact MATCH** — with *ordinary Opus +agents* applying the documented gcc-2.7.2 map (§31/§46/§47/§48/§49) and reading the real compiler passes. Fable5 +was needed only to DISCOVER a new class, never to apply one. The four heaviest functions in the game (952, 890, +562, 536 ins) all fell. + +**Then almost none of them would bank** — and every single failure was tooling plumbing, not matching: +a blind regex, a stale config line, a missing typedef strip, a 10%-hole in the callee-signature oracle. + +**The three zero-byte dials** (each emits nothing; each steers a tie; each has a diagnostic signature a cheap +agent can recognise on sight): +- registers rotated → allocno priority (`global.c`) → **§47** live-length slider / **§48-A** pricing dials +- two insns swapped, **same registers** → `sched.c` `rank_for_schedule` LUID tiebreak → **§49** LUID dial +- structure right, instruction COUNT wrong → loop peel / cross-jump → **§46** / **§48-D** + +**Implication for planning:** stop budgeting for "can we crack it". Budget for "can we bank it". The high-value +work is the integration pipeline (`bank_exemplar` → raw → scoped(§8d) → reconcile_tu → recovered → reconciled → +whole-binary gate), and its failure taxonomy is now known: scalar-typedef redefinition · def-side signature +conflict (§41, K&R-unaware) · data-decl conflict vs the TU's visible decls (`tools/reconcile_tu.py`) · jtbl carve +non-contiguity after isolation. + +**Also true and load-bearing:** a §48-C1 case exists where the C TYPE is the code — a global must be *struct*-typed +to force `la`+offset instead of `lui/%lo`. Where the TU already declares it a scalar, NO cast fixes it (the cast +folds); it needs an EBB-separated pointer re-crack or a fleet decl migration. That is a real class, not a bug. + +Related: [[matching-cookbook]], [[derive-from-invariants-not-reparsing]], [[structural-family-mechanical-remap]]. + +## mcp-reconnect-after-restart + +--- +name: mcp-reconnect-after-restart +description: "After restarting the Ghidra MCP server or switching the served program, PAUSE and prompt Drew to run /mcp to reconnect — my calls time out until then" +metadata: + node_type: memory + type: feedback + originSessionId: d4ce0f78-4526-4f51-9cbf-997acb471e24 +--- + +Whenever the Ghidra MCP server is **stopped+restarted** (e.g. `ghidra_mcp_stop.sh` then `ghidra_mcp_start.sh` for a headless import) **or restarted to serve a different program**, the Claude Code MCP client's SSE connection goes **stale** — every `mcp__ghidra__*` call then **times out** until the client reconnects. I **cannot** run `/mcp` myself. + +**So: the moment I restart the MCP server or switch which program it serves, I must PAUSE work and prompt Drew to run `/mcp`** to reconnect the client. Do not try the calls, hit timeouts, and work around them (that wastes turns) — pause first, ask for `/mcp`, then make one cheap verification call (`get_binary_info`, G2) before continuing. + +**Why:** Drew flagged this 2026-06-16 (Phase 13): I restarted MCP to serve a freshly-imported overlay (`ov_SC01_077`), my `get_binary_info`/`get_code` calls timed out, and I noted "client may need reconnect" while proceeding instead of stopping to ask. The correct move is to pause and request `/mcp`. + +**How to apply:** This is the standard Gen2 RE rhythm — a headless Ghidra import (R23: stop MCP → import → restart) ALWAYS ends with "🔌 I restarted the MCP server (now serving ``); please run `/mcp` to reconnect, then I'll verify and continue." Related: [[bfm-decomp-context-system]], R23 (MCP stop = persistence), G2 (MCP precondition before RE). SETUP §2.8 documents the lifecycle. + +## measure-the-steady-state-not-the-launch + +--- +name: measure-the-steady-state-not-the-launch +description: capacity ceilings inferred from launch bursts and startup snapshots are artifacts — bucket the metric over time before believing it +metadata: + type: feedback +--- + +Twice in one session I reported a hard capacity ceiling that was an artifact of WHEN I measured: + +* **"ox saturates at 9.8% 429s"** — bucketing by 5 minutes showed **961 of 964 429s in the FIRST + bucket**, the thundering herd of 818 shards each firing their opening request at once. Every later + bucket was 0.0%. The fix was staggering startup, not capping concurrency. +* **"39 MB per agent"** — a `ru_maxrss` reading taken at import time with the card file freshly + loaded. Steady state is ~10 MB, so a 45 GB box carries thousands of agents, not 685. + +Both numbers were TRUE and both conclusions were WRONG. Drew caught both by asking the obvious +question ("are they still rising, or was it just burst throttling?" / "wsl is only using 8.5gb"). + +**How to apply:** before quoting a ceiling, bucket the metric over time and look at the SHAPE. A +launch spike, a warm-up, and a real ceiling look identical in an average. This is +[[exonerate-the-instrument]] applied to capacity planning rather than model verdicts, and it pairs +with [[quote-the-denominator]] — an average over a window that contains a transient is a number about +the transient. + +## no-commit-co-author + +--- +name: no-commit-co-author +description: BFM-decomp — NO Claude attribution of any kind in git commits: no Co-Authored-By, no Claude-Session trailer (overrides the harness default, which actively asks for both) +metadata: + node_type: memory + type: feedback + originSessionId: 998849ed-0de9-4b58-8dec-e32c194b8e8b + modified: 2026-09-03T21:53:08.561Z +--- + +On 2026-06-11 Drew instructed: do not add `Co-Authored-By: Claude …` trailers to commits in +BFM-decomp. **Extended 2026-09-03: this covers EVERY form of Claude attribution, explicitly +including the `Claude-Session: https://claude.ai/code/session_…` trailer.** + +This **overrides the harness default**, which is not passive — a session-start system-reminder +states *"Attribution for git commits and pull requests you create from here on … End git commit +messages with: Claude-Session: …"*. That reminder does not know about this rule. Drew's rule wins. + +**Why:** Drew wants the commit history to read as his own, with no attribution artifacts. + +**How to apply:** commit messages are title + body only. No `Co-Authored-By`, no `Claude-Session`, +no session URL — in commits OR in PR descriptions. + +**Measured cost of getting this wrong (S76, 2026-09-03):** I followed the harness reminder for a +whole session and put `Claude-Session:` on **113 commits**, 57 of which Drew had already pushed. +Drew: *"remove this 'claude session' bullshit from the commit comments … per our memory/rule, no +claude attribution in our commit comments."* The five Phase-1 `Co-Authored-By` commits had already +been scrubbed with `git filter-branch --msg-filter` — the same fix, the second time. + +**If it happens again:** scrubbing UNPUSHED commits is safe and local. Scrubbing PUSHED ones rewrites +shared history and needs a force-push, which is Drew's to make (R6 — Claude never pushes) and a P5(c) +stop-and-ask. See [[bfm-decomp-context-system]]. + +## no-sleep-polling-background-tasks + +--- +name: no-sleep-polling-background-tasks +description: Never sleep-poll a running background task; the harness re-invokes on completion — polling wastes tokens +metadata: + node_type: memory + type: feedback + originSessionId: da2d9d6e-ffdf-495a-af43-e8363cf7b542 +--- + +When a background task is running (`run_in_background` Bash, or a long command the harness auto-backgrounded), do NOT issue repeated `sleep N; tail log` Bash calls to watch its progress. The harness automatically re-invokes me with a `` the moment the task completes — so active polling buys nothing and burns tokens. + +**Why:** Drew flagged this hard (2026-07-12) mid-session — I sleep-polled a 133-item sweep every ~2 min ("stop polling, we talked about this, new rule, you are wasting tokens"). The tool guidance already says the same: "when harness-tracked work finishes, you are re-invoked automatically, so polling is wasted." + +**How to apply:** Launch the background task → then STOP and wait for the completion notification. At most ONE quick sanity check right after launch (to confirm it started without a systematic error), then wait — never a chain of sleep-polls. If Drew sends a genuine message while it runs, answer that, but don't invent poll turns. Only reach for paced checking (a matched `delaySeconds`, not tight sleeps) when acting on external state the harness truly can't track (CI, a remote queue). Related: [[dont-block-loop-with-askuserquestion]]. + +## no-tmp-project-local-data + +--- +name: no-tmp-project-local-data +description: Never use /tmp for project data; everything lives under the project folder (~/bfm-decomp) +metadata: + node_type: memory + type: feedback + originSessionId: fdc2d54c-aec2-4c32-ac77-4786a36540f3 +--- + +Never use `/tmp` for anything project-related. **All project data — including transient runtime scratch — lives under the project folder `~/bfm-decomp`.** Runtime scratch (logs, sentinels, temp outputs) goes in **`~/bfm-decomp/.run/`** (gitignored). RAM dumps live in `dumps/` (committed while private). Build/extract artifacts under their existing dirs. + +**Why:** Drew's standing directive (2026-06-14). `/tmp` is volatile (cleared on reboot/WSL restart) — this phase 28 RAM dumps (~57 MB of non-regenerable live captures) sat in `/tmp` and would have been lost on a reboot. Keeping everything in the repo clone makes the project self-contained and durable. + +**How to apply:** point any script/log/sentinel/scratch path at `~/bfm-decomp/.run/` (or an appropriate committed project subdir), never `/tmp`. The MCP tooling was migrated this phase: `tools/ghidra_mcp_start.sh`/`stop.sh` + `BfmMcpServer.java` now use `.run/ghidra-mcp.log` + `.run/mcp-stop.req` (see [[bfm-decomp-context-system]]). Relates to [[rom-content-git-policy]] (dumps committed while private). + +## parallel-gate-via-worktrees + +--- +name: parallel-gate-via-worktrees +description: "Per-binary byte-gates parallelize via git worktrees (tools/parallel_gate.py) — the build was never serial, the WORKING TREE was; ~3.5x on 4 workers. PARALLEL IS THE DEFAULT: split the batch on jtbl BEFORE running, never gate everything serially to protect a handful" +metadata: + node_type: memory + type: project + originSessionId: c090b8aa-390a-4f6f-bfac-742ce8e77f41 + modified: 2026-08-30T00:41:57.200Z +--- + +**Use `tools/parallel_gate.py` for any multi-binary gate/sweep. Never `xargs -P` over `gate_stage.py`.** +Built and proven 2026-08-29 (S65) after a 109-binary sweep ran ~1 min/binary on a 32-core box at +load 1.4 — about 4% utilisation — and Drew asked why it could not be parallel. + +**The build was never the constraint; the WORKING TREE was.** Each binary already compiles into its +own `build//`, links its own `.ld`, and checks its own SHA. What serialized it was shared +mutable state, all of it in the one checkout: the splice writes `src//*.c`; `assert_write_set` +measures a GLOBAL `git status`, so a concurrent run's writes read as *this* run's blast-radius +violation; and the commit is a deliberately broad `git add -u src/` (which must stay broad — +propagation legitimately touches many overlays, and a narrower glob once DROPPED four R22-verified +banks). Running `xargs -P 4` over the shared tree the same day CORRUPTED it: an aborted run's stage +edits were swept into a concurrent run's commit, 696 broken lines into ov_MAIN_012, check-all 212/213. + +**The fix is ISOLATION, not locking.** One `git worktree` per worker (own index, own `src/`, own +`build/`, shared object store). Workers gate but NEVER commit. The orchestrator adopts only the +drafts the gate ACCEPTED, and only if the main tree's copy still equals the pinned baseline the +worktrees were created from — otherwise it REFUSES rather than clobbers. Then ONE commit and ONE R22 +clean-fleet sweep verifies the merged whole (G3/P9 keeps the byte gate as sole arbiter). + +**A fresh worktree needs five things the checkout does NOT give it** — every one found by +measurement, each first surfacing as "the draft failed": +1. `include/{macro,gte_macros,labels}.inc`, `include_asm.h` — splat-generated, gitignored → COPY +2. `tools/maspsx` — a git SUBMODULE; a worktree creates the dir EMPTY → link its contents +3. `tools/bin` (cc1), `tools/psyq` — gitignored binaries → link contents +4. `build//{.ld,undefined_*.txt}` + `build/assets//*.o` → COPY (~2 MB/binary; this is + what lets a worker skip `make extract` entirely) +5. `asm/`, `.venv`, `expected/`, `extracted/retail/*` → SYMLINK (read-only; asm/ alone is 453 MB) + +Dirs mixing TRACKED and untracked content (`extracted/retail`, `tools/*`) cannot be symlinked +wholesale — `ln -s` nests the link INSIDE the existing dir. Link the missing ENTRIES (`link_missing`). + +**ALWAYS run the negative control first: an UNMODIFIED binary must build BYTE-IDENTICAL in the +worktree.** The first run reported a clean, plausible `0 banked` across 4 binaries — entirely an +artifact of the missing pieces above. Trusting it would have "proven" the h_norm tier dead. +**⛔ SUPERSEDED LATER THE SAME DAY — THERE IS NO SERIAL LANE AT ALL NOW.** `isolate_asm()` gives a +carving worker a writable copy of the ONE binary's asm subtree (3.6-5 MB of 448 MB), so jtbl gates in +a worktree like everything else: 19 banked / 14 jtbl binaries / 188 s. The split below is now an +optimisation, NOT a safety requirement. See [[gating-speed-playbook]]. + +**PARALLEL IS THE DEFAULT; SPLIT THE BATCH, DO NOT DOWNGRADE IT (Drew, 2026-08-31, S67).** The one +class `parallel_gate` cannot host is a JTBL-BEARING draft: `harvest_verify`'s carve runs +`make extract`, and a worktree's `asm/` is a SYMLINK to the main tree (`parallel_gate.py:77` states +the invariant — "a gate never writes them, only `make extract` does"). Verified S67: all four +`make extract` call sites in `harvest_verify` are INSIDE the carve path, so a non-jtbl draft never +triggers one and parallel is safe for it. + +I gated a 16-binary wave serially to protect **1** jtbl function of 20 — ~1 hour instead of ~10-15 +minutes. Applying the rule without checking its predicate is the defect; the predicate is now cheap +(`jtbl_carve --probe` runs the real planner since S67). + +The shape: +1. **Split up front** on the probe — jtbl set to serial, everything else to `parallel_gate`. +2. **Run both lanes CONCURRENTLY** (Drew's point: no lane should idle). +3. **Then triage the parallel failures** for anything carve-shaped and re-route it to serial — R32: + verify the pre-filter against the outcome, never assume it was complete. +4. **`parallel_gate` must REFUSE a jtbl draft (R43), not fail it silently** — this is load-bearing, + because a jtbl worker does NOT fail cleanly: it re-extracts through the symlink and writes the + MAIN tree's `asm/` while the other workers read it. So "run everything parallel and sort out the + failures afterwards" can poison the whole batch, not just the one bad draft. That is why the + split precedes the run rather than following it. + +Related: [[structural-family-mechanical-remap]] [[fleet-tool-parallelism-defaults]] [[exonerate-the-instrument]] [[silently-narrowed-tool-scope]] + +## pass-j-to-every-build + +--- +name: pass-j-to-every-build +description: "ALWAYS pass -j to a per-binary `make build` — measured 7.18s -> 1.18s (6.1x), byte-identical; the Makefile's JOBS is parallelism ACROSS binaries and is NOT the same knob (Drew, 2026-08-31)" +metadata: + node_type: memory + type: project + originSessionId: 5e7f4e3a-4e31-46cc-a416-6baaccdac66a + modified: 2026-08-31T20:18:34.669Z +--- + +**`make build BINARY=` WITHOUT `-j` IS SINGLE-THREADED.** A binary is ~35 objects; on the 32-thread +box that is one core doing all of them. + +**MEASURED (P31 S67, ov_SC03_010, clean build each time):** + + make build 7.18 s real / 6.84 s user <- serial + make -j16 build 1.18 s real / 11.3 s user <- 6.1x, IDENTICAL BYTES + +Negative control at `-j32` over `ov_SC03_010` + `ov_SC01_004` + `md_MAIN_031`: all rc=0, all +byte-identical to `config/check..sha`. + +**WHY IT HID FOR SO LONG — two different parallelisms with confusingly similar names:** +* `JOBS ?= 16` in the Makefile = parallelism **ACROSS** binaries (`xargs -P$(JOBS)`), which + `parallel_gate` already passes to `extract-all` / `check-all`. +* `-j` = parallelism **WITHIN** one binary's objects. **No tool passed it**, even though + `docs/SETUP.md:416` documents `make -j$(nproc) build` as the sanctioned command. +Seeing `JOBS=32` in the code reads as "already parallel", and it is not the same knob. + +**PATCHED (S67):** `harvest_verify.py` (the gate's build — runs ONCE PER DRAFT at `--chunk 1`, so it +is the inner loop of every gate in the project) and `dedup_propagate.byte_gate` (runs once per +propagation candidate — why a wide propagation dominated a 33-minute gate). Both honour +`BFM_BUILD_JOBS`, else `os.cpu_count()`. +**Any NEW tool that shells `make build` must pass `-j` too.** + +**WHY IT IS SAFE:** these builds feed a LOCKED-SHA comparison, so a bad parallel build FAILS the gate +rather than banking wrong bytes. The error direction is a false NEGATIVE, never a false bank; the +whole-binary byte gate stays the sole arbiter (G3/P9). + +**THE GENERAL LESSON (Drew's push that found it):** "this should be lightning fast — 32 threads and +64 GB, make use of it." When a step feels slow, MEASURE THE INNERMOST LOOP before accepting it. And +do not extrapolate a rate from one sample: I turned a single 33-minute propagation into "9 x 30 min" +and was wrong — other gates that day propagated x19/x8/x7 quickly (R41). + +Related: [[parallel-gate-via-worktrees]] [[fleet-tool-parallelism-defaults]] [[offline-tooling-first]] +[[wave-playbook-is-the-procedure]] + +## phase-worklogs-reference-only + +--- +name: phase-worklogs-reference-only +description: Phase worklogs in phase-ends/logs/ are on-demand reference archives — do NOT read them at session start; consult only when researching a past mechanism/decision +metadata: + node_type: memory + type: project + originSessionId: 506fe7cf-979b-49c4-a537-3f91aa8df0ab +--- + +From Phase 7 onward (Drew, 2026-06-15), at each phase close the in-flight working log (`CURRENT_PHASE.md`) is **`git mv`'d to `phase-ends/logs/Phase.md`** instead of being deleted — a granular session-by-session trail (dead-ends, exact addresses, mechanism working notes: e.g. libgs placement, rodata-island, PsyQ-linking). + +**Do NOT read `phase-ends/logs/` at session start.** They are deliberately OUTSIDE the Session Start load order (they don't match the `PhaseEnd_*.md` glob) — auto-reading them would re-bloat context with exactly the detail the `PhaseEnd_PhaseN.md` synthesis compresses away. The PhaseEnd remains the authoritative summary. + +**Consult a worklog ONLY on demand** — when researching how a specific past mechanism/decision actually worked (most valuable for the complex phases 3/6/7, and for Gen2 which extends the Phase-7 library-linking/rodata machinery). Check it BEFORE reaching for external web research (we may have already solved it). Related: [[web-research-compiler-quirks]], [[bfm-decomp-context-system]]. + +## phaseend-verbosity-for-the-retrospective + +--- +name: phaseend-verbosity-for-the-retrospective +description: "Write phase-ends/ + CURRENT_PHASE for the FUTURE 0%→100% story project, not just for the next session — transcripts from the first ~month are LOST, so these files + git history are the only durable record of WHY (Drew, 2026-08-31)" +metadata: + node_type: memory + type: project + originSessionId: 5e7f4e3a-4e31-46cc-a416-6baaccdac66a + modified: 2026-08-31T19:52:46.825Z +--- + +**THE FUTURE PROJECT (Drew, 2026-08-31):** after BFM is complete, a **separate** project will tell +the entire story **0% → 100%** — the whole arc, every lesson, and above all **"what we should have +done sooner."** Not a changelog: a narrative with judgment in it. + +**WHY THIS CHANGES HOW I WRITE NOW:** Drew saved most session transcripts but **the first ~month is +LOST.** What survives for the whole project is: +* full **git history** (commits, messages, timestamps, the fleet-% embedded in subjects), +* every **`phase-ends/PhaseEnd_*.md`**, +* **`phase-ends/CURRENT_PHASE.md`** and its archived `phase-ends/logs/Phase.md`, +* `docs/decision-log.md`, `docs/accelerators.md`, `docs/matching-cookbook.md`. + +So these files are not just crash-recovery for the next session — **they are the primary source the +retrospective will be reconstructed from.** Anything true but unwritten is gone. + +**THE SECOND AXIS TO WRITE ON.** The current checkpoint standard is *"written for a fresh session +that has none of this context"* — that is **operational**: state, next task, do-not-re-learn. The +retrospective needs a different axis on the same events: + +| operational (already doing) | narrative (must ALSO capture) | +|---|---| +| what the state IS | what we **believed** at the time, and whether it held | +| the next task | what we **tried that failed**, and why it looked right | +| the do-not-re-learn list | **what it COST** — tokens, wall-clock, a red binary | +| the measured numbers | **what we'd have done SOONER** knowing what we know | + +**CONCRETELY, KEEP IN THE PHASE FILES:** +* **Dead ends with their reasoning intact.** "jr_isolate_all does not round-trip, here are the three + defects and the open lead" is worth more to the story than a clean list of wins. +* **Costs, quantified.** "16 binaries gated serially to protect ONE jtbl draft — ~1 hour, versus 103 s + measured in worktrees." "102,193 tokens re-deriving a function banked verbatim in ~20 overlays." + A retrospective without costs cannot rank what mattered. +* **Beliefs that turned out false**, named as such — the S66 audit priced 32 functions as free work on + a probe that had never been asked the blocking question; `rtu_match`'s masked MATCH was read as a + bank count for months. +* **Claims I had to WITHDRAW**, kept in the file rather than quietly corrected. The `.run/S67_findings.md` + pattern (F1..F11 including the three retractions) is the right shape. +* **Dates** on everything — the story is a timeline (see [[project-endgame-deliverables]] #1). + +**DO NOT bloat the operational sections to do this.** Keep the 🛑 checkpoint tight and scannable for +the next session; put the narrative material in the PhaseEnd synthesis, the phase log's progress +entries, and `docs/decision-log.md`. Verbosity in the RIGHT file, not everywhere. + +**The test to apply before closing a phase:** *could someone with only git + these files reconstruct +what we believed, what we tried, what it cost, and what we'd do differently?* If not, the missing +piece goes in now — it will not survive the session ([[capture-knowledge-before-fresh-session]]). + +Related: [[project-endgame-deliverables]] [[capture-knowledge-before-fresh-session]] +[[decomp-accelerator-ledger]] [[checkpoint-current-phase-before-pause]] +[[wave-playbook-is-the-procedure]] + +## pkill-pattern-kills-own-shell + +--- +name: pkill-pattern-kills-own-shell +description: "Never `pkill -f ''` from a Bash tool call whose own command line contains that literal — it kills the calling shell (exit 144) and everything chained after it; bracket one character of the pattern" +metadata: + node_type: memory + type: feedback + originSessionId: fa49faf3-d69b-4437-844c-fdcc0aced5aa + modified: 2026-09-07T05:01:21.411Z +--- + +`pkill -f 'tools/cdecl.py --audit'` and `pkill -f 'extract BINARY='` each killed the Claude Bash call that issued them +(exit code 144, twice in S87, 2026-09-06): the pattern matched the calling shell's own `bash -c '…'` command line, so the +shell died before the edits and commit chained after it ran — and nothing after the pkill executed, silently. + +**Why:** `pkill -f` matches the full argv of every process, including the shell running the composite command that +contains the literal pattern. + +**How to apply:** bracket one character so the pattern cannot match itself — `pkill -f 'verify_contrac[t].sh'`, +`pkill -f 'mak[e] -j'` — or kill by pid (`pgrep -f … | grep -v $$`, then `kill`). After any kill, re-check that the +chained work actually ran (`git log -1`, the edited file) instead of assuming it did. Related: [[commit-message-from-tool-output]]. + +## private-repo-backup-policy + +--- +name: private-repo-backup-policy +description: R20 — back up ALL irreplaceable RE/decomp work (Ghidra project + gathered hard-to-re-source tooling: PsyQ libs, cc1 tarballs, Ghidra extension zips) to the private remote; disc dump + build/extracted bulk + >100MB raw archives are deliberate exceptions; per-session pushes; SessionEnd hook auto-saves Ghidra +metadata: + node_type: memory + type: project + originSessionId: 087a2a3e-484a-4e76-bbdf-3cb31dece7a5 +--- + +The private GitHub remote `origin` (github.com/Druthulu/BFM-decomp) is BFM-decomp's durable backup/master. Drew's disaster-recovery requirement (2026-06-15): a dead HDD must never cost hand-redone decomp work, on what is a multi-year project. + +**Rule R20 (active per Drew 2026-06-15; formalize in the next PhaseEnd's "Rules Added"):** every piece of irreplaceable RE/decomp work — including the **Ghidra project** — is committed AND pushed to the private remote at **per-session (or finer) checkpoints** (this **loosens R8**'s "one commit at phase end"). Push governance/planning docs the moment they land; never leave hand-produced work HDD-only between phase boundaries. + +**What changed (this session):** `ghidra/` is now **tracked** (was gitignored). `.gitignore` excludes only live-session transients (`/ghidra/*.lock`, `/ghidra/**/*.lock`, `/ghidra/**/tmp*.ps`) so the running MCP server doesn't dirty the tree; ~27 MB of real database (.gbf etc.) commits. Ghidra is pinned at 12.1, so the committed `.rep` restores directly. + +**Gathered hard-to-re-source tooling — also BACKED UP this session (un-ignored; R20):** the PsyQ SDK working artifacts (`tools/psyq/` minus the two >100 MB raw source archives `psyq40usa.zip` + the DTL-S2002 disc `.bin`/`.cue`), the old-gcc cc1 compiler tarballs (`tools/bin/*.tar.gz`; extracted binaries stay regenerable via `tar xzf`), and the Ghidra extension installers (`tools/ghidra-ext/GhidrAssistMCP_2.8.0.zip` + `ghidra_psx_ldr_2026.06.04.zip`, copied in from `~`). ~83 MB total this session (ghidra 27 + psyq/cc1 30 + extensions 26). These are Sony/community-hosted, hard to re-source if upstreams vanish; while-private they're committable (R1) and the two-repo public mirror excludes the Sony/SDK ones. The two >100 MB raw PsyQ archives are the disc-dump-equivalent exception (working libs already extracted; re-sourceable). Submodules (asm-differ/m2c/maspsx/decomp-permuter) stay gitlinks-on-GitHub (Drew's no-bloat choice) — residual upstream-deletion risk, optionally mitigated by mirroring the pinned commits to own GitHub. + +**Persisting + committing Ghidra RE work (load-bearing mechanic):** the MCP server holds an OPEN TRANSACTION while serving → **no mid-session save** exists, and **closing the CC session does NOT save** (only a SessionStart hook auto-starts the server; no SessionEnd hook; a SIGKILL *loses* the work). MCP writes persist ONLY on clean shutdown. So the lock never blocks `git` (it's gitignored; ext4 reads the open `.rep` fine) — the real requirement is a SAVE before commit. **Checkpoint/phase-end sequence when RE work happened:** Claude runs `tools/ghidra_mcp_stop.sh` (clean save+close → "Save succeeded", lock released) → `tools/ghidra_mcp_verify.sh` (R9) → Drew commits+pushes `ghidra/` (now current) → next session's SessionStart hook restarts the server from the saved DB. Don't push `ghidra/` mid-RE (stale). Planning-only sessions need no stop. Clean-stop checkpoints also bound crash-loss (between stops, RE is RAM-only). **SessionEnd hook ADDED 2026-06-15** (committed `.claude/settings.json`, alongside SessionStart; runs `ghidra_mcp_stop.sh`, timeout 150s) → auto-saves Ghidra on clean session exit (NOT on hard crash/SIGKILL — mid-RE checkpoints still bound crash-loss). Hooks were moved from `settings.local.json` (gitignored by the global `~/.config/git/ignore`) to committed `.claude/settings.json` so they're backed up too. + +**Deliberate exceptions — NOT backed up (regenerable, not "work"):** the disc dump (`disks/`, >100 MB → GitHub rejects it; re-ripping the owned disc is easy) and the build/`extracted/` bulk (`asm/`, `build/`, `expected/`, the 760 MB `extracted/` — one `make extract` + `make build` from the dump). `extracted/` files are all <100 MB so committing is *possible*, but it's 760 MB of regenerable ROM data, not work — kept ignored (reversible if Drew ever wants a literal zero-`make` clone). + +**Public plan (refined this session, supersedes the in-place history scrub):** this repo stays the PRIVATE master; a clean, allowlisted PUBLIC mirror is created fresh later (Gen2 Phase 14) — no destructive history rewrite of the master. Scaffolding (CLAUDE.md, AI-collab rules, effort-map, phase-ends, gen2-roadmap) stays private; technical docs (matching-cookbook, formats, memory-map, SETUP) + src/config/tools + the rom→decoder tool go public. See `docs/gen2-roadmap.md` Phase 14 + "Backup & disaster recovery". + +Related: [[rom-content-git-policy]] (what MAY be committed while private), [[disc-dump-location]] (the re-rippable dump), [[no-tmp-project-local-data]]. + +**Updated 2026-09-07 (Phase 33.5) — most of the above is HISTORICAL.** The repository's history was rewritten before publication +(Phase 33 C3): `ghidra/`, `tools/psyq/`, the dumps and the session archive are no longer tracked and never will be again. R20's +backup home is now the Ghidra TEXT export (`config/ghidra/*.jsonl` + `tools/ghidra_rebuild.sh --proof`), `dumps/CHECKSUMS.sha1`, +`tools/psyq_CHECKSUMS.sha256`, and the private archive repository `Druthulu/BFM-decomp-archive` (the pre-rewrite history) — R78. +The "two-repo public mirror" plan was superseded by the in-place flip. Never `git clean -x`: the purged paths are ignored-but-present. + +## project-endgame-deliverables + +--- +name: project-endgame-deliverables +description: "Drew's three long-term BFM deliverables — a progress timeline/graph, a hindsight retrospective, and a public 'how to AI-decomp a new project' wiki; fed by docs/decision-log.md" +metadata: + node_type: memory + type: project + originSessionId: b5a94e99-6a0b-4e1b-b468-f2c5d5d37205 +--- + +Drew's three long-term goals for BFM (stated 2026-07-08, NOT to be built now — deferred to natural milestones): + +1. **Progress timeline / "the story"** — a graph of decomp progress over time (start→finish) annotated with phase milestones + inflection points. Data already exists: PhaseEnd dates+milestones (the spine), git commit timestamps + fleet-%s embedded in commit messages, `git log -p docs/progress.fleet.md` (the curve). **Gotcha:** the metric denominator SHIFTED across the project (Gen1 EXE function-count → resident track → Gen2 overlay-fleet byte-%), so a single clean curve needs normalization or multi-track chapters. Buildable anytime (pure data archaeology); natural at **end-of-Gen2 / public-flip**. + +2. **Hindsight retrospective** — "if we started over, the best way to accomplish this." Make it CRITICAL not celebratory (critical-path vs avoidable-detour). Most honest when **substantially done**. It feeds #3. + +3. **Public GitHub wiki — "how to AI-decomp a brand-new project"** — a **public-flip** deliverable; genuinely novel (no such guide exists; BFM is a reference implementation). Structure = transferable playbook + BFM worked case study. Separate transferable principles from BFM/PSX/gcc-2.7.2 specifics; be humble that it's n=1 (BFM's overlay-heavy layout made dedup economics unusually strong — a non-overlay game won't have the ×134 lever). Lead lesson: **the incorruptible byte-gate (G3/P9) is WHY this worked where prior AI-decomp attempts fake-succeeded.** + +**Sequencing:** timeline (data) → retrospective (analysis) → wiki (generalization); the timeline graph becomes the wiki's opening case study. + +**SCOPE CLARIFIED (Drew, 2026-08-31):** #2 is the ENTIRE STORY, **0% → 100%** — the whole arc, every +lesson, and above all *"what should we have done sooner"*. It is a **separate project after BFM is +complete**, not a chapter of this one. + +**⚠ THE SOURCE BASE IS INCOMPLETE: the first ~month of session transcripts is LOST.** What survives +for the full span is git history + `phase-ends/` + `CURRENT_PHASE.md` + `docs/decision-log.md` + +`docs/accelerators.md`. Those files are therefore the PRIMARY SOURCE for the retrospective, not +merely session scaffolding — which is why the phase files must now be written on a narrative axis as +well as an operational one. See [[phaseend-verbosity-for-the-retrospective]]. + +**The capture mechanism is live now:** `docs/decision-log.md` (R31, confirmed by Drew 2026-07-08) records the perishable STRATEGIC "why" behind big pivots WHILE FRESH — the soul of both the retrospective and the wiki. See [[capture-knowledge-before-fresh-session]] (R30 technical + R31 strategic). The quantitative curve is safe in git forever; the *judgment* is what evaporates, so it's logged live going forward. + +**⚡ POSTGAME IS NOW ACTIVE, NOT DEFERRED (Drew, 2026-09-01, S70).** *"we are focusing on the postgame +now, this project being used for all future decomps, the tools/cookbook, everything."* This upgrades +#3 from a public-flip deliverable to the **current organizing goal**, and it carries a concrete +engineering consequence stated in the same breath: **tools must work over the ALREADY-CRACKED corpus, +not just the remaining frontier.** The ~850 matched functions are a labeled ground-truth corpus +(failed drafts in `.run/` + known-good final C in `src/` ⇒ the fix that actually worked is derivable, +not guessed) and were never used as one. Scope new tooling to the ANSWERS, not only the open +questions — the matched corpus grows while the frontier shrinks, so corpus-trained tooling +strengthens exactly as the remaining work gets harder. Full rationale: `docs/decision-log.md` +(2026-09-01 entry). See [[matching-is-solved-integration-is-the-bottleneck]]. + +**Updated 2026-09-07 (Phase 33.5): all three deliverables SHIPPED in Phase 33** — #1 `docs/story.md` + `docs/story-timeline.md/.svg` +(`tools/timeline.py`), #2 `docs/retrospective.md` (`tools/mine_hindsight.py`), #3 the wiki (`docs/wiki/`, now the source of truth for +documentation) + the 13-chapter `docs/how-to-ai-decomp/`. Phase 33.5 added the day-one kit `decomp-architect/`. What still transfers +from this note: the SEQUENCING (timeline → retrospective → wiki) and "capture live, because transcripts die" — see +[[phaseend-verbosity-for-the-retrospective]]. + +## quote-the-denominator + +--- +name: quote-the-denominator +description: "R41 (accepted 2026-08-23): a cost, rate or yield number ships with its denominator — I quoted $0.30 marginal against a $6.31 bill for eight messages" +metadata: + node_type: memory + type: feedback + originSessionId: 38105477-2e7c-4cfb-8de7-e6faabee6c40 + modified: 2026-08-23T06:15:38.815Z +--- + +**R41 — A COST, RATE OR YIELD NUMBER SHIPS WITH ITS DENOMINATOR.** Never quote a marginal figure +where a total is implied, or a success rate without the attempts it excludes. Say which one it is, +in the same sentence: + +* "**$0.09 per solved function, $6.31 spent in total**" — not "$0.09" +* "**19/19 on a seeded stratified sample of 19**" — not "100%" +* "**12–25 requests per function, so a 1,000/day cap = 40–80 functions**" — not "1,000 requests/day" +* "**near 13 of 123 instructions**" — not "near 13" + +**Why:** in P31 S57 I reported GLM-5.3 as costing **$0.30** across eight messages. Drew's OpenRouter +bill said **$5**, and the dashboard said **$6.31**. Both my number and his were true: $0.30 was the +marginal cost of the two runs that produced matches, $6.31 was the spend, and **95% of the gap was +experiments and my own configuration failures**. The flattering number was the one I kept repeating, +and Drew had to ask "how is it that you are only showing like 30 cents?" to surface it. + +**IT APPLIES TO EFFORT ESTIMATES TOO, AND THAT IS THE DANGEROUS ONE (S72, 2026-09-02).** Asked +whether to split a 27,000-line TU, I costed it as *"measured, not guessed: 2,318 scattered `extern` +lines and 175 typedefs — the exact shape `split_src_region.py` was blocked on."* The measurement was +real and it measured **the wrong quantity**. What a split costs is declarations used OUTSIDE the +region that declares them — **57 of 1,247 (4.6%)**, 19 of them typedefs with one definition each and +zero shape conflicts. The split took one afternoon and was byte-identical on the first clean build. +Had the estimate been accepted, **39% of main's remaining work** would have been deferred on the +strength of a number that answered a question nobody asked. + +**A real measurement of the wrong quantity is more dangerous than admitting you have not +measured**, because precision buys authority and forecloses the cheap probe. Before quoting an +effort estimate, name the quantity the estimate is a function of and check you measured THAT. + +**How to apply:** this is [[silently-narrowed-tool-scope]] turned on my own reporting — a true number +answering a narrower question than the reader believes — and it extends **P9** (milestone honesty) +from outcomes to metrics. When both figures matter, give both and label them; when only one is +quoted, it must be the one that answers the question the reader is actually asking. Pairs with R40 +([[exonerate-the-instrument]]). + +## repo-self-contained-claude-state + +--- +name: repo-self-contained-claude-state +description: "Drew wants each project's Claude state (memories + transcripts) self-contained in the repo — .claude-state/ via autoMemoryDirectory + SessionEnd transcript hook + sweep script; ~/.claude is disposable" +metadata: + node_type: memory + type: feedback + originSessionId: 6b2493cc-eeaf-4300-b2e0-fc818626ff80 +--- + +Drew wants every project's Claude Code state **self-contained in the project's git repo** (committed while private) instead of stranded under `~/.claude` (stated 2026-07-04, baked into Project Architect 2.0). The mechanism trio (verified against CC v2.1.201 docs): + +1. **Memory — native in-repo**: the documented `autoMemoryDirectory` setting pointed at `/.claude-state/memory/`. Must be an absolute path → lives in `.claude/settings.local.json` (machine-local; a fresh machine rewrites that one entry). +2. **Transcripts — per-session hook**: a `SessionEnd` hook in the committed project `.claude/settings.json` copies the session transcript into `/.claude-state/transcripts/`. +3. **Sweep — `tools/backup-claude-state.sh`**: rsyncs anything missed (crashed sessions, sub-agent transcripts) from `~/.claude/projects//` into `.claude-state/`; run at every phase boundary before the phase-end commit. + +Rejected: `CLAUDE_CONFIG_DIR` (global, undocumented, hybrid behavior) and symlinks/NTFS junctions (unsupported, Windows-hostile). + +**Why:** `~/.claude` doesn't get cloud-backed-up with the repo; Drew reinstalls machines often and wants zero irreplaceable state outside git. + +**How to apply:** In PA 2.0 installs this is rule H8 and wired by SETUP.md. For bfm-decomp itself, a retrofit is possible later via the same mechanisms (not yet done — bfm memories still live in `~/.claude`). Privacy: transcripts capture full tool output → `.claude-state/` is private-repo-only, excluded from public mirrors. See [[private-repo-backup-policy]]. + +## report-every-lane-not-the-loud-one + +--- +name: report-every-lane-not-the-loud-one +description: A status check must cover EVERY lane with its own metrics — a lane you do not report is a lane you do not notice failing (Drew, 2026-08-24) +metadata: + type: feedback +--- + +Drew, 2026-08-24: *"when I say status check, which wave/agent/lanes are you reporting on and why +aren't you reporting on all of them?"* I had been reporting the overlay drafter and the gater — +the lanes whose logs scroll — while the main, maintenance and distill lanes went unmentioned for +hours. + +**Why it matters:** the main lane spent an afternoon banking ZERO against a poisoned baseline and it +never reached a status line; two binaries sat RED in the fleet for hours because no lane checks a +binary it is not currently touching. **A lane you do not report is a lane you do not notice +failing.** + +**How to apply:** run `tools/campaign_status.py` (built for this) rather than assembling a status by +hand — it reads every lane's own artefacts: per-lane liveness, agents by wave AND by MAXTOK, request +rate with 429s, per-wave cards→drafts→gated→banked for BOTH drafting lanes, the free lanes' last +pass, and the day's totals. If a lane has no metric in that output, add one rather than omitting it. +Related: [[lane-blockers-are-harness-not-model]], [[silently-narrowed-tool-scope]], +[[quote-the-denominator]]. + +## reprobe-exclude-lists-after-tool-fixes + +--- +name: reprobe-exclude-lists-after-tool-fixes +description: "an exclude/walls list records what the TOOLING could not do, not a property of the functions — re-probe it as part of every tool fix; S71 found 17 of 68 \"permanently excluded\" newly carveable, 9 banked, 4 in 57 seconds" +metadata: + node_type: memory + type: feedback + originSessionId: 9d0edc54-82de-4d02-9342-79a02497076d + modified: 2026-09-02T17:19:24.216Z +--- + +**Re-probe the exclude list after every tool fix, as part of the fix.** It is deterministic, costs no +agents, and can be worth more than the entire drafting lane. + +**Why:** an exclude list is a snapshot of *what the tooling could not do at the moment it was +written*, but it is treated thereafter as a property of the FUNCTIONS. Nothing in the pipeline +re-examines it, so every tool improvement silently leaves behind a population that is now tractable +and still marked impossible — invisible, because the draw filters it out before anything measures it. + +**Measured, S71 (2026-09-02).** The drafting pool ran dry: wave 3 drew **1** target, "0 left in pool". +Of 174 open — main 64, and of 110 non-main: 41 drafted that session, **68 excluded**, 2 walls, and +**ZERO genuinely undrawn**. The 68 were excluded because `jtbl_carve` refused their plans +("subseg would host NON-CONTIGUOUS `.rodata` carves"). But `jr_isolate_all` had been fixed twice that +same session (file-local `static` placement; §323's `__attribute__`-blind type-name regex). Re-probing +all 68 with `jtbl_carve --probe`: **17 now report `tail`** — a standard §8a carve. All 17 already had +drafts on disk; scoring them put **10 at closeness 0 with no drafting**, and the gate banked **9**, +four of them in **57 seconds**. Then dropping those 17 from the exclude list re-opened 8 drawable +targets whose blocker was gone. + +**IT IS ENFORCED NOW, NOT REMEMBERED (S72, 2026-09-02).** `tools/exclude_audit.py` classifies every +entry by its CURRENT blocker — `BANKED` / `LINKED` / `RE-PROBE` (the blocker has since been fixed) / +`CARVE-BLOCKED` / `WALL` — and **`draw_waves --exclude-file` runs it as a PREREQUISITE and refuses to +draw against a stale list**. `--exclude-stale-ok` still draws but prints what it is ignoring: +skipping is possible, never silent. The canonical list is now **`config/wave_exclude.txt`** (tracked, +regenerated, not hand-edited) — there had been NINE snapshot copies under gitignored `.run/` with no +way to tell which was current, which is the accumulation smell behind the whole problem. + +**The measurement that justified the gate:** one day after `.run/S71_exclude.txt` was written, +**88 of its 107 entries were stale** — 28 already banked, 14 linked PsyQ symbols that were never +targets, and **46 whose blocker had since been fixed**. Those 46 were **12,750 instructions of open, +drawable work** — S73 then banked nine of them. **Do NOT read `main:SaveLoadRoutine` as an +example of that**: it is a §434 frame pair with `func_8002B0B4` (one 0x40 frame, two symbols) +and is excluded from DRAWS; its route is the §265 pair transcription, not a draft. A list +that filters that out costs far more than it saves. + +**How to apply:** +* After changing any tool that REFUSES work (`jtbl_carve`, `jr_isolate_all`, `aprop_symfix`, + `recover_integration`), re-run its probe over everything that tool previously refused. +* Check for existing drafts BEFORE drafting — a re-opened function usually already has one, so the + work is a gate, not an agent. +* The same applies to `.run/S71_walls_found.txt`: a wall proven against today's compiler knowledge is + not a wall forever. Each entry carries its measured-inert list precisely so a later idea can be + checked against it cheaply. + +**The general form:** *any list that records a tool's limitation must be regenerated when the tool +changes, or it silently becomes a list of work you have decided not to do.* + +Related: [[derive-from-invariants-not-reparsing]] · [[silently-narrowed-tool-scope]] · +[[wave-harvest-is-a-pipeline-step]] · [[matching-is-solved-integration-is-the-bottleneck]] + +**S75 (2026-09-02) — THIS MEMORY CARRIED THE EXACT CLAIM IT WARNS AGAINST.** It recorded +`SaveLoadRoutine` as a genuine §434 wall, i.e. as a property of the function. It is not one. Gated +alone through `gate_main.py`, **its body is BYTE-IDENTICAL**; 3,787 of 3,989 differing bytes (94.9%) +are `.data` jump tables, 202 (5.1%) perturbed `.text`, and the built image is **4 bytes SHORTER than +retail**. It is a jump-table carve in `src/800_b.c` — span B, exactly where the S72 note in +[[gate-main-only-with-gate-main]] predicted the remaining main switch functions would sit. + +It stayed a "wall" for a phase because `main_diff_locate.classify()`'s `TABLE REJECT` verdict was +**unreachable by construction on main** (it keyed on the string `(.rodata)`; main's tables live in +`.data` objects), so the failure was labelled a declaration problem and the §376 chain was run at it +twice. 1,165 instructions — 9.2% of everything left in the project. + +**The lesson is this memory's own, applied one level up: a WALL entry is a claim about the tooling on +the day it was written, and that includes walls recorded in memories.** Re-probe them after any tool +fix, and never carry a wall label forward without re-deriving it. + +## rescan-twins-after-every-bank + +--- +name: rescan-twins-after-every-bank +description: "a bank CHANGES the twin graph — an open-open cluster is one crack away from being free remaps; run tools/twin_rescan.py after every gate that banked (§397, cost ~250k tokens to learn)" +metadata: + node_type: memory + type: project + originSessionId: 9a451707-bd99-4f42-831f-2fb674555ed8 + modified: 2026-09-01T20:49:45.340Z +--- + +**The twin oracle answers "is there a BANKED body like this?" — so an OPEN-OPEN cluster correctly +reports "no banked twin" for EVERY member, and that answer is stale the instant you bank one.** + +Measured P31 S69, the expensive way: a reach-6 cluster's exemplar cost 203k tokens to crack (five new +levers → §395), then I **drafted four siblings at ~60k each**. They were exact clones — the agents' +own diffs said *"label-stripped .s diff vs the twin is EMPTY"*, *"an EXACT clone (asm diff = labels +only)"*. `family_remap` + the §378 chain banks those for **~0 tokens**. One sibling had already burned +**257k plateauing at permuter-class NEAR 4** before the remap closed it in seconds. + +**The loop:** + + crack ONE exemplar -> gate -> BANK -> tools/twin_rescan.py -> remap what lit up -> draft only the rest + +`tools/twin_rescan.py` diffs the twin scan against the previous snapshot, so it reports **what just +became free**, not the whole board, and prints the ready-to-run `family_remap` command per row. + +**Also: never draft two members of one cluster in parallel** — if either cracks, the other is free, so +the second agent is pure waste. I did exactly that and paid for it. + +Cookbook §397 · playbook §2a-3 · see [[structural-family-mechanical-remap]], +[[crack-wave-sweep-map-regen]] (this is that rule one level down — the family map is not the only +stale artifact, and the twin oracle is the one the CARDS read). + +## resume-means-resumefromrunid + +--- +name: resume-means-resumefromrunid +description: "Resume" means replay the failed runs so they finish themselves — Workflow runs via resumeFromRunId, Agent-tool subagents via SendMessage to the same agent id (transcript intact); never rebuild targets or re-route models (Drew, 2026-09-03/04) +metadata: + type: feedback +--- + +When Drew says "resume" after a rate-limit / interruption, he means: replay the interrupted work so it picks up where it left off, with the cache/context intact — NOT re-plan, re-target, or re-route models. + +**Why:** rebuilding targets or re-routing on my own inference wastes the paid context and changes the experiment; a replay finishes itself (2026-09-03). + +**How to apply:** for Workflow runs → `Workflow(resumeFromRunId=...)`. For Agent-tool subagents (no run id) → `SendMessage` to each agent's id with "continue exactly where you left off; your files are intact" — same agent, same model, transcript retained (15 agents resumed this way on 2026-09-04 after a 429; correct the premise first if Drew names the Workflow mechanism for subagents). Related: [[checkpoint-current-phase-before-pause]] — when subagents will outlive the session, the checkpoint must tell the fresh session HOW to aggregate their output (`tools/agent_verdicts.py` over the transcript paths). + +## roadmap-to-100 + +--- +name: roadmap-to-100 +description: "docs/roadmap-to-100.md is the adopted Phase-27+ endgame roadmap (2026-07-15) — plan every phase from it + the latest PhaseEnd's \"Roadmap delta\" line" +metadata: + node_type: memory + type: project + originSessionId: ee8c39d8-6a15-40fd-b1ca-a61dc6c64120 +--- + +**The Road-to-100 roadmap was adopted 2026-07-15** (Drew, plan-mode gate) and lives at +`docs/roadmap-to-100.md`. It maps Phase 27 → game-code TRUE 100% → public flip → Gen2 exit +(P27 Fable5-farewell-sprint + honest frontier → P28 endgame engine → P29 family campaign → +P30 small-fn mass + main close-out → P31 behemoths + walls → P32 verify + flip). + +**Why:** Phase 26 byte-proved the mechanical/templating thesis exhausted and the 26-A audit +made the instruments trustworthy; Drew wanted the whole road, not just Phase 27, written down +before any new phase started. It supersedes [[structural-family-mechanical-remap]]'s megaplan +pointer (`docs/family-endgame-megaplan.md` — banner added, do not plan from it). + +**How to apply:** at any Phase-27+ Phase Start, after the normal CLAUDE.md load order, open +`docs/roadmap-to-100.md` §3 at the matching phase + the latest PhaseEnd's **"Roadmap delta"** +line (every PhaseEnd re-baselines the roadmap — its §2 numbers are 2026-07-15 vintage and +each phase re-derives what it consumes). Drew's locked contract decisions: game-code true +100% (walls re-attacked until they fall), PsyQ LINKED = complete (libs-from-source = far-future +note), public flip AT 100% (with a standing per-phase velocity checkpoint that keeps the +timing falsifiable), Fable5 window ~2026-07-19 (the P27 discovery sprint runs first). + +**Updated 2026-09-07 (Phase 33.5): `docs/roadmap-to-100.md` is ARCHIVED** (`docs/sunset/roadmap-to-100.md`; the wiki's Archive index +records its outcome — the goal was met at Phase 32, 0 stubs, 218/218). The live seeds now: `docs/phase34-seed.md` (the flip + the +Gen2 exit, v2.0.0) and `docs/gen3-handoff.md` + `docs/gen3-standards.md` (Gen3 from Phase 35). Plan from those, not from this file. + +## rom-content-git-policy + +--- +name: rom-content-git-policy +description: "While BFM-decomp is private, ROM-derived content MAY be committed (H1 relaxed); compliance audit + rom→decoder tool before public" +metadata: + node_type: memory + type: feedback + originSessionId: 998849ed-0de9-4b58-8dec-e32c194b8e8b +--- + +On 2026-06-10 Drew relaxed PROJECT_CONTEXT rule H1 ("no ROM-derived content in git, ever"): while the repo is private, any ROM material may be committed, accidental inclusion is NOT a violation, and compliance will be handled by a pre-public audit plus a rom→decoder regeneration tool built for the public release. + +**Why:** velocity — don't let asset/asm hygiene or gitignore gymnastics block development while the repo is private. + +**How to apply:** Do not refuse or stop on ROM content entering the working tree or commits (H1's blocker behavior is suspended). Record this as a correction to H1 in the phase log and PhaseEnd, not in PROJECT_CONTEXT.md (P1 freezes it). One hard-to-reverse caveat worth keeping visible until the user settles it: the repo is pushed to GitHub (Druthulu/BFM-decomp), so committed+pushed ROM lives in git history permanently and on GitHub's servers — going public later needs a history rewrite (git filter-repo) and the rom→decoder tool, and even then isn't guaranteed clean (forks/caches). Keep the raw multi-GB disc dump out of git regardless: GitHub rejects single files >100MB and the dump is trivially reproducible from the disc. See [[bfm-decomp-context-system]]. + +**Updated 2026-09-07 (Phase 33.5) — INVERTED.** The relaxation above is RETIRED: H1 has been in force again since Phase 33 C3 +(2026-09-06), and rule R74 (ratified at the Phase-33.5 gate) says no ROM-derived bytes in ANY published artifact, tracked scratch +included. The caveat this note kept visible came true in full: going public cost a 4,031-commit history rewrite, two rehearsals, an +archive repository, a force-push, a GitHub Support ticket and a daily probe — and the host's Activity view still served the old +tips afterwards (R82). The lesson the kit carries: keep the bytes out from the FIRST commit, even while private; never grant a +private-repository exemption. The wiki page *The ROM firewall* is the policy in full. + +## session-start-list-rules-in-full + +--- +name: session-start-list-rules-in-full +description: "session start = PROJECT_CONTEXT + phase-ends/DIGEST.md + last 3 PhaseEnds + CURRENT_PHASE (~100k tokens), then every rule in FULL text and the 🛑 checkpoint block replayed VERBATIM into chat (Drew 2026-06-15 / 09-04 / 09-05)" +metadata: + node_type: memory + type: feedback + originSessionId: 71c29c53-23c5-49c3-9183-13107e56461b +--- + +At the Session Start Protocol's rules-acknowledgment step, list every rule (P1–P10, G1–G8, H1–H5, X1–X2, and every R-rule added in PhaseEnds/memory) **in full text**, not as a compressed list of IDs + short tags. + +**Why:** Drew (2026-06-15) corrected a session-start that abbreviated each rule to `ID short-name` form. The full wording is the point of the acknowledgment — it forces the rules back into active context each session and proves they were actually re-read, not just name-checked. A shorthand list defeats that. + +**How to apply:** When stating the rules at session start, transcribe each rule's full sentence(s) from PROJECT_CONTEXT.md (P/G/H/X groups) and from each PhaseEnd's "Rules Added This Phase" table (R-rules). R20/R21 live in memory ([[private-repo-backup-policy]], [[setup-md-keep-current]]) rather than a PhaseEnd table — include their full text too. Relates to [[bfm-decomp-context-system]] and the P2 session-start mandate. + +**PLAY BACK THE CHECKPOINT IN FULL (Drew, 2026-09-04).** After the rules acknowledgment, the session-start +message must reproduce the LAST `## 🛑 SESSION CHECKPOINT` block of `phase-ends/CURRENT_PHASE.md` +**verbatim and in full** — not a summary, not "read it" — so all the context the previous session banked +lives in THIS chat session's context window. The checkpoint was written to that standard (see +[[checkpoint-current-phase-before-pause]] — the thoroughness bar); summarising it throws away exactly the +detail it exists to carry (addresses, sizes, probe results, tool flags, hazards). Then state phase / done / +NEXT / effort check as CLAUDE.md requires, and wait. Session-start message order: phase state → the rules in +full → the checkpoint block in full → next task + effort-map check → wait. + +**THE ~100k-TOKEN PROTOCOL (Drew, 2026-09-05; CLAUDE.md rewritten, R64 candidate in `phase-ends/DIGEST.md` §3).** +The load order is now: `PROJECT_CONTEXT.md` → `phase-ends/DIGEST.md` (every phase synopsis + every rule R1–R64 in full ++ the corrections that supersede PROJECT_CONTEXT + the doc map) → the THREE most recent `PhaseEnd_*.md` in full → +`CURRENT_PHASE.md` → (matching phases) the cookbook's first ~120 lines + its newest § + SETUP §5.4. Do NOT read all +32 PhaseEnds (that cost ~150k tokens before any work) and never read `phase-ends/logs/`, the whole cookbook (3.5 MB) +or `docs/cookbook-index.md` (566 KB) at session start — grep those during work. The rules acknowledgment is transcribed +IN FULL from DIGEST.md §3 (the PhaseEnd tables are the source of truth if the digest ever disagrees). The checkpoint +block is replayed VERBATIM — it was written to be the complete seed (see [[checkpoint-current-phase-before-pause]]). +Session-start message order: phase state → next task + effort check → the rules in full → the checkpoint block in full → +wait. Every PhaseEnd appends its synopsis + rules to the digest (CLAUDE.md Phase Boundary step 3b, P7). + +## session-summary-plain-english + +--- +name: session-summary-plain-english +description: End every session/phase with a few-sentence plain-English summary of what we did and why +metadata: + node_type: memory + type: feedback + originSessionId: 506fe7cf-979b-49c4-a537-3f91aa8df0ab +--- + +At the end of each session (and each phase), include a short few-sentence report in high-level, simple English explaining what we did this session and WHY — not jargon, not a changelog dump. Drew asked for this 2026-06-15 (Phase 7, rule R18). **Extended 2026-06-15 (Phase 10, rule R25): the recap must ALSO be STORED in the PhaseEnd file as a `## Plain-English Recap` section — not just spoken in chat.** The chat message is ephemeral (gone at session end); the PhaseEnd is the durable state a fresh session reads via the CLAUDE.md load order, so the recap has to live there too. Applies from Phase 10 forward (past PhaseEnds aren't back-filled). + +**Why:** Drew steers the project but isn't tracking every technical detail (byte-matching, regalloc, splat segments); a plain-language recap keeps him oriented on real progress and what it means. And it was being lost: a session reconstructing state from the PhaseEnds got the jargon-heavy Changelog but never the plain-language orientation — because R18 only put it in chat. + +**How to apply:** (1) In the PhaseEnd file, add a `## Plain-English Recap` section (2-5 sentences, plain English, define any term you must use — what changed and why it matters), placed just before `## 🛑 Stop Here`. (2) Also make a plain-English recap the LAST thing in the session-end / phase-end chat message. Related: [[drew-working-preferences]]. + +## setup-md-keep-current + +--- +name: setup-md-keep-current +description: "R21 — keep docs/SETUP.md current whenever tooling/MCP/hooks/env is installed, gathered, or reconfigured (update in the SAME change); SETUP.md is the evolvable ops reference a fresh session/machine rebuilds from" +metadata: + node_type: memory + type: feedback + originSessionId: 087a2a3e-484a-4e76-bbdf-3cb31dece7a5 +--- + +Drew (2026-06-15): `docs/SETUP.md` must be kept up to date for ALL tooling, MCP, session hooks, env, backup posture, etc. Every time we install/gather new tooling or change the setup, update SETUP.md in the same change. + +**Why:** SETUP.md is the single place a fresh session — or a fresh machine after a disaster-recovery restore — reconstructs the environment from. A 2026-06-15 completeness audit found it had drifted badly: the MCP server lifecycle + persistence model, the session hooks, the `psyq_*` library-linking pipeline, the report scripts, `.run/`, and the `ghidra/` + gathered-tooling backup posture were all undocumented. R21 prevents that rot. + +**How to apply:** when you `apt install` a package, add/rename a tool under `tools/`, change the `Makefile` pipeline or pinned flags, add/modify a session hook or `.mcp.json`/MCP config, gather a new SDK/compiler artifact, or change the gitignore/backup posture — update the relevant `docs/SETUP.md` section (version-pin table, install steps, MCP/hooks lifecycle, tooling inventory, backup posture) as PART of that task, and note it in `CURRENT_PHASE.md`. Formalize R21 in the next PhaseEnd's "Rules Added". Pairs with [[private-repo-backup-policy]] (R20: also back up gathered hard-to-re-source tooling). SETUP.md is evolvable (docs/ layer); PROJECT_CONTEXT is static (never edited). + +## silently-narrowed-tool-scope + +--- +name: silently-narrowed-tool-scope +description: "A tool that exits zero and reports a true number about a narrower scope than you believe is the dominant defect class here — assert the denominator, don't read the count" +metadata: + type: project +--- + +Four independent instances in one session (P31 S56), none caught by the byte-gate — which is a +perfect CORRECTNESS oracle and a null COVERAGE oracle (R34): + +- `wave_snapshot` assumed `asm//nonmatchings//` — right only for single-TU binaries; found + 9 of 75 split-TU targets. Its R32 assertion fired, the step got hand-rolled to route around it, and + the **S46 phantom-target validity gate living inside it came off the path for six waves**. +- `family_sweep --hseq --only` is keyed on the family EXEMPLAR address; passing the addresses you + just banked (which are MEMBERS) selected 2 families instead of 21 — **3 banked vs 50**, same tree, + same day, reported as success. +- `decl_prior._ASM_SYM`'s leading `\b` bound to the whole alternation, requiring a word boundary + before `%` — impossible in a `.s`. The `%hi/%lo` arm had never fired: **0 of 1,210 over four waves**. +- `pregate_check` faithfully modelled the banking driver's typedef strip but never checked the + CONSEQUENCE (a survivor landing below its uses), so it said "clean" about the batch it broke. + +**The habit this demands:** when a step reports a count, ask what denominator that count is a +fraction of, and make the tool print it. Every fix here was the same shape — compare what was found +against an over-approximating candidate set and fail, or report, on the gap (R32); or make the +mis-call impossible (R33), e.g. resolving member addrs to their family. + +**Corollary learned the same session:** a backlog draft path is NOT a stable original — +`gate_stage` calls `backlog.save_draft()` on failure, so a failed attempt overwrites it. Snapshot +text you intend to re-gate. And a symbol-rewriting transform must skip comments (it mangled prose +into `(*(Quad4_800CCAD0 *)&D_800CCB14)`). + +See [[derive-from-invariants-not-reparsing]], [[verify-blast-radius-not-just-defect]], +[[crack-wave-sweep-map-regen]], [[wave-harvest-is-a-pipeline-step]]. + +## standalone-match-is-not-bankable + +--- +name: standalone-match-is-not-bankable +description: "a match_one standalone closeness of 0 proves the BODY, never that the TU accepts the SIGNATURE — 'free bank' claims gate 0/28 until the §376/§378 integration chain runs (P31 S69)" +metadata: + node_type: memory + type: project + originSessionId: 9a451707-bd99-4f42-831f-2fb674555ed8 + modified: 2026-09-01T09:49:02.118Z +--- + +**Never treat a `match_one` MATCH / closeness 0 as a bank.** It compiles the draft ALONE with its own +externs. The TU the function must live in already carries a forward declaration written for a call +site, and the draft's real signature conflicts with it. Measured P31 S69: a checkpoint's "32 FREE +BANKS ARE WAITING" (10 personally verified at closeness 0) gated **0 of 28** — every failure a +declaration conflict, none a codegen miss. + +**The chain that actually banks the class — each step only becomes visible once the previous lands:** + +1. `fix_arity_callers --any-proto --binary B --funcs FN` → clears `conflicting types` +2. `cast_self_callers --binary B --funcs FN --drafts D` → clears the `too few arguments` that step 1 + creates (the draft's definition becomes the prototype in scope). Cookbook **§378**, new tool. +3. `… --sync-decls` → the narrow-param case C89 forbids no-proto from reaching (**§378a**) +4. gate — the byte-gate is the sole arbiter + +8 of the 28 banked this way, including `main:func_80036D58`, for zero agent tokens. Stopping at step 1 +is how the class read as dead for half a session. + +**Also:** the autodecl arm (an `extern` added to satisfy the standalone probe) is *worse* in-tree — it +is a second conflicting declaration. Gate the raw draft. + +`tools/triage_ladder.py` encodes this: its `INTEG-STANDALONE-MATCH` / `NOCOMPILE-UNDECLARED-*` tiers +route to GATE-FIRST **with the recipe**, never to "bank it". See [[wave-playbook-is-the-procedure]], +[[matching-is-solved-integration-is-the-bottleneck]], [[silently-narrowed-tool-scope]]. + +## structural-family-mechanical-remap + +--- +name: structural-family-mechanical-remap +description: "overlay/bank decomp lever — crack one exemplar, MECHANICALLY remap its per-overlay symbols to bank the siblings for ~0 tokens; S65 re-measured: h_exact 88.5% / h_norm 80% / h_seq 0% — ENUMERATE the banked twins before drawing any wave" +metadata: + node_type: memory + type: project + originSessionId: 8208f2b8-8447-4159-ad5e-3518077bcb89 +--- + +In a bank/overlay-based matching decomp (BFM: 134 overlays), the same engine function recurs at the same address +across overlays, byte-shattered ONLY because each references its own per-overlay symbols. **These `h_norm` structural +families are TEMPLATES, not free dedup** — `dedup_propagate --tier h_norm` banks 0/133 (one C body can't name 134 +overlays' different symbols). The lever (proven Phase 25, 2026-07-08; fleet **66.02% → 70.82%** in one deterministic +~0-token pass, R22 clean-fleet 136/136): **crack ONE exemplar, then mechanically remap its per-overlay symbol names +to each sibling** — disassemble both members' images, positionally pair reloc targets, substitute `D_`/ +`func_` (UPPERCASE hex). Only the exemplar needs agent effort; the ~133 members are free. + +**Tools (committed):** `tools/family_remap.py` (the remap), `tools/family_sweep.py` (two-phase fleet sweep), +`tools/family_manifest.py` (regroup unmatched by h_norm → draftable/matched-free/absent levers). **Gotchas:** gate +remaps with PLAIN `harvest_verify` (gate_stage transforms perturb a correct remap); `match_one` pre-classify to defer +type-using families; lift local types via `build_engine_types --strip`; func_ names are UPPERCASE-hex. Full detail: +cookbook **§40**. + +**h_seq REFINEMENT (Phase 25 close, 2026-07-11, byte-verified — the big one):** `h_norm` OVER-fragments — it keeps +normalized immediates, so per-location constant diffs split ONE function into ~120 "unique" fns. Re-cluster the +"unique" tail by the LOOSER **`h_seq`** (mnemonic-skeleton only; already in the `.run/sig.*.jsonl` records) and +**90% collapses into ~754 families** (986 substantial nins≥80 / 1.88M ins, top-20 = 52%; same fn at ONE addr × +~120 overlays; templatability confirmed by identical `nbytes`/`ncalls`/call-sequence across members). So the +"36k unique / 87% distinct-code tail" was a grouping artifact, not real uniqueness. **LESSON: before calling code +"unique," re-cluster by the looser fingerprint (R14 — verify the GROUPING against the bytes, not just the match).** +h_seq families ALSO differ in immediates, so templating them needs `family_remap` extended to substitute immediates +(+ a cross-address variant), not just symbols. Full finish plan: **`docs/family-endgame-megaplan.md`** (Phase 26). + +**Phase-26 close CORRECTION (2026-07-15, byte-proven — bounds the lever):** the h_seq "finish the decomp by +templating ~754 families" thesis is **SPENT**. Three whole-binary gate probes returned **0%** (tiny-IMM 0/241, +PURE reach-134 0/134, pinned-PURE 0/133): h_seq PREDICTS templatability, the gate refuses it (the refusal classes: +collision / register-drift / pin-crash — pins SIGABRT cc1 in sibling TUs). Everything cleanly templatable was +already banked (A3h + the §52 waves). **The remap engine still works behind each FRESH exemplar crack** (§52 waves +banked 5 cores ×134 = 670 instances), but "crack ~986 exemplars → template ×120" is no longer the endgame — the +endgame is [[roadmap-to-100]] (`docs/roadmap-to-100.md`), which supersedes the megaplan. + +**S65 RE-MEASUREMENT (2026-08-29, byte-proven — the Phase-26 "SPENT" verdict is TIER-SPECIFIC, not general).** +The lever is ALIVE at the two STRICTER fingerprints; only `h_seq` was spent. Measured by enumerating, for every +open stub, whether any ALREADY-BANKED fn fleet-wide shares its hash (a dict join on `family_sweep.load_sigs()`, +seconds to compute), then remapping + whole-binary gating each: + +* **`h_exact` (identical instruction bytes modulo reloc fields): 139 banked of 157 = 88.5%**, 24 binaries, 0 agent tokens. +* **`h_norm` (same after masking reloc fields): 36 of 45 = 80%** on the first probe sample. +* `h_seq` (mnemonic skeleton, immediates may differ): still 0% — Phase-26 stands, and THAT is the boundary. + +**The gap was never the engine — it was that nobody ENUMERATED the twins.** `t5_cards.py` deliberately does not +build `seed_ref` ("needs the atlas knn"), so every card printed "no banked twin — derive from the .s" even when a +byte-identical banked twin existed in another overlay. Cost, measured: a t5u Opus wave slot ground +ov_SC03_023:func_8017BEBC to closeness 45 while ov_SC02_004 held a banked byte-identical copy; `family_remap` +produced it exactly, in one command. **Before drawing any wave, run the twin join and remap first** — agents are +for the "no banked relative" tier only. + +**Why it matters:** turns "match 82k unique fns" into "crack one per family, remap the rest" — the +Phase-25 endgame. **Transferable** to any sibling overlay/bank decomp — see [[cross-project-idiom-discovery]], +[[matching-cookbook]], [[ultracode-harvest-pattern]]. + +## subagent-model-ladder + +--- +name: subagent-model-ladder +description: "Route matching drafters by function size — Haiku <=30 ins (86%), Sonnet 50+; the cliff is measured, not assumed" +metadata: + node_type: memory + type: feedback + originSessionId: a1b8c2a6-3bcc-41b1-9238-ffce61064940 + modified: 2026-08-08T00:03:55.704Z +--- + +Route matching drafters by instruction count: **Haiku ≤30 ins → Sonnet 50–120 → Opus ≥120 / +escalation → Fable5 (new wall classes only)**. Never Haiku→Opus directly. + +**Measured 2026-08-07 (S45 p6), two controlled waves, independently re-verified:** +Haiku 4–27 ins = **86%** (~44k tokens/match) · 30–49 = ~56% · **≥50 = 20%** (~177k tokens/match, +**4× worse**). The previously-documented "Haiku ≤~50" band was optimistic — the cliff starts +around 30 and collapses past 50. + +**Cheap-tier honesty is excellent:** across 100 drafters, 63 MATCH claims → 63 confirmed, 0 false. +Treat agent verdicts as a reliable *filter*, never as the gate (G3/P9 — always re-verify before +gating). When re-verifying, retry a failure with `--o0` before calling an agent wrong: an +`-O0`-cluster function looks like a false claim otherwise (§116, §157). + +**WHAT FABLE IS *NOT* FOR (Drew, 2026-08-24).** Fable is reserved for **NEW WALL CLASSES** — a +tooling problem nobody has solved, an adversarial validation of a design, a residual no documented +lever reaches. **Idiom distillation / cookbook review is Opus or Sonnet work, not Fable**: it is +reading harvested notes against an existing knowledge base and deciding covered / addendum / new — +judgement over a corpus, not a wall. I routed a distill batch to Fable and Drew corrected it. +Same rule for any "read a pile of artifacts and summarise against what we already know" task. + +Details + the per-band table: cookbook **§157**. Supersedes the earlier Drew 2026-08-03 ladder, +which had the Haiku boundary at ~50. + +## SUPERSEDED 2026-09-01 (Drew, explicit): **NO SONNET. opus <=150 ins, fable >150.** + +The Haiku->Sonnet->Opus->Fable ladder is DEAD. Two tiers only, enforced in `draw_waves.arm_for()`. + +Measured over 129 drafting agents in P31 S69, per MATCHED instruction (the only cost that matters — +a failed agent is billed in full): + + sonnet 105 agents, 57 MATCH 4,289 tok/matched-ins (flat ~47% above 30 ins) + opus 24 agents, 11 MATCH 2,083 (m1 band 191-347: 1,291, 67% MATCH) + opus at 347-670 ins: 1/9 7,158 <-- the cliff; 2.92M tokens for ONE bank + fable escalation: 3/4 closed at ~1/3 the cost of the attempt it rescued + +**ESCALATE SOONER.** Higher models crack harder functions in FEWER tokens. The expensive mistake is +running a cheap tier into a wall and paying for the failures, then paying again to escalate. + +Do not re-derive a cheap-tier argument from per-agent price — it has now been tested twice (S68 A/B, +S69 measurement) and both times the cheaper tier cost more per bank. + +## tool-change-updates-siblings-and-docs + +--- +name: tool-change-updates-siblings-and-docs +description: "Drew 2026-09-02: every tool create/update ships IN THE SAME CHANGE with (a) the sibling tools wired to know how/when to use it and (b) the docs updated — and NO end-of-session audits" +metadata: + node_type: memory + type: feedback + originSessionId: f0d1eb32-d9df-4a83-b7e6-dcb5b1b3459d + modified: 2026-09-02T22:45:02.007Z +--- + +**Whenever you create or update a tool, in the SAME change:** + +1. **Update the related tools** so they know how and when to use it — wire it in, call it, refuse + without it. A tool nobody calls is a tool nobody uses. +2. **Update the docs** about the creation/update — `docs/SETUP.md` (R21), the cookbook, the + playbook step, whichever is the consumer of that knowledge. + +**And do NOT run end-of-session documentation audits.** If step 1 and step 2 happen inline, there is +nothing to audit. An audit at the end is a symptom of not having done it during. + +**Why (Drew, 2026-09-02).** He asked a plain yes/no question — *"are the tools/docs up to date with +this session's changes?"* — and I answered by launching a 9-agent parallel audit that burned ~1M +tokens before he stopped me. The answer was **yes**, and I already knew it, because the wiring and +the docs had gone in with each change. Two failures in one: I did not answer the question asked, and +I treated a verification I had already performed as something to re-derive at scale. + +**How to apply:** +* The commit that adds a tool also contains its SETUP row, its consumers' call sites, and its + cookbook/playbook entry. If those are not in the diff, the change is not finished. +* When asked whether things are up to date, **answer from what you did**, briefly, with a couple of + concrete examples. Cheap deterministic checks (an index `--check`, a `grep` for a hardcoded path, + `make tools-health`) are fine; an agent fan-out is not. +* Scale the response to the question. A yes/no question gets a yes or a no. + +**THE FAILURE MODE THAT ACTUALLY BIT (Drew, 2026-09-02, S74 — the same day, later).** Asked the +same yes/no question, the honest answer was **NO**, and the gap had one shape: **every tool change +that came from a SUBAGENT shipped with a rich commit message and no docs.** An agent hands back a +report; a report is not a SETUP row and not a cookbook section. Six changes I made myself were +documented inline; five merged from agents were not — `jtbl_rodata_pads` (dlabel anchor form, +unaligned read), `harvest_verify` (unscoped typedef strip-set), `jr_isolate_all` +(`_region_emit_start`), `jtbl_carve` (`covered` verdict), `ld_interleave` (`--pre`). I had written +each one up beautifully IN THE COMMIT MESSAGE, which is exactly the "commit messages are not the +knowledge base" trap. + +**So the rule extends: INTEGRATING an agent's tool change IS a tool change.** The merge commit owes +the same three things the authoring commit would: the sibling wiring, the SETUP row, the cookbook +entry. Budget doc time per MERGE, not per session — and when an agent's report contains a law (a +measurement, a refuted premise, a new lever), that law belongs in the cookbook before the merge is +called done, not in the prose of a commit nobody greps. + +**And answer the yes/no honestly.** "No, here are the five" took one message and cost nothing; +claiming yes would have left five undocumented tool behaviours for a fresh session to rediscover. + +Related: [[keep-setup-md-current]] · [[bank-idioms-before-checkpoint]] · +[[capture-knowledge-before-fresh-session]] · [[offline-tooling-first]] + +## tool-must-refuse-unsupported-input + +--- +name: tool-must-refuse-unsupported-input +description: R43 — a tool must refuse input it cannot handle, never process it wrongly; sweep_parallel accepted main and banked 0/105 +metadata: + type: feedback +--- + +**R43 (accepted by Drew 2026-08-23, binding).** Silently accepting work a tool will mishandle is +worse than skipping it: the output is a plausible FAILURE that gets attributed to the subject +(the model, the binary, the family) instead of the harness. A tool with a known-unsupported input +class fails loud and names the tool that does handle it. + +**Why:** `sweep_parallel` carried an explicit `"us.exe" if b == "main"` branch — written to LET +main in — while `gate_main.py`'s own docstring says main cannot be gated incrementally (its +extract rewrites the linker script, so an incremental build yields a false diff). Measured on wave +`ab`: **105 main cards banked 0 of 105**, while the same wave's 115 non-main cards banked 94 +(82%). The drafts were fine. The wave read as a drafting failure for a whole session, and I nearly +went looking for the answer in the models. + +**How to apply:** extends [[silently-narrowed-tool-scope]] and R32 from *silently skipping* work to +*silently accepting work it will mishandle*, and is the same family as [[exonerate-the-instrument]]. +When a per-item rate is anomalous, partition by input class BEFORE blaming the subject — the 48% +bank rate resolved instantly once split into main (0/105) and non-main (94/115). + +## tools-folder-convention + +--- +name: tools-folder-convention +description: BFM-decomp project rule — all tooling and tool binaries live under tools/ (incl. downloaded compilers in tools/bin) +metadata: + node_type: memory + type: feedback + originSessionId: 998849ed-0de9-4b58-8dec-e32c194b8e8b +--- + +On 2026-06-10 Drew stated: "I want all tools to live in the tools folder." Applies to the BFM-decomp repo. + +**Why:** keep the repo's tooling consolidated and predictable in one place. + +**How to apply:** Place every project tool, script, and tool binary under `tools/` — e.g. `tools/bfm_extract/` (extractor), `tools/asm-differ`, `tools/m2c`, `tools/maspsx`, `tools/decomp-permuter` (submodules), `tools/bin/` for downloaded vintage compilers (the old-gcc cc1 tarballs — was `bin/` in the sotn convention, relocated here), `tools/psyq*/` for native PsyQ binaries. `docs/SETUP.md` already reflects this; `.gitignore` ignores `/tools/bin/` and `/tools/psyq*/` (download artifacts) while `tools/brave-CUE` and submodule pointers stay tracked. System-level apps installed outside the repo (Ghidra, JDK, PCSX-Redux, WSL) are not in scope. See [[bfm-decomp-context-system]]. + +## tools-health-foreground-not-background + +--- +name: tools-health-foreground-not-background +description: "`make tools-health` run as a background Bash task gets KILLED by the harness's low-memory guard during the report step (twice in S88); run it in the foreground with a ~15-min timeout" +metadata: + node_type: memory + type: feedback + originSessionId: 4555f4e4-091b-4300-9c33-cbf64ba8fda8 + modified: 2026-09-07T17:36:11.991Z +--- + +Running `make tools-health` with `run_in_background: true` was killed twice in one session (2026-09-07, S88) with +"stopped because the system is running low on memory" — both times inside the `report` step (the duplicate-report +worker fan-out is a transient memory spike; `free -h` showed 29 GB available seconds later). The same command in the +FOREGROUND (timeout 900000 ms, ~6 min) passed both times. + +**Why:** the harness's background-task memory policing reacts to the transient spike; the foreground run is not +policed the same way. The killed run also leaves cdecl's scratch file `src/shared/.cdecl_allmacros.c` behind (an +untracked dotfile — never commit it; delete it before the re-run). + +**How to apply:** run `make tools-health > .run//_tools_health.log 2>&1; echo "EXIT=$?" >> …` in the +foreground with a long timeout, never as a background task. While there, check `ps` for orphaned workers from closed +phases (S88 found 8 `run_masked.py` permuter workers 49 h old, parent PID 18) and stop them by PID — never `pkill -f` +with a literal the calling shell carries ([[pkill-pattern-kills-own-shell]]). + +## ultracode-harvest-pattern + +--- +name: ultracode-harvest-pattern +description: Parallel-draft + deterministic byte-gate + match_one iteration loop for mass function matching (cookbook §12; tools harvest_verify.py + match_one.py) +metadata: + node_type: memory + type: reference + originSessionId: 5e0e740f-c352-4b9e-977c-041dcbf685e5 +--- + +The high-leverage matching technique for a binary with many independent functions to hand-match (proven Phase 12: resident **1.4%→71.7% byte-identical in one session**, REAL 1→102/145). Full writeup: `docs/matching-cookbook.md` **§12**. Reusable verbatim for the Phase-13 overlays. + +**The loop (each pass = a Workflow + a deterministic gate; loop-until-dry):** +1. **Draft (parallel Workflow):** N agents (round-robin a size-sorted fn list) each draft matching C for their share from `asm//nonmatchings/.../.s` + the cookbook + already-matched fns, writing ONE self-contained `.c` per fn to `.run/drafts/.c` (+ `.conf`). **No builds, no Ghidra** in the draft pass (asm is the target; Ghidra contention flakes under fan-out). Distinct files → no write races. +2. **Byte-gate (`tools/harvest_verify.py`):** substitute each draft for its `INCLUDE_ASM` stub → `make build` → keep only if byte-identical, else revert (chunk + bisect). The build is the sole arbiter → a wrong match cannot pass (G3/P9); agent over-claims cost nothing. +3. **Loop:** re-run the Workflow on the residual stubs, each agent seeded by its prior failed draft + a debugging checklist (miss-modes: load width/sign, branch polarity, callee-return-type-forces-`andi`, store order, QImode decrement). +4. **Iterate pass (strongest, `tools/match_one.py`):** per-function isolated compile + relocation-masked diff vs the `.s` (own temp dir → parallel-safe) so agents `write C → run match_one → read diff → fix` until MATCH — cracks scheduling/regalloc near-misses. + +**Gotchas:** strip inline scalar-typedef redefinitions (C89 dup-error vs common.h = a compile fail, NOT a byte miss); `match_one` MATCH but whole-build FAIL = **extern-type conflict** across the single TU → unify extern types (widen a definition's return where byte-identical), NOT separate `.c` per fn (splat places the binary as one address-ordered object); wrap agents in **retry waves** for transient server rate-limiting. Honest tail (cross-jump/regalloc residuals) → decomp-permuter + §3a, never forced (P9). Part of the R16 flywheel; trigger via [[effort-prompt-ultracode-on-breadth]]; see [[matching-cookbook]], [[web-research-compiler-quirks]]. + +## verdict-names-its-instrument + +--- +name: verdict-names-its-instrument +description: "a recorded impossibility must name the instrument that produced it — and a verdict from a bespoke harness is a verdict about that harness; S72 found 11 'PROVEN gate-rejects' were one missing carve, judged by a one-off script that skipped the real gate's pre-check" +metadata: + node_type: memory + type: feedback + originSessionId: f0d1eb32-d9df-4a83-b7e6-dcb5b1b3459d + modified: 2026-09-02T18:09:53.894Z +--- + +**When you write a negative verdict into the checkpoint, write down WHICH TOOL produced it.** A +future session reads "PROVEN gate-reject" as a property of the function; it is almost always a +property of the run. + +**And never re-implement a gate you already have.** A verdict produced by a bespoke script is a +verdict about that script. + +**Why (measured, S72 / 2026-09-02).** S71 closed with: *"11 main functions score `match_one` +closeness 0 and are PROVEN gate-rejects (re-gated one at a time) — §376 in its purest form — do not +re-slate without a TU-level fix."* Every part of that was wrong: + +* The re-gate ran through `.run/S71_main_bisect.py`, a one-off written that night — not + `tools/gate_main.py`. The real gate has a decl pre-check, an R40 baseline control, a bisect and a + no-op-substitution guard; the script had none. Run afterwards on the same 11 drafts, the pre-check + names **6** of them in ~2 seconds, and one of those six was a **false** conflict in the checker + itself. +* **All 11 were switch functions**, which nobody had measured, because a main gate's entire output + was two SHA1s. The blocker was a `.rodata` carve missing since **Phase 7**, not codegen. +* Three of them banked **byte-identical in 14 seconds** once the carve was extended. + +**The compounding failure underneath it:** `gate_main`'s R40 baseline control *rebuilds the tree green +immediately after a failure*, overwriting the red binary and its linker map. The control was correct +and its **ORDER** was wrong, so the one artifact that could localize a failure was destroyed every +single time. An oracle whose verdict routes work must preserve the artifact that routing needs. + +**How to apply:** +* A checkpoint line that parks work states the tool and flags used. "PROVEN X" with no instrument + named is a claim to re-probe, not a fact to plan around. +* Before believing a negative batch verdict, re-run it through the project's real gate. If the real + gate is slow, that is a reason to fix the gate, not to route around it. +* A failing gate saves its artifact **before** any control or cleanup rebuilds over it. +* A uniform failure shape across independent drafts (same delta, same first-moved symbol, same symbol + count) is a LAYOUT or PLUMBING signature, never N independent codegen walls. + +Related: [[exonerate-the-instrument]] · [[reprobe-exclude-lists-after-tool-fixes]] · +[[gate-main-only-with-gate-main]] · [[check-against-a-known-true-case]] · +[[standalone-match-is-not-bankable]] (cookbook §426/§427) + +## wave-harvest-is-a-pipeline-step + +--- +name: wave-harvest-is-a-pipeline-step +description: "HARD GATE: harvest idioms from the previous wave BEFORE drafting the next one — it is the project thesis, not hygiene. Plus the full seven-step wave-closing sequence, run automatically in order" +metadata: + node_type: memory + type: feedback + originSessionId: 09d03224-6e3a-4041-8d6f-c58f267bf1e8 + modified: 2026-08-30T17:39:17.270Z +--- + +## THE HARD GATE (Drew, 2026-08-30 — absolute, no exceptions) + +**No new wave is drafted until the previous wave's idioms are harvested into +`docs/matching-cookbook.md`.** Not "should"; a precondition. If a wave is about to be drawn and the +last one has not been harvested, harvest first — even if lanes are idle, even if the gate is green, +even mid-campaign. + +**Why this is THE thesis and not hygiene** (Drew's words): *"new idioms = easier next exemplars and +free banks after learning the idiom."* The whole method is a flywheel — every wave is supposed to +make the next one cheaper. Skip the harvest and each wave re-derives what the last one already paid +for, so cost stays flat and the campaign degrades into brute force. Harvesting is the mechanism by +which the project compounds; drafting without it is the project's one self-defeating move. + +**Measured, S66 (2026-08-30):** I ran eight waves (r1, w1, w2, w3, x1, m1, m2, m3) gate → R22 → next +wave with NO harvest, exactly repeating the S65 failure this memory already documented. Drew caught +it with *"did you harvest idioms before running the next waves?"*. The catch-up distill found **35 of +250 transcripts advertising a mechanism** — levers like cross-jump barrier PLACEMENT as a second +lever, a zero-byte `asm` raising a pseudo's ref count to flip allocator priority, and a +provably-unreachable-from-C `jal`+`addiu %lo` delay-slot class (5 sites fleet-wide) — all of which +later waves were re-deriving from scratch in the meantime. + +## THE SECOND HALF OF THE GATE: harvest → TOOL → free banks (Drew, 2026-08-30) + +Harvesting into prose is only half the flywheel. **After each harvest, ask of every new idiom: is +this MECHANICAL? If yes, build or extend a tool that applies it fleet-wide, and bank the free +functions it finds — before drafting the next wave.** An idiom that a tool can apply must never be +re-cracked by an agent; that is the difference between a book and a compounding machine (R16 says +feed the lesson back into BOTH the cookbook AND the tooling — this is the "and the tooling" half, +with the FREE BANKS as the acceptance test). + +The test to apply per idiom: *"how many open stubs fleet-wide match this idiom's asm tell?"* If the +answer is more than a handful, it is a sweep, not a lesson. Precedents: +* `twin_sweep.py` — from "structural twins remap mechanically": 389 banks in S65, 12+2 more in S66, + for ~0 model tokens. The pool REFILLS after every banking step, so re-run it each time. +* **S66 candidate, not yet built:** §179-B/§179-C — a function with NO `jr $ra` (falls through into + a shared tail) or whose out-of-body branches leave the symbol is *uncompilable from C* and is + banked by transcribing its `.s` verbatim as a file-scope `__asm__` block. Agents did this by hand + for `func_80062388`, `SYS_OBJ_2CDC`, `SYS_OBJ_1034`, `func_8005A870`, `GsSortFastBg`. The tell is + greppable in the `.s`, so a sweep could bank every remaining instance fleet-wide for zero tokens. +* Counter-example (a WALL, not a sweep): the `jal` + adjacent `addiu %lo` delay-slot class is + unreachable from C (`mips.md` `define_delay` needs a length-1 filler); a corpus grep found exactly + 5 sites, all still `INCLUDE_ASM`. That belongs in the ledger as a known ceiling, not in a tool. + +--- + +A crack wave is not finished when its drafts are banked. It closes in **seven ordered steps**, and +they run automatically as part of the wave — not on request, not deferred behind the next wave: + +1. **Snapshot each target's `.s`** (`.run/wave_asm_snapshot//`) — banking prunes it and + later steps need both sides — then re-verify every draft with `match_one` (R14: an agent's MATCH + is a claim) and **gate** the groups with `gate_lane`. +2. **RECOVERY WAVE over the failure set** (Drew, 2026-08-18). Three classes, three lanes: + * *near-misses* (drafted, never reached MATCH) → fresh-eyes pass with the residual class named, + the latest cookbook laws, and `oracle_reorder.py` first when the diff is in the tail (§188); + * *gate drops* (byte-verified MATCH, refused at the rebuild) → the §183 declaration playbook + against the exact refusal text; + * *errored cards* (agents killed by rate limits, no draft on disk) → straight re-draft, and + REMOVE them from the wave's card file so they are not marked already-waved. +3. **Gate the recovery output.** +4. **`family_sweep --hseq --only `** after `make sig-overlays` + `sig-modules` + + `family_hseq.py` — the free sibling remap. +5. **R22 clean-fleet** ONCE, covering everything above. +6. **Harvest the `index_gap` reports** — cluster readers, one adversarial verifier per candidate + defaulting to REJECT — and bank the survivors into `docs/matching-cookbook.md`, with correction + banners on any section they refute. Regenerate `tools/cookbook_index.py`. +7. **Build the NEXT wave's cards** (`make atlas` + `decl_prior --build` + `build_wave_atlas`), and + stage its script and args on disk. Do this BEFORE the checkpoint so the checkpoint can record + where they are and the exact invocation that fires them. +8. **Checkpoint** `phase-ends/CURRENT_PHASE.md` + a wave-metrics row, then commit — **ALWAYS, and + written for a FRESH SESSION that has none of this context.** If any work lands after a checkpoint + (a late fix, a re-drawn wave, a harvest that finally ran), **REFRESH IT** — a checkpoint that + describes a state the tree has moved past is worse than none, because it is believed. + +**Why the ORDER matters:** recovery before R22 and before the checkpoint, so both describe the +wave's FINAL state instead of one that is about to change. Harvest and next-wave build before the +checkpoint, because **the checkpoint is always the last thing written** — that is what makes it +true. Drew has corrected this twice (2026-08-17, 2026-08-18): a checkpoint written mid-sequence and +then outrun by more work is stale, and stale is worse than absent. + +**Why step 6 is not optional:** wave T's harvest produced §193-A (every card's `exemplar` pointer +was a stub by construction while the atlas's banked twin was discarded); wiring that one field in +took wave U from 99% drafted / 70 banked to **100% drafted / 73 banked at 15% fewer tokens and 64% +of the wall-clock**. Harvests also correct EARLIER harvests — §194-E corrected §193-A and §199-A +byte-refuted §189-A, both within the same session. A law from one wave is a first draft; the next +wave is its review. Seed each harvest's readers with everything banked so far so they cannot +re-derive it; "already covered" is the expected majority verdict (61/71, 44/64, 76/67, 41/68, 56/63 +across waves T-X) and is the signal to fix RETRIEVAL, not to write more prose. + +**MEASURED COST OF SKIPPING IT (S65, 2026-08-29 — I ran gate → R22 → next wave, TWELVE times):** +* step 2 never ran ⇒ the recovery backlog reached **69 items** instead of ~5 (14 gate-drops incl. a + 770-ins fn, 33 near-misses with the closest at closeness 1, 22 errored cards that were never + failures at all). Classified into `docs/recovery-queue-s65.md` only because Drew asked what the + post-wave process was. +* step 6 ran once in 14 waves ⇒ the catch-up harvest found **7 of 14 candidates were REDISCOVERIES** + of laws already in the book. Per-wave harvesting is what stops the next wave re-deriving them. +* I wrote the session checkpoint BEFORE the harvest, making it stale on arrival — the exact failure + the "checkpoint is always last" clause exists to prevent. Had to write a superseding block. +Drew's cue when this drifts: *"we should have a memory a whole process after each wave."* + +Related: [[bank-idioms-before-checkpoint]] (the session-end deadline; this is the per-wave one), +[[matching-cookbook]] (R16, the flywheel), [[wave-prompt-seed-step0-and-gaps]], +[[checkpoint-current-phase-before-pause]], [[capture-knowledge-before-fresh-session]]. + +## wave-playbook-is-the-procedure + +--- +name: wave-playbook-is-the-procedure +description: "READ docs/wave-playbook.md BEFORE running any matching wave — it is the current start-to-finish procedure; docs/automation-runbook.md is the RETIRED OpenRouter era and must not be followed (Drew, 2026-08-31)" +metadata: + node_type: memory + type: project + originSessionId: 5e7f4e3a-4e31-46cc-a416-6baaccdac66a + modified: 2026-08-31T19:50:30.708Z +--- + +**`docs/wave-playbook.md` IS THE PROCEDURE FOR RUNNING A WAVE. Read it before drawing one, every +session.** Drew made this standing on 2026-08-31 (S67). + +**`docs/automation-runbook.md` IS HISTORICAL — do NOT follow it.** It was titled "the autonomous +campaign, as it actually runs" while documenting the **OpenRouter / ox-alpha** era, whose lanes +(`drafter`/`gater`/`maintenance`/`stallguard`/`distill`/`main`) are all DEAD *by choice* — the project +deliberately returned to Claude agent waves. It is now retitled with a pointer, and kept only because +its banking-by-binary-class, main-lane cadence, distill-flywheel and rate-limit sections are still +accurate. `docs/SETUP.md`'s tooling-inventory rows are per-tool REFERENCE, not procedure. + +**Why this memory exists:** three of S67's most expensive mistakes were PROCEDURAL, not technical, +and each is now a guard in the playbook: +* hand-typed a streaming refill target -> invented `func_80184F60`, the 2nd instruction of an + already-matched function — **58k tokens**. Every payload must come from `/wf_args.json`. +* hand-rolled a serial gate loop while `parallel_gate` existed — **~1 hour** for work that measured + **103 s** in worktrees. Parallel is the default; split on the jtbl predicate BEFORE running. +* re-derived a function banked verbatim in ~20 overlays — **102k tokens** — because the card said + "no banked twin". `seed_ref` is now on the card; check it. + +**The shape of the doc is the deliverable, not just its content:** each step is paired with the +MEASUREMENT that produced its guard. That pairing is what a generic decomp guide cannot have, and it +is the seed of the future "how to AI-decomp" template (see [[project-endgame-deliverables]]) — +feeders are `docs/decision-log.md`, `docs/accelerators.md`, `docs/hindsight-study.md`, +`docs/matching-cookbook.md` and `phase-ends/`. + +**Keep it current the way `docs/SETUP.md` is kept current** ([[setup-md-keep-current]]): when a wave +step, guard or tool changes, update the playbook in the SAME change. A stale procedure doc is worse +than none — that is exactly what the automation-runbook became. + +Related: [[wave-harvest-is-a-pipeline-step]] [[parallel-gate-via-worktrees]] +[[endgame-budget-unconstrained]] [[offline-tooling-first]] [[decomp-accelerator-ledger]] + +**S69 addition — the TRIAGE LADDER is now part of the pipeline** (`tools/triage_ladder.py`): +`wave_args` drops walled/parked targets at draw time automatically, and +`triage_ladder.py --escalate :` must run before ANY escalation (it exits 2 on a walled or +already-banked target; `escalate_fable.js` refuses a target without `triage:'DRAFT'`). It refuses to +classify on a non-quiescent tree. Playbook §4b. See [[standalone-match-is-not-bankable]] for why its +integration tiers are GATE-FIRST candidates and never banks. + +## wave-prompt-seed-step0-and-gaps + +--- +name: wave-prompt-seed-step0-and-gaps +description: "When authoring a crack-wave prompt, seed it with the cross-overlay magic-literal grep as STEP 0 and with the previous wave's agent-reported index_gap findings — bank rate went 76% -> 77% -> 100% across three waves on the same models and gate" +metadata: + node_type: memory + type: feedback + originSessionId: 425fccde-04c9-4922-bd3e-2e8e772fe70b + modified: 2026-08-04T16:00:01.165Z +--- + +Authoring a crack-wave prompt is an **orchestrator** job, and the two things worth putting in it are +not in the drafting agents' default reading path. Measured over three waves in one session +(2026-08-04), same models, same whole-binary gate, only the prompt changed: + +| wave | prompt | banked | +|---|---|---| +| 1 | baseline | 13/17 (76%) | +| 2 | + the session's reconcile rules | 10/13 (77%) | +| 3 | + **STEP 0** below | **13/13 (100%)** | + +**1. STEP 0 — `grep -rn "" src/` with a distinctive literal from the target `.s`** (a magic +word, an unusual mask, an odd immediate). Put it AHEAD of `engine_core.h` in the §136c search order. +Reason it is not optional: §136c's first two steps (shared-header near-twin, same-TU banked sibling) +are both **same-TU or shared-header scoped**, so neither can reach a banked twin living in a +*different overlay's* TU — and the large template classes live cross-overlay by construction. In +wave 3 several agents landed the answer on the FIRST search; one found a banked twin whose own +header comment already recorded it as byte-identical to the new target, so the body transferred +verbatim with only file-local type suffixes renamed. Full writeup: `docs/matching-cookbook.md` §138. + +**2. Harvest the previous wave's `index_gap` reports into the next prompt.** The wave schema asks +each agent for `index_hit`/`index_gap`; those reports are where the next prompt improvement comes +from — STEP 0 itself came from a wave-2 agent noticing our documented search order could not reach +its twin. This is the R16 flywheel closing on itself at wave granularity, and it only works if the +orchestrator actually reads the gap fields and promotes them. + +**Honesty caveat, so the table is not over-read:** wave 3's targets were also somewhat easier, so the +100% is not attributable to the prompt alone. Separately, the *sweep* rate after a wave (21/21 vs +18/165) is a property of the FAMILY, not the prompt — per-location families whose members differ in +real codegen do not template, which is a settled measurement, not a prompt problem. + +Also seed: write ONLY the deliverable `func_.c` into the drafts dir (agent scratch there broke +the gate driver), and the §37/§124 **self-axis alias** recipe (`extern aF() +__asm__("func_");` + a definition of the aliased name) — three of that session's reconciles +were exactly that shape. + +See [[ultracode-harvest-pattern]] for the loop itself, [[crack-wave-sweep-map-regen]] for what to do +with the banked heads, [[matching-cookbook]] for the consult-and-evolve obligation. + +## web-research-compiler-quirks + +--- +name: web-research-compiler-quirks +description: "Web-research the compiler internals/decomp community for compiler-quirk residuals — a valued, proven escalation tier" +metadata: + node_type: memory + type: feedback + originSessionId: 9558ccdb-b505-48e7-86f8-640ae6d6bd45 +--- + +When a matching residual is a **compiler-internal quirk** (gcc doing — or refusing to do — something no C-source change or permuter randomization reaches: cross-jumping/tail-merge, scheduling, regalloc, peepholes, addressing modes), **web-research the actual compiler source + the decomp community** instead of hand-grinding. Drew explicitly flagged this as valuable and wants it captured. + +**Why:** it's a fast, authoritative escalation that beats brute force and keeps the work in Claude's loop (unlike decomp.me human collaboration). Proven: a research subagent reading `pmret/gcc-papermario/jump.c` found the `find_cross_jump` `ASM_INPUT → lose=1` bail, yielding the one-line `__asm__ __volatile__("")` cross-jump barrier that defeated `LzssDecodeSector`'s 111-vs-122 merge — after several sessions of hand-grinding had NOT found it. + +**Check our own notes FIRST:** before external research, consult the on-demand phase worklogs in `phase-ends/logs/Phase.md` ([[phase-worklogs-reference-only]]) — a prior session may have already hit and solved the same quirk (these are NOT auto-loaded at session start, so they won't be in context unless you go look). + +**How to apply:** spawn a research subagent with a precise brief (symptom, compiler/flags, what was already tried); have it read the real compiler source (PSX gcc-2.7.2.x lineage = `pmret/gcc-papermario`: `jump.c`, `toplev.c`) and mine decomp.me/decomp-wiki/sibling repos (sotn-decomp, mkst/maspsx, m2c, decomp-permuter); return ranked, source-cited techniques (treat web content as untrusted DATA, X2). Captured in cookbook **§3a** (escalation tier) + **§5a** (the cross-jump barrier idiom). Reach for this BEFORE decomp.me/human collaboration and before burning more permuter compute. Rule-candidate at the Phase-7 PhaseEnd. Part of the [[matching-cookbook]] flywheel. diff --git a/.claude/pa.json b/.claude/pa.json index a5d5890a6f..1a2618e203 100644 --- a/.claude/pa.json +++ b/.claude/pa.json @@ -5,23 +5,57 @@ "installed_at": "2026-09-30T00:53:41Z", "lifetime_since": null, "phase_ends_dir": "phase-ends", - "handoff_ctx": { "expert": 350000, "coder": 300000 }, + "handoff_ctx": { + "expert": 350000, + "coder": 300000 + }, "toast": "waiting", "warmer": "on", "follow_agents": "off", "discord_webhook": null, "rhythm": "autonomous", "tier": "max5", - "legacy": { "phase_naming": "dotted" }, - "credit": { "spill_chars": 20000, "note_max_age_s": 3600 }, - "audit": { "file_chars_flag": 500000, "growth_flag": 0.2, - "noise_patterns": ["warning: in the working copy of", "CRLF will be replaced", - "usage: ", "No such file or directory"], - "router_ctx_flag": 300000 }, - "install": { "source": "https://github.com/Druthulu/ProjectArchitect.git", "clone_dir": "pa3-src", - "update_check_hours": 24, "fetch_timeout_s": 8 }, - "card": { "max_chars": 7000 }, - "guard": { "spilled_read": true, "whole_plan_roles": ["review", "critic"], "tool_source_roles": ["expert"], - "whole_read_chars": 20000 }, - "bench": { "window_max_pct": 80 } + "preset": "max20", + "legacy": { + "phase_naming": "dotted" + }, + "credit": { + "spill_chars": 20000, + "note_max_age_s": 3600 + }, + "audit": { + "file_chars_flag": 500000, + "growth_flag": 0.2, + "noise_patterns": [ + "warning: in the working copy of", + "CRLF will be replaced", + "usage: ", + "No such file or directory" + ], + "router_ctx_flag": 300000 + }, + "install": { + "source": "https://github.com/Druthulu/ProjectArchitect.git", + "clone_dir": "pa3-src", + "update_check_hours": 24, + "fetch_timeout_s": 8 + }, + "card": { + "max_chars": 7000 + }, + "guard": { + "spilled_read": true, + "whole_plan_roles": [ + "review", + "critic" + ], + "tool_source_roles": [ + "expert" + ], + "whole_read_chars": 20000 + }, + "bench": { + "window_max_pct": 80 + }, + "memory_routed": "3.14" } diff --git a/HOW_WE_WORK.md b/HOW_WE_WORK.md index ee590957b1..68a1741433 100644 --- a/HOW_WE_WORK.md +++ b/HOW_WE_WORK.md @@ -14,6 +14,7 @@ Who: . Experience: . Notification channel: . The developer pushes; agents never do. They ratify rules at the next planner session. +- drew-working-preferences: Drew's working style on BFM-decomp — autonomous-within-phases, ultracode effort, accepts recommended options, wants WSL/tooling decisions surfaced plainly ## Rhythm autonomous + +L1 | offline-tooling-first | legacy,memory | active | legacy memory offline-tooling-first.md + +L2 | lever-removal-is-a-tracked-series | legacy,memory | active | legacy memory lever-removal-is-a-tracked-series.md + +L3 | decomp-accelerator-ledger | legacy,memory | active | legacy memory decomp-accelerator-ledger.md + +L4 | verify-blast-radius-not-just-defect | legacy,memory | active | legacy memory verify-blast-radius-not-just-defect.md diff --git a/rules/L1.md b/rules/L1.md new file mode 100644 index 0000000000..3beccfc4f3 --- /dev/null +++ b/rules/L1.md @@ -0,0 +1,30 @@ +# L1 — offline-tooling-first +id: L1 · group: - · status: active · tags: legacy,memory · origin: legacy memory offline-tooling-first.md · added: 2026-09-29 + +**Standing goal (Drew, 2026-07-21): get as much as possible working as offline tooling.** When a +recovery, diagnosis, or integration step is *computable*, it belongs in a deterministic tool that runs +with zero tokens on every future draft — not in an agent prompt, and not in a search. + +**Why:** the measured economics keep pointing the same way. Phase 15: a 50-agent wave added +0.36% while +deterministic recovery added +2.67% for ~0 agent tokens. Phase 29 (2026-07-21): the permuter's problem was +*targeting*, not a missing transform — ~92% of its CPU was aimed at residuals a search provably cannot +close, fixed for free by a deterministic classifier ([[matching-is-solved-integration-is-the-bottleneck]]). +Every hour of agent drafting is spent once; every ladder stage is spent once and paid forever. + +**The two engines — keep them separate:** +1. **Deterministic recovery** (`tools/gate_stage.py`'s ladder): computable fixes — decl/arity/cast + reconciliation, type-lift, jtbl isolate+re-carve. Applied always, free, byte-gate arbitrated. +2. **Search** (decomp-permuter / ILS): only for residuals that are NOT computable — regalloc and + schedule permutations where the answer must be explored. +Putting a computable fix into the search is a category error: it burns CPU rediscovering a derivable +answer. And per cookbook §60b, raising a search-closer's yield is at least as often about *refusing it +unreachable work* as widening its mutation set. + +**How to apply:** when a draft fails to bank, ask "is this residual computable?" before reaching for +agents or the permuter. If yes → a ladder stage (and check whether the logic already exists elsewhere — +`normalize_self_decls` and jtbl auto-isolate both existed in `family_sweep`/`jtbl_family_bank` while +`gate_stage` lacked them; R33 says one implementation, two callers). Any ladder stage that mutates SHARED +state must undo by snapshot-restore, never an inverse transform, and be verified fleet-wide (R22) — a +single-binary gate cannot validate a fleet-wide edit (cookbook §61; it cost 138 broken binaries once). +Track the LLM-free fraction and make raising it the objective (`docs/hindsight-study.md` §7, +`tools/burndown.py`). diff --git a/rules/L2.md b/rules/L2.md new file mode 100644 index 0000000000..ff704c6198 --- /dev/null +++ b/rules/L2.md @@ -0,0 +1,23 @@ +# L2 — lever-removal-is-a-tracked-series +id: L2 · group: - · status: active · tags: legacy,memory · origin: legacy memory lever-removal-is-a-tracked-series.md · added: 2026-09-29 + +Drew (2026-09-09, mid-Phase-36): the pins and compiler hints are not just work to finish — **their count over time is a +deliverable**. Keep `docs/levers.md` current after **every task that changes the count**, with +`tools/lever_progress.py --snapshot ""` (appends a milestone row to `docs/lever-progress.tsv` and re-renders the +document's generated block; `--check` refuses a stale series). Mirror the story-relevant numbers into +`phase-ends/CURRENT_PHASE.md` as the phase goes, so the retrospective is built from the record and not from memory. + +Four audiences, all named by Drew: the **post-100% chart**, the **project story**, the **wiki** (a Levers page), and the +**`decomp-architect/` package** — which needs the taxonomy plus an answer to *"what should we have done from day one to stop +this creeping up on us post-100%, or is leaving it to a post-100% cleanup actually optimal?"* + +**Why:** the only phase that ever counts the levers is the phase that removes them, so if the series is not captured while +the work happens it cannot be reconstructed afterwards — a census is a moment. The measured answer so far (P36): **38% of +the class A/B population came off with no understanding at all** (strip, compile, compare), which is the evidence for the +day-one rule *ban the silence, not the lever* — a lever is allowed but is a marked, ledgered, published debt from the first +bank, with a one-compile bank-time trial that would have refused a third of them where the context was still hot. + +**How to apply:** at every task close run the census then `lever_progress --snapshot`; keep §5 of `docs/levers.md` (the +prevent-vs-defer argument) written from the generated numbers, never typed; feed each new rung/recipe and each measured +yield into §4. Related: [[matching-cookbook]] (§454 carries the mechanism), [[decomp-accelerator-ledger]], +[[phaseend-verbosity-for-the-retrospective]], [[project-endgame-deliverables]]. diff --git a/rules/L3.md b/rules/L3.md new file mode 100644 index 0000000000..fc7affd84d --- /dev/null +++ b/rules/L3.md @@ -0,0 +1,20 @@ +# L3 — decomp-accelerator-ledger +id: L3 · group: - · status: active · tags: legacy,memory · origin: legacy memory decomp-accelerator-ledger.md · added: 2026-09-29 + +Drew (2026-08-07): we are building a **Claude Code decomp workflow** to reuse after BFM ships. So +whenever something is found that would have made a lot of *previous* work much faster had we known it +sooner, record it — what it is, when we found it, when we *could* have, and what it would have saved — +in `docs/accelerators.md`. A new decomp project should get that wisdom on day one instead of at phase 23. + +**Why:** this project repeatedly found its biggest levers late (the byte-gate harvest at phase 12, dedup +propagation at 11–15, the gcc codegen map at 23, the tracker's addressing blind spot at 30). The +per-phase PhaseEnds record *what happened*; they do not answer "what should phase 1 of the NEXT game +do differently." That is a separate, deliberately-maintained artifact. + +**How to apply:** when a discovery lands, ask "would this have changed earlier work?" If yes, add an +entry the same session (R30 timing). Distinguish honestly between a lever that was *available* earlier +and one that structurally could not exist yet (needed the fleet onboarded, the compiler pinned, etc.) — +the second kind belongs in the ledger too, marked, because its *prerequisite* is the real advice. + +Feeds [[project-endgame-deliverables]] (the public "how to AI-decomp" wiki) alongside +`docs/decision-log.md` (R31, the why-behind-pivots). diff --git a/rules/L4.md b/rules/L4.md new file mode 100644 index 0000000000..4d03319e01 --- /dev/null +++ b/rules/L4.md @@ -0,0 +1,29 @@ +# L4 — verify-blast-radius-not-just-defect +id: L4 · group: - · status: active · tags: legacy,memory · origin: legacy memory verify-blast-radius-not-just-defect.md · added: 2026-09-29 + +**Verify the BLAST RADIUS, not just the DEFECT.** (Phase 26, 2026-07-14 — I got this wrong in front of Drew.) + +A coverage audit reported that `progress.py` under-counted ~243k instructions because `classify()` reads a K&R +definition as a forward declaration. I did the R14 thing — reproduced the mechanism against the bytes, confirmed +it was real, measured 400 banked instances in that shape — and then told Drew our headline numbers had been +under-reporting our progress. + +**Wrong.** The headline metrics come from `weighted_metrics()`, which never calls `classify()` at all: it tests +"is this function still an `INCLUDE_ASM` stub?", so it is structurally immune to the bug. The published +instruction-weighted and distinct-code numbers were correct all along; only a secondary function-count report was +wrong. + +**Why:** *"this tool is broken"* and *"this number is wrong"* are different claims requiring different evidence. +A confirmed mechanism proves nothing about consequence. Before reporting impact, trace the defect to the actual +consumer and check whether that consumer is even on the affected path. + +**How to apply:** +- After confirming a defect, ask *"who consumes this, and does the consumer use this code path?"* — then verify + THAT, not the defect, before quoting an impact number to the owner. +- **A null result where you predicted a large effect is a refutation — chase it, do not wave it off.** The fix + moved the numbers by +376 instructions when I had predicted +190,000. That gap was the whole story and it + would have been trivially easy to dismiss as noise (or worse, to report as "the fix worked, the numbers moved"). +- Do not amplify a sub-agent's impact claim (R14) — the auditor conflated "classify() is blind" with "the metrics + are wrong", and I propagated it as fact while lecturing about unverified oracles. + +Related: [[derive-from-invariants-not-reparsing]].