From c7772bc452023713c0feaa4ccbb2dd9f965f436a Mon Sep 17 00:00:00 2001 From: Drew T <50529377+Druthulu@users.noreply.github.com> Date: Tue, 14 Jul 2026 10:40:04 -0600 Subject: [PATCH] =?UTF-8?q?docs(phase-26a):=20SESSION-9=20CLOSE=20?= =?UTF-8?q?=E2=80=94=20cookbook=20=C2=A751=20(the=20tooling-integrity=20la?= =?UTF-8?q?ws)=20+=20handoff?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit R30/R16: the context-dependent artifacts, written while the context is live. cookbook §51 — the SILENT SKIP: the bug class, why the byte-gate cannot see it, the over-approximating-detector method, and FOUR LAWS: 1. Derive, don't re-derive — the best outcome is a DELETED SCANNER (28 findings -> one derived oracle + ~10 deleted scanners). A derived fact cannot rot; a hand-maintained copy of it is a liability that grows with every structural change. 2. Assert your COVERAGE, not merely your correctness. *** A LOUD FAILURE THAT NOBODY COUNTS IS EXACTLY AS INVISIBLE AS A SILENT ONE *** — build_engine_types printed '[overlap] handle manually' every single time for four phases while dead on 81% of its own corpus. This CORRECTS the first draft of R32 ('fail loud'), which was not enough. 3. When an oracle is structurally blind to a class of error, add a SECOND ORACLE THAT CAN DISAGREE WITH IT — not a better assertion inside it. We had two all along and never made them argue. (And scope the comparison to where the second oracle is genuinely independent: the same check run outside its domain reports 914 slices when the truth is 193.) 4. A rule that needs a human to remember it is not a gate. Make it structural. + the FALSE-WALL PIPELINE (a silent skip -> a wasted draft -> a backlog 'matching failure' -> reserved_walls() PERMANENTLY blacklists a function that was never attempted), and a checklist for any new corpus-scanning tool. CURRENT_PHASE: session-9 handoff — what is done, what remains (each with its spec on disk), and the R32-corrected / R33 / R34-new rule candidates for P10 ratification. --- docs/matching-cookbook.md | 141 ++++++++++++++++++++++++++++++++++++ phase-ends/CURRENT_PHASE.md | 52 +++++++++++++ 2 files changed, 193 insertions(+) diff --git a/docs/matching-cookbook.md b/docs/matching-cookbook.md index 9b07786e9..62283ba8c 100644 --- a/docs/matching-cookbook.md +++ b/docs/matching-cookbook.md @@ -3470,3 +3470,144 @@ keep same-page stores apart. **Never "tidy up" the store order of a matched func and **both are blocked by sched.c**, which always fills the load-use stall between `lh` and its consumer, and places byte-free (reload-deleted) copies *only* in such stalls. The one productive angle left: a lever that injects a **reload-deleted no-op reg copy** inside `[lhu 4($a2) … sh %lo(D_801152AC)]`. That is the whole delta. + +--- + +## §51 — TOOLING INTEGRITY: the silent skip, and how to hunt it + +> Phase 26-A (the inserted tooling-integrity audit). 68 findings across ~36 tools. This section is the +> METHOD and the LAWS; the findings themselves are in `docs/tooling-audit.md`, the strategic why in +> `docs/decision-log.md`. **Read this before writing any tool that scans the corpus.** + +### §51a — The bug class + +> **A scanner extracts N items from a corpus. The true count is M > N. Nobody ever compared N to M.** + +That is the whole class. It is not a typo, it is a *structural* blind spot, and it produced every one of +the seven bugs found in Phase-26 session 8 and the twenty-eight found in the audit. Its signature: + +* the tool reports success; +* the number it reports is *smaller than reality* and self-consistent; +* nothing downstream can tell, because **a target that is never nominated produces silence, not an error.** + +Measured consequences in this project: **91.6% of all remaining work was invisible to target selection** +(a 3-file allowlist against a 14-file tree); the byte-gate could reach **4.9%** of the canonical overlay; +**62% of the endgame plan's byte-weight was already-matched phantom targets**; and one 10% hole in a +callee-signature oracle made **nine byte-exact functions look like an intrinsic compiler wall.** + +### §51b — Why the byte-gate cannot save you + +**The whole-binary byte-gate is a perfect CORRECTNESS oracle and a NULL COVERAGE oracle.** It has never +once accepted a wrong match. It is also blind *by construction* to work never attempted: it has been +green since Phase 5, when 0% was decompiled, because `INCLUDE_ASM` pastes the ORIGINAL assembly. + +> **A green byte-gate is compatible with ANY decomp percentage.** + +So the instrument the project trusts absolutely cannot see this class at all. Do not reach for it here. + +### §51c — THE METHOD (do not audit by reading the regex) + +Reading regexes is the failure mode that *wrote* these bugs. For every scanner: + +1. Build a deliberately **OVER-APPROXIMATING** candidate detector for what it is *supposed* to find. +2. Run **both** the detector and the real scanner over the **real corpus** (production data, not a toy). +3. `gap = candidates − parsed`. +4. **Classify EVERY item in the gap** — real silent skip, or justified exclusion. "I sampled a few" is + not acceptable; if the gap is large, classify by shape and count each shape. +5. **Measure the blast radius against the corpus.** Not "this could affect X" — go count how many + functions/members/banks are *actually* affected today. Distinguish LIVE from LATENT (armed but not + firing). Both are real; conflating them is not. +6. Pair each finding with an **adversarial skeptic** told to REFUTE it. In this audit the skeptics killed + 4 of 32 findings outright and corrected magnitudes in **both** directions. + +### §51d — THE LAWS + +**LAW 1 — Derive, don't re-derive (R33).** Where a proven invariant answers the question, derive the +answer from it rather than re-parsing the source. + +* The invariant here: *the fleet builds byte-identical, and `INCLUDE_ASM` pastes the ORIGINAL assembly, + therefore a function NOT wrapped in `INCLUDE_ASM` is byte-exact.* +* `progress.py` is the proof, in one file: `weighted_metrics()` **derived** from the invariant and was + correct; `classify()` **re-parsed C** and inherited a bug. Same question, two tools, and *the one that + refused to re-derive was the one that was right.* +* **The best outcome of an audit is a DELETED SCANNER, not a fixed regex.** 28 findings collapsed to one + defect — *a hand-maintained model of the corpus layout sitting on top of a filesystem that already + answers the question* — and the fix was **one derived oracle (`tools/corpus.py`) and ~10 deleted + scanners.** A dict literal is strictly worse than the filesystem AND it **fails OPEN** (silently yields + a plausible wrong answer) instead of closed. +* Corollary: **a derived fact cannot rot; a hand-maintained copy of it is a liability that grows with + every structural change.** `.run/fuel_manifest.json` recorded 130 live stubs on 2026-07-08; the same + code returned **30** six days later, because a TU split moved ~100 stubs out from under a dict literal. + +**LAW 2 — Assert your COVERAGE, not merely your correctness (R32).** A tool that scans the corpus must +compare what it found against an over-approximating candidate set, and fail on the gap. + +> ⚠️ **The sharpest lesson of the audit, and a correction to the first draft of R32:** +> `build_engine_types` **was never silent.** It printed `[overlap] … handle manually` *every single +> time*, for four phases, while hard-exiting on **81% of its own corpus** — the type-heavy tail's only +> sanctioned unblocker, unable to run on the corpus that tail lives in. It went unfixed because the +> message reads like a rare edge case rather than a four-fifths coverage failure, **so nobody ever +> counted it.** +> **A loud failure that nobody counts is exactly as invisible as a silent one.** "Fail loud" is not the +> rule. **"Assert your coverage" is the rule.** + +**LAW 3 — When an oracle is structurally blind to a class of error, add a SECOND ORACLE THAT CAN +DISAGREE WITH IT.** Not a better assertion inside the first one. + +* `config/symbols.us.txt` declared a main-EXE RAM symbol at an address that is *live code* in every + overlay. splat cut **97 real functions in half** and **invented 96 phantoms** — 193 slices unmatchable + *by construction* (one phantom's `.s` literally begins `lw $ra,0x10($sp)` / `addiu $sp,$sp,0x18` / + `jr $ra`: splat cut a function immediately before its **epilogue** and called the epilogue a function). +* **The byte-gate stayed green throughout and always would have**, because the `.s` halves are pasted + back verbatim in original order. One phantom even got **banked** as a real match. +* What exposed it: `sig_image` computes function boundaries from the ORIGINAL bytes *without splat*, and + **disagreed**. `make audit-corpus` is now that second oracle, standing. +* We had both oracles all along and never made them argue. **Redundancy is only worth what you spend + comparing it.** +* ⚠️ **Scope a cross-oracle check to the domain where the second oracle is genuinely independent.** Run + naively over all 136 binaries the same check reports **914** slices; the truth is **193**. main/resident + are signed by the *Ghidra* dumper, whose boundaries are shorter by design — so the comparison measures + *Ghidra's* limits, not splat's errors. **A check applied outside its valid domain does not become more + thorough; it becomes noise.** + +**LAW 4 — A rule that needs a human to remember it is not a gate. Make it structural.** + +* `.o ← .s` is **not** a dependency `make` can see: assembly arrives via `INCLUDE_ASM`, expanded to a + `.include` consumed by maspsx/as *after* cpp, while `-MMD` tracks headers only. Re-extract, build + incrementally, and make links a **stale object**. +* This is not merely slow. `INCLUDE_ASM` pastes the ORIGINAL bytes, so a stale object still yields the + original image: **SHA1 goes GREEN while the split you just changed is never exercised.** A broken + `config/` change can be "verified" by an incremental build. +* R22/H3 already legislate this ("clean rebuild"; "`make clean` after any `config/` change"). They are + right, and they were broken anyway — by me, mid-audit. So `extract` now **deletes the objects that + include what it just rewrote**. Structural, not advisory. + +### §51e — The false-wall pipeline (why this is not just hygiene) + +A silent skip does not stay quiet. It **compounds into a false wall**: + +1. `wave_targets` hands a drafter an asm path that does not exist (78 of 87 targets). +2. The drafter drafts against nothing and fails. +3. The failure is booked into the backlog as a **matching** failure. +4. `reserved_walls()` reads the backlog and **permanently blacklists a function that was never attempted.** + +Same shape with a lying closeness oracle: `masked_diff` left `R_MIPS_PC16` unmasked, so **155 functions +scored a phantom non-zero** against an unresolved placeholder that can never compare equal. An agent +grinds forever at a wall that is not there, and the result is filed as an intrinsic compiler residual. + +> **Before you write up a wall as intrinsic, prove your instruments could have seen the alternative.** +> How many of the walls "byte-proven" across 26 phases were lookup misses wearing a wall's clothes? + +### §51f — Checklist for any new corpus-scanning tool + +* [ ] Does the **filesystem** already answer this? Then glob it — never keep a second copy (LAW 1). +* [ ] Does a **proven invariant** already answer this? Then derive it — never re-parse (LAW 1). +* [ ] Over-approximating candidate set + `assert found == candidates`, failing with the unparsed items (LAW 2). +* [ ] Does it print a count nobody checks? Then it is not asserted — it is decoration (LAW 2). +* [ ] Symbol regexes: **any C identifier**, not `func_[0-9A-Fa-f]{8}` — curated names exist, and curated + naming *increases* as RE quality improves, so a `func_`-only oracle **rots by design**. +* [ ] File lists: **glob**, never a suffix allowlist — the next split kind re-opens the hole. +* [ ] Definition detectors: handle **K&R** (`f(a)` / `int a;` / `{`), **multi-line signatures**, and + **single-line bodies**. K&R is this project's house style for exactly the biggest, highest-reach + functions. And find the signature's closing paren with a real **paren-walk** — `line.count('(')` and + `split(')')[-1]` both land on the wrong paren for a one-line body containing a call. diff --git a/phase-ends/CURRENT_PHASE.md b/phase-ends/CURRENT_PHASE.md index 425112284..a31bca66c 100644 --- a/phase-ends/CURRENT_PHASE.md +++ b/phase-ends/CURRENT_PHASE.md @@ -169,6 +169,58 @@ bytes (R14) and still got the conclusion wrong because **I did not verify its BL --- +## 🔁 SESSION-9 CLOSE — HANDOFF (2026-07-14). Read this, then `docs/tooling-audit.md`. + +**State: `check-all` 136/136 BYTE-IDENTICAL · `make audit-corpus` 0 unmatchable slices · dedup 1823/0 · +fleet instr-weighted 66.5% → 66.7% · 0 NON_MATCHING. Tree clean, all work committed.** + +### DONE (A0–A8 + two unplanned finds) + +| | what | outcome | +|---|---|---| +| **A1** | `dedup_integrate` — a fail-closed gate that printed **false greens** | 3 paths closed w/ negative controls; 7 ghost groups purged | +| **A2** | THE FULL AUDIT (18 tools, 38 agents, 2.24M tok) | **32 raised → 28 survived**, 4 refuted, 40 scanners measured clean | +| **A3** | **`tools/corpus.py`** — ONE derived oracle + **`make audit-corpus`** (a *second oracle that can disagree*) | targets **30→263** · reach-134 **10→127** · gain **83k→994,633 ins** · byte-gate reach **4.9%→100%** | +| **A4** | the **`listCdBuffer`** corpus defect | **193 unmatchable slices → 0**; 4 real functions un-hidden; a **banked phantom** removed | +| **A5** | the closeness oracle (`masked_diff` PC16) | **150 lies → 4** (coverage-asserted over 2,741 fns) | +| **A6** | `dedup_propagate` (glob + K&R `find_site`) | **17 fns banked ×134 free**, incl. all 4 the registry lied about | +| **A7** | family engine (`extract_unit`/`symbol_map`/`gather_externs`/`stub_map`) | **96 phantom exemplars → 0** (216/216 real) | +| **A7** | `build_engine_types` | ran on **73–81%** of its corpus for the first time | +| **A8** | `jr_isolate_all._SAFE_TYPE` | **683 dropped prototypes** — a **latent BYTE-CHANGER** — fixed + coverage-asserted | +| ➕ | **stale objects can produce a FALSE PASS** | `extract` now invalidates them. **Structural, not advisory.** | + +### REMAINING — all fully specified on disk; nothing lives only in a dead session's context + +1. **`tools/cdecl.py`** — the ONE coverage-asserting C-decl parser. The same char-class bug (`[\w\s\*]` cannot hold `(`, `,`, `[N]`) is re-implemented in **6+ scanners**, and two tools in ONE pipeline already disagree about what a data decl *is*. ~15 of the round-1 findings. **Spec + evidence: `docs/tooling-audit.md`, GROUP data-decls / callee-sigs / recovery-rest.** +2. **Wire `tools/reconcile_tu.py`** (written + validated at `commit:0580`, still **NOT WIRED**) into `bank_exemplar`/`jtbl_family_bank`/`gate_stage`; retire `reconcile_decls`' fleet-majority oracle (**3,717 actively-wrong decls**; 21.7% of (TU,symbol) pairs). Unblocks `func_8017A4AC` (287 KB), `func_8013F350`, `func_80131340`. **Then DELETE `census_conflict_callees`** (already marked; `reconcile_tu` answers its question from the build). +3. **`lint_symbol_refs`** — RED (43 false positives) and UNWIRED. Fix the 4 blind spots → green on HEAD → **wire into `make report`**. It is the only detector for the R22 rename-drift failure mode. +4. **`overlay_src_split.scan_construct`** `force_decl` latch (swallows 2 real defs **in the exemplar overlay**, while its selftest passes green — a *serialisation* check masquerading as a *coverage* check) + `jr_inventory`'s banked-roster read from an **ephemeral gitignored scratch file** (R33 violation). +5. **🏆 A10 — RE-TEST THE WALLS.** *This is the payoff and the reason the audit was gated ahead of matching.* + - **"The permuter's fuel is exhausted" (Phase 22) is UNSAFE.** `grinder` banks through `harvest_verify`, which could see ONE TU — **1,290 of its own 1,298 queued fns could never have banked.** "0 banks since Phase 21" is *equally consistent* with *the tool could not bank*. **Re-run it against the fixed gate before repeating that conclusion.** + - The def-side loose-typing wall (§20/§41, "triple-confirmed"); the 159 arity conflicts; the 3,098 type-heavy tail + 9 zero-bank type-using families (`build_engine_types` can now RUN); the **780 h_seq rejections** against the repaired callee oracle. + - Any wall whose closeness came from the **155 wrong scores**, or whose target was one of the **193 listCdBuffer slices** (unmatchable *by construction* — no C exists for them). +6. **A11 — distill + close**: `docs/tooling-audit.md` DIAGNOSIS→ledger (partly done), PhaseEnd, **then resume Phase 26 at Task 7**. + +### RULE CANDIDATES for Drew's ratification (P10) + +- **R32 — Assert your COVERAGE.** A tool that scans the corpus must compare what it found against an + over-approximating candidate set and fail on the gap. + > ⚠️ **This is a CORRECTION to the first draft** ("fail loud on unparsed input"). `build_engine_types` + > **failed loud every single time for four phases** while hard-exiting on 81% of its own corpus — and was + > still invisible, because the message read like an edge case and **nobody counted it**. + > **A loud failure that nobody counts is exactly as invisible as a silent one.** +- **R33 — Derive, don't re-derive.** Where a proven invariant answers the question, derive from it rather + than re-parse. **The best outcome is a DELETED SCANNER, not a fixed regex.** (28 findings → one derived + oracle + ~10 deleted scanners.) +- **R34 (new) — A second oracle, not a better assertion.** When an oracle is *structurally* blind to a class + of error, no assertion inside it can help. Add an independent oracle that can **disagree** with it, and + make them argue. (The byte-gate is a perfect correctness oracle and a **null coverage oracle**; `sig_image` + disagreeing with splat is what exposed the 193 slices. We had both all along and never compared them.) + +**Reusable method + laws: `docs/matching-cookbook.md` §51.** Strategic why: `docs/decision-log.md`. + +--- + ## ▶ SESSION-8 RESULTS (2026-07-13/14) ### 📊 SESSION-8 SCOREBOARD