docs(phase-26a): SESSION-9 CLOSE — cookbook §51 (the tooling-integrity laws) + handoff

R30/R16: the context-dependent artifacts, written while the context is live.

cookbook §51 — the SILENT SKIP: the bug class, why the byte-gate cannot see it, the
over-approximating-detector method, and FOUR LAWS:
  1. Derive, don't re-derive — the best outcome is a DELETED SCANNER (28 findings -> one
     derived oracle + ~10 deleted scanners). A derived fact cannot rot; a hand-maintained
     copy of it is a liability that grows with every structural change.
  2. Assert your COVERAGE, not merely your correctness. *** A LOUD FAILURE THAT NOBODY
     COUNTS IS EXACTLY AS INVISIBLE AS A SILENT ONE *** — build_engine_types printed
     '[overlap] handle manually' every single time for four phases while dead on 81% of its
     own corpus. This CORRECTS the first draft of R32 ('fail loud'), which was not enough.
  3. When an oracle is structurally blind to a class of error, add a SECOND ORACLE THAT CAN
     DISAGREE WITH IT — not a better assertion inside it. We had two all along and never made
     them argue. (And scope the comparison to where the second oracle is genuinely independent:
     the same check run outside its domain reports 914 slices when the truth is 193.)
  4. A rule that needs a human to remember it is not a gate. Make it structural.
  + the FALSE-WALL PIPELINE (a silent skip -> a wasted draft -> a backlog 'matching failure'
    -> reserved_walls() PERMANENTLY blacklists a function that was never attempted), and a
    checklist for any new corpus-scanning tool.

CURRENT_PHASE: session-9 handoff — what is done, what remains (each with its spec on disk),
and the R32-corrected / R33 / R34-new rule candidates for P10 ratification.
This commit is contained in:
Drew T
2026-07-14 10:40:04 -06:00
parent 2086b15b48
commit c7772bc452
2 changed files with 193 additions and 0 deletions
+141
View File
@@ -3470,3 +3470,144 @@ keep same-page stores apart. **Never "tidy up" the store order of a matched func
and **both are blocked by sched.c**, which always fills the load-use stall between `lh` and its consumer, and
places byte-free (reload-deleted) copies *only* in such stalls. The one productive angle left: a lever that
injects a **reload-deleted no-op reg copy** inside `[lhu 4($a2) … sh %lo(D_801152AC)]`. That is the whole delta.
---
## §51 — TOOLING INTEGRITY: the silent skip, and how to hunt it
> Phase 26-A (the inserted tooling-integrity audit). 68 findings across ~36 tools. This section is the
> METHOD and the LAWS; the findings themselves are in `docs/tooling-audit.md`, the strategic why in
> `docs/decision-log.md`. **Read this before writing any tool that scans the corpus.**
### §51a — The bug class
> **A scanner extracts N items from a corpus. The true count is M > N. Nobody ever compared N to M.**
That is the whole class. It is not a typo, it is a *structural* blind spot, and it produced every one of
the seven bugs found in Phase-26 session 8 and the twenty-eight found in the audit. Its signature:
* the tool reports success;
* the number it reports is *smaller than reality* and self-consistent;
* nothing downstream can tell, because **a target that is never nominated produces silence, not an error.**
Measured consequences in this project: **91.6% of all remaining work was invisible to target selection**
(a 3-file allowlist against a 14-file tree); the byte-gate could reach **4.9%** of the canonical overlay;
**62% of the endgame plan's byte-weight was already-matched phantom targets**; and one 10% hole in a
callee-signature oracle made **nine byte-exact functions look like an intrinsic compiler wall.**
### §51b — Why the byte-gate cannot save you
**The whole-binary byte-gate is a perfect CORRECTNESS oracle and a NULL COVERAGE oracle.** It has never
once accepted a wrong match. It is also blind *by construction* to work never attempted: it has been
green since Phase 5, when 0% was decompiled, because `INCLUDE_ASM` pastes the ORIGINAL assembly.
> **A green byte-gate is compatible with ANY decomp percentage.**
So the instrument the project trusts absolutely cannot see this class at all. Do not reach for it here.
### §51c — THE METHOD (do not audit by reading the regex)
Reading regexes is the failure mode that *wrote* these bugs. For every scanner:
1. Build a deliberately **OVER-APPROXIMATING** candidate detector for what it is *supposed* to find.
2. Run **both** the detector and the real scanner over the **real corpus** (production data, not a toy).
3. `gap = candidates − parsed`.
4. **Classify EVERY item in the gap** — real silent skip, or justified exclusion. "I sampled a few" is
not acceptable; if the gap is large, classify by shape and count each shape.
5. **Measure the blast radius against the corpus.** Not "this could affect X" — go count how many
functions/members/banks are *actually* affected today. Distinguish LIVE from LATENT (armed but not
firing). Both are real; conflating them is not.
6. Pair each finding with an **adversarial skeptic** told to REFUTE it. In this audit the skeptics killed
4 of 32 findings outright and corrected magnitudes in **both** directions.
### §51d — THE LAWS
**LAW 1 — Derive, don't re-derive (R33).** Where a proven invariant answers the question, derive the
answer from it rather than re-parsing the source.
* The invariant here: *the fleet builds byte-identical, and `INCLUDE_ASM` pastes the ORIGINAL assembly,
therefore a function NOT wrapped in `INCLUDE_ASM` is byte-exact.*
* `progress.py` is the proof, in one file: `weighted_metrics()` **derived** from the invariant and was
correct; `classify()` **re-parsed C** and inherited a bug. Same question, two tools, and *the one that
refused to re-derive was the one that was right.*
* **The best outcome of an audit is a DELETED SCANNER, not a fixed regex.** 28 findings collapsed to one
defect — *a hand-maintained model of the corpus layout sitting on top of a filesystem that already
answers the question* — and the fix was **one derived oracle (`tools/corpus.py`) and ~10 deleted
scanners.** A dict literal is strictly worse than the filesystem AND it **fails OPEN** (silently yields
a plausible wrong answer) instead of closed.
* Corollary: **a derived fact cannot rot; a hand-maintained copy of it is a liability that grows with
every structural change.** `.run/fuel_manifest.json` recorded 130 live stubs on 2026-07-08; the same
code returned **30** six days later, because a TU split moved ~100 stubs out from under a dict literal.
**LAW 2 — Assert your COVERAGE, not merely your correctness (R32).** A tool that scans the corpus must
compare what it found against an over-approximating candidate set, and fail on the gap.
> ⚠️ **The sharpest lesson of the audit, and a correction to the first draft of R32:**
> `build_engine_types` **was never silent.** It printed `[overlap] … handle manually` *every single
> time*, for four phases, while hard-exiting on **81% of its own corpus** — the type-heavy tail's only
> sanctioned unblocker, unable to run on the corpus that tail lives in. It went unfixed because the
> message reads like a rare edge case rather than a four-fifths coverage failure, **so nobody ever
> counted it.**
> **A loud failure that nobody counts is exactly as invisible as a silent one.** "Fail loud" is not the
> rule. **"Assert your coverage" is the rule.**
**LAW 3 — When an oracle is structurally blind to a class of error, add a SECOND ORACLE THAT CAN
DISAGREE WITH IT.** Not a better assertion inside the first one.
* `config/symbols.us.txt` declared a main-EXE RAM symbol at an address that is *live code* in every
overlay. splat cut **97 real functions in half** and **invented 96 phantoms** — 193 slices unmatchable
*by construction* (one phantom's `.s` literally begins `lw $ra,0x10($sp)` / `addiu $sp,$sp,0x18` /
`jr $ra`: splat cut a function immediately before its **epilogue** and called the epilogue a function).
* **The byte-gate stayed green throughout and always would have**, because the `.s` halves are pasted
back verbatim in original order. One phantom even got **banked** as a real match.
* What exposed it: `sig_image` computes function boundaries from the ORIGINAL bytes *without splat*, and
**disagreed**. `make audit-corpus` is now that second oracle, standing.
* We had both oracles all along and never made them argue. **Redundancy is only worth what you spend
comparing it.**
* ⚠️ **Scope a cross-oracle check to the domain where the second oracle is genuinely independent.** Run
naively over all 136 binaries the same check reports **914** slices; the truth is **193**. main/resident
are signed by the *Ghidra* dumper, whose boundaries are shorter by design — so the comparison measures
*Ghidra's* limits, not splat's errors. **A check applied outside its valid domain does not become more
thorough; it becomes noise.**
**LAW 4 — A rule that needs a human to remember it is not a gate. Make it structural.**
* `.o ← .s` is **not** a dependency `make` can see: assembly arrives via `INCLUDE_ASM`, expanded to a
`.include` consumed by maspsx/as *after* cpp, while `-MMD` tracks headers only. Re-extract, build
incrementally, and make links a **stale object**.
* This is not merely slow. `INCLUDE_ASM` pastes the ORIGINAL bytes, so a stale object still yields the
original image: **SHA1 goes GREEN while the split you just changed is never exercised.** A broken
`config/` change can be "verified" by an incremental build.
* R22/H3 already legislate this ("clean rebuild"; "`make clean` after any `config/` change"). They are
right, and they were broken anyway — by me, mid-audit. So `extract` now **deletes the objects that
include what it just rewrote**. Structural, not advisory.
### §51e — The false-wall pipeline (why this is not just hygiene)
A silent skip does not stay quiet. It **compounds into a false wall**:
1. `wave_targets` hands a drafter an asm path that does not exist (78 of 87 targets).
2. The drafter drafts against nothing and fails.
3. The failure is booked into the backlog as a **matching** failure.
4. `reserved_walls()` reads the backlog and **permanently blacklists a function that was never attempted.**
Same shape with a lying closeness oracle: `masked_diff` left `R_MIPS_PC16` unmasked, so **155 functions
scored a phantom non-zero** against an unresolved placeholder that can never compare equal. An agent
grinds forever at a wall that is not there, and the result is filed as an intrinsic compiler residual.
> **Before you write up a wall as intrinsic, prove your instruments could have seen the alternative.**
> How many of the walls "byte-proven" across 26 phases were lookup misses wearing a wall's clothes?
### §51f — Checklist for any new corpus-scanning tool
* [ ] Does the **filesystem** already answer this? Then glob it — never keep a second copy (LAW 1).
* [ ] Does a **proven invariant** already answer this? Then derive it — never re-parse (LAW 1).
* [ ] Over-approximating candidate set + `assert found == candidates`, failing with the unparsed items (LAW 2).
* [ ] Does it print a count nobody checks? Then it is not asserted — it is decoration (LAW 2).
* [ ] Symbol regexes: **any C identifier**, not `func_[0-9A-Fa-f]{8}` — curated names exist, and curated
naming *increases* as RE quality improves, so a `func_`-only oracle **rots by design**.
* [ ] File lists: **glob**, never a suffix allowlist — the next split kind re-opens the hole.
* [ ] Definition detectors: handle **K&R** (`f(a)` / `int a;` / `{`), **multi-line signatures**, and
**single-line bodies**. K&R is this project's house style for exactly the biggest, highest-reach
functions. And find the signature's closing paren with a real **paren-walk** — `line.count('(')` and
`split(')')[-1]` both land on the wrong paren for a one-line body containing a call.
+52
View File
@@ -169,6 +169,58 @@ bytes (R14) and still got the conclusion wrong because **I did not verify its BL
---
## 🔁 SESSION-9 CLOSE — HANDOFF (2026-07-14). Read this, then `docs/tooling-audit.md`.
**State: `check-all` 136/136 BYTE-IDENTICAL · `make audit-corpus` 0 unmatchable slices · dedup 1823/0 ·
fleet instr-weighted 66.5% → 66.7% · 0 NON_MATCHING. Tree clean, all work committed.**
### DONE (A0–A8 + two unplanned finds)
| | what | outcome |
|---|---|---|
| **A1** | `dedup_integrate` — a fail-closed gate that printed **false greens** | 3 paths closed w/ negative controls; 7 ghost groups purged |
| **A2** | THE FULL AUDIT (18 tools, 38 agents, 2.24M tok) | **32 raised → 28 survived**, 4 refuted, 40 scanners measured clean |
| **A3** | **`tools/corpus.py`** — ONE derived oracle + **`make audit-corpus`** (a *second oracle that can disagree*) | targets **30→263** · reach-134 **10→127** · gain **83k→994,633 ins** · byte-gate reach **4.9%→100%** |
| **A4** | the **`listCdBuffer`** corpus defect | **193 unmatchable slices → 0**; 4 real functions un-hidden; a **banked phantom** removed |
| **A5** | the closeness oracle (`masked_diff` PC16) | **150 lies → 4** (coverage-asserted over 2,741 fns) |
| **A6** | `dedup_propagate` (glob + K&R `find_site`) | **17 fns banked ×134 free**, incl. all 4 the registry lied about |
| **A7** | family engine (`extract_unit`/`symbol_map`/`gather_externs`/`stub_map`) | **96 phantom exemplars → 0** (216/216 real) |
| **A7** | `build_engine_types` | ran on **73–81%** of its corpus for the first time |
| **A8** | `jr_isolate_all._SAFE_TYPE` | **683 dropped prototypes** — a **latent BYTE-CHANGER** — fixed + coverage-asserted |
| ➕ | **stale objects can produce a FALSE PASS** | `extract` now invalidates them. **Structural, not advisory.** |
### REMAINING — all fully specified on disk; nothing lives only in a dead session's context
1. **`tools/cdecl.py`** — the ONE coverage-asserting C-decl parser. The same char-class bug (`[\w\s\*]` cannot hold `(`, `,`, `[N]`) is re-implemented in **6+ scanners**, and two tools in ONE pipeline already disagree about what a data decl *is*. ~15 of the round-1 findings. **Spec + evidence: `docs/tooling-audit.md`, GROUP data-decls / callee-sigs / recovery-rest.**
2. **Wire `tools/reconcile_tu.py`** (written + validated at `commit:0580`, still **NOT WIRED**) into `bank_exemplar`/`jtbl_family_bank`/`gate_stage`; retire `reconcile_decls`' fleet-majority oracle (**3,717 actively-wrong decls**; 21.7% of (TU,symbol) pairs). Unblocks `func_8017A4AC` (287 KB), `func_8013F350`, `func_80131340`. **Then DELETE `census_conflict_callees`** (already marked; `reconcile_tu` answers its question from the build).
3. **`lint_symbol_refs`** — RED (43 false positives) and UNWIRED. Fix the 4 blind spots → green on HEAD → **wire into `make report`**. It is the only detector for the R22 rename-drift failure mode.
4. **`overlay_src_split.scan_construct`** `force_decl` latch (swallows 2 real defs **in the exemplar overlay**, while its selftest passes green — a *serialisation* check masquerading as a *coverage* check) + `jr_inventory`'s banked-roster read from an **ephemeral gitignored scratch file** (R33 violation).
5. **🏆 A10 — RE-TEST THE WALLS.** *This is the payoff and the reason the audit was gated ahead of matching.*
- **"The permuter's fuel is exhausted" (Phase 22) is UNSAFE.** `grinder` banks through `harvest_verify`, which could see ONE TU — **1,290 of its own 1,298 queued fns could never have banked.** "0 banks since Phase 21" is *equally consistent* with *the tool could not bank*. **Re-run it against the fixed gate before repeating that conclusion.**
- The def-side loose-typing wall (§20/§41, "triple-confirmed"); the 159 arity conflicts; the 3,098 type-heavy tail + 9 zero-bank type-using families (`build_engine_types` can now RUN); the **780 h_seq rejections** against the repaired callee oracle.
- Any wall whose closeness came from the **155 wrong scores**, or whose target was one of the **193 listCdBuffer slices** (unmatchable *by construction* — no C exists for them).
6. **A11 — distill + close**: `docs/tooling-audit.md` DIAGNOSIS→ledger (partly done), PhaseEnd, **then resume Phase 26 at Task 7**.
### RULE CANDIDATES for Drew's ratification (P10)
- **R32 — Assert your COVERAGE.** A tool that scans the corpus must compare what it found against an
over-approximating candidate set and fail on the gap.
> ⚠️ **This is a CORRECTION to the first draft** ("fail loud on unparsed input"). `build_engine_types`
> **failed loud every single time for four phases** while hard-exiting on 81% of its own corpus — and was
> still invisible, because the message read like an edge case and **nobody counted it**.
> **A loud failure that nobody counts is exactly as invisible as a silent one.**
- **R33 — Derive, don't re-derive.** Where a proven invariant answers the question, derive from it rather
than re-parse. **The best outcome is a DELETED SCANNER, not a fixed regex.** (28 findings → one derived
oracle + ~10 deleted scanners.)
- **R34 (new) — A second oracle, not a better assertion.** When an oracle is *structurally* blind to a class
of error, no assertion inside it can help. Add an independent oracle that can **disagree** with it, and
make them argue. (The byte-gate is a perfect correctness oracle and a **null coverage oracle**; `sig_image`
disagreeing with splat is what exposed the 193 slices. We had both all along and never compared them.)
**Reusable method + laws: `docs/matching-cookbook.md` §51.** Strategic why: `docs/decision-log.md`.
---
## ▶ SESSION-8 RESULTS (2026-07-13/14)
### 📊 SESSION-8 SCOREBOARD