mirror of
https://github.com/Druthulu/BFM-decomp
synced 2026-09-29 15:18:24 -04:00
docs(phase-26a): SESSION-9 CLOSE — cookbook §51 (the tooling-integrity laws) + handoff
R30/R16: the context-dependent artifacts, written while the context is live.
cookbook §51 — the SILENT SKIP: the bug class, why the byte-gate cannot see it, the
over-approximating-detector method, and FOUR LAWS:
1. Derive, don't re-derive — the best outcome is a DELETED SCANNER (28 findings -> one
derived oracle + ~10 deleted scanners). A derived fact cannot rot; a hand-maintained
copy of it is a liability that grows with every structural change.
2. Assert your COVERAGE, not merely your correctness. *** A LOUD FAILURE THAT NOBODY
COUNTS IS EXACTLY AS INVISIBLE AS A SILENT ONE *** — build_engine_types printed
'[overlap] handle manually' every single time for four phases while dead on 81% of its
own corpus. This CORRECTS the first draft of R32 ('fail loud'), which was not enough.
3. When an oracle is structurally blind to a class of error, add a SECOND ORACLE THAT CAN
DISAGREE WITH IT — not a better assertion inside it. We had two all along and never made
them argue. (And scope the comparison to where the second oracle is genuinely independent:
the same check run outside its domain reports 914 slices when the truth is 193.)
4. A rule that needs a human to remember it is not a gate. Make it structural.
+ the FALSE-WALL PIPELINE (a silent skip -> a wasted draft -> a backlog 'matching failure'
-> reserved_walls() PERMANENTLY blacklists a function that was never attempted), and a
checklist for any new corpus-scanning tool.
CURRENT_PHASE: session-9 handoff — what is done, what remains (each with its spec on disk),
and the R32-corrected / R33 / R34-new rule candidates for P10 ratification.
This commit is contained in:
@@ -3470,3 +3470,144 @@ keep same-page stores apart. **Never "tidy up" the store order of a matched func
|
||||
and **both are blocked by sched.c**, which always fills the load-use stall between `lh` and its consumer, and
|
||||
places byte-free (reload-deleted) copies *only* in such stalls. The one productive angle left: a lever that
|
||||
injects a **reload-deleted no-op reg copy** inside `[lhu 4($a2) … sh %lo(D_801152AC)]`. That is the whole delta.
|
||||
|
||||
---
|
||||
|
||||
## §51 — TOOLING INTEGRITY: the silent skip, and how to hunt it
|
||||
|
||||
> Phase 26-A (the inserted tooling-integrity audit). 68 findings across ~36 tools. This section is the
|
||||
> METHOD and the LAWS; the findings themselves are in `docs/tooling-audit.md`, the strategic why in
|
||||
> `docs/decision-log.md`. **Read this before writing any tool that scans the corpus.**
|
||||
|
||||
### §51a — The bug class
|
||||
|
||||
> **A scanner extracts N items from a corpus. The true count is M > N. Nobody ever compared N to M.**
|
||||
|
||||
That is the whole class. It is not a typo, it is a *structural* blind spot, and it produced every one of
|
||||
the seven bugs found in Phase-26 session 8 and the twenty-eight found in the audit. Its signature:
|
||||
|
||||
* the tool reports success;
|
||||
* the number it reports is *smaller than reality* and self-consistent;
|
||||
* nothing downstream can tell, because **a target that is never nominated produces silence, not an error.**
|
||||
|
||||
Measured consequences in this project: **91.6% of all remaining work was invisible to target selection**
|
||||
(a 3-file allowlist against a 14-file tree); the byte-gate could reach **4.9%** of the canonical overlay;
|
||||
**62% of the endgame plan's byte-weight was already-matched phantom targets**; and one 10% hole in a
|
||||
callee-signature oracle made **nine byte-exact functions look like an intrinsic compiler wall.**
|
||||
|
||||
### §51b — Why the byte-gate cannot save you
|
||||
|
||||
**The whole-binary byte-gate is a perfect CORRECTNESS oracle and a NULL COVERAGE oracle.** It has never
|
||||
once accepted a wrong match. It is also blind *by construction* to work never attempted: it has been
|
||||
green since Phase 5, when 0% was decompiled, because `INCLUDE_ASM` pastes the ORIGINAL assembly.
|
||||
|
||||
> **A green byte-gate is compatible with ANY decomp percentage.**
|
||||
|
||||
So the instrument the project trusts absolutely cannot see this class at all. Do not reach for it here.
|
||||
|
||||
### §51c — THE METHOD (do not audit by reading the regex)
|
||||
|
||||
Reading regexes is the failure mode that *wrote* these bugs. For every scanner:
|
||||
|
||||
1. Build a deliberately **OVER-APPROXIMATING** candidate detector for what it is *supposed* to find.
|
||||
2. Run **both** the detector and the real scanner over the **real corpus** (production data, not a toy).
|
||||
3. `gap = candidates − parsed`.
|
||||
4. **Classify EVERY item in the gap** — real silent skip, or justified exclusion. "I sampled a few" is
|
||||
not acceptable; if the gap is large, classify by shape and count each shape.
|
||||
5. **Measure the blast radius against the corpus.** Not "this could affect X" — go count how many
|
||||
functions/members/banks are *actually* affected today. Distinguish LIVE from LATENT (armed but not
|
||||
firing). Both are real; conflating them is not.
|
||||
6. Pair each finding with an **adversarial skeptic** told to REFUTE it. In this audit the skeptics killed
|
||||
4 of 32 findings outright and corrected magnitudes in **both** directions.
|
||||
|
||||
### §51d — THE LAWS
|
||||
|
||||
**LAW 1 — Derive, don't re-derive (R33).** Where a proven invariant answers the question, derive the
|
||||
answer from it rather than re-parsing the source.
|
||||
|
||||
* The invariant here: *the fleet builds byte-identical, and `INCLUDE_ASM` pastes the ORIGINAL assembly,
|
||||
therefore a function NOT wrapped in `INCLUDE_ASM` is byte-exact.*
|
||||
* `progress.py` is the proof, in one file: `weighted_metrics()` **derived** from the invariant and was
|
||||
correct; `classify()` **re-parsed C** and inherited a bug. Same question, two tools, and *the one that
|
||||
refused to re-derive was the one that was right.*
|
||||
* **The best outcome of an audit is a DELETED SCANNER, not a fixed regex.** 28 findings collapsed to one
|
||||
defect — *a hand-maintained model of the corpus layout sitting on top of a filesystem that already
|
||||
answers the question* — and the fix was **one derived oracle (`tools/corpus.py`) and ~10 deleted
|
||||
scanners.** A dict literal is strictly worse than the filesystem AND it **fails OPEN** (silently yields
|
||||
a plausible wrong answer) instead of closed.
|
||||
* Corollary: **a derived fact cannot rot; a hand-maintained copy of it is a liability that grows with
|
||||
every structural change.** `.run/fuel_manifest.json` recorded 130 live stubs on 2026-07-08; the same
|
||||
code returned **30** six days later, because a TU split moved ~100 stubs out from under a dict literal.
|
||||
|
||||
**LAW 2 — Assert your COVERAGE, not merely your correctness (R32).** A tool that scans the corpus must
|
||||
compare what it found against an over-approximating candidate set, and fail on the gap.
|
||||
|
||||
> ⚠️ **The sharpest lesson of the audit, and a correction to the first draft of R32:**
|
||||
> `build_engine_types` **was never silent.** It printed `[overlap] … handle manually` *every single
|
||||
> time*, for four phases, while hard-exiting on **81% of its own corpus** — the type-heavy tail's only
|
||||
> sanctioned unblocker, unable to run on the corpus that tail lives in. It went unfixed because the
|
||||
> message reads like a rare edge case rather than a four-fifths coverage failure, **so nobody ever
|
||||
> counted it.**
|
||||
> **A loud failure that nobody counts is exactly as invisible as a silent one.** "Fail loud" is not the
|
||||
> rule. **"Assert your coverage" is the rule.**
|
||||
|
||||
**LAW 3 — When an oracle is structurally blind to a class of error, add a SECOND ORACLE THAT CAN
|
||||
DISAGREE WITH IT.** Not a better assertion inside the first one.
|
||||
|
||||
* `config/symbols.us.txt` declared a main-EXE RAM symbol at an address that is *live code* in every
|
||||
overlay. splat cut **97 real functions in half** and **invented 96 phantoms** — 193 slices unmatchable
|
||||
*by construction* (one phantom's `.s` literally begins `lw $ra,0x10($sp)` / `addiu $sp,$sp,0x18` /
|
||||
`jr $ra`: splat cut a function immediately before its **epilogue** and called the epilogue a function).
|
||||
* **The byte-gate stayed green throughout and always would have**, because the `.s` halves are pasted
|
||||
back verbatim in original order. One phantom even got **banked** as a real match.
|
||||
* What exposed it: `sig_image` computes function boundaries from the ORIGINAL bytes *without splat*, and
|
||||
**disagreed**. `make audit-corpus` is now that second oracle, standing.
|
||||
* We had both oracles all along and never made them argue. **Redundancy is only worth what you spend
|
||||
comparing it.**
|
||||
* ⚠️ **Scope a cross-oracle check to the domain where the second oracle is genuinely independent.** Run
|
||||
naively over all 136 binaries the same check reports **914** slices; the truth is **193**. main/resident
|
||||
are signed by the *Ghidra* dumper, whose boundaries are shorter by design — so the comparison measures
|
||||
*Ghidra's* limits, not splat's errors. **A check applied outside its valid domain does not become more
|
||||
thorough; it becomes noise.**
|
||||
|
||||
**LAW 4 — A rule that needs a human to remember it is not a gate. Make it structural.**
|
||||
|
||||
* `.o ← .s` is **not** a dependency `make` can see: assembly arrives via `INCLUDE_ASM`, expanded to a
|
||||
`.include` consumed by maspsx/as *after* cpp, while `-MMD` tracks headers only. Re-extract, build
|
||||
incrementally, and make links a **stale object**.
|
||||
* This is not merely slow. `INCLUDE_ASM` pastes the ORIGINAL bytes, so a stale object still yields the
|
||||
original image: **SHA1 goes GREEN while the split you just changed is never exercised.** A broken
|
||||
`config/` change can be "verified" by an incremental build.
|
||||
* R22/H3 already legislate this ("clean rebuild"; "`make clean` after any `config/` change"). They are
|
||||
right, and they were broken anyway — by me, mid-audit. So `extract` now **deletes the objects that
|
||||
include what it just rewrote**. Structural, not advisory.
|
||||
|
||||
### §51e — The false-wall pipeline (why this is not just hygiene)
|
||||
|
||||
A silent skip does not stay quiet. It **compounds into a false wall**:
|
||||
|
||||
1. `wave_targets` hands a drafter an asm path that does not exist (78 of 87 targets).
|
||||
2. The drafter drafts against nothing and fails.
|
||||
3. The failure is booked into the backlog as a **matching** failure.
|
||||
4. `reserved_walls()` reads the backlog and **permanently blacklists a function that was never attempted.**
|
||||
|
||||
Same shape with a lying closeness oracle: `masked_diff` left `R_MIPS_PC16` unmasked, so **155 functions
|
||||
scored a phantom non-zero** against an unresolved placeholder that can never compare equal. An agent
|
||||
grinds forever at a wall that is not there, and the result is filed as an intrinsic compiler residual.
|
||||
|
||||
> **Before you write up a wall as intrinsic, prove your instruments could have seen the alternative.**
|
||||
> How many of the walls "byte-proven" across 26 phases were lookup misses wearing a wall's clothes?
|
||||
|
||||
### §51f — Checklist for any new corpus-scanning tool
|
||||
|
||||
* [ ] Does the **filesystem** already answer this? Then glob it — never keep a second copy (LAW 1).
|
||||
* [ ] Does a **proven invariant** already answer this? Then derive it — never re-parse (LAW 1).
|
||||
* [ ] Over-approximating candidate set + `assert found == candidates`, failing with the unparsed items (LAW 2).
|
||||
* [ ] Does it print a count nobody checks? Then it is not asserted — it is decoration (LAW 2).
|
||||
* [ ] Symbol regexes: **any C identifier**, not `func_[0-9A-Fa-f]{8}` — curated names exist, and curated
|
||||
naming *increases* as RE quality improves, so a `func_`-only oracle **rots by design**.
|
||||
* [ ] File lists: **glob**, never a suffix allowlist — the next split kind re-opens the hole.
|
||||
* [ ] Definition detectors: handle **K&R** (`f(a)` / `int a;` / `{`), **multi-line signatures**, and
|
||||
**single-line bodies**. K&R is this project's house style for exactly the biggest, highest-reach
|
||||
functions. And find the signature's closing paren with a real **paren-walk** — `line.count('(')` and
|
||||
`split(')')[-1]` both land on the wrong paren for a one-line body containing a call.
|
||||
|
||||
@@ -169,6 +169,58 @@ bytes (R14) and still got the conclusion wrong because **I did not verify its BL
|
||||
|
||||
---
|
||||
|
||||
## 🔁 SESSION-9 CLOSE — HANDOFF (2026-07-14). Read this, then `docs/tooling-audit.md`.
|
||||
|
||||
**State: `check-all` 136/136 BYTE-IDENTICAL · `make audit-corpus` 0 unmatchable slices · dedup 1823/0 ·
|
||||
fleet instr-weighted 66.5% → 66.7% · 0 NON_MATCHING. Tree clean, all work committed.**
|
||||
|
||||
### DONE (A0–A8 + two unplanned finds)
|
||||
|
||||
| | what | outcome |
|
||||
|---|---|---|
|
||||
| **A1** | `dedup_integrate` — a fail-closed gate that printed **false greens** | 3 paths closed w/ negative controls; 7 ghost groups purged |
|
||||
| **A2** | THE FULL AUDIT (18 tools, 38 agents, 2.24M tok) | **32 raised → 28 survived**, 4 refuted, 40 scanners measured clean |
|
||||
| **A3** | **`tools/corpus.py`** — ONE derived oracle + **`make audit-corpus`** (a *second oracle that can disagree*) | targets **30→263** · reach-134 **10→127** · gain **83k→994,633 ins** · byte-gate reach **4.9%→100%** |
|
||||
| **A4** | the **`listCdBuffer`** corpus defect | **193 unmatchable slices → 0**; 4 real functions un-hidden; a **banked phantom** removed |
|
||||
| **A5** | the closeness oracle (`masked_diff` PC16) | **150 lies → 4** (coverage-asserted over 2,741 fns) |
|
||||
| **A6** | `dedup_propagate` (glob + K&R `find_site`) | **17 fns banked ×134 free**, incl. all 4 the registry lied about |
|
||||
| **A7** | family engine (`extract_unit`/`symbol_map`/`gather_externs`/`stub_map`) | **96 phantom exemplars → 0** (216/216 real) |
|
||||
| **A7** | `build_engine_types` | ran on **73–81%** of its corpus for the first time |
|
||||
| **A8** | `jr_isolate_all._SAFE_TYPE` | **683 dropped prototypes** — a **latent BYTE-CHANGER** — fixed + coverage-asserted |
|
||||
| ➕ | **stale objects can produce a FALSE PASS** | `extract` now invalidates them. **Structural, not advisory.** |
|
||||
|
||||
### REMAINING — all fully specified on disk; nothing lives only in a dead session's context
|
||||
|
||||
1. **`tools/cdecl.py`** — the ONE coverage-asserting C-decl parser. The same char-class bug (`[\w\s\*]` cannot hold `(`, `,`, `[N]`) is re-implemented in **6+ scanners**, and two tools in ONE pipeline already disagree about what a data decl *is*. ~15 of the round-1 findings. **Spec + evidence: `docs/tooling-audit.md`, GROUP data-decls / callee-sigs / recovery-rest.**
|
||||
2. **Wire `tools/reconcile_tu.py`** (written + validated at `commit:0580`, still **NOT WIRED**) into `bank_exemplar`/`jtbl_family_bank`/`gate_stage`; retire `reconcile_decls`' fleet-majority oracle (**3,717 actively-wrong decls**; 21.7% of (TU,symbol) pairs). Unblocks `func_8017A4AC` (287 KB), `func_8013F350`, `func_80131340`. **Then DELETE `census_conflict_callees`** (already marked; `reconcile_tu` answers its question from the build).
|
||||
3. **`lint_symbol_refs`** — RED (43 false positives) and UNWIRED. Fix the 4 blind spots → green on HEAD → **wire into `make report`**. It is the only detector for the R22 rename-drift failure mode.
|
||||
4. **`overlay_src_split.scan_construct`** `force_decl` latch (swallows 2 real defs **in the exemplar overlay**, while its selftest passes green — a *serialisation* check masquerading as a *coverage* check) + `jr_inventory`'s banked-roster read from an **ephemeral gitignored scratch file** (R33 violation).
|
||||
5. **🏆 A10 — RE-TEST THE WALLS.** *This is the payoff and the reason the audit was gated ahead of matching.*
|
||||
- **"The permuter's fuel is exhausted" (Phase 22) is UNSAFE.** `grinder` banks through `harvest_verify`, which could see ONE TU — **1,290 of its own 1,298 queued fns could never have banked.** "0 banks since Phase 21" is *equally consistent* with *the tool could not bank*. **Re-run it against the fixed gate before repeating that conclusion.**
|
||||
- The def-side loose-typing wall (§20/§41, "triple-confirmed"); the 159 arity conflicts; the 3,098 type-heavy tail + 9 zero-bank type-using families (`build_engine_types` can now RUN); the **780 h_seq rejections** against the repaired callee oracle.
|
||||
- Any wall whose closeness came from the **155 wrong scores**, or whose target was one of the **193 listCdBuffer slices** (unmatchable *by construction* — no C exists for them).
|
||||
6. **A11 — distill + close**: `docs/tooling-audit.md` DIAGNOSIS→ledger (partly done), PhaseEnd, **then resume Phase 26 at Task 7**.
|
||||
|
||||
### RULE CANDIDATES for Drew's ratification (P10)
|
||||
|
||||
- **R32 — Assert your COVERAGE.** A tool that scans the corpus must compare what it found against an
|
||||
over-approximating candidate set and fail on the gap.
|
||||
> ⚠️ **This is a CORRECTION to the first draft** ("fail loud on unparsed input"). `build_engine_types`
|
||||
> **failed loud every single time for four phases** while hard-exiting on 81% of its own corpus — and was
|
||||
> still invisible, because the message read like an edge case and **nobody counted it**.
|
||||
> **A loud failure that nobody counts is exactly as invisible as a silent one.**
|
||||
- **R33 — Derive, don't re-derive.** Where a proven invariant answers the question, derive from it rather
|
||||
than re-parse. **The best outcome is a DELETED SCANNER, not a fixed regex.** (28 findings → one derived
|
||||
oracle + ~10 deleted scanners.)
|
||||
- **R34 (new) — A second oracle, not a better assertion.** When an oracle is *structurally* blind to a class
|
||||
of error, no assertion inside it can help. Add an independent oracle that can **disagree** with it, and
|
||||
make them argue. (The byte-gate is a perfect correctness oracle and a **null coverage oracle**; `sig_image`
|
||||
disagreeing with splat is what exposed the 193 slices. We had both all along and never compared them.)
|
||||
|
||||
**Reusable method + laws: `docs/matching-cookbook.md` §51.** Strategic why: `docs/decision-log.md`.
|
||||
|
||||
---
|
||||
|
||||
## ▶ SESSION-8 RESULTS (2026-07-13/14)
|
||||
|
||||
### 📊 SESSION-8 SCOREBOARD
|
||||
|
||||
Reference in New Issue
Block a user