docs(cookbook): 3 byte-proven addenda from the t5b-t5d distill — §211's guard-hoist INVERTS on a reg-reg copy (LENGTH-DRIFT -1; use the §164-36a fence instead, func_8017DD80); §194-M third instance + new residual shape (func_80182420); §176-B B1 sharpened (func_80180B44). 2 more verdicts COVERED. Index green at 914 (P31 S63 T5.5)

This commit is contained in:
Drew T
2026-08-26 21:42:47 -06:00
parent 926fc3f1c0
commit 1180748502
2 changed files with 1256 additions and 1236 deletions
+1236 -1236
View File
File diff suppressed because it is too large Load Diff
+20
View File
@@ -17690,6 +17690,12 @@ else { s32 *ptr2 = ...; s32 field2 = ptr2[1]; field2 &= 0x7FFFFFFF; ptr2[1]
---
**Addendum (P31 S63 t5b-t5d, func_80180B44):** B1's parenthetical — *"naively fixing that by reusing one local restores the schedule via a WAR dependency but merges two live ranges into one pseudo that then conflicts with the `$a2` arg setup → `$a3`"* — is **not universal**, and on the very same alias family it runs the other way. After B1's own §37 asm-label alias for `D_80126B5E`/`D_80126B62` (closeness 158 → 25) the residual is not a register perm at all but **OPCODE-MIXED with `move`/`nop` swapped at one slot pair**: a reg-to-reg copy of the aliased load stranded immediately *after* the `subu` that overwrites its source register, one instruction too early to be sunk into the following conditional branch's delay slot (which holds `nop`). That is a register-identity residual wearing a scheduling costume — a fresh single-purpose `s32 d` for the tail-guard difference gets an isolated allocno whose register the intervening `subu` is free to clobber, so reorg has nothing to sink, and ~15 naming/statement-order/early-return variants (v4–v11, w1–w3, b2e–b2l, x1–x3, q3/q4) all plateau at 21 with the same signature.
**The two-step fix, and WHICH variable is the dial.** Writing **both** tail-guard differences into ONE variable took 21 → **4** (`REGALLOC-PERM/$v0>$a0` — shape already right, only the hard register left); writing them into the loop's **already-live `rand`-result local** instead of a fresh name merged both defs into the allocno that lands in `$a0`, the register the target actually uses, and the leftover `move` sank into the delay slot on its own: **MATCH 170/170**, whole-binary byte-gate accepted, law-1c re-walk clean (25 external relocations, both masked internal `j` pairs 0x16C×2 / 0x28C×2). So it is not merely *whether* you merge (§45 Lever A / §76 / §162b1) but *which existing allocno you merge into* — §156's donor law ("the preference travels with the allocno; reuse is what carries it") read at delay-slot granularity, and the same "reuse before pins or densities" route.
⚠ **The pin is the wrong tool here, as in §176-B4 / §72 / §156 / §176-C:** `register s32 d __asm__("$4")` on that same variable **regressed 4 → 22** and back to OPCODE-MIXED.
**New row for §164-20's delay-slot table** (whose three rows are all *copy MISSING*): a copy **present but one slot early**, sitting directly under an instruction that clobbers its register, with the branch's slot left `nop` ⇒ do not reorder statements and do not pin — fix the register identity by reusing an already-live local. *(Honest bound: why THAT local yields `$a0`, and why the pin disturbs downstream allocation instead of reproducing the effect, are established empirically by cc1-asm probing (z1→z2 regression), not read out of `local-alloc.c`/`global-alloc.c`.)*
### §176-C — 🔴 WALL REFUTATION: a hard-register pin CANNOT schedule around a call, because of a genuine gcc-2.7.2 bug
**This refutes the universality of `sched.md` S11's protocol step (1) ("pin the callee-saved homes FIRST so scheduling levers can't cascade the alloc") and of the standing "don't conclude unsteerable — try register pins" reflex.** There is a class where the pin *is* the defect, and the correct move is the inverse: **unpin.**
@@ -19606,6 +19612,9 @@ against the banked source shape `ret = func_8012C1B8(); *(s32*)(a0+0x20) = ret;
Population: `.run/harvest_u/scan_slots.py` over all 14,899 `asm/**/*.s` → **2,482** conditional-branch-slot store sites.
**Addendum (P31 S63 t5b-t5d, func_80182420):** THIRD INSTANCE, and a **new residual face for the triage note** — this class also presents as plain **LENGTH-DRIFT +2**, not only as the equal-length `ADDRESSING/move!=sw profile=cse` of the triage note or the equal-length `DELAY-SLOT` of bound 3's `B_below`. `func_80182420` (ov_SC04_011, 132 ins, banked `src/ov_SC04_011/ov_SC04_011_jr_8017D494.c:5618`): a `u16` counter RMW (`cnt = a0->unkF2 + 1; a0->unkF2 = cnt;`) whose `sh $v1,0xF2($s0)` sits in the slot of the *neighbouring* clamp branch `bgtz $v0,.L80182530` (`lim = a0->unkF6 - 1; if (lim <= 0) lim = 1;`). Written BELOW the clamp `if` — bound 3's `B_below` direction — the draft is `LENGTH-DRIFT/2`, close 62/58, 134 vs 132 ins, delta at idx 15; hoisting the two RMW statements ABOVE the `if`, the only edit made, gives **MATCH 132/132, close 0** on the very next `match_one`. **So a length drift does not exonerate statement placement**: bound 6 already says the cost direction is unbounded, and this is its first measured **+2** and the first time `match_one` stamps the class with a *count* class instead of a schedule/addressing one — exactly the misdirection the triage note exists to head off.
**Two claims from this card NOT banked.** (a) Its reading that the statement must sit *immediately* before the branch in the C — bound 4 already states the invariant is **dominance**, not lexical adjacency, and nothing here re-opens that. (b) Its proposed tell that "the same `sll $v0,$a0,16` appearing TWICE, straddling the branch, is the fingerprint of an unfilled slot patched by rematerialization" — the card traced no RTL for it (self-flagged as unconfirmed), and both byte-read duplicated-shift sightings in this file read a duplicated `sll` the OTHER way, as a slot that got **filled**: the cross-BB combine law (L2489, `fill_slots_from_thread` steals the top-BB `sll` into the `bnez` slot and redirects the label) and the §74 counter/clamp addendum (L27920, dbr duplicates an idempotent promotion head into a mutually-exclusive arm's slot). Read the slot itself, per §194-M's own reading rule; do not read the duplicate.
### §194-N — §193-D's C dial is misstated: the lever is a SURVIVING CODE_LABEL (a label with a real incoming edge), not "a label between the block and the call" — a bare label, or a `goto L; L:` pair whose target is the next active insn, is deleted by jump1 (jump.c:663-669 → delete_insn → jump.c:3458-3461, and jump.c:243 for the bare case) long before sched1/local-alloc, and costs exactly zero bytes
**§193-D CORRECTION (replaces the "THE C DIAL" paragraph at L18546 and precondition 1 at L18549).**
@@ -22395,6 +22404,17 @@ dependence** on the address between the two statements (`first = base[0];` place
when the guard/branch actually has a fillable slot and the two induction pseudos actually contend.
On a loop with no dominating guard there is nothing to hoist above.
**Addendum (P31 S63 t5b-t5d, func_8017DD80):** §211's guard-hoist is a lever for a **materialisation**, and it INVERTS on a **copy**. Every card in §211 and in its S58b addendum hoists an insn that computes a fresh value (`addiu $a1,$a0,0xE4`, `addiu $v1,$zero,4`, `i = 0`); when the target's guard slot instead holds a bare `addu $rD,$rS,$zero` copy of a pseudo defined **one statement earlier**, hoisting that copy above the guard does not relocate it into the slot — it lands the copy adjacent to its source's own def inside one basic block, where §46-L2/§48-B's EBB rule ("a source-level `fp = q;` ALWAYS dies **unless defined in a guard block and used in the loop**") deletes it and folds its destination straight into the producing ALU op. **An instruction is eliminated, not moved: `LENGTH-DRIFT/-1` with the whole tail shifted one slot early.** *(Which pass performs the fold was not traced — empirical, not internals-verified; the deletion law it instantiates is byte-proven at §46-L2/§48-B/§162p.)*
**THE TELL.** Target: `addu $a1,$a2,$v0` (address calc) ; `blez $v1,.L` ; **slot** `addu $a0,$a1,$zero`. The hoisted draft has no copy at all and the address calc writes the copy's register directly — `addu $a0,$a2,$v0` — at `nins_mine = nins_tgt - 1`. **Read the slot insn's SHAPE before applying §211: `rD,rS,$zero` ⇒ do NOT hoist; a computed operand pair ⇒ §211 applies as written.**
**THE FIX FOR THIS SHAPE.** Leave the copy inside the guarded block and park §164-36a's zero-byte `__asm__ __volatile__("");` **between the address calc and the guard** — `reorg.c:675 stop_search_p` halts `fill_simple_delay_slots`' backward scan on any asm, so the addr calc can no longer be stolen into the slot and `fill_eager_delay_slots` takes the in-block copy instead. The pre-fence residual names itself: the addr calc and the `blez` appear **transposed** against the target (mine `blez` then `addu`, target `addu` then `blez`), which is the backward scan having won.
**BYTE EVIDENCE — `func_8017DD80` (ov_SC06_025, 55 ins, banked `src/ov_SC06_025/ov_SC06_025_jr_8017BEBC.c:3719`; target `8017ddc8`-`8017ddd0`).** v3 (pins + a `$2`-pinned integer-space temp for `param_1[4]*4`, §239) = closeness **2**, residual exactly `[[18,"blez v1,78","addu $a1,$a2,$v0"],[19,"addu a1,a2,v0","blez $v1,.L8017DDF8"]]`. v4/v5 hoisting `puVar3 = puVar7;` above the `if` — the §211-prescribed edit — **regressed to closeness 39, `LENGTH-DRIFT/-1`, 54 vs 55**. v6/v7 reverted the hoist and added the one fence line ⇒ `MATCH`, closeness **0**, residual `[]`. The fence is solo-proven (§266): it is the only delta between the closeness-2 and the closeness-0 build.
**DISCRIMINATE against the two near neighbours before reaching for this.** §164-82 has the *same* surface tell (`LENGTH-DRIFT −1`, a guard branch whose slot holds `addu $sD,$v0,$zero`) but its cause is a hoisted **LOAD** and its fix is the test-then-re-read spelling — that one applies when the slot copy's source is a memory value, this one when it is a just-computed address. And §156's `sub = pct` bullet prescribes the **opposite** placement ("survives cse ONLY placed in the LOAD's bb before the branch; in an arm every spelling dies") — that regime has the copy's use two fall-through branches down, so cse's follow-jumps table reset saves it; here the use is in the same block as the def and the guard block is the only surviving boundary.
## §212 — THE WALKING CURSOR IS COUNTABLE: `*wp++` emits one `addiu` PER STORE, `wp[0..2]` emits one (P31 S58)
**THE TELL.** Three stores at displacement **0** off a base that steps between them: