From 57ff04c1ac539ec1dbacf78962d1ea9ea5c03b0b Mon Sep 17 00:00:00 2001 From: Drew T <50529377+Druthulu@users.noreply.github.com> Date: Wed, 12 Aug 2026 18:55:05 -0600 Subject: [PATCH] =?UTF-8?q?docs(phase-30=20S48):=20=C2=A7167=20=E2=80=94?= =?UTF-8?q?=20wave-5/6=20harvest,=20and=20the=20saturation=20signal?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 27 note-sets, 197 claims, one skeptic each, against a cookbook already holding §162-§166 from this campaign: NEW 5 · SHARPENS 43 · COVERED 126 · UNSOUND 23 byte-probed 92 · single-instance 75 · asserted 30 COVERED+UNSOUND: 57% (§164) -> 64% (§165) -> 76% (here). The duplicate rate rises monotonically as the base grows. FIVE genuinely new laws out of 197 claims is the signal that the idiom well for this class of function is approaching dry — future waves should spend tokens on cracks, not on mining notes for idioms, and harvest only what a skeptic grades byte-probed. The skeptics ran their own A/Bs this round. Best example: a crack agent claimed "the source STATEMENT BOUNDARY decides whether the scheduler hoists a far-consumed load". The vetter built that spelling and got .text BYTE-IDENTICAL to the inline form, then swept eight POSITIONS and got five distinct objects — showing the lever is statement position (the already-banked INSN_LUID tie-break), not the boundary. Plausible mechanism, refuted by measurement, true lever named in its place. 46 entries banked as §167-01..46; §167z records the 23 refutations. cookbook_index.py: 508 sections. --- docs/cookbook-index.md | 10 +- docs/matching-cookbook.md | 1256 +++++++++++++++++++++++++++++++++++++ 2 files changed, 1263 insertions(+), 3 deletions(-) diff --git a/docs/cookbook-index.md b/docs/cookbook-index.md index 849ef506cf..57549c4dbb 100644 --- a/docs/cookbook-index.md +++ b/docs/cookbook-index.md @@ -2,7 +2,7 @@ > **Generated by `tools/cookbook_index.py` — do not hand-edit** (R33). Regenerate after adding a cookbook section. > -> `docs/matching-cookbook.md` is ~716 KB / 506 sections. Grepping it blind is how three P30 wave-1 agents each "discovered" an idiom that was already written down. **Start here, then read the section.** A section appears under every symptom it addresses. +> `docs/matching-cookbook.md` is ~716 KB / 508 sections. Grepping it blind is how three P30 wave-1 agents each "discovered" an idiom that was already written down. **Start here, then read the section.** A section appears under every symptom it addresses. **How to use:** name what you SEE in the diff (a stolen delay slot, an extra `la`, a swapped register pair, a `conflicting types` error), find that symptom below, read those sections first. If nothing fits, THEN grind — and add a section when you win. @@ -435,7 +435,7 @@ - **§162** — THE LICM PAIR: what makes an address a movable AT ALL, and why the preheader order is the body order (P30 S47, `ov_MAIN_012`) L11140 - **§166** — THE DESTINATION-TU ORACLE (P30 S48): the seven-attempt bug that was never codegen L14899 -### build graph, splat & the harness (108) +### build graph, splat & the harness (109) - **§4** — Flag/toolchain gotchas L190 - **Build** — mechanism — per-file opt override (splat resegmentation) L288 @@ -545,8 +545,9 @@ - **§163** — S48 WAVES 2-3 HARVEST (P30, 2026-08-11/12): the five that were byte-probed and are actionable L11757 - **§165** — S48 WAVE-4 HARVEST (P30, 2026-08-12): banked the same day the wave landed L13686 - **§166** — THE DESTINATION-TU ORACLE (P30 S48): the seven-attempt bug that was never codegen L14899 +- **§167** — S48 WAVE-5/6 HARVEST (P30, 2026-08-12): the saturation point L14955 -### process, measurement & doctrine (71) +### process, measurement & doctrine (72) - **§8e** — The jtbl ALIGNMENT LAW + the pad-spec filter — multi-table .rodata spans (Phase 29, byte-proven; `.run/probe_jtbl/verdict.md`) L530 - **§3-The** — mechanism: game-code dedup is SOURCE-LEVEL, not an object swap (R-D1, the key lesson) L926 @@ -619,6 +620,7 @@ - **§163z** — THE UNVETTED REMAINDER (do not cite as law; each needs a dedupe pass) L11825 - **§164z** — REFUTED CLAIMS: do NOT re-derive these L13620 - **§165z** — REFUTED THIS WAVE: do NOT re-derive L14855 +- **§167z** — REFUTED IN WAVES 5/6: do NOT re-derive L16151 ### (unbucketed — title matched no symptom vocabulary) (166) @@ -1298,3 +1300,5 @@ - **§3-The** — ADDRESS-CLASS TABLE: which load/store pairs even REACH the `/s` clause (P30 S48 wave 4, `func_80185B44`, ov_SC03_014) L14339 - **§165z** — REFUTED THIS WAVE: do NOT re-derive L14855 - **§166** — THE DESTINATION-TU ORACLE (P30 S48): the seven-attempt bug that was never codegen L14899 +- **§167** — S48 WAVE-5/6 HARVEST (P30, 2026-08-12): the saturation point L14955 +- **§167z** — REFUTED IN WAVES 5/6: do NOT re-derive L16151 diff --git a/docs/matching-cookbook.md b/docs/matching-cookbook.md index c8636f6342..66b2d6ea4f 100644 --- a/docs/matching-cookbook.md +++ b/docs/matching-cookbook.md @@ -14948,3 +14948,1259 @@ pinned triple end-to-end, and masked-diff your function out of the WHOLE-TU obje TUs: 279/279, 0 masked diffs, cpp/cc1 clean; collateral check 71/71 other sized symbols identical. The same agent corrected the family reach to **×2** (only two `.s` exist, byte-identical modulo the overlay name) against a map that claimed ×5. + + +--- + +## §167 — S48 WAVE-5/6 HARVEST (P30, 2026-08-12): the saturation point + +27 note-sets, **197 distinct laws claimed**, one independent skeptic each, vetted against a cookbook +already holding §162-§166 from this same campaign: + +| verdict | count | | evidence | count | +|---|---:|---|---|---:| +| NEW | 5 | | byte-probed | 92 | +| SHARPENS | 43 | | single-instance | 75 | +| COVERED | 126 | | asserted | 30 | +| UNSOUND | 23 | | | | + +**COVERED+UNSOUND: 57% (§164) → 64% (§165) → 76% (here).** The duplicate rate rises monotonically as +the base grows — **the idiom well for THIS class of function is approaching dry.** Five genuinely new +laws out of 197 claims is the signal to stop mining waves for idioms and spend the tokens on cracks. + +**THE SKEPTICS RAN THEIR OWN A/Bs THIS TIME** — several refutations are byte-measured, not +argued. The best example: a crack agent claimed "the SOURCE STATEMENT BOUNDARY decides whether the +scheduler hoists a far-consumed load". The vetter built that exact spelling and got `.text` +**byte-identical** to the inline form (12 mismatched either way) — then swept eight *positions* and +got five distinct objects, showing the lever is the statement's POSITION (the already-banked +`INSN_LUID` tie-break), not the boundary. A plausible mechanism, refuted by measurement, with the +true lever named in its place. + +*(NEW; evidence: byte-probed; from `func_8017EEEC`)* + +**§167-01 — A JUMP-TABLE ENTRY THAT LANDS IN THE *INTERIOR* OF ANOTHER ARM'S BODY IS A C CASE FALL-THROUGH — AND `match_one` CANNOT SEE THE DIFFERENCE.** *(NEW. §163c reads the entry values only for the `jtbl[k] == default label` fingerprint; §162a1/§162a2 rule the two EDGES; §165-42 rules an entry at `arm_start+4` (the reorg copy-steal) and says outright it is 'not a case'. This is the interior case, it IS a case, and it is the positive half §161a's corollary (L10989) never stated. It also adds a FIFTH row to §87's blindness ladder — §81's row is a duplicated table at the wrong ADDRESS; this is the right table at the right address with the WRONG WORDS.)* + +**Target shape** (`func_8017EEEC`, ov_SC06_011, 108 ins, reach 6): + + lhu $v1,0x34($s0) ; sltiu $v0,$v1,0x5 ; beqz $v0,.L8017F074 ; jr $v0 <- 5-entry dispatch + jtbl_801AA880 = [ 0x8017EFB0, 0x8017F074, 0x8017EFF0, 0x8017F074, 0x8017EF3C ] + ^interior ^epi ^arm ^epi ^FIRST body + 8017EF3C: la $s1,D_801202A0 … 0x60-iteration sweep … + 8017EFAC: addiu $s1,$s1,0x10C <- loop ends, NO `j` — falls through + 8017EFB0: lw $v0,%lo(D_80126CC8)($v0) <- table entry[0] lands HERE, 0x74 into the arm + 8017EFE8: j .L8017F074 <- *this* arm's `break` + +**THE LAW.** Every word in the `ADDR_VEC` is a case LABEL. `expand_end_case` emits one word per value in `[minval,maxval]` and the bodies in SOURCE order (§55a), so: +* the **lowest** body address in the table names the arm written **FIRST**; +* an entry that lands strictly between two other entries' addresses splits that block into **two arms joined by a `/* fall through */`** — the earlier-addressed case's body flows into the later one. Confirm it with the `j`: gcc emits a `j ` for every `break` **except** the last-emitted arm, so **an arm that ends without a `j` and is not the last arm is a fall-through arm.** +Read the shape above as `case 4: {sweep} /* fall through */ case 0: {if (D_80126CC8==a0) …} break;` — case 4 first, because 0x8017EF3C < 0x8017EFB0. + +**WHY IT IS INVISIBLE — the fifth blindness row.** Moving the sweep from `case 0` to a fall-through `case 4` changes **zero instructions**. Both spellings, pinned triple, masked vs `asm/ov_SC06_011/nonmatchings/ov_SC06_011_jr_8017BEBC/func_8017EEEC.s`: + +| source | ins | masked diffs | emitted `.rdata` | +|---|---|---|---| +| `case 0: {sweep; if(D_…)…} … case 4: default: break;` | **108** | **0** | `[$L3, $L2, $L12, $L2, $L2]` — entry[0]=sweep, entry[4]=default | +| `case 4: {sweep} /*fall through*/ case 0: {if(D_…)…} break;` | **108** | **0** | `[$L10, $L2, $L13, $L2, $L3]` — entry[0]=interior, entry[4]=sweep — **the target** | + +Two functionally different programs, one instruction stream. `match_one` masks relocations and never looks at `.rodata` at all; `rtu_match` masked-diffs the function out of the whole-TU object and is blind for the same reason. **Both report MATCH on the wrong one.** The error surfaces only after `jtbl_carve` puts the compiled table at the table's real address — as a 2-word whole-binary DIFF with a spotless per-function gate (§52b, with no residual to read). + +**THE DIAGNOSTIC TELL — three lines, before writing any C for a `jr` function.** +1. Take the table's DISTINCT non-epilogue entries and sort them by address. Lowest = the arm written first (§55a). +2. For each entry, ask where it lands: `== an arm's start` ⇒ ordinary arm; `== arm_start+4` ⇒ §165-42's stolen head insn, **not** a case; `== the epilogue` ⇒ empty case (§162a2/§162a1 at the edges); **strictly interior** ⇒ **case fall-through, and the enclosing arm's case value is the one written first.** +3. Cross-check with the `j`s: count arms ending in `j `. It must equal (number of arms with a `break`) − (1 if the last-emitted arm breaks). A missing `j` at an arm boundary that carries a table entry IS the fall-through. + +**BYTE EVIDENCE.** `func_8017EEEC` (ov_SC06_011, 108 ins). Target table `asm/ov_SC06_011/data/tail18.data.s:21-27`. Both C variants compiled on the pinned triple (`tools/bin/gcc-2.7.2-psx/cc1 -quiet -O2 -G0 -mips1 -mcpu=3000 -mgas -msoft-float -fgnu-linker`, maspsx `--aspsx-version=2.56 --expand-div`), `masked_diff.structured_diff` = 0 for both. The wave-5 draft `.run/wave6/func_8017EEEC/func_8017EEEC.c` (sha1 `baae6aba…`) is the first row and was blessed as a bankable MATCH by two independent agents and by `rtu_match`. *Scope: one function, both directions byte-measured; the mechanism (ADDR_VEC words are case labels, bodies in source order) is structural, not statistical.* + +*(NEW; evidence: byte-probed; from `func_801805D4`)* + +**§167-02 — `match_one` DECIDES WHETHER TO PREPEND `common.h` BY GREPPING THE RAW DRAFT TEXT, SO THE DIRECTIVE SPELLED INSIDE A COMMENT SUPPRESSES THE REAL ONE.** *(NEW. §40a (L2610) and §42a (L2960) both state that isolation compiles are `-Iinclude` + a prepended `common.h`; each describes the prepend as a fact of the environment and neither says it is CONDITIONAL, let alone on what. §96/§104 establish the class — a scan that reads RAW text counts prose as code — for `reconcile_tu` and `gather_externs`; this is that class inside the gate proxy every drafter runs, and §165-13 is its sibling on the other side of the same file.)* + +The check, verbatim (`tools/match_one.py:71-75`): + + src = masked_diff.strip_scalar_typedefs(src) + if '#include "common.h"' not in src: + src = '#include "common.h"\n' + src + +`strip_scalar_typedefs` → `cdecl.strip_provided_typedefs` drops typedef STATEMENTS only; it never masks comments, so comment text reaches the test intact. + +**THE LAW.** A draft that spells `#include "common.h"` anywhere — including inside a `/* */` header block — satisfies the test, the prepend is suppressed, and cc1 compiles a TU with no scalar types at all. The diagnostic is a `CC1 FAIL` in which `s32`/`s16`/`u32` **and every parameter and local** are "undeclared (first use this function)" — which reads as an `engine_types` / §163a declaration-conflict problem and has nothing whatever to do with the body. + +**BYTE EVIDENCE** (vet-time, pinned triple, `func_801805D4`, ov_SC07_006 / `jr_8017BEBC`). Shipped draft `.run/wave6/func_801805D4/func_801805D4.c`, which paraphrases the directive → **MATCH (212 ins)**. The SAME file with one comment word changed — `(via its common.h include ->` → `(via its #include "common.h" ->`, **no code touched** → **CC1 FAIL**: `t.c:131: parse error before 't'`, `:133: 's32' undeclared`, `:135: 't' undeclared`, `:136: 'u' undeclared`, `:136: 's16' undeclared`, `:136: 'arg1' undeclared`. + +**BLAST RADIUS — the same defect is on a second rung of the ladder.** `tools/reloc_verify.py:71` carries the identical `if '#include "common.h"' not in src:` inside `build_object()` (the §88f relocation gate), so a draft that trips `match_one` trips `reloc_verify` the same way and for the same reason. + +**DIAGNOSTIC TELL.** A `CC1 FAIL` in which the SCALAR TYPES THEMSELVES are undeclared is never a body problem and never a decl conflict — run `grep -n 'common\.h' ` before reading a single line of C. If the hit is inside a comment, paraphrase it ("via its common.h include"). Tool-side fix is §104's two-text discipline: run the test against `cdecl._mask`ed text, never the raw source. + +*(NEW; evidence: byte-probed; from `func_801832F8`)* + +**§167-03 — A `short / CONSTANT` DIVISION IS COMPUTED IN HImode, SO ITS QUOTIENT CARRIES AN EXTENSION — AND THE *DESTINATION'S* DECLARED WIDTH DECIDES WHERE THAT `sll 16 ; sra 16` LANDS.** *(§1-I3 gives only the authoring rule for the magic multiply and one ÷10 pair; §164-04 covers pure powers of two, where the shortening leaves no residue. §164-37/§165-29 own the same `get_unwidened`/promotion axis for a SWITCH operand only. Nothing in the file states the division case, the position rule, or the −2 lever.)* + +Target shape — a magic-multiply quotient joined to a running total, with a 16-bit round trip on ONE side of the add: + + sra $v0,$a3,3 + subu $v0,$v0,$v1 <- the quotient, NOT truncated + addu $v0,$s1,$v0 <- + pan + addu $s1,$v0,$zero + sll $v0,$v0,16 + sra $v0,$v0,16 <- the SUM is truncated + bgez $v0,... + +**THE LAW.** `c-typeck.c:2023-2033` (`build_binary_op`, the TRUNC_DIV_EXPR arm) sets `shorten = 1` whenever the divisor is an `INTEGER_CST != -1`; `c-typeck.c:2352-2356` then calls `get_narrower` on both operands. So `*(s16 *)p / 20` is evaluated in **HImode**, the magic-multiply quotient is a 16-bit value, and it must be sign-extended before it can join an SImode sum. **Where the `sll 16 ; sra 16` pair lands is decided by the width of the variable the sum is assigned to, not by the division:** + +| destination of `pan + field/K` | emitted | +|---|---| +| `s16` | quotient bare, extension rides the **SUM** (after the `addu`) — *the target* | +| `s32` | extension rides the **QUOTIENT** (before the `addu`), sum bare | + +**THE −2 LEVER, AND THE CAST THAT IS INERT.** Only a real `s32` local kills the shortening — `get_narrower` strips a cast but cannot strip a VAR_DECL: + + pan + out1.x / 20; -> 32 ins (shortened, quotient extended) + pan + (s32)out1.x / 20; -> 32 ins (IDENTICAL — the cast is inert) + s32 xv = out1.x; pan + xv / 20; -> 30 ins (no shortening at all) + +**Byte evidence.** All on the pinned triple. `func_801832F8` (ov_SC02_041), the agent's own isolators re-run: `.run/wave6/func_801832F8/dbg/isolate.c` (`s32 pan`, pinned) 35 ins with `sll/sra` on the quotient; `isolate2.c` (`s16 pan`, same pin) 35 ins with `sll/sra` on the sum; `isolate3.c` (`s32 pan`, unpinned) 32 ins, same shape as isolate — **the pin is not the variable.** Controls `.../vet832F8/cA,cB,cC,cD` give the table above; `cD` (`/16`) confirms the power-of-two path (§164-04) carries no residue because the `sra` already leaves the value in range. Target anchor: `asm/ov_SC02_041/nonmatchings/ov_SC02_041_jr_8017BEBC/func_801832F8.s` at 801833C4-801833D8 (`pan`) and 80183454-80183464 (`vol`, the same shape with `sll` alone for a sign test) — **both accumulators in this function are `s16`.** + +**THE DIAGNOSTIC TELL.** A magic-multiply quotient with a `sll 16 ; sra 16` pair on ONE side of the following `addu`/`subu`, count-neutral. **Read which side.** Extension on the SUM ⇒ the accumulator is a `short`; declare it `s16` and assign to it. Extension on the QUOTIENT ⇒ the accumulator is an `int`. Neither ⇒ the dividend reached the division through a named `s32` local. Do not reach for pins, `(s32)` casts or the permuter — the cast is provably inert and the width is one character of source. + +**Bound.** Constant divisors only (`shorten` needs `INTEGER_CST`); a *variable* divisor takes the `divu`+break path (§1-I4) and never shortens. Unsigned dividends shorten unconditionally (`TREE_UNSIGNED (orig_op0)`). + +*(NEW; evidence: byte-probed; from `func_801832F8`)* + +**§167-04 — A register-pinned struct base's field reads need explicit per-field temps to keep gcc's natural multi-register** + +**§16N+1 — A `register __asm__` PIN ON A STRUCT BASE SERIALISES THAT BASE'S OWN FIELD READS INTO ONE REGISTER; NAMED PER-FIELD TEMPS BUY THE PARALLEL FORM BACK. UNPINNED, THE TWO SPELLINGS ARE BYTE-IDENTICAL.** *(a sixth channel for §164-49's class — a pin whose damage is not the pinned register but the code AROUND it. §72/§74 price a pin as a preference and a call-crossing hazard; §165-47 prices singleton pins on one insn; §45-Lever-B/§76 own the REUSED-temp serialisation via `local-alloc.c:472`. None covers a pinned base's own field loads, and none names the cure.)* + +Target shape — three halfword fields copied off one base into a stack struct, all three loads ABOVE all three stores, in three different registers: + + lhu $v0,0x6($s3) + lhu $v1,0xA($s3) + lhu $a2,0xE($s3) + addu $a1,$s0,$zero + sh $v0,0x20($sp) + sh $v1,0x22($sp) + jal func_8012EFB8 + sh $a2,0x24($sp) + +**THE LAW.** With the base declared `register s32 base __asm__("$19")`, writing the copies inline — +`in2.x = *(u16 *)(base + 0x6);` ×3 — emits a strictly serial `lhu $v0 / sh / lhu $v0 / sh / lhu $v0 / sh` chain: each load's destination dies into its own store, so one scratch is reused and sched1 has nothing independent to interleave. Naming the three values first — + + { u16 fx = *(u16 *)(base + 0x6), fy = *(u16 *)(base + 0xA), fz = *(u16 *)(base + 0xE); + in2.x = fx; in2.y = fy; in2.z = fz; } + +gives three pseudos that are all live before the first store, forcing three registers and letting all three loads rise above all three stores — the target's shape. + +**⚠ THE LEVER IS CONDITIONAL ON THE PIN, AND IT IS NOT A LENGTH LEVER.** 2×2 on `func_801832F8` (ov_SC02_041), `match_one` vs `asm/ov_SC02_041/nonmatchings/ov_SC02_041_jr_8017BEBC`: + +| base | field reads | result | +|---|---|---| +| pinned `$19` | per-field temps | 111 ins, **40** mismatched | +| pinned `$19` | inline | 111 ins, **54** mismatched | +| unpinned | per-field temps | 109 ins, 101 mismatched | +| unpinned | inline | 109 ins, 101 mismatched — **byte-identical to the row above** | + +Unpinned, the edit is a no-op: spend it only when the base carries a pin. And **do not price it as instructions** — the pin costs +2 in both spellings and the temps recover none of it; they buy register identity only. (Files: `.../vet832F8/w6_{base,inlinefields,nopin_temps,nopin_inline}.c`.) + +**THE DIAGNOSTIC TELL.** The target reads N fields off one base into N DIFFERENT registers and stores them afterwards; your draft emits N `load ; store` pairs through one scratch, same instruction count — **and the base is pinned.** Give each field a named local. If the base is NOT pinned, this residual is somewhere else entirely; do not buy the edit. + +*Scope, honestly: one function, one base, three fields, four cells. The mechanism (dying-into-store scratch reuse vs three overlapping live ranges) is reconstructed from the emitted register pattern, not from a `-dl`/`-dg` join; the cheap confirmation nobody has run is `cc1 -da` on the pinned pair and a diff of `.lreg`.* + +*(NEW; evidence: byte-probed; from `func_801874C0`)* + +**§167-05 — QUALIFYING *BOTH* MEMs OF A DISAMBIGUATED PAIR `volatile` IS THE ONLY WAY TO PUT A DEPENDENCE EDGE BACK. A LONE `volatile` IS BYTE-INERT.** *(NEW. Completes §16Z's ADDRESS-CLASS TABLE, whose rows 1-3 record "**0** — not needed" in the escape column and therefore leave every disambiguated cell with no lever at all. §37, §16Xy, §136d-3, §162q and §165-46 are all edge-DROPPING levers; this is the file's first edge-CREATING one. §164-79/§165-44 name the two-base barrier only as a diagnostic. §16Z states the both-volatile conjunction only for `read_dependence`, i.e. load-vs-load.)* + +Target shape — a store and a later load through ONE base pseudo at two constant offsets, where the target keeps them in source order and every draft hoists the load: + + TARGET (both volatile) MINE (every non-volatile spelling) + sh $v1,0xA($s0) addiu $a0,$sp,0x10 + lhu $v0,0x6($s0) <- stays lhu $a1,0x6($s0) <- hoisted ABOVE the store + addiu $a0,$sp,0x10 … + sh $v1,0xA($s0) + +**THE LAW.** All three dependence predicates are one disjunction (`tools/reference/gcc-2.7.2/sched.c:817-838` `true_dependence`, `:844-864` `anti_dependence`, `:866-882` `output_dependence`): + + (MEM_VOLATILE_P (x) && MEM_VOLATILE_P (mem)) + || (memrefs_conflict_p (…) && ! ) + +`memrefs_conflict_p` (`:613`) is the first conjunct of the **second** term only. So on §16Z's rows 1-3 — two distinct symbols, `$sp+K` vs a symbol, and **the same base register at two constant offsets whose byte ranges are disjoint** — the second term is dead outright, and **no `/s` grant, no statement order, no register pin and no scope/reuse edit can create the edge**: the `/s` clause only ever *removes* one. The first term is the only remaining door, and it is a **CONJUNCTION**. `sched.c`'s own header comment says it (`:768-771`): *"If both memory references are volatile, then there must always be a dependence between the two references, since their order can not be changed. A volatile and non-volatile reference can be interchanged though."* + +**BYTE EVIDENCE — the full 2×2, re-verified at vet time from the objects on disk** (`func_801874C0`, ov_SC03_014, 241 ins; `.run/wave4/func_801874C0/w_*/func_801874C0/t.o`, pinned triple, every variant 241 ins): + +| spelling | `objdump` lines differing from the no-volatile baseline | +|---|---| +| neither (`w_v6`) | — baseline, 8 mismatched vs target | +| **store only** `*(volatile s16*)(a0+0xA) = …` (`w_pA1`) | **0 — byte-IDENTICAL** | +| **load only** `*(volatile u16*)(a0+0x6)` (`w_pA2`) | **0 — byte-IDENTICAL** | +| **both** (`w_pA`) | **16** — the `lhu 0x6` drops below the `sh 0xA`, the `addiu $a0,$sp,0x10` follows it: **MATCH 241/241** | + +**Two controls that pin the mechanism to the alias oracle and to nothing else.** +* **`w_pA4` — volatile on the store to `0xA` *and* on the load from `0xA` (the SAME address): 0 differing lines.** A pair `memrefs_conflict_p` already conflicts on gains nothing. The qualifier pays **only** where the oracle had disproved the conflict. +* **A lone `volatile` also defeats cse-hashing for that access, and it moves zero bytes** — which rules out a cse story and leaves the scheduler conjunction. + +**Widening is not free — qualify the ONE pair, not the object.** `w_pA3` (store + all three loads) and `w_pA5` (7 accesses) are byte-identical to the minimal pair; `w_pV3` (14) = 246 ins (**+5**); `w_pV1` (whole entity, 38) = **258 ins (+17)**. + +**THE DIAGNOSTIC TELL.** A load and a store through the same base pseudo at two constant offsets, the target holding them in source order, your draft hoisting the load, and **every** structural respelling returning the *identical* mismatch count — the flat plateau that means you are not steering the pass that decides. Before writing an "unsteerable" verdict, evaluate `memrefs_conflict_p` by hand for the pair against §16Z's table: **if it returns 0 you are not on a scheduling tie at all — you are missing an edge, and only the volatile pair can supply it.** + +**⚠ HONEST SCOPE.** `volatile` here is a compiler-steering construct, near-certainly not what the original author wrote; some non-volatile spelling yielding two *distinct base pseudos* would create the same edge (§164-79) but costs an address materialisation. 20 single-base variants — shared and per-field scratch temps (§162b), `s16*`/`u16*` pointer locals (combine folds them back), struct-vs-array buffers, a single `V8 v[3]`, operand re-association, statement order, arg-address hoisting, pointer-typed prototypes, a copy of the parameter — all sat at exactly 8 mismatched. Unlike a `register __asm__` pin (§165-47/§37), a `volatile` cast is ordinary C and is not special-cased anywhere in `tools/dedup_propagate.py`, so it should not forfeit the h_seq family — untested on a real sweep. + +**⚠ TWO LIST CORRECTIONS.** (1) This law is the *explanation* of the §164z `func_8017CA18` entry's observation half ("`volatile` applied to the RMW side is a no-op") — one-sided `volatile` is **supposed** to be inert. (2) The §165z `func_801874C0` entry ("the residual is unreachable by respelling") is **byte-superseded**: the mechanism half was right and already banked as §164-79, but the *unreachable* verdict is now refuted by the `w_pA` object. + +*(NEW; evidence: byte-probed — 2×2 ablation plus two controls, re-verified at vet time from the on-disk objects; mechanism source-cited to the pinned `tools/reference/gcc-2.7.2/sched.c`; from `func_801874C0`)* + +*(SHARPENS — sharpens §165-03 (L13716-13740), §163e (L11813-11821), §164-72 (L13412), §164-53; evidence: byte-probed; from `func_8017C294`)* + +**§167-06 — EVERY RELOAD SPILL SLOT IS 8 BYTES, IN ANY MODE, BECAUSE `alter_reg` PASSES `align = -1`.** *(Derives the constant §165-03 measures and builds its frame formula on; BOUNDS §163e's 'a scalar `s32` gets only 4-byte alignment' to DECLARED locals — carrying that sentence across to a SPILL mis-prices the frame by 4 per slot.)* + +**Target shape.** `vars=` exceeds the sum of your declared locals by a multiple of 8, none of the excess is `$sp`-addressed, and you are trying to price a spilled `s16`/`s32`. + +**THE LAW.** `reload1.c:657-658` walks `for (i = LAST_VIRTUAL_REGISTER+1; i < max_regno; i++) alter_reg (i, -1);` — ascending pseudo NUMBER (§163e), `from_reg == -1` — and `alter_reg` allocates with `x = assign_stack_local (GET_MODE (regno_reg_rtx[i]), total_size, -1)` (`reload1.c:2349`; the slot-reuse path at `:2382` passes `-1` too). `assign_stack_local`'s `align == -1` arm (`function.c:681-685`) sets `alignment = BIGGEST_ALIGNMENT / BITS_PER_UNIT` and, **uniquely among its three arms, rounds the size**: `size = CEIL_ROUND (size, alignment)`. MIPS sets `BIGGEST_ALIGNMENT 64` (`config/mips/mips.h:1080`). ⇒ **a spilled HImode pseudo, a spilled SImode pseudo and a combine-orphan all cost exactly 8 bytes, 8-aligned. The mode never reaches the frame.** Only DECLARED locals see their own alignment — `assign_stack_temp` takes the `align == 0` arm (`function.c:675-680`), which is the case §163e measured. + +**Corollary — the frame is a two-ended pseudo-BIRTH oracle.** Ascending regno + `expand` minting in source order ⇒ the LOWEST spill slot belongs to the earliest-born pseudo (in an arg-spilling function, the incoming-argument pseudo) and the HIGHEST to whatever pass minted last; after expand only `loop.c` mints (§147-CORRECTED A). *A value the target spills at the TOP of the block therefore cannot be a declared local at all.* + +**BYTE EVIDENCE** (`.run/wave6/func_8017C294/`, pinned `cc1 -O2 -G0 -mips1 -mcpu=3000`; two drafts, one formula): +- `dmp_fas/`: declared aggregates `32+32+32+8+8 = 112` (every scalar is a register candidate and contributes 0), `grep -c 'ST_REGS or none' t.i.lreg` = **16**, `t.cc1.s` = `.frame $sp,312,$31 # vars= 256` = `112 + 8×(16 orphans + 2 real spills)`, exact. +- `dmp_base/` (same body plus `s32 pad[4]`): declared **128**, orphans **14**, same `vars= 256` = `128 + 8×(14+2)`. Different split, same arithmetic. + +**DIAGNOSTIC TELL.** **Never price a spill by its C type.** Count orphans with `grep -c 'ST_REGS or none' t.i.lreg` (§165-03) and real spills from `.greg`'s `regs to allocate` minus the placed ones, then multiply **both** by 8 — `vars = Σ(declared locals) + 8×(orphans + real spills)`. + +*(SHARPENS — sharpens §165-03 (L13738, the flagged ⚠ open detail), §165-02 (L13914), §147-B CORRECTED (L10117); evidence: byte-probed; from `func_8017C294`)* + +**§167-07 — ONLY A DELETION THAT HAPPENS *AFTER* `life_analysis` CAN ORPHAN A PSEUDO — AND AN ORPHAN LEAVES *TWO* DIFFERENT RESIDUES IN `-dc`.** *(Closes §165-03's flagged ⚠ open detail and corrects its causal order; supplies the necessary condition behind §165-02's 'must come from MEMORY'; reconciles §147-B-corrected's `(use)` account with §165-03's 'present in no insn'.)* + +**THE PASS ORDER, from `toplev.c`.** cse (`:2865`) → loop → cse2 (`:2926`) → **`flow_analysis` / `life_analysis` (`:2983`)** → **combine (`:3004`)** → sched1 (`:3033`) → `regclass` + `local_alloc` (`:3051-3052`) → `global_alloc` (`:3080`) → reload. `reg_n_refs` is filled by `life_analysis` and is **never recomputed** before `alter_reg` reads it (`reload1.c:2330`, gated `reg_n_refs[i] > 0`). + +**THE LAW.** A pseudo whose insns die in **cse / loop / cse2 / jump** is recounted to **zero** by the later flow pass and can never take a slot. A pseudo whose insns die in **combine** keeps its pre-combine count forever and is slotted. **The recipe's real precondition is not 'memory' — memory is what makes the deletion a COMBINE deletion** (there is a load for the widening to fold into). A redundant copy, an alias of a parameter, or a register-sourced short is killed by cse, *upstream of the count*, which is why no amount of retyping buys frame from it. + +**TWO RESIDUES, ONE OUTCOME.** §165-03 says the orphan 'appears in no insn in `-dc`'s combine dump'; §147-B-corrected says combine leaves `(use (reg)) + REG_DEAD`. **Both forms exist in one function** (`.run/wave6/func_8017C294/dmp_fas/`, 16 orphans): +- **`(use)` form, 14 of 16** — `distribute_notes` has a `REG_DEAD` note with no home and plants a bare `(use)` (`combine.c:10835-10847`, §36). `t.i.combine:396` = `(insn 719 191 193 (use (reg:SI 147)) -1 (nil)` with `(expr_list:REG_DEAD (reg:SI 147)` on the next line; `t.i.lreg` = `Register 147 used 4 times across 1 insns in block 4; ST_REGS or none.` +- **Ghost form, 2 of 16** — nothing survives at all. Pseudos 103/104 (the head `ws`/`hs`) appear as `(set (reg:SI 103) (ashift …))` + `(ashiftrt (reg:SI 103) 16)` in `t.i.cse2:200-206` and `t.i.jump:65-71`, appear **nowhere** in `t.i.combine`, and still print `Register 103 used 2 times across 2 insns in block 0; dies in 0 places; ST_REGS or none.` — a count for two insns that no longer exist. + +⇒ **the `(use)` is not what keeps the count alive; flow banked the count before combine ran.** The `(use)` matters only in that it carries no constraints, so `regclass` (after combine) records no class for either form, the all-zero cost vector converges on `ST_REGS`, and §165-03's allocation story runs unchanged. + +**BYTE EVIDENCE.** Five dumps under `.run/wave6/func_8017C294/`: `grep -c 'ST_REGS or none' t.i.lreg` = 14 / 16 / 16 / 14 / 16 (`dmp_base`, `dmp_fas`, `dmp_m2`, `dmp_s9`, `dmp_zbss`) against `grep -c '(use (reg' t.i.combine` = 28 / 30 / 30 / 28 / 30 — a **constant offset of 14** (the call-argument and return `(use)`s) ⇒ **+1 orphan ⇔ +1 planted `(use)`**, five bodies, one function. + +**DIAGNOSTIC TELL.** You added a narrow memory value and `vars` did not move: find out **which pass ate it**, not which type it had. `grep -n 'reg:SI N' t.i.cse2 t.i.combine t.i.lreg` — absent from `t.i.cse2` onward ⇒ cse killed it, the count went with it, no slot exists at any spelling; present in `t.i.cse2`, absent from `t.i.combine`, still printed in `t.i.lreg` ⇒ ghost form, the slot is already yours. + +*(SHARPENS — sharpens L2266-2270 (§ giant-crack recipe item 2) — 'Declare each callee to match the ACTUAL call site — count the `$a0–$a3` (+ stack) set before each `jal`, NOT the canonical sig; evidence: byte-probed; from `func_8017EEEC`)* + +**§167-08 — AN ARGUMENT REGISTER THAT IS *READ* BEFORE THE `jal` IS SCRATCH, NOT AN ARGUMENT — COUNT THE DEFS, NOT THE MENTIONS.** *(sharpens the giant-crack recipe's arity rule at L2267, 'count the `$a0–$a3` set before each `jal`', which is right about DEFS and is routinely misread as 'count the $aN mentions'; the complement of §164z's `func_80186C4C` refutation, which killed the inverse inference — a `nop` delay slot does NOT prove a 0-arg callee.)* + +**Target shape** (`func_8017EEEC` @8017EFC4, ov_SC06_011) — four `$a`-register mentions around one call, and the callee takes nothing: + + lhu $v1,0x88($s0) ; lhu $a0,0x8A($s0) ; lhu $a1,0x8C($s0) + sh $v0,0x34($s0) ; sh $zero,0x5C($s0) ; sh $v1,0x6($s0) + sh $a0,0xA($s0) <- $a0 READ as a store SOURCE, before the jal + jal func_8017F364 + sh $a1,0xE($s0) <- $a1 READ in the DELAY SLOT + +**THE LAW.** `$a0–$a3` are ordinary caller-saved registers; local-alloc hands them to any pseudo whose live range does not cross the call. `expand_call` materialises real arguments as the **last defs before the `jal`** (a `move`/`lw`/`li` into `$aN` with no intervening use of that register as a source). So: **an `$aN` whose last event before the `jal` is a READ — it is a store's value operand, a compare operand, an address — died before the call and was never an argument.** A store in the delay slot reads its operand *before* the callee runs, so it is not evidence either way on its own; the def/use direction is. + +**BYTE EVIDENCE — the counterfactual costs two `move`s and permutes the file.** Same body, only the callee's arity changed, pinned triple: + +| spelling | ins | masked diffs | +|---|---|---| +| `extern void func_8017F364(void); … func_8017F364();` | **108** | **0 — MATCH** | +| `extern void func_8017F364(u16,u16); … func_8017F364(t1,t2);` | 110 | **53**, LENGTH-DRIFT | + +The 2-arg build is unmistakable: gcc allocates the stored halfwords to `$a2`/`$v1` (`lhu $a2,138($s0) ; lhu $v1,140($s0)`) and emits `move $a0,$a2 ; move $a1,$v1` immediately before the `jal` — **it will not source a store from a register it is about to load an argument into.** The target has no such `move` pair, so the call is 0-arg. Corroborated independently by §166a's oracle: the destination TU `src/ov_SC06_011/ov_SC06_011_jr_8017BEBC.c` defines `void func_8017F364(void)` itself, ~60 lines below the splice point. + +**THE DIAGNOSTIC TELL.** Before declaring a callee, walk backwards from the `jal` and mark each `$aN`: **DEF with no later use before the call ⇒ argument. USE (source operand of a store/compare/ALU op) ⇒ scratch.** Then stop at the first `$aN` that is neither — argument registers are filled contiguously from `$a0`. If the TU (or a sibling TU) defines the callee, that definition outranks the register read (§166a). + +*(SHARPENS — sharpens §164-12, §136d-2, §164-57, §164-73; evidence: byte-probed; from `func_8017F17C`)* + +**§167-09 — A `?:` IN AN `if`'s CONTROLLING EXPRESSION IS EXPANDED ONE CONDITIONAL BRANCH PER ARM. HOIST IT ANYWHERE — TERNARY-INTO-A-LOCAL, `if/else`, OR CONDITIONAL OVERWRITE; ALL THREE ARE BYTE-IDENTICAL.** *(sharpens §164-12, which reaches the SAME `do_jump` COND_EXPR case only via `fold-const.c:3276-3333`'s distribution over an ENCLOSING comparison and is stated for a `?:` used as a comparison OPERAND — the `?:` as the condition itself needs no fold and is not covered. Distinct from §136d-2 / §164-57 / §164-73 / §164-74 / §165-21, every one of which is about the select's VALUE DESTINATION — reg, MEM, call — and none about it being the test.)* + +Target shape — TWO branches, the one-instruction arm living entirely in the first branch's delay slot, and no `j` to a join: + + andi $v0,$v1,0x2000 + bnez $v0,.L8017F330 + andi $v0,$v1,0x8000 <- the whole ELSE arm, in the delay slot + jal func_8012BEE8 + addu $a0,$s0,$zero <- the THEN arm, falling through + .L8017F330: beqz $v0,.L8017F354 <- ONE consumer test + +**THE LAW.** `if (c ? A : B)` reaches `do_jump`'s COND_EXPR case (`expr.c:9124-9151`) directly: each arm gets its own `do_jump` aimed at the CONSUMER's true/false labels, so you pay **one conditional branch per arm plus the drop-through** — three conditional branches where the target has two, **+2 ins**. Take the select out of the controlling position and give it a name; both arms then expand as VALUES into one pseudo and a single test sits below the join. + +**The dial is POSITION, not spelling.** Pinned cc1 (`-quiet -O2 -G0 -mips1 -mcpu=3000 -mgas -msoft-float`), one function, `r` from a preceding call so the prologue is already emitted: + +| spelling | ins | conditional branches | +|---|---:|---:| +| `if ((r & 0x2000) ? (r & 0x8000) : h(a0))` | **20** | **3** | +| `c = (r & 0x2000) ? (r & 0x8000) : h(a0); if (c)` | 18 | 2 | +| `if ((r&0x2000)==0) c = h(a0); else c = r & 0x8000; if (c)` | 18 | 2 | +| `c = r & 0x8000; if ((r&0x2000)==0) c = h(a0); if (c)` | 18 | 2 | + +The last three are byte-identical up to label numbering. **The +2 is the isolated figure**; the crack note reported +1 from an in-function count — §165-22's precedent, do not use the in-function number as the tell. + +**⚠ DO NOT RE-BUY (byte-refuted at vet time).** *"The CALL must be the FALLTHROUGH arm — invert the source condition to choose."* Putting the call in the `else` arm gives the **same 18 instructions and the same 2 branches**, differing only in label numbers: gcc inverts the test for you. (Arm order does move one byte in the degenerate shape with NO preceding call — 16 vs 17 — but that is the prologue's own `sw $31` competing for the slot, not the select.) Likewise do not reach for §2-T4's polarity invert; the polarity is a consequence of the hoist, not a dial. + +**BYTE EVIDENCE.** `func_8017F17C` (ov_SC02_026, 133 ins, banked at `src/ov_SC02_026/ov_SC02_026_jr_8017C180.c:3723`) uses the if/else-into-`c` form; the isolated 5-spelling A/B above reproduces the target's exact branch/delay-slot shape from the hoisted form and only from it. + +**DIAGNOSTIC TELL.** THREE conditional branches where the target has two, **+2 ins**, and your extra branch tests a value one arm has just computed. Hunt for a `?:` sitting in an `if`'s controlling expression and bind it to a local — any binding will do. Reading the target from the other side: a branch whose **delay slot holds a complete one-instruction arm**, whose fall-through is the other arm, with **no `j` to a join** and a single test after the label, is a two-armed select assigned to a temp — not a threaded condition and not two independent tests. + +*(SHARPENS — sharpens §164-20, §160d, §165-02, §16Xy; evidence: byte-probed; from `func_8017F17C`)* + +**§167-10 — A SIGNED-NARROW LVALUE COMPARED AND THEN RE-READ COSTS ONE `move` **AND** 8 BYTES OF FRAME — ONE EDIT, TWO SYMPTOMS.** *(extends §164-20's second-read law off the DELAY SLOT — §164-20's three-way tell is keyed on the slot's owner and has no route for a copy that sits between a load and a compare; gives that law a SIGNEDNESS gate; and supplies §165-02 / §16Xy the construct row they lack, joining the copy to the orphan slot. BOUNDS §162i1's "Δ multiple of 8 ⇒ declare `s32 pad[Δ/4]`".)* + +Target shape — no delay slot in sight; the copy sits between the load and a compare that clobbers the load's register: + + lh $v0,0x102($s0) + addu $a0,$v0,$zero <- the copy IS the second source read + slti $v0,$v0,0x20 <- clobbers the load's register + ... + addiu $v0,$a0,1 + sh $v0,0x102($s0) + +**THE LAW.** Reading ONE signed narrow (HI/QI) memory lvalue in a guard **and again inside the guarded arm** mints two things at once. cse rewrites the second load as `(set P2 P1)` and it survives as a bare `move` (§164-20's mechanism, off the delay slot). And the SImode extension `combine` cannot fold — a narrow use, the `sh` store-back, survives — is stranded as an orphan pseudo that `alter_reg` hands **exactly 8 bytes** of `vars` no instruction ever references (§165-02's rule, §16Xy's `.greg` mechanism). **Hoisting the read into one `s32` local deletes BOTH. Never diagnose them as two residuals.** + +**The gate is the load opcode.** Isolated on the pinned cc1 (`-quiet -O2 -G0 -mips1 -mcpu=3000 -mgas -msoft-float`), one 4-line function, `if (*(T*)(p+0x102) < 0x20) { *(T*)(p+0x102) = *(T*)(p+0x102) + 1; g(p); }`: + +| T | emitted | vars | ins | +|---|---|---:|---:| +| `short` | `lh` + **`move`** + `slt` | **8** | 12 | +| `signed char` | same shape | **8** | 12 | +| `unsigned short` | `lhu` + `sltu`, **no copy** | 0 | 11 | +| `int` | **no copy** (cse collapses it) | 0 | 11 | +| `short`, read hoisted into `int n` | **no copy** | 0 | 11 | +| `unsigned short` + a `(short)` cast on the compare | `lh` + **`move`** | **8** | 12 | + +The last row is the discriminator: the **signed compare** is the trigger, not the pointer's spelling. + +**Bounds, each byte-measured.** An `(s32)` cast on the second read buys nothing (still 8/12 — §21, gcc re-derives). The guard must compare the SAME lvalue the arm re-reads: compare `x` / store from `y` ⇒ 0; two reads passed as call ARGS with no store-back ⇒ 0 (cse commons them); the read-modify-write with no compare at all ⇒ 0. Ordered vs equality compare is irrelevant (`!= 0x20` ⇒ 8/12). A call in the arm is not required. Two independent sites ⇒ `vars= 16`, additive (§16Xy says the same of its own construct). + +**BYTE EVIDENCE.** `func_8017F17C` (ov_SC02_026, 133 ins, banked at `src/ov_SC02_026/ov_SC02_026_jr_8017C180.c:3723`). Real `cpp | cc1` pipeline: banked body `.frame $sp,48 # vars= 16, regs= 4/0, args= 16`, 118 cc1-insns. Hoist the `0x102` read into `s32 n` ⇒ `vars= 8`, 117 insns — one `move` and 8 bytes of frame from one character of source. Neutralise the guarded block entirely (`if (0)`) and `vars` drops 16 → 8, locating the 8 bytes. + +**⚠ DO NOT RE-BUY — the 8-byte hole above an align-1 local is NOT the struct's slot being rounded.** Measured, 7 spellings: an 8-byte `struct { u8 c[8]; }`, address-taken and assigned from a global, reserves `vars= 8` — identical to `char s[8]`, `int s[2]`, `struct {int a,b;}`; only a 16-byte struct reserves 16. And "the hole absorbed my `s32 pad[1]`" proves nothing: §162i1 byte-proves a one-element array is INERT **everywhere**. + +**DIAGNOSTIC TELL.** A bare `move` / `addu rD,rS,$zero` between a narrow LOAD and a compare that CLOBBERS the load's register — no delay slot involved, so §164-20's three-way tell never fires — is a SECOND SOURCE READ, not an allocator artifact. **Check the frame in the same breath:** if you are also 8 bytes light, it is the same edit, and §162i1's "Δ multiple of 8 ⇒ declare `s32 pad[Δ/4]`" will send you to invent a dead local you do not need (measured here: `s32 pad[2]` ⇒ 0x38, FAIL). Then route by the second read's opcode: a second `lw`/`lbu` ⇒ §160d (keep the re-read, the load survives); `addu …,$zero` in a delay slot ⇒ §164-20; `addu …,$zero` before a clobbering compare ⇒ here. + +*(SHARPENS — sharpens §166a, §165-36 rule 2, §162g; evidence: byte-probed; from `func_8017F2D4`)* + +**§167-11 — THE DESTINATION-TU ORACLE, STATED EXACTLY: the LAST component of the `INCLUDE_ASM` subdir is the TU stem, and the stem ALONE resolves the file.** + +§166a says *"the third path component **is** the TU stem"* and draws one shape. Both need fixing before an agent can apply it mechanically. + +**THE TWO SHAPES** (they differ by the presence of an overlay component, not by the rule): + + asm//nonmatchings//.s -> src//.c (398 of 476 live TU dirs) + asm/nonmatchings//.s -> src/.c ( 78 of 476 — the main exe) + +**THE LAW, ordinal-free.** The TU stem is the **last component of the `INCLUDE_ASM` subdir** = the **parent directory of the `.s` file**. Never count from the left: in the overlay form the third component is `nonmatchings`, and §166a's `src//.c` template is wrong for every main-executable stub (`asm/nonmatchings/libgte21/…` is `src/libgte21.c`, not `src/nonmatchings/libgte21.c`). + +**WHY IT IS STRUCTURAL, not a convention that could drift.** A splat `c` subsegment is a SINGLE token: `- [0x56c04, c, ov_SC01_005_jr_8017ED5C]` (`config/splat.ov_SC01_005.yaml:125`). That one token generates both `/.c` and `/nonmatchings//`. The paths cannot disagree unless the config is edited between the two. + +**BYTE EVIDENCE (whole tree, current HEAD).** +* **13,630 / 13,630** `INCLUDE_ASM` rows in `src/`: subdir's last component == the containing file's stem. Zero exceptions. +* **13,630 / 13,630**: the overlay component (when present) == the containing directory; the 2,002 rows with no overlay component all live in `src/` root. +* **476 / 476** `asm/**/nonmatchings//` dirs holding `.s` resolve to an existing `src` file. The only non-conforming `asm` dirs are `*/data/` (rodata carves) and `asm/header.s` — neither holds a function stub. +* **4,153 / 4,153** `src/**/*.c` stems are unique repo-wide ⇒ **the stem alone is a sufficient key.** You do not need the overlay component to resolve the TU, which is what makes the ordinal-free phrasing safe across both shapes. + +**⚠ AUTHORITATIVE ≠ DURABLE (see §165-36 rule 2).** The oracle binds the **current** `asm/` tree to the **current** `src/` tree. A subdir recorded in a note, a backlog row or a `.run/*_ready.json` is a *stale* path, not an oracle — splat moved `func_8017F2D4` from `…_jr_8017C340` to `…_jr_8017ED5C` and `.run/jr48/wave1_ready.json` still carries the old one. Re-glob `asm/**/nonmatchings/*/.s` before deriving anything. + +**THE DIAGNOSTIC TELL.** You are about to splice, and the path in the prose has a *different* stem from the `--asm-subdir` you were handed ⇒ the prose is wrong, every time (a name grep hits callers and prototypes in sibling TUs and reads identical to a destination hit). One-line check, no build: `ls src/$(dirname )`— or simply `find src -name "$(basename ).c"`, which is unambiguous because stems are unique. + +*(SHARPENS — sharpens §164-62, §80 (R7), §158, §36 (KEEPALIVE KILLS THE DYING-HARD-REG SUGGESTION, L2498); evidence: byte-probed; from `func_8017F3C8`)* + +**§167-12 — THE KEEPALIVE'S ANCHOR CAN BE `volatile` INSTEAD OF A SECOND OPERAND — and §36's bare single-input form is inert only because it is NON-volatile.** *(SHARPENS §164-62, which supplies exactly one anchor — "name a value whose def is at or below the consuming insn" — and never names this one, though its own mechanism paragraph contains the fact that makes it work; widens §80-R7, stated only at behemoth/pin scale; and identifies §158's range-extender as the same construct seen from the `allocno_compare` side rather than the `qty_phys_sugg` side.)* + +Target shape — §164-62/§80's in-place-reuse tell on an ordinary leaf, the value held in a callee-saved pin: + + mine: sll $s1,$s1,16 <- destructive, in place; $s1 "dies" at this insn + target: sll $v1,$s1,16 <- preserving copy into a fresh scratch + +**THE LAW.** `local-alloc.c:1795-1834` records the dying hard reg in `qty_phys_sugg` unconditionally, with no death guard (§80), and `find_free_reg` (`:2150`) honours the suggestion only while that reg is free across the quantity's range — i.e. only while it dies at that insn. §80-R7's cure (keep it alive past the temp) is right, and §164-62 is right that a NON-volatile input-only `asm` is an ordinary schedulable insn carrying only its operands' deps, so written after the consuming insn while naming only the dying source it floats ABOVE and anchors nothing. §164-62 fixes that with a second operand. **`__asm__ __volatile__` is the other fix and needs only the one operand:** `stmt.c:1502` sets `MEM_VOLATILE_P` from the keyword and `sched.c:1957` then takes the clobber-everything barrier path, so the asm cannot float above the insn it was written after, the death moves onto it, and the suggestion never fires. Materialise the comparison into a named boolean first so there is a statement boundary to sit on: + + s16 cmp = sVar2 > sVar3; + __asm__ __volatile__("" :: "r"(sVar2)); /* the `volatile` is the whole anchor */ + if (cmp) { … } + +**BYTE EVIDENCE** — `func_8017F3C8` (ov_SC06_011, 66 ins, **NEAR 5/66, not banked**). Eight A/B pairs on one base, one variable each: dropping `__volatile__`; placing the keepalive before the compare; inside either arm after the `if`; and wrapped as a GNU statement-expression evaluated before the compare — **all regress to the in-place `sll $s1,$s1,16` form at 11-13 mismatches**, against 5 for the form above. Both the keyword and the position are load-bearing, independently. + +**DIAGNOSTIC TELL.** §164-62's tell unchanged (one-register diff, target's dest a fresh `$v0`/`$v1`, yours the operand's own register, pseudo NUMBER identical across every `-da` dump and only its COLOUR differing). **Before reaching for §164-62's second operand, check whether you wrote `volatile`.** Free confirmation either way: `grep -n '#APP' t.s` — if the block sits ABOVE the insn you wrote it after, the asm is non-volatile and inert; adding `volatile` pins it there at zero bytes. Prefer the volatile form when there is no value defined at or below the consuming insn to name; prefer §164-62's second operand when you are inside a region where a new barrier would move a schedule (§47's rule). + +*Scope: one function, eight controlled pairs, measured as mismatch counts against the target rather than a gated MATCH — the function was still 5-off when the probes were taken. Re-confirm on a banked exemplar before treating the position rule as exact.* + +*(SHARPENS — sharpens L2465 (§34 'Statement-position / type levers': "a leading `pb = param_3;` rides sched.c:3191-3215's 'don't delay getting parameters' pin so combine folds the parm-save in; evidence: byte-probed; from `func_8017F76C`)* + +**§167-13 — A LEADING `p1 = a1;` DOES NOT *RIDE* THE bb0 PARAMETER PIN — IT **TERMINATES** IT, AND ONLY FOR THE COPIES AFTER IT. The dose is "break the run as EARLY as the target needs".** *(SHARPENS L2465 (§34's statement-position lever list), whose one line — "a leading `pb = param_3;` rides `sched.c:3191-3215`'s 'don't delay getting parameters' pin so combine folds the parm-save into it (prologue order)" — names the pin but has the direction, the pass and the dose wrong for this shape; and BOUNDS `docs/gcc-2.7.2-map/sched.md` S8, "anything before the first non-param-copy insn is immovable — don't fight it".)* + +Target shape — the call's argument setup WOVEN INTO the prologue save/copy pairs: + + target: sw $s3 / move $s3,$a0 / li $a0,0x3B / sw $s6 / move $s6,$a1 / + move $a1,$s3 / sw $s5 / move $s5,$a2 / sw $s4 / move $s4,$a3 / jal + mine: sw $s3 / move $s3,$a0 / sw $s6 / move $s6,$a1 / sw $s5 / + move $s5,$a2 / sw $s4 / move $s4,$a3 / li $a0,0x3B / move $a1,$s3 / jal + +Same instruction count, same register letters, same everything else — the arg setup just sits BELOW every pair. + +**THE LAW.** `schedule_block` (`tools/reference/gcc-2.7.2/sched.c:3186-3213`, guarded `reload_completed == 0 && b == 0`) walks the head of bb0 and sets `INSN_REF_COUNT = 1` — *"Keep this insn from ever being scheduled"* — on the **leading run** of `(set pseudo hardreg)` parameter copies. The loop demands `GET_CODE (head) == INSN`, so **any NOTE inside the run ends it**, and a deleted insn is a NOTE. The prologue saves are threaded after global-alloc (`toplev.c:3103`) and weave at **sched2**, where every candidate here ties at priority 1 and at class 3 against the last-scheduled `sw $ra`, so `rank_for_schedule` falls all the way through to `INSN_LUID (tmp) - INSN_LUID (tmp2)` = stream order. Unlevered, all four parm copies are pinned, permanently out-LUID the arg setup, and win every tie — **no statement order anywhere in the body reaches this.** + +**THE LEVER.** A leading `s32 p1 = a1;` makes cse DELETE the original a1 parm copy (insn 6 → NOTE) and rewrite the later `p1 = a1` to read `(reg:SI 5 a1)` directly, so it survives as a HIGHER-UID insn further down the stream. That NOTE terminates the pin run **after the a0 copy alone**; a1/a2/a3 become schedulable; sched1's `adjust_priority` birthing boost (`birthing_insn_p`, `REG_N_SETS == 1`) lifts those single-set copies to `0x7f000001` while the arg setups stay at 1; sched1 is BACKWARD, so boosted = picked first = emitted LAST. Post-sched1 the stream is `4 (s3=a0) / 19 ($a0=0x3B) / 16 (s6=a1) / 21 ($a1=s3) / 8 (s5=a2) / 10 (s4=a3) / call` = the target's bb0 verbatim, and because sched2's LUIDs are sched1's output order (sched.md S9) there is nothing left to undo. + +**THE DOSING RULE (the operational half, and not what you would guess).** A leading copy frees only the parm copies **after** it. A/B, same file, one line each, pinned triple: + +| leading copy | residual | +|---|---| +| *(none)* | 8 mismatched | +| `p3 = a3` | 8 — frees nothing | +| `p2 = a2` | 7 | +| `p1 = a1` | **MATCH 154/154** | +| all three | MATCH | + +So the rule is **"break the pin run as EARLY as the target needs"**, not "launder the parameter you happen to use". Also byte-inert here: retyping parameter 1 to `u8 *` (8). + +**DIAGNOSTIC TELL.** A residual that is a pure SCHEDULE-REORDER **confined to bb0** — instruction count exact, every register letter already correct — with the call's argument setup (`li $aN,K` / `move $aN,$sN`) sitting BELOW all the `sw $sN`/`move $sN,$aN` prologue pairs the target weaves them into. Do **not** reach for a `register __asm__` pin (§36: pins wreck the save-birthing order) or §67's asm launder: this is a plain C statement, costs zero instructions, and keeps `compiles_standalone` so the draft still propagates to its family (§37). + +**⚠ BOUND ON L2465, recorded because the one-liner mispredicts this shape.** The levered copy does **not** ride the pin in prologue order — in the `cc1 -dS` stream it lands BELOW the arg setup — and the deletion is **cse**'s, not combine's. Read L2465 as naming the pin, not as describing the mechanism. + +**BYTE EVIDENCE.** `func_8017F76C` (ov_SC02_026, 154 ins, MATCH, clean standalone compile, no pins, no permuter; draft + dumps at `.run/wave4/func_8017F76C/`, header L38-L92). The source half was re-verified at vetting time against `tools/reference/gcc-2.7.2/sched.c:3186-3213`. Second natural instance: the §160g STEP-0 sibling `func_80181F88` (`src/ov_SC03_098/ov_SC03_098_jr_8017D898.c:5153`) carries the same odd-looking leading copy as its first line — **reading a sibling's weird first line beat every scheduler analysis on this function.** + +*(SHARPENS — sharpens §164-24 / §16x 'THE PLUS/MINUS CONSTANT MIRROR' (L12384-12405), incl. "A MINUS whose constant is FIRST comes back through the `varsign == -1` code flip (:3766) to the sam; evidence: byte-probed; from `func_8017F76C`)* + +**§167-14 — A TWO-TERM `CON - VAR` IS NOT THE SAME TREE AS `-VAR + CON`: `fold`'s `associate:` PATH CANNOT FIRE ON A BARE CONSTANT OPERAND, SO THE SPELLING SURVIVES TO RTL AND PICKS THE OPCODES.** *(BOUNDS §164-24 / "THE PLUS/MINUS CONSTANT MIRROR" (L12384-12405), whose "A MINUS whose constant is FIRST comes back through the `varsign == -1` code flip (`:3766`) to the same tree as the PLUS form" is byte-true for its THREE-term exemplars (`a -= K - J` ≡ `a += J - K`) and byte-FALSE for the two-term form; and bounds §164-16's "parentheses, operand order and same-mode casts are all inert", which holds only once `associate:` can run at all.)* + +Target shape — a negate feeding an add-immediate, with the negate in the branch delay slot: + + bnez $v0,L + negu $v0,$s0 <- head of the FALL-THROUGH arm, taken into the slot + addiu $v0,$v0,0x400 + j ... + L: addiu $v0,$s0,0x400 + +**THE LAW.** `-base + 0x400` is `PLUS (NEG (base), 1024)` → `subu $2,$0,$4 ; addu $2,$2,1024` (the `negu`/`addiu` pair). `0x400 - base` stays `MINUS (1024, base)`: `split_tree` (`fold-const.c:882-950`) fires only on a PLUS/MINUS node with a decomposable operand, and a bare `INTEGER_CST` beside an opaque `VAR_DECL` offers nothing to split, so `associate:` never runs, the tree reaches RTL as a reverse subtract, and MIPS — which has no reverse-subtract-immediate — must materialise the constant first: `li $2,0x400 ; subu $2,$2,$4`. **Identical instruction COUNT, different opcodes, and a different insn at the HEAD of the arm** — which is exactly what `reorg` moves into the delay slot. + +**BYTE EVIDENCE** (re-measured at vetting time, pinned triple `cc1 -O2 -G0 -mips1 -mcpu=3000`, two files differing in one token): + + s = -base + 0x400; -> bne $5,$0,$L2 / subu $2,$0,$4 / j $L3 / addu $2,$2,1024 / $L2: addu $2,$4,1024 + s = 0x400 - base; -> bne $5,$0,$L2 / li $2,0x400 / j $L3 / subu $2,$2,$4 / $L2: addu $2,$4,1024 + +Consumer: `func_8017F76C` (ov_SC02_026, 154 ins, MATCH) — `if ((rand() & 1) == 0) spd = -base + 0x400; else spd = base + 0x400;` (the polarity half is §3-T4, read off the target's `bnez`). + +**DIAGNOSTIC TELL.** A `li $vN,K` you never asked for occupying a branch delay slot where the target has `negu $vN,$sM`. Reading a target back to source: **`negu` immediately followed by `addiu +K` ⇒ the original wrote `-x + K`; `li K` followed by `subu` ⇒ it wrote `K - x`.** Do not sweep parenthesisations (§164-16: inert) and do not apply §164-24's mirror — with only one splittable term there is nothing to mirror. + +*(SHARPENS — sharpens §164-40 (byte evidence, L12714), §165-36 rule 2 (L14646); evidence: byte-probed; from `func_8017F9AC`)* + +**§167-15 — PROVENANCE — the cookbook's cite for morph_lerp (…jr_8017BEBC.c:4018) is one line past the helper; the definit** + +**⚠ Provenance (R37) — §164-40's `morph_lerp` cite is one line long, and every `func_8017F9AC` entry is gate-outstanding.** L12714 cites the helper at `src/ov_SC07_006/ov_SC07_006_jr_8017BEBC.c:4018`; the definition is **:3994-4017** (:4018 is the blank line after its closing brace) and `func_8017F9AC`'s `INCLUDE_ASM` splice site is **:4086**, not the inherited notes' "~:4089". Re-read against the live tree 2026-08-12. This is §165-36 rule 2 firing on THIS campaign's own entries — coordinates decay inside a single phase, not just across phases; carry the reasoning, re-resolve the line. **And the status line:** `func_8017F9AC` (ov_SC07_006, 275 ins) is **not banked** — backlog row 64, `match_one` MATCH / gate rejected on TU plumbing, `INCLUDE_ASM` still at :4086 — so §164-05, §164-06, §164-07, §164-18, §164-40, §164-41, §165-14 and §165-38 all rest on a MATCH that is in-situ-proven (below) but has never faced the whole-binary arbiter. ⚠ The `func_8017F9AC` *definition* that does exist in `src/` is `ov_SC02_028_jr_8017D898.c:3586` — a different overlay, an unrelated body at the same address. Do not read it as this function banked. + +*(SHARPENS — sharpens §88f (L6718-6724), §136 self-verification (L9333-9339), the two offline oracles (L9395-9405), §42c / tools/rtu_match.py (L3033-3043); evidence: byte-probed; from `func_8017F9AC`)* + +**§167-16 — THE IN-SITU PROOF: KEEP `INCLUDE_ASM` **LIVE** IN THE BASELINE AND THE WHOLE-TU `.text` sha1 BECOMES THE ORACLE.** *(sharpens §88f — this IS the "missing rung between `match_one` and the binary" it asks `tools/` for, obtained without writing a relocation resolver; sharpens §136's self-verification (L9333-9339) and offline-oracle 2 (L9403-9405), whose collateral check must ALLOW a `j`-addend shift 'of exactly your function's size'; and BOUNDS §42c / `tools/rtu_match.py`, which NEUTRALIZES the stub (`-DINCLUDE_ASM(a,b)=`, rtu_match.py:7,137) and therefore cannot produce this baseline at all.)* + +**Target situation.** `match_one`/`rtu_match` say MATCH. That is a candidate, not a bank (§52b) — relocations are masked (§81/§84/§87) and file-scope decl damage is invisible. You want the strongest verdict obtainable *without* a gate cycle. + +**THE PROCEDURE.** Scratch-copy the host TU and build **two** objects through the pinned triple (`cpp → cc1 -O2 → maspsx --aspsx-version=2.56 --expand-div → as`): +* **baseline** — the TU exactly as committed, `INCLUDE_ASM` **left live**. `include/include_asm.h:7` expands to `.include "/.s"`, so `as` assembles the ORIGINAL game asm into the object at the function's exact size. +* **candidate** — the same TU with the draft body spliced **over** that one `INCLUDE_ASM` line. + +Then compare **the entire `.text` section byte-for-byte** and set-compare `objdump -r`. + +**THE LAW.** Because the splice replaces the stub *in place*, the candidate's function occupies the same byte range the original asm did — **no size shift exists, so the collateral check collapses to exact `.text` equality with no addend allowance to reason about.** And because the baseline's bytes ARE the shipped code as assembled, the body comparison is against the game, not against the `.s` text, with relocation ENTRIES (symbol + type) directly comparable — closing the wrong-`jal`-target (§81), wrong-`%lo`-symbol/addend (§84) and wrong-`D_`-symbol (§87) classes in one diff. Neutralising the stub throws all of that away. + +**BYTE EVIDENCE.** `func_8017F9AC` (ov_SC07_006, 275 ins), artefacts at `.run/wave6/func_8017F9AC/insitu/`: `base.text.bin` and `tu.text.bin` are both **31,064 bytes, sha1 `0cacfe12d5774ea408d822fec2e923ecda49bb19`** — the whole TU `.text` identical spliced-vs-baseline, all 44 neighbour functions included; `relocs.txt` shows the function's 30 relocations agreeing symbol-for-symbol (11 `D_` globals as HI16/LO16 pairs, 4× `R_MIPS_26 func_8004787C`, 2× internal `R_MIPS_26 .text`); `cc1.err` carries **only** the TU's pre-existing `conflicting types for built-in function memcpy` at :2305 — §165-13's warning-visibility check satisfied, zero new diagnostics. + +**DIAGNOSTIC TELL — and it discriminates the failure, not just detects it.** One extra TU compile. Read the two shapes apart: **`.text` differs only INSIDE your function's range** ⇒ ordinary codegen, go back to the diff. **`.text` differs OUTSIDE it** ⇒ your probe layer changed the declaration environment (§8d/§161c) and the function's own bytes may be perfect — the class L9405 says 'the target function's own bytes cannot show'. Spend it on any draft where a wrong `%lo` or a file-scope collision would cost a gate cycle. + +⚠ **Still not the arbiter.** The whole-binary SHA1 (G3/P9) is (§52b). `func_8017F9AC` passed every clause above and is *still* backlog row 64, unbanked. + +*Honest scope: n=1 function, procedure verified from the artefacts on disk; PROCESS, not a compiler law.* + +*(SHARPENS — sharpens §165-40, §165-39, §165-14/§NNN (L13999-14029), §42; evidence: byte-probed; from `func_8017FE38`)* + +**§167-17 — A READ IS PLACED BY ITS STATEMENT'S POSITION, NOT BY HAVING A STATEMENT OF ITS OWN: the own-statement-but-late spelling is BYTE-IDENTICAL to the inline one.** *(BOUNDS §165-40, which reads as 'sched1 hoists it to the top of the block regardless'; instantiates §165-14's LUID law on a GLOBAL load whose consumer is ~200 instructions away.)* + +Target shape — a global `lhu` consumed ~200 instructions later, yet emitted at insn 6-7, with the `addiu` that competes for `$a0` pushed below the store that frees it: + + /* 6 */ lui $v1,%hi(D_800B9A02) + /* 7 */ lhu $v1,%lo(D_800B9A02)($v1) + ... + /* 14 */ sw $a0,4($a3) <- the 0x808080 constant KEEPS $a0 + /* 15 */ addiu $a0,$a1,0xDC <- $a0 reused only after the store frees it + /* 16 */ sll $v1,$v1,0xE + +**THE LAW.** Free the read from its consumer's expression **and** put the statement at the head of the block. cc1 emits each statement where it stands and `rank_for_schedule` falls through to `INSN_LUID` on the all-ties case (`sched.c:2425`, §165-14), so source position IS the schedule: an inline sub-expression inherits its enclosing statement's position, and a statement written beside the consumer inherits the same one. **The boundary is not the lever; the POSITION is** — and it is a plateau with a monotone tail, so sweep it (§67). + +| `d = *(u16 *)&D_800B9A02;` written … | mismatched vs the MATCH | +|---|---| +| head of body — 1st, 2nd or 3rd statement | **0 — MATCH** | +| after `*(u32 *)(pkt + 4) = 0x808080;` | 7 | +| after `va = (u8 *)(p + 0xDC);` · after `*(u8 *)(pkt + 7) = 0x2C;` | 11 | +| its own statement, immediately above `ot = …[d << 14];` | 12 | +| inline: `ot = (u32)&D_800A6610[(*(u16 *)&D_800B9A02) << 14];` | 12 — **byte-identical `.text` to the row above** | + +**BOUNDS §165-40.** §165-40 (same function) says sched1 hoists an independent symbol-address pair to the TOP of its block and prescribes a bare `__asm__ volatile("")` fence to deny it. That holds only up to the LUID tie-break: an independent load moves ahead of *dependent* work but keeps source order against other *independents*, which is why eight positions give five distinct objects. **Sweep the statement before you spend the fence.** + +**BYTE EVIDENCE.** `func_8017FE38` (ov_SC07_001, 239 ins, banked at `src/ov_SC07_001/ov_SC07_001_jr_8017BEBC.c:3976`). One-statement A/Bs against the banked body on the pinned triple (`cc1 -O2 -G0 -mips1 -mcpu=3000`, `maspsx --aspsx-version=2.56 --expand-div`), masked-diffed against the banked object; ladder above. The inline and own-statement-late objects differ only in the embedded source filename. + +**THE DIAGNOSTIC TELL.** A global `lui`/`lhu` pair the target places in the prologue region while your draft emits it beside its far-away consumer — with an unrelated `addiu $aN,$aM,K` sitting in the slot the target gives the load, and a materialised constant displaced out of `$a0`. Do not fence and do not pin: walk the read's STATEMENT up one statement at a time and take the plateau. + +*(SHARPENS — sharpens §78 third block (L6203-6209), §164-74 ⚠ (L13469), §67 refuted hypothesis (L5390-5397), §76; evidence: byte-probed; from `func_8017FE38`)* + +**§167-18 — THE TARGET NAMES THE VARIABLE TO REUSE: a value sitting in a register that a DIFFERENT value just vacated was the same SOURCE variable reassigned.** *(§78 and §164-74 both prescribe 'reuse an already-busy variable' and neither says WHICH one; both leave 'maybe it is declaration order' open. This closes both.)* + +Target shape — count-neutral, one register, on a mask of a still-live operand: + + mine: andi $a1,$a1,0xFFFF ; … ; addiu $a1,$a1,-0x100 ; sb $a1,0xD($a3) + target: andi $v0,$a1,0xFFFF ; … ; addiu $v0,$v0,-0x100 ; sb $v0,0xD($a3) + +`$a1` is `y`'s own register — the value being masked. `$v0` is the register the PREVIOUS temp (`tp`, stored two insns earlier at `sh $v0,0x16($a3)`) has just released. + +**THE LAW.** A fresh local for the masked value is tied by local-alloc to the source it reads, so it lands in the operand's register. Reassigning the temp whose live range has just ended — `tp = y & 0xFFFF;` written after `x -= tp << 6;` — makes the allocator re-use the register that temp vacated, which is the target's. **Read the target's register choice as a statement about variable IDENTITY: when the target's value sits in a register a now-dead value released, rather than in its own operand's register, the original reassigned THAT variable.** This is the selector §78 and §164-74 do not supply. + +**DECLARATION ORDER IS NOT THE KNOB — negative control.** Six spellings of a separate `sv` local (declared first / mid / last in the block, assigned at the load, assigned after the `u0` store, block-scoped inside the `if`) give **`.text` byte-identical across all of them** — 4 mismatched every time. Only the reuse closes it, 4 -> 0. Second instance of §67's 'declaration order is likewise inert', now at HImode and in a call-free leaf. (A seventh probe that moves the ASSIGNMENT rather than the declaration costs 27 — that is §42 statement placement, a different axis; do not conflate them.) + +**BYTE EVIDENCE.** `func_8017FE38` (ov_SC07_001, 239 ins, banked at `src/ov_SC07_001/ov_SC07_001_jr_8017BEBC.c:3976`). Drop-one ablation against the banked body on the pinned triple: fresh `sv` = **4 mismatched at 239 ins**; `tp` reuse = **MATCH**. Permutation set `.run/wave4/func_8017FE38/S1…S6.c` (S1/S2/S3/S5/S6 identical `.text`). + +**THE DIAGNOSTIC TELL.** Count-neutral, one register, on a value produced by masking a still-live operand: if YOUR register is the operand's and the TARGET's is one a nearby store has just freed, you manufactured an allocno. Do not permute the declaration block and do not pin — rename the assignment onto the dead temp. + +*Scope: n=1 function, no `.lreg` dumped; the measurement is the law.* + +*(SHARPENS — sharpens §166a (THE DESTINATION-TU ORACLE — 'prove it in situ … splice into a private copy of the real TU, one directory deep, with a `shared` symlink … run the pinned triple end-; evidence: byte-probed; from `func_8017FFD0`)* + +**§167-19 — THE WHOLE-TU SYMBOL-LAYOUT ORACLE: keep `INCLUDE_ASM` LIVE, assemble from the repo root, and check every sized symbol's object offset against its SHIPPED VRAM delta.** *(SHARPENS §166a's in-situ recipe, which proves the function's BYTES and the collateral symbols' BYTES but never checks WHERE anything landed; and §8a, whose 'rtu_match … excludes the §8 jtbl rodata, so it MATCHes a body whose switch is subtly wrong' false-MATCH class this closes offline. Distinct from §137a oracle 2, which compares with-splice against without-splice — a SELF-comparison; this one compares against the SHIP.)* + +**The recipe — three lines on top of §166a.** +1. Splice the draft over its own `INCLUDE_ASM` and leave **every OTHER `INCLUDE_ASM` live**. Do **not** neutralize them with `-DINCLUDE_ASM(a,b)=` — that is §42c's `rtu_match` trick and it is the opposite choice. Run `as` with **cwd = repo root** so the stubs' `.include "asm//nonmatchings/…"` paths resolve; the object then holds the TU's whole shipped text, not just your function. +2. `nm -S` the object and drop zero-size symbols (§166a's `.NON_MATCHING` alias trap, restated — an unsized alias has no offset worth checking). +3. Require, for every survivor, `st_value == VRAM(sym) − BASE(section)`, with `BASE(.text)` = the TU's first function VRAM and `BASE(.rodata)` = its first jtbl VRAM. Both bases and every symbol VRAM are already in the `.s` headers / `symbols.us.txt`. + +**THE LAW.** `.text` offsets are a running sum of sizes, so ONE wrong-length body shifts every symbol after it. Requiring the whole offset VECTOR to equal the shipped deltas therefore proves, in one command and with no target `.s` diff at all: your function's LENGTH, its POSITION (i.e. that nothing above it drifted), the length of every already-banked C sibling in the TU, and — because the `.rodata` base catches them — that the **jump tables are the shipped size**. That last clause is the point: §8a's false-MATCH class (`func_80159C84`'s second jtbl 5 words instead of 6) shifts the following jtbl by 4 and fires here, offline, where `rtu_match`/`match_one` are structurally blind and the file's only prescribed remedy is a full whole-binary gate. + +**Byte evidence.** `func_8017FFD0` (ov_SC03_108, 196 ins) spliced over `src/ov_SC03_108/ov_SC03_108_jr_8017F83C.c:2917`, one directory deep with a `shared` symlink so `../shared/engine_core.h` resolves, pinned triple end-to-end (`cpp -Iinclude` → `cc1 -O2 -G0 -mips1 -mcpu=3000 -mgas -msoft-float` → `maspsx --aspsx-version=2.56 --expand-div` → `as -march=r3000 -O1 -G0`): masked diff **196/196, 0 mismatches**, and **21/21 symbols at their exact shipped offsets, layout drift 0** — `func_8017FFD0` at `0x794` = `0x8017FFD0 − 0x8017F83C`. The 21 reconciles exactly against the tree: 15 `.s` under `asm/ov_SC03_108/nonmatchings/ov_SC03_108_jr_8017F83C/`, 4 col-0 C definitions in the TU (`func_8017F83C`:2759, `func_801803E0`:2932, `func_80180814`:3096, `func_8018087C`:3109), and 2 jtbls (`jtbl_801A0510`, `jtbl_801A05D8`). + +**⚠ Three bounds — this is a LENGTH/POSITION oracle, not a byte gate.** +* It cannot see wrong bytes at the right length. §166a's masked diff is still the body check and the whole-binary SHA is still the arbiter (§52b). +* It is **blind to the last symbol of each section** — nothing follows it to shift. Close that by also comparing each section's total size against `VRAM(last) + size(last) − BASE`. +* Its live surface is your function + the TU's already-banked C functions + the jtbls. An `INCLUDE_ASM`'d sibling is verbatim shipped assembly and *cannot* drift, so in a 100%-stub TU the oracle gives you back only your own function's length. + +**Evidence scope (honest).** n = 1 and **green-only**: it was run on a MATCHing draft and never on a known-bad object. The negative control existed in the same note-set and was never fed to it — the barrier-stripped build of this very function is 189 ins, −28 bytes (§164-09's first table row), which would shift `func_801802E0` and everything below. That the oracle *fires* on drift is arithmetic, not a measurement; run it once on a bad object before quoting it as a gate. + +**THE DIAGNOSTIC TELL.** Offsets exact up to symbol K and uniformly off by a constant Δ from K+1 onward ⇒ **symbol K is Δ bytes wrong**, and K is the function to look at — even when K is not the one you spliced. + +*(SHARPENS — sharpens §152 (BYTE SIZE is the family key that name- and h_seq-grouping both miss — 'Byte size is an allocator-independent, name-independent, cache-independent family key … One c; evidence: byte-probed; from `func_8017FFD0`)* + +**§167-20 — §152's BYTE-SIZE FAMILY KEY IS BAND-DEPENDENT: exact for WHALES, ~92% garbage under 0x100.** *(BOUNDS §152, whose 'one command, exact, no false positives on the case measured' was measured on a 0xECC / 947-instruction body — the one band where it is in fact clean. Does not refute it: size is still the right candidate GENERATOR.)* + +§152 keys structural families on the `.s` header size (`grep -rl 'nonmatching .*, 0x' asm/*/nonmatchings/*_jr_/`). Measured over **all 11,681** `asm/*/nonmatchings/*/*.s`, with identity taken as the SHA1 of the opcode-only stream (the mnemonic after each `*/`, delay slots included): + +| size band | size-groups (≥2 files) | files | groups holding >1 body | files outside the group's dominant body | +|---|---|---|---|---| +| < 0x100 | 62 | 9,103 | 62 | **8,367 (91.9%)** | +| 0x100–0x200 | 64 | 1,897 | 64 | 1,641 (86.5%) | +| 0x200–0x400 | 91 | 460 | 89 | 291 (63.3%) | +| 0x400–0x800 | 21 | 67 | 15 | 21 (31.3%) | +| ≥ 0x800 | 4 | 14 | **0** | **0 (0.0%)** | + +**THE LAW.** Size is a *sieve*, and its selectivity IS the size. Above ~0x800 it is an exact key — which is why §152's whale exemplar had no false positives, and why that result must not be carried down. At ordinary overlay-function sizes it must be composed with a structural oracle before any remap is planned or any reach is reported. + +**Byte evidence — the case that found it.** `func_8017FFD0`'s family (ov_SC03_108, 196 ins, 0x310). The §152 command on `0x310` returns **5** files. Four share opcode-stream SHA1 `9afa3e4a…` — ov_SC03_108 `func_8017FFD0`, ov_SC03_110 `func_8018035C`, ov_SC03_112 `func_80181F74`, ov_SC05_001 `func_80183C9C` — and a full instruction-text diff with only symbol names and `.L` labels normalized is drift 0 across all 196 lines, which is what makes that remap a pure 4-symbol substitution. The fifth, `asm/ov_SC04_015/nonmatchings/ov_SC04_015_jr_8017AE2C/func_8017E520.s`, is an unrelated body: identical for four instructions (`addiu;sw;addu;sw` — every -O2 leaf prologue looks like this), then `lhu` where the family has `lh`, then `addiu;sll;sra;sltiu` where the family has `slti;bnez` — a §163b minval-biased `s16` switch dispatch, not a guard. + +**THE CHEAP COMPOSITION (keep §152, add two steps).** Size grep as the candidate generator → group candidates by the opcode-only stream → confirm each survivor with a full instruction-text diff, symbols and `.L` labels normalized. Steps 2 and 3 are grep-and-diff, zero tokens, and they turn a band-dependent sieve into an exact key at any size. §152's caution 1 still binds afterwards: `match_one`/`rtu_match` mask exactly the fields a remap rewrites, so only the whole-binary SHA gates the result. + +**DIAGNOSTIC TELL.** A size-keyed reach whose members' FIRST branch tests different struct offsets, or where one member extends its switch index and the others do not ⇒ you have a size collision, not a family. Run §164-45's 10-second check across the whole candidate set, not just against the Ghidra seed. + +*(SHARPENS — sharpens §45 Lever A (merged accumulator variables; gcc-2.7.2 global-alloc has no coalescing, K8), §164-64 (a local-alloc'd expression temp can claim $a0 and evict parameter 1), §; evidence: byte-probed; from `func_801805D4`)* + +**§167-21 — A REASSIGNED PARAMETER CARRIES ITS INCOMING ARGUMENT REGISTER THROUGH THE WHOLE FUNCTION: give the recomputed value its OWN local, and expect NO length change.** *(SHARPENS §45 Lever A, whose law — "gcc-2.7.2 global-alloc has no coalescing (K8), so one hard reg spanning disjoint regions can only come from one reused source variable" — is stated and byte-proven only in the MERGE direction; and §164-64, which supplies the `set_preference`/`prune_preferences` chain but at the opposite polarity, an anonymous temp EVICTING the parameter. Neither prices a reused PARAMETER, and neither warns that the residual is COUNT-NEUTRAL. Distinct from §161b/§162n2, which price an ALIAS of a parameter, not a reassignment of it.)* + +**Target shape** — one value recomputed between N inlined blocks, where the FIRST block reads the raw argument register and the later blocks read a callee-saved one: + + mult $v0,$a0 <- block 1: the raw incoming parameter + ... + jal func_8004787C ; addu $s0,$v0,$zero <- left operand parked across call 2 + jal func_8004787C ; sra $v0,$v0,4 + addu $s0,$s0,$v0 <- the recomputed factor lands in $s0 + ... + mult $v0,$s0 <- blocks 2 and 3 + +**THE LAW.** The parameter's allocno is seeded with a copy preference for its incoming hard register off the prologue's `(set pseudo (reg 4))` (§164-64's citation, `global.c:1535-1619`). Reassign the parameter and both roles share ONE pseudo, so it holds `$4` for its entire life and the recomputed value can never reach `$s0`. Two named variables give two allocnos, and the second is then free to take the register of the anonymous left-operand temp that must survive call 2 — which is what puts the accumulation's result in the same `$s0` the park used. + +**BYTE EVIDENCE** (vet-time, pinned triple; `func_801805D4`, ov_SC07_006, 212 ins; base `.run/wave6/func_801805D4/func_801805D4.c`). Baseline with a separate `s32 u` → **MATCH 212**. Delete `u` and reassign the parameter `t` instead, nothing else changed → **mine=212, target=212, 76 mismatched, class `REGALLOC-PERM`, sig `$a0>$s0,$a1>$a0,$a2>$a1,$a3>$a2,$t0>$a3,$t1>$t0,$t2>$t1,$t3>$t2`**. Objdump of both builds at offset `0x128` isolates the single causal instruction: baseline `addu s0,s0,v0`, reassigned `addu a0,s0,v0`; everything downstream is that one register shift propagating through first-fit. + +⚠ **The producing note's own mechanism is byte-refuted — do not repeat it.** It reads "forced into a callee-saved reg for the WHOLE function → prologue `addu $s0,$a0,$zero` and a first loop `mult $v0,$s0`." The fused build emits **no prologue parameter copy at all**, its count is **identical**, and **all three** of its loops read `mult $v0,$a0`: the fused pseudo keeps `$a0`, it does not migrate to `$s0`. The prescription's direction is right; the register and the phantom extra instruction are not. + +**DIAGNOSTIC TELL.** Equal instruction counts, no prologue copy, and the whole caller-saved file shifted by one slot (`$a0>$s0,$a1>$a0,…`), with the target reading a callee-saved register where you read the raw argument register **in the later blocks only**. That is one variable doing two jobs. Split it before touching `register __asm__` pins, §148-C density sliders or the permuter — it is a one-line declaration edit. §45 Lever A is the same knob run the other way; read the target's FIRST block to decide which direction you need. + +*(SHARPENS — sharpens §164-05 (THE INLINE-EXPANSION FRAME ORACLE: frame - args - 4*regs COUNTS THE EXPANSIONS, L11944), §162i1 (the unreferenced-local frame oracle); evidence: byte-probed; from `func_801805D4`)* + +**§167-22 — THE INLINE-EXPANSION FRAME ORACLE MUST SUBTRACT THE ALIGNMENT PAD BEFORE DIVIDING: an ODD SAVED-REGISTER COUNT BREAKS THE "exact integer" STEP.** *(SHARPENS §164-05, whose procedure — divide `(frame − args − 4×regs)` by the body count, "an exact integer is your answer" — is byte-proven on two functions whose saved-register areas are 0 bytes (`func_8017DC1C`) and 16 bytes (`func_8017F9AC`). Both are already multiples of 8, so no rounding can occur in either row and the oracle's failure mode is invisible in its own evidence.)* + +**Target shape** — the same `static inline` helper, in a caller that saves an ODD number of registers: + + addiu $sp,$sp,-0x68 + sw $s1,0x5C($sp) ; sw $ra,0x60($sp) ; sw $s0,0x58($sp) <- 3 saved regs = 12 bytes + ... and NOTHING else in the function touches $sp + +**THE LAW.** gcc-2.7.2/MIPS rounds the total frame up to 8. When `args + vars + 4×regs` is not already 8-aligned the frame carries a 4-byte pad, §164-05's residual absorbs it, and the quotient misses an integer by `4/N` — which reads as "this is not an inline expansion" and sends you to §162i1's dead-local pad, the exact misroute §164-05 exists to prevent. Compute `p = (8 − ((args + 4×regs) mod 8)) mod 8` — equivalently try `p ∈ {0,4}` — and divide `(frame − args − 4×regs − p)`. + +**BYTE EVIDENCE** (`func_801805D4`, ov_SC07_006 / `jr_8017BEBC`, 212 ins, MATCH). Frame **0x68 = 104**; outgoing-args area 16 (the function makes four `jal`s); saved regs 3 × 4 = **12** at `0x58`/`0x5C`/`0x60`; the only `$sp` references in the entire function are the 8 prologue/epilogue words (`grep -c '\$sp'` = 8). §164-05 as written gives 104 − 16 − 12 = **76**, and 76/3 = **25.33 — no integer**. The true vars region is `0x10..0x57` = **0x48 = 72 = 3 × 0x18**, i.e. the same 24-byte helper local area measured across 25 expansions on `func_8017DC1C` and 4 on `func_8017F9AC`, with **4 bytes of alignment pad** sitting above the register saves. Same helper, same 24, N = 3 — the law itself reproduces exactly; only the arithmetic recipe needed the pad term. + +**DIAGNOSTIC TELL.** The quotient misses an integer by exactly `4/N`. Subtract 4 and re-divide **before** concluding the body is not an inline expansion; an odd saved-register count (3 or 5 `sw $sN`, `$ra` included) is the fingerprint that the pad is present. + +*(SHARPENS — sharpens §164-61 / §16Xd (L13164-13181) — the section being bounded, §48-B THE EBB RULE (L3468-3480) — 'cse resets its hash table at a label'; stated for copies that must SURVIVE,; evidence: byte-probed; from `func_80181500`)* + +**§167-23 — THE ARG-COPY ARITY TELL IS LABEL-SCOPED, NOT BASIC-BLOCK-SCOPED.** *(bounds §164-61/§16Xd; reads §48-B's EBB boundary for a DELETED copy instead of a surviving one)* + +§164-61 scopes its "the `nop` carries no arity information" exemption to "the parameter's own basic block" and tells you to "ask which block the `jal` is in." Read literally that mis-predicts, because **a conditional branch does not end the block that matters here.** + +Target shape — one function, one callee, one argument, the two halves 0x18 bytes apart: + + /* 80181508 */ move $s0,$a0 <- the parm save + /* 8018151C */ bne $v0,$zero,.L8018153C + /* 80181524 */ jal func_80131C78 <- fall-through of a conditional branch: + /* 80181528 */ nop <- NO arg copy, delay slot is a real `nop` + .L8018153C: <- the FIRST CODE_LABEL in the function + /* … */ jal func_80131C78 <- same callee, same argument, past the label + /* … */ move $a0,$s0 <- and now the copy appears + +**THE LAW.** The copy is deleted by cse, and **cse's block runs label to label**: `cse_end_of_basic_block` scans `while (p && GET_CODE (p) != CODE_LABEL)` (`tools/reference/gcc-2.7.2/cse.c:8039`). A `JUMP_INSN` is not a terminator, so the **not-taken side of every conditional branch is still inside the parameter's own cse block** — which is sound, that side has exactly one predecessor path. Inside it the incoming hard reg still carries the value, the arg-setup `(set (reg $4) (parm-pseudo))` is redundant, and nothing is emitted. At the first `CODE_LABEL` the table resets (§48-B's EBB rule, stated there for copies you need to *survive* — this is the same boundary read for copies that get *deleted*) and every later call must materialise `move $a0,$sN`. **§164-61's exemption is therefore wider than its own prose: it covers every call between function entry and the first label, however many conditional branches sit in between.** + +**BYTE EVIDENCE** (`func_80181500`, ov_SC03_107, 212 ins, `match_one` MATCH ×3; `.run/wave6/func_80181500/{func_80181500.c,cc1.s}`). The draft declares `extern void func_80131C78(s32 a0);` and calls it with one argument at **four** sites. cc1 emits the first (`cc1.s:33` → `0x80181524`) with **no arg setup at all**, and all three later ones (`cc1.s:222`, `:352`, and `func_8012B030` at `:359`) with `move $4,$16` in the `jal` delay slot. Same function, same callee, same argument — the only variable is which side of `$L2`/`$L25` the call sits on. This is a **second overlay and a second callee** for §164-61, whose evidence was `func_80186C4C`/`func_8012C218` alone; and note §164-61's own listing already shows `bgtz v0,0x38` at idx 8 with the exempt `jal` at idx 10 — its exemplar was always the fall-through case its prose calls "the entry block." + +**DIAGNOSTIC TELL.** Before reading arity out of a `jal` delay slot, find the **first `CODE_LABEL`** in the target `.s` — the lowest address any branch or jump names. A bare-`nop` `jal` **above** that address is arity-SILENT: take the arity from the destination TU (§166), a sibling overlay, or a later call site, and do **not** buy the §17a-1 `((void (*)(void))f)()` cast to explain it. Only a bare `nop` on a `jal` **below** the first label is evidence, and only if `$a0` has not been re-materialised in between. + +*(SHARPENS — bounds §164-61/§16Xd (L13164-13181, "THE ARG-COPY ARITY TELL IS BASIC-BLOCK SCOPED"), extends §48-B (L3468-3480) to deleted copies, guards §161c/§17a-1; evidence: byte-probed; from `func_80181500`)* + +*(SHARPENS — sharpens §55a / §49-variant (L4099-4102), §162b1 (L11069, incl. tell #2), §162k1 (L11421), §164-48 (L12861); evidence: byte-probed; from `func_80181948`)* + +**§167-24 — THE TEMP YOU ADD TO KILL THE BIRTHING BOOST IS ALSO A WIDTH DECLARATION: share it at the VALUE'S OWN WIDTH, or the boost-kill stops being zero-byte.** *(sharpens §55a's §49-variant (L4099-4102) — "route the load through a temp assigned in BOTH halves of a branch ⇒ `reg_n_sets==2` ⇒ boost dead, **at zero byte cost**" — which names no TYPE for that temp; composes it with §162k1 (L11421) / §164-48 (L12861), which price a local's declared width but never as a PRECONDITION of a scheduling lever; and bounds §162b1's tell #2 (L11090), whose mask-cost route sends you to SPLIT the shared temp — which here re-arms the boost.)* + +**Target shape** — a jr-switch whose arms bump the same halfword state word, with a 2-insn transposition around a call and NO mask anywhere in the residual: + + case 0: … sh $v0,0x34($s0) <- `state + 1` written back + case 1: … lhu/addiu/sh 0x34($s0) <- the same bump, other arm + residual: your `beqz` slot holds `move $a0,$s0`; the target's holds `addiu $v0,$a1,1` + +**THE LAW.** The §49-variant S2 kill is zero-byte **only when the shared pseudo's MODE matches the value's**. Declared `u16`, every set is HImode, combine's `reg_nonzero_bits` union (`combine.c:718-788`, §162b1) stays inside 0xFFFF and nothing re-widens or re-ranks; declared `s32` — one token, nothing else touched — the function repermutes. **Take the width off the lvalue the value is stored back into, never off the arithmetic.** The correction §162b1 would have you make (split it per arm) is the one edit you must not make: splitting restores `reg_n_sets == 1` and the boost with it. + +**BYTE EVIDENCE** (`func_80181948`, ov_SC01_077, 132 ins, `match_one` MATCH; probes and harness in `.run/wave6/func_80181948/`, generator `mk3.py`). One function-scope `u16 t` written in case 0 (`t = state + 1; *(u16*)(a0+0x34) = t;`) and case 1 (`t = *(u16*)(a0+0x34) + 1; …`) ⇒ two sets ⇒ no boost ⇒ **132/0 MATCH**, and `state` moves `$a2`→`$a1` for free. `t3_s32temp.c`, identical but for `u16 t` → `s32 t`: **131 ins, 113 mismatched.** Reusing the function's existing `s32 v0` for the same bump (`t2_reusev0.c`): the same **131/113**. Measured NULLs on the same base, all 132/6 — per-arm block scope, memory re-read, decl-order swap, an extra dummy local, `+=`, `<= 0x170`, a `$2` pin. `register u16 state __asm__("$5")` buys the REGISTER only and leaves the 2-insn swap (132/2), consistent with §164-50: the pin was on the mis-allocated value, not on the boosted insn's dest. + +**⚠ THE MECHANISM IS A HYPOTHESIS — do not repeat the originating note's version.** It reads "the s32 temp forces the zero-extend back and wrecks the dispatch". The length went **down** by one, so a re-materialised `andi` is not what happened; 113 mismatched at LENGTH-DRIFT −1 is §164-48's declared-width **allocno re-ranking** signature. No `-dl`/`-dS` was taken of the `s32` build. Ship the precondition, not the story. + +**THE DIAGNOSTIC TELL.** You introduced a shared multi-set temp to kill an S2 boost and the mismatch count EXPLODED instead of dropping, with a length drift of ±1. Do not split it back and do not reach for a pin — **re-declare it at the width of the lvalue it is written into, and re-gate.** + +*(SHARPENS — sharpens §1-I3, §1-I4, §164-04, §28's div-by-constant magic table (0x66666667→/10, 0x55555556→/3); evidence: byte-probed; from `func_801832F8`)* + +**§167-25 — Division idioms confirmed: `x/20` uses gcc's `/5` magic 0x66666667 with shift 1+2=3 (20 = 5·2^2); `(excess*127** + +**§16N+2 — THE MAGIC CONSTANT NAMES ONLY THE *ODD PART* OF THE DIVISOR; THE `sra` COUNT CARRIES THE POWER OF TWO. Read them together or you will guess the wrong divisor.** *(sharpens §1-I3, which gives the authoring rule and exactly one pair (0xCCCCCCCD / shift 3 = ÷10) and never says the constant is reusable; and corrects the §28-imported table's `0x66666667 → /10` row, which is a divisor-ladder entry misfiled as a unique mapping. Disjoint from §164-04, which owns the pure-power-of-two `bgez ; addiu 2^k-1 ; sra k` path.)* + +Target shape — a reciprocal multiply whose magic you recognise but whose shift you do not: + + lui $a0,(0x66666667 >> 16) ; ori $a0,$a0,(0x66666667 & 0xFFFF) + mult $v0,$a0 + sra $v1,$v1,31 + mfhi $a3 + sra $v0,$a3,3 <- the shift is the whole message + subu $v0,$v0,$v1 + +**THE LAW.** `expand_divmod` factors the divisor as `odd × 2^k`, synthesises the magic for the ODD part only, and folds `k` into the post-multiply shift. So **divisor = odd × 2^(shift − base_shift(odd))**, and one magic covers a whole ladder. + +| magic | odd part | shift → divisor | +|---|---|---| +| `0x55555556` | 3 | 0 → /3 | +| `0x66666667` | 5 | 1 → /5 · 2 → /10 · 3 → **/20** · 4 → /40 · 5 → /80 | +| `0x51EB851F` | 25 | 5 → /100 · 6 → /200 · 7 → /400 · 9 → /1600 · 13 → **/25600** | + +**Byte evidence — 11 divisors on the pinned triple** (`cpp / cc1-2.7.2 -O2 -G0 -mips1 -mcpu=3000 / maspsx 2.56 / as`), `s32 f(s32 x){ return x / D; }`, D ∈ {3,5,10,20,40,80,100,200,400,1600,25600}: the table above is read straight off the objdumps (`.../vet832F8/dv.c`). Target anchor: `func_801832F8` (ov_SC02_041) uses **both** rows — `0x66666667 / sra 3` for `out1.x / 20` at 801833B0-801833C4 and `0x51EB851F / sra 13` for `(dist-0x100)*127 / 25600` at 80183438-80183454. Neither is a ÷10 or a ÷100. + +**THE DIAGNOSTIC TELL.** You recognise the magic, so you write the divisor the table told you and the magic comes out right while the `sra` is off by k. **Multiply the divisor by `2^k` and write it literally** — never hand-craft a magic, never respell as `(x/5) >> 2` (that rounds differently and emits an extra bias). Read it the other way too: an unfamiliar magic plus a large shift is a compound divisor, and the shift alone tells you how many factors of two to strip before you look the odd part up. + +*(SHARPENS — sharpens §165-21 (L14195) — TWO *COMPUTED* ARMS WANT ONE SHARED TRAILING STORE; already prescribes this exact cure for this exact target shape (arithmetic in the bnez delay slot, ; evidence: byte-probed; from `func_80183834`)* + +**§167-26 — WHEN THE DUPLICATED STORE'S TAIL FALLS THROUGH, THE DUPLICATE-vs-JOIN DIAL IS FREE AND COSTS ONLY THE SAVED-REGISTER ORDER. §165-21's "+2 instructions" is the two-`j` case.** *(sharpens §165-21, whose only measured cost for the duplicated spelling on COMPUTED arms is a length drift; and §48-A1 / §164-32, which own the duplicate↔join allocno dial for a LOCAL's init and a switch arm's BODY but never for a STORE's BASE POINTER. The discriminator is §50-B's cross-jump floor.)* + +Target shape — two computed arms updating ONE memory cell by ±K, the tail already merged, **one arm falling through**: + + jal rand + … + andi $v0,$v0,1 + bnez $v0,.L1 + addiu $v0,$sX,-0x400 <- TRUE arm, in the slot + addiu $v0,$sX,0x400 <- FALSE arm, falls through + .L1: sh $v0,0x12($sY) <- ONE `sh`, reached by FALL-THROUGH, not by a `j` + +**THE LAW.** Both spellings emit this identical 135-instruction stream. The only thing that moves is which of {base pointer, loaded value} gets `$s1` and which gets `$s2`: + + if (c) *(u16*)(obj+0x12) = ang-0x400; else *(u16*)(obj+0x12) = ang+0x400; -> BASE = $s1, VALUE = $s2 + *(u16*)(obj+0x12) = c ? ang-0x400 : ang+0x400; -> VALUE = $s1, BASE = $s2 <- target + if (c) d = ang-0x400; else d = ang+0x400; *(u16*)(obj+0x12) = d; -> same as the ternary + +Duplicating the store hands the BASE POINTER one extra `REG_N_REFS` at allocation time: `jump_optimize (…, JUMP_CROSS_JUMP, …)` runs after reload (pass order, §45-B), so the second `sh` is real to `global.c:594 allocno_compare` and is refunded before `final`. This is **§48-A1's rule applied to a store's BASE instead of a local's VALUE**, and **§164-32's zero-byte ref-boost run subtractively** — collapse to the join to LOWER the base, duplicate into the arms to RAISE it. *(The refs reading is INFERENCE: no `-dl`/`-dg` dump was taken. What is measured is the direction and the zero length cost.)* + +**BYTE EVIDENCE** — `func_80183834` (ov_SC01_077, 135 ins, banked `src/ov_SC01_077/ov_SC01_077_jr_80183324.c:3357`, TU has 0 remaining `INCLUDE_ASM`). Every probe preserved under `.run/wave6/func_80183834/`; **`base.c` → `varO.c` is a one-hunk diff whose only change is the store spelling**: + +| spelling | file | result | +|---|---|---| +| ternary, ONE store | `varO.c` (= the banked body) | **MATCH 135/135** | +| named `d`, store after the join | `varL.c` | **MATCH** | +| duplicated store, plain | `base.c` | 5 mismatched | +| duplicated store, nested-CSE load order | `varK.c` | 5 mismatched | +| duplicated store, inverted condition + swapped arms | `varN.c` | 5 mismatched | +| duplicated store through a `u16 *` base (`*obj`) | `varR.c` | 5 mismatched | + +**IT IS NOT A TIE-BREAK, AND PINS ARE STRICTLY WORSE.** Declaration order (`varA.c`) and block scope (`varB.c`, `varI.c`) are **inert** — so §158 / §162o1's decl-order tie-break is not the dial, and the refs inequality is strict, not an exact tie. Pinning BOTH (`varC.c`: `obj`→`$18`, `ang`→`$17`) and pinning `ang` alone (`varG.c`) each cost **+2 instructions and flipped a branch sense**; pinning `obj` alone (`varH.c`) does MATCH but fails `dedup_propagate.compiles_standalone` (§37) and forfeits this exemplar's ×3 reach (§162p rung 3). **Take the spelling, not the pin.** + +**THE DIAGNOSTIC TELL — count the `j`s into the shared store, not the instructions.** A two-arm ±const update of one memory cell where your instruction COUNT is exact and the residual is a clean saved-register swap between the store's BASE and the VALUE it stores ⇒ re-spell the store; do not reach for a pin, an §47/§158 slider or the permuter. Which direction to try is read off the tail: +- **one arm FALLS THROUGH into the `sh`** ⇒ §50-B's `minimum=1` path merges the duplicate for free, the two spellings are length-identical, and the dial is *purely* the register order — **try both directions**; +- **both arms reach the store by `j`** ⇒ §165-21, the duplicate does not merge and you pay its **+2 instructions**, so only the join spelling is live. + +**⚠ Do NOT cite §164-69 as the opposite pole.** Its seven losing join spellings are a **call-bearing, multi-instruction merged TAIL** whose mechanism is sched1's basic-block SCOPE; nothing there is a register permutation on a one-instruction store tail. The two-way framing already belongs to §48-A1, which states both directions outright. + +*(SHARPENS — sharpens §136g-1 (L9214-9218), §164-55 (L13035-13063), §164-47 (L12840-12851), §3-T4 (L90-103); evidence: byte-probed; from `func_80187130`)* + +**§167-27 — WHEN *EVERY* ARM RETURNS, gcc-2.7.2 PHYSICALLY SWAPS THE TWO ARMS AND INVERTS THE BRANCH (`jump.c:1806`). THE ARM YOU WRITE **FIRST** IS THE ONE THAT LANDS NEXT TO THE EPILOGUE.** *(⚠ BOUNDS §164-55 and §164-47, whose "RTL block order == source statement order" / "only the last-written body can fall into the epilogue" is stated with no precondition; settles the internal contradiction in §3-T4 — its clause (a) "put the target's fall-through block in the `if`" and its clause (b) "`if (ok){…return good;} return 0;` makes `return 0` the fall-through" cannot both hold, and (b) is the one that governs here; supplies the FORWARD-direction lever for the transform §136g-1 already names but only teaches how to DISARM.)* + +**Target shape** — a 2-way test where both sides `return`, the branch is taken *forward into* the block that sits last, and that block falls into the epilogue with no `j`: + + bnez $v0,.L801871XX <- branch-if-condition-TRUE, forward + li $a0,1 <- delay slot stolen from the block it branches TO + j .Lepi <- the OTHER arm, inline, pays the jump + move $v0,$zero <- …its return value, in the j's slot + .L801871XX: … li $v0,1 <- falls straight into the epilogue, no `j` + .Lepi: lw $ra,0x6C($sp) … + +**THE LAW.** `jump.c:1800-1875` — `/* Look for if (foo) bar; else break; */` — is a real **block-swapping** transform in gcc-2.7.2. It matches `condjump label1 / range1 / jump label2 / label1: / range2 / / label2:`, calls `invert_jump (insn, label1)`, and then **splices range1 and range2 past each other** with raw `NEXT_INSN`/`PREV_INSN` surgery (`jump.c:1866-1875`). Its preconditions ARE the "every arm returns" shape: +* `label1` = the if-join label, with `LABEL_NUSES (label1) == 1`; +* `range1end` (last insn of the then-arm) is a **simplejump** targeting `label2 = next_label (label1)` — i.e. **the then-arm ends in `return` and the label after the if-join IS the return/epilogue label**; +* `range2end` is a JUMP_INSN followed by a BARRIER — i.e. **the other arm also ends in `return`**; +* `! first` (a later `jump_optimize` round) and `reload_completed ? ! flag_delayed_branch : 1`. + +So **`if (C) { P; return X; } Q; return Y;` emits Q-then-P**, branch-if-`C` to P, and **P — the arm you wrote FIRST — is the epilogue-adjacent block with no `j`.** Put the target's no-jump arm **inside the `if`**, never after it. **The spelling is inert:** `… return Y;` and `… else { return Y; }` are byte-identical, both ways round; only WHICH arm is the then-arm moves anything. + +**Byte evidence — `func_80187130` (ov_SC06_018, 86 ins, banked `commit:1713`, `src/ov_SC06_018/ov_SC06_018_jr_80186270.c:3243`).** Five variants through the pinned triple, re-run at vetting (`cpp / cc1-2.7.2 -O2 -G0 -mips1 -mcpu=3000 / maspsx / as`; probe artifacts `.run/match/func_80187130.246967/` and `.242624/`): + +| spelling | cc1 `.s` layout | verdict | +|---|---|---| +| `if (c != 0) { BODY; return 1; } return 0;` | `bne …,$L2` · `j $L3 ; move $2,$0` · `$L2: BODY … li $2,1` · `$L3:` | **MATCH — the banked body** | +| `if (c != 0) { BODY; return 1; } else { return 0; }` | **`.s`-identical modulo label numbers** | MATCH | +| `if (c == 0) { return 0; } BODY; return 1;` | `beq …,$L2` · `BODY … j $L3 ; li $2,1` · `$L2: move $2,$0` | DIFF, **86 vs 86 ins**, 11 mismatched | +| `if (c == 0) { return 0; } else { BODY; return 1; }` | **`.s`-identical to the row above** | DIFF | +| then-arm falls to a shared tail (no `return` in the arm) | `beq …,$L2` · **BODY inline** · `$L2: ` — source order, **no swap** | regime control | + +The last row is the precondition failing: with the arm falling into an interior join, `range1end` is not a jump to the return label, `jump.c:1806` never fires, and §164-55/§164-47's source-order law holds unmodified. **The two laws are disjoint, not contradictory — the discriminator is whether the label after the if-join is the RETURN label, i.e. whether EVERY arm returns.** + +**⚠ This also bounds a §164z verdict.** §164z's `func_80184944` entry reads *"No relocation occurred … gcc-2.7.2 has no block-reordering pass; describing plain source-order emission as the compiler 'physically relocating' a block will send the next agent hunting for a pass that doesn't exist."* That verdict is right **for `func_80184944`** (interior join ⇒ precondition unmet) and its refutation stands, but the pass is not imaginary: it is `jump.c:1806`, it splices insn chains, and `func_80187130` is the byte-proven instance. Read that sentence as "no *bb-reorder*" (gcc-3.x), not as "no block motion in 2.7.2". + +**THE DIAGNOSTIC TELL.** A **count-neutral** residual — *no* LENGTH-DRIFT, which is what separates it from §164-55's +1 — in which ONE conditional branch flips `beqz`↔`bnez` **and a 2-instruction `j ; ` pair changes sides of it**, sliding between the function tail and the slot immediately after the branch. `match_one` files it SHIFT-DRIFT / BRANCH-POLARITY and the index routes to §3-T4: **do not invert the condition** (the polarity you can see is `invert_jump`'s output and carries no source information, cf. §165-23) and **do not add a fence** (`jump.c` runs long before `reorg`, §136g-1). Swap which arm is the `if` body: the target's no-jump, epilogue-adjacent arm goes INSIDE the `if`, and the condition is spelled so the branch tests it TRUE. + +**Cross-link — the disarm direction.** §136g-1 needs this swap NOT to fire and buys that by putting any label between the if-join and the return label (wrap the body in an `if` with ONE trailing `return`), which also breaks `LABEL_NUSES (label1) == 1`. Same pass, opposite goal — read the target's block order and match it. + +*(SHARPENS — bounds §164-55 (L13035), §164-47 (L12851), §3-T4 (L90-103) and §164z's func_80184944 "no relocation" verdict (L13677); completes §136g-1 (L9214); evidence: byte-probed (5 variants, cc1 `.s` + relocation-masked `.o`, re-run at vetting) + source-read `jump.c:1800-1875`; from `func_80187130`)* + +*(SHARPENS — sharpens §49 (L3528-3576, THE LUID DIAL) — incl. its Method note: '-dS -dR ... the ready lists, the computed priorities, and the chosen order ... if the priorities are equal, you ; evidence: byte-probed; from `func_801874C0`)* + +**§167-28 — WHEN THE TARGET PICKS A *LOWER*-PRIORITY INSN THAN YOUR DRAFT, PRIORITY IS NOT THE DIAL: THE HIGHER-PRIORITY INSN WAS NOT READY, SO YOUR DRAFT IS MISSING A DEPENDENCE EDGE.** *(sharpens §49's Method note, which reads the ready lists only for the EQUAL-priority cell and prescribes the LUID dial; and the `gcc-2.7.2-map/sched.md` §4 recipe, whose ladder is "equal-pri ⇒ S1 / pri differs via a load-or-mul chain ⇒ S3 = intrinsic, route to the permuter" — this is the missing third rung, and it reaches the opposite verdict.)* + +**THE LAW.** `schedule_block` traverses each bb **backward** (`sched.c:3144`, *"we are traversing the instructions backwards"*), so *picked earlier* = *placed later*. READY = every **successor** already scheduled (`schedule_insn`, `:2557`). `rank_for_schedule` (`:2385`) returns `INSN_PRIORITY(y) − INSN_PRIORITY(x)` **first**, so the class rule and the LUID tie-break never run across a priority difference. And `priority()` (`:1425`) is `max(1, priority(pred) + insn_cost − 1)` over LOG_LINKS — **a C edit can only ever RAISE a priority, never lower one** (`sched.md` §1.3). Therefore: *a target that places an insn your draft's ready list ranked BELOW another one cannot be explained by any priority you can reach from C.* The higher-priority candidate was **absent from the ready list** — it still owed a successor. Re-read the residual as an alias/dependence question, not as a scheduling tie. + +**⚠ BOUND — rule out the one other way a high-priority insn gets passed over.** `schedule_select` (`sched.c:2616-2650`) works the ready list in **equal-priority groups**: it `queue_insn`s every member of the top group that `actual_hazard` blocks, and if the whole group is queued it falls through to the **next, lower-priority group**. That is a lower-priority pick with the higher-priority insn present *and* ready. The dumps separate the two cases: **`;; blocking insn N for K cycles` at that tick ⇒ §165-45's memory-unit story, not this one; no such line ⇒ non-readiness.** + +**THE READING, step by step** (worked on `func_801874C0`, ov_SC03_014, 241 ins). At the step after `sh ,0x18($sp)` the matching build picks `addiu $a0,$sp,0x10` (priority 1) over `sh 0xA($s0)` (priority 2 — the `lhu 0xDE` load edge into the `subu` costs 2). Backward ⇒ the `addiu` is *placed last*, and in the matching object it is (`w_pA`: `sh 0xA` → `lhu 0x6` → `addiu $a0,$sp,0x10`); in the 8-mismatch draft it sits four insns earlier (`w_v6`). `sh 0xA($s0)` can only have been unready if something it must precede was still unscheduled, and the only insn between them is `lhu 0x6($s0)` ⇒ the original carried a store→load edge at `0xA`/`0x6` that our alias oracle refuses ⇒ §16Z-vol. + +**⚠ THE INSTRUMENT IS NOT NEW — do not re-introduce it.** §49's Method note (`cc1 -dS -dR` → the `.sched`/`.sched2` traces: *"the ready lists, the computed priorities, and the chosen order"*), `sched.md` §3 *Dump tells* and the cookbook at L12914 already name the `;; ready list at T-N` lines, the `INSN_PRIORITY` column, `(7f000001)` = birthing boost, `;; blocking insn N for K cycles` and `;; insn N has a greater potential hazard`. **What is new is the inference across a priority difference**, and the verdict flip it produces: `sched.md` §4 currently sends the unequal-priority case to the permuter as intrinsic. + +*(SHARPENS — sharpens §49 (Method note), `gcc-2.7.2-map/sched.md` §4 recipe and §S3, §165-45, §25; evidence: byte-probed on one function (the `w_v6`/`w_pA` object pair), mechanism source-cited to the pinned `sched.c`; from `func_801874C0`)* + +*(SHARPENS — sharpens §136g rule 15 (L8921-8925) — split an RMW into `v = *p + 1; … *p = v;` so its load RISES, §165-06 / §164-XX (L13829-13876) — `(*p)++` vs `*p += 1`, the seven-spelling tab; evidence: byte-probed; from `func_801874C0`)* + +**§167-29 — A MEMORY RMW WHOSE RESULT IS A CALL ARGUMENT WANTS THE *HYBRID*: RE-READ THE FIELD AT EVERY USE SITE, AND KEEP THE RESULT IN A LOCAL.** *(sharpens §136g rule 15, which splits an RMW into `v = *p + 1; … *p = v;` precisely to make its load RISE — this is the same dial at the opposite polarity, for a target whose load sits mid-block; and §165-06/§164-XX, whose seven-spelling table is about the STORE's `SET_SRC` and never about where the LOAD lands. §49 supplies the argument half. **The drafter's citation of §162b/§76 is wrong — those are scope→allocno/`combine` levers and do not reach load placement.**)* + +Target shape — ONE narrow load mid-block, an if/else that bumps the field by two different constants, and the result passed to a call: + + lh $v0,0xDC($s0) <- one load, MID-block, not at its head + slti … ; beqz … ; nop <- the compare's delay slot is a real nop + addiu $v0,$v0,0x30 / 0x50 + sll $a0,$v0,16 ; sra $a0,$a0,16 ; jal … ; sh $v0,0xDC($s0) + +**THE LAW, as a three-cell decision** (all measured on `func_801874C0`, ov_SC03_014, pinned triple): + +| spelling | result | +|---|---| +| **hoist to a block-head local** — `s16 ang = *(s16*)(a0+0xDC); if (ang<0x400) ang+=0x30; else ang+=0x50;` | 240 ins, **201 mismatched**. The `lh` now has no LOG_LINKS predecessor ⇒ `priority()` = 1 ⇒ under backward traversal it is picked LAST and **placed FIRST**; the `sw 0xC` then falls into the `beqz` delay slot the target leaves `nop`, and the whole block cascades. | +| **fully in memory** — `*(s16*)(a0+0xDC) += 0x30;` + `f(*(s16*)(a0+0xDC))` | 241 ins, **12**. The load lands correctly, but the argument becomes `sh; lh` (store then re-load) where the target has `sll; sra` off the register — and the `sh` takes the `jal` delay slot. | +| **HYBRID** — read the field in the condition **and** in each arm, assign the sum to a local, store the local | 241 ins, **8**. cse commons the three reads into the one `lh` at the FIRST read site (= the target's position), and the local keeps the argument a register sign-extension. | + +**DIAGNOSTIC TELL — two symptoms that pull in opposite directions.** +* (a) A `nop` in a compare's delay slot that your draft fills, with a single narrow load at the **block head** that the target has **mid-block** ⇒ you hoisted the RMW's load into a local. Delete the local; re-read the field at each use site. +* (b) `sh` in a `jal` delay slot followed by an `lh` of the same field, where the target has `sll; sra` ⇒ you left the result in memory. Bind it to a local (§49; at the lvalue's own width, §165-06). + +**This shape needs BOTH fixes at once**, which is exactly why the single-axis sweeps stall: §136g-15's hoist alone gives you (a), §165-06's operator table alone gives you (b). + +*(SHARPENS — sharpens §136g rule 15, §165-06/§164-XX, §49; evidence: byte-probed, three measured cells on one function (variants under `.run/wave4/func_801874C0/`); the 201/12 cells were **not** re-run at vet time; from `func_801874C0`)* + +*(SHARPENS — sharpens L1808-1815 register-resident `short` truthiness bullet (sll 16 + branch; contrasts s16 vs int ONLY — no u16 arm), §164-30 (L12507) `sll $aN,$aN,16` in place at a truthine; evidence: byte-probed; from `func_80187F2C`)* + +**§167-30 — Field a0+0xFE must be declared s16, not u16, even though the machine load is `lhu`: MIPS LOAD_EXTEND_OP==ZERO_** + +**§16?-NN — THE MEMORY-FIELD SIGNEDNESS ARM: `--field == 0` EMITS `sll 16` FOR `s16` AND `andi 0xffff` FOR `u16`, AND THE MISS ALSO PERMUTES THE LOAD SCHEDULE.** *(sharpens §165-11 / §136 type-form rule 12 (L8904), which state the signed/unsigned narrowing pair only on the RETURN axis; sharpens L1808's register-resident-`short` truthiness bullet, which contrasts `s16` against `int` and has no `u16` arm; sharpens §164-30, which owns the lone-`sll 16` shape for a PARAMETER; and BOUNDS §164z's `func_80180128` refutation — 'the byte gate is BLIND to the signedness of that field' — by naming the precondition under which it is NOT blind.)* + +**Target shape** — an in-place decrement of a narrow field, tested against zero: + + lw $v1,0xC($s0) + lhu $a0,0xFE($s0) <- `lhu` EITHER WAY (§136g: LOAD_EXTEND_OP==ZERO_EXTEND, mips.h:1163) + addu $v1,$v1,$v0 + addiu $a0,$a0,-1 + sh $a0,0xFE($s0) + sll $a0,$a0,0x10 <- the ONLY signedness carrier. NO `sra`. + bnez $a0,… + +**THE LAW.** The load opcode carries **zero** signedness information for a HImode field — `lhu` is emitted for `s16` and `u16` alike. The whole declaration is readable at the surviving HImode→SImode use, and for a *zero-equality* test it is one instruction either way: + + *(s16 *)(p+K) -> sll $r,$r,16 (`sra` deleted: `simplify_comparison` case ASHIFTRT, combine.c:7422-7433 — + "If this is an equality comparison with zero, we can do this as a logical shift") + *(u16 *)(p+K) -> andi $r,$r,0xffff (the high bits are load-bearing; nothing collapses it) + +**And the miss is not opcode-only.** The `u16` spelling also hoists the `lhu` above the neighbouring `lw` and swaps the two `addu`s with it — **~5 mismatched instructions at LENGTH-DRIFT 0**, which reads like a §135-4/§49 scheduling residual and is not one. + +**BYTE EVIDENCE** (`func_80187F2C`, ov_SC03_118, 56 ins, banked at `src/ov_SC03_118/ov_SC03_118_jr_801863CC.c:3937`). Target read off the SHA1-green object, `mipsel-linux-gnu-objdump -d build/ov_SC03_118/ov_SC03_118.elf` @ `0x80187F80-0x80187F98`. Controlled A/B on the pinned triple, the two files differing in exactly one character sequence (`.run/wave4/func_80187F2C/experiments/v4.c` = `--*(u16 *)(a0+0xFE)`, `v5.c` = `--*(s16 *)(a0+0xFE)`), 110 cc1 lines each: + + v5 (s16, = target): lw $3,12($16) ; lhu $4,254($16) ; addu $3,$3,$2 ; addu $4,$4,-1 ; sh $4,254($16) ; sll $4,$4,16 + v4 (u16): lhu $4,254($16) ; lw $3,12($16) ; addu $4,$4,-1 ; addu $3,$3,$2 ; sh $4,254($16) ; andi $4,$4,0xffff + +**DIAGNOSTIC TELL.** `lhu ; addiu -1 ; sh ; sll 16 ; beqz/bnez` on a FIELD ⇒ write the lvalue `*(s16 *)`. The same run ending `andi $x,$x,0xffff` ⇒ `*(u16 *)`. Do not read the `lhu` — it is the same in both. **Bound:** the missing `sra` is the *equality-with-zero* rule; a target that shows `sll 16 ; sra 16` before a signed relational compare is the same `s16` field with a non-equality test (§164-37/§163b), not a different width. And §164z's blindness result still holds for the shape it was measured on — a field that is only added to and stored back with `sh` has no surviving HImode→SImode use, so no spelling is observable. + +*(SHARPENS — sharpens §165-11 (L13956), §136 rule 12 (L8904), L1808, §164-30 (L12507); bounds §164z `func_80180128` (L13680); evidence: byte-probed by the vet (objdump + pinned cpp|cc1 A/B); from `func_80187F2C`)* + +*(SHARPENS — sharpens §164-56 (L13067-13086) — THE LAW's 'the subtract and compare happen in the OPERAND'S OWN MODE, so a `short` operand emits `andi 0xffff` + `sltiu` and an `int` operand emi; evidence: byte-probed; from `func_80188348`)* + +**§167-31 — THE `range_test` MASK IS A FUNCTION OF THE BIAS SIGN, NOT OF THE OPERAND'S WIDTH. A `short` OPERAND WITH `LO ≥ 0` FOLDS TO A BARE `addu −LO ; sltu HI−LO`, AND §164-56'S "COUNT THE `andi 0xffff`" TELL NEVER FIRES.** + +Target shape — a `short` counter that is BOTH range-tested and passed as a call argument: + + lh $v1,0x70($s0) <- SIGNED load + slt $v0,$v1,8 ; bne $v0,$0,.Lskip + slt $v0,$v1,14 ; beq $v0,$0,.Lskip <- two independent signed `slt`; no bias, no mask + +Yours, from `if (cnt >= 8 && cnt < 14)`: + + lhu $a1,0x70($s0) <- load flipped to UNSIGNED + addu $v0,$a1,-8 ; sltu $v0,$v0,6 <- the fold; NO `andi 0xffff` anywhere + beq $v0,$0,.L ; sll $a1,$a1,16 + li $a0,0x1EA ; sra $a1,$a1,16 <- the sign-extension REMATERIALISED for the call arg + +**THE LAW (two halves; the first bounds §164-56).** +1. **The mask is conditional on the sign of `LO`, not on the operand's mode.** `range_test` does rewrite `VAR >= LO && VAR < HI` to `(unsigned TYPEOF(VAR))(VAR − LO) < (HI − LO)` in the operand's own mode (§164-56) — but combine then DELETES the HImode truncation whenever it cannot change the answer. With **`LO < 0`** the bias is a positive `addu`: `x + |LO|` carries above 0xFFFF for `x` near 0xFFFF and the 32-bit result disagrees with the HImode one on the in-range side ⇒ the `andi 0xffff` is load-bearing and survives. With **`LO ≥ 0`** the bias is a negative `addu`: every `x < LO` wraps to a huge unsigned that is still ≥ `HI − LO` ⇒ the truncation is provably redundant and is removed. **A `short` operand with a non-negative low bound therefore emits exactly the same two instructions as an `int` operand.** +2. **On a `short` operand the fold is visible in the LOAD and in a rematerialised extension instead.** The fold wants the value zero-extended, so `lh` becomes `lhu`; any OTHER use of the same value that needs it signed (a call argument, a narrow store) then pays a fresh `sll 16 ; sra 16` pair. That pair, not the `andi`, is the tell in this family, and it costs **+1 instruction** over the split form. + +**BYTE EVIDENCE** — six-way isolated A/B on the pinned triple (`-quiet -O2 -G0 -mips1 -mcpu=3000 -mgas -msoft-float`), same body both sides: + +| operand | bounds | spelling | compare region emitted | +|---|---|---|---| +| `short` local | `>= 8 && < 14` | `&&` | `lhu ; addu −8 ; sltu 6` — **no mask**, 21 ins | +| `short` local | `>= 8 && < 14` | nested `if`s | `lh ; slt 8 ; slt 14` — **the target**, 20 ins | +| `int` local | `>= 8 && < 14` | `&&` | `lh ; addu −8 ; sltu 6` — no mask, 19 ins | +| `short` local | `>= −0x101 && < 0xF2` | `&&` | `lhu ; addu 257 ; **andi 0xffff** ; sltu 499` | +| `short` local | `< −0x101 \|\| >= 0xF2` | `\|\|` | identical + `bne` — §164-56's own shape | +| `int` local | `>= −0x101 && < 0xF2` | `&&` | `lh ; addu 257 ; sltu 499` — no mask | + +The mask tracks the **bias sign** in every row; the declared width tracks only `lh` vs `lhu`. §164-56's evidence (`func_80185EF8`, biases +0x101 / +0x4AA) sits entirely in the `LO < 0` half — that is why it read the mask as the mode's signature. + +**⚠ THE SPLIT FORM DOES NOT DECLARE THE WIDTH.** `short cnt` and `int cnt` compile byte-identically once the chain is split (20 ins, same instructions). The `lh`/`lhu` distinction exists ONLY in the folded form; do not read the target's `lh` as evidence for an `s16` local here. + +**THE DIAGNOSTIC TELL — use this one when `LO ≥ 0`; §164-56's does not apply.** You are **LENGTH-DRIFT +1**, the target has **two `slt`/`slti` on one register**, and you have **one `addu −K ; sltu M`**. There is no `andi` to count. Confirm on the LOAD: target `lh`, yours `lhu`, and yours grows an `sll 16 ; sra 16` pair on that same register straddling the branch. Fix is §164-56's — nested `if`s (§165-22's lever; reach for §164-19(b)'s `goto` ladder only if a 0/1 flag with a body is involved). + +**⚠ DO NOT CONFUSE WITH §164-37 / §163b.** Both shapes contain `addu −K`, `sll 16 ; sra 16` and an `slt*` on one register. **Position discriminates them absolutely:** the switch dispatch is `addiu −MINVAL ; sll 16 ; sra 16 ; sltiu` — extension **between** the subtract and the compare, one value with one use. The `range_test` residual is `addu −LO ; sltu HI−LO ; … sll 16 ; sra 16` — extension **after** the compare, because it belongs to a DIFFERENT use of the value. `sll/sra` after the `sltu` ⇒ this section, and no switch is involved. + +*Provenance:* `func_80188348` (ov_SC03_118, 52 ins, S48 wave-6 `confirmed`, draft `.run/wave6/func_80188348/func_80188348.c:21-29`, sha1 `8effd252f70f5d475300df0839e6db4b55e3e388`) supplied the shape; the six-row A/B was run at vetting time on the pinned cc1 and is reproducible from that body. §164-56's core law is unchanged — only its mask rule and its tell are bounded. + +*(SHARPENS — sharpens §164-02 (L11888) — states the precondition as 'every register loaded with `la sym`' but never says what happens one add downstream, §164-01 (L11871) — the pointer_int_sum; evidence: byte-probed; from `func_8018C3F8`)* + +**§167-32 — `qty_const` DECAYS AFTER ONE ADD: A SYMBOL ADDRESS THAT ALREADY CARRIES A RUNTIME TERM IS AN ORDINARY REGISTER, AND FIX A1 GOVERNS IT.** *(BOUNDS §164-02; supplies the discriminator against §164-01. Retires the 'a spill/reload kills the constant equivalence' story offered from `func_8018C3F8` — see below.)* + +Target shape — **two `addu`s off the same symbol in one function, with opposite operand orders**: + + lhu $2, D_800B9A02($2) ; runtime index + la $3, D_800A6610 ; BARE symbol -> CONSTANT_P + sll $2, $2, 14 + addu $2, $2, $3 <- INDEX first (cse swapped it; §164-02) + sw $2, 0x30($sp) + ... + lw $10, 0x30($sp) ; same value, ONE add downstream + sll $17, $17, 2 + addu $17, $10, $17 <- BASE first (no swap; plain source order) + +**THE LAW.** §164-02's precondition is `qty_const`, and `insert` records it only from a class member whose `elt->is_const` is true — `CONSTANT_P (x) || (RTX_UNCHANGING_P (x) && REG && REGNO >= FIRST_PSEUDO_REGISTER) || FIXED_BASE_PLUS_P (x)` (`cse.c:1313-1319`; `FIXED_BASE_PLUS_P`, `cse.c:575-586`, is frame/arg-pointer only). A pseudo set from `(plus (reg ) (reg ))` has a class holding exactly that PLUS — none of the three. The second arm of `insert` (`cse.c:1383-1397`) scans the class for `p->is_const && GET_CODE (p->exp) != REG`, finds nothing, and leaves `qty_const` unset; `equiv_constant` (`cse.c:5699-5710`) then returns 0 and the commutative swap at `cse.c:5282-5289` never fires. **The equivalence dies at the first add.** `p = &SYM[i];` makes `p` an ordinary register, and the next `p + x` is §10-Residual-A territory: write the operand you want in `rs` FIRST. + +**IT IS NOT THE SPILL — and it can never be.** `cse_main` runs at `toplev.c:2865` (cse1) and `:2926` (cse2); `reload` runs inside `global_alloc` at `:3080`, with `reload_completed = 1` at `:3096`. `grep -rn reload_cse_regs tools/reference/gcc-2.7.2/` ⇒ **0 hits** — 2.7.2 has no post-reload cse. Operand order is frozen before a stack slot exists, so **no reload-time observation (`lw 0x30($sp)`, register pressure, 9 busy callee-saveds) can ever explain a cse-time operand swap.** If your story about an `addu`'s operand order mentions a spill, the story is wrong. + +**BYTE EVIDENCE — one function, one symbol, both outcomes.** `func_8018C3F8` (ov_SC06_018, 236 ins, MATCH, banked; `src/ov_SC06_018/ov_SC06_018_jr_80187AEC.c:4854`). In the shipped `build/ov_SC06_018/ov_SC06_018.elf`: +* `8018c464: addu v0,v0,v1` — v0 = `sll` index, v1 = `la D_800A6610`. **Index first**, although the C at `:4877` is written pointer-first (`&D_800A6610[(*(u16*)&D_800B9A02) << 14]`, which `pointer_int_sum` also canonicalises pointer-first per §164-01). Both source-level orderings lose: this is §164-02 firing, in the shipped bytes. +* `8018c554: addu s1,t2,s1` — t2 = `lw 0x30($sp)` (that same value), s1 = `z << 2`. **Base first**, matching the C at `:4897`, `(s32) ot + (z << 2)`. The drafter's A/B is the dial: `(z << 2) + (s32) ot` gave `addu $s1,$s1,$t2`, the one and only mismatch in a 236/236 body; swapping the addends made it MATCH. Two builds, one character. +* Sibling control, same TU: `func_8018F694` (`:6776`) spells the second add the OTHER way, `(z << 2) + (s32) ot` (`:6810`), and gets `8018f7f0: addu s0,s0,a3` — index first — while **its `ot` is stack-resident too** (`8018f7d8: lw a3,0x38($sp)`). Both functions spilled, both obeying source order, opposite results. Spilling is not the variable. + +**THE DIAGNOSTIC TELL.** Before reaching for §164-02's self-set asm launder, ask **how the base got into the register**, not where it currently lives: +* `la SYM` materialised in the same expression ⇒ §164-02. Source order and pins are inert; the `__asm__("" : "=r"(e) : "0"(e))` launder is the only lever. +* A pointer local **assigned** a symbol-plus-runtime-index earlier ⇒ ordinary register. **Try Fix A1 first — it is one character and costs one build.** +* And spell that second add in `s32`, not pointer, arithmetic. `(s32) base + off` keeps it a plain binop where source order reaches RTL; `ptr + off` funnels through `pointer_int_sum` and §164-01 pins the pointer to operand 0 before RTL, making Fix A1 inert for a second, unrelated reason. + +*(SHARPENS — bounds §164-02, discriminates against §164-01; evidence: byte-probed in the shipped object, 2 functions + a 2-build A/B; from `func_8018C3F8`)* + +*(SHARPENS — sharpens §164-08 (L12007) — the ASCENDING N-site form: `&SYM[k*STRIDE]` off one array symbol, tells `LENGTH-DRIFT/+2` (N symbols) and `-5` (walked pointer), §136 rule 5 (L8872) — ; evidence: byte-probed; from `func_8018D654`)* + +**§167-33 — `use_related_value` ALSO RUNS DESCENDING, AND AT TWO SITES THE MISTAKE IS LENGTH-NEUTRAL: §164-08's LENGTH-DRIFT TELL CANNOT FIRE.** *(BOUNDS §164-08 (L12007), whose law, both refutations and both tells are stated for N ASCENDING sites off one array symbol; and resolves its "Tension to know about" §136-5 note in the OPPOSITE direction — at two offsets the pointer local is REQUIRED, not refuted.)* + +**Target shape** — ONE `%hi/%lo` pair built for the FAR offset, a zero-displacement load off it, and the plain `&SYM` later reached by a single NEGATIVE bump: + + 8018d698 lui $s2,%hi(D_80126B64) ; = D_80126B58 + 0xC + 8018d69c addiu $s2,$s2,%lo(D_80126B64) + 8018d6a4 lw $a0,0x0($s2) <- the far read: THREE instructions + ... + 8018d738 addiu $a0,$s2,-0xC <- &D_80126B58: ONE instruction + +**THE LAW.** `use_related_value` (`cse.c:1781`, called at `:6535`) is symmetric in the sign of the offset: with `SYM+K` already live in a pseudo, cse rewrites a later plain `SYM` as `reg − K`. Two preconditions, and both invert §164-08's ascending case: +(a) **the FAR offset must go through a POINTER LOCAL** so the full address lands in a pseudo (§136-5, L8872). Spelled as a bare read of the higher symbol it is `(mem (const (plus SYM K)))`, which this MIPS port accepts as a legal address and folds to `lui/%lo` — the register never exists and cse has nothing to relate. (§162l, L11482, is the same fold read from the MEM side.) +(b) **the near address must be spelled as the same symbol** (`&SYM`), or the related_value chain does not link. + +**WHY IT IS INVISIBLE — the drift cancels at N=2.** Spelling the far offset as its own extern (`extern s32 D_80126B64;` — exactly what splat's symbol map hands you) costs **−1** on the read (`lui/lw %lo` instead of `lui/addiu/lw 0(reg)`) and **+1** on the address (`lui/addiu` instead of one `addiu`). **Net 0.** §164-08's tells — `+2` for N separate symbols, `−5` for a walked pointer — are counts over N≥3 ascending sites; at two sites, one read and one address, they sum to zero. The gate reports the right length with a scattered OPCODE-MIXED residual and nothing that names the mistake. + +**BYTE EVIDENCE.** `func_8018D654` (ov_SC06_018 / `_jr_80187AEC`, 135 ins, banked MATCH 135/135; `src/ov_SC06_018/ov_SC06_018_jr_80187AEC.c:5411`). The target's own bytes settle it without an A/B — `8018d698-8018d6a4` spends three instructions reading `D_80126B64`, `8018d738` spends one on `&D_80126B58`; the bare-extern spelling cannot produce the first. + +**DIAGNOSTIC TELL.** A bare `addiu $rD,$rS,-K` feeding a call argument or an address, where `$rS` was built by a `lui/addiu` pair of a symbol exactly K bytes HIGHER and is read at displacement 0 ⇒ ONE symbol, the far offset in a pointer local, the near one as `&SYM`. **Do not read the length — this class has none.** Read the address-instruction COUNT instead: a symbol whose read costs three instructions lives in a pointer local; two means it is a bare global. + +*Companion, already law — do not re-derive:* each cse REGION needs its OWN pointer local. The target rematerialises the same base at `8018d77c` (`lui/addiu $a1`) because `.L8018D77C` carries four jump refs and starts a fresh cse table (§164-52, L12943); one shared C variable would stay live in `$s2` across the join. That is §44-Lever-3 (L3263) / §76 / §136-1 / §162b1's split-vs-share law applied to an address. + +*(SHARPENS — bounds §164-08; evidence: byte-probed (banked MATCH + target read); from `func_8018D654`.)* +*Honest scope: the isolated single-variable A/B on the spelling was never run — the only bare-extern draft (`.run/wave6/func_8018D654/v1.c`) also carried the wrong struct layout, so its "135 ins, OPCODE-MIXED" number is confounded. The length-neutrality above is arithmetic plus the target read, not a gated pair. Two minutes to close it: respell the banked body's two `lim` locals as `extern s32 D_80126B64;` and gate.* + +*(SHARPENS — sharpens §49 (L3527) — THE LUID DIAL: materialise a temp to shift an insn's expand-stream position and split a `rank_for_schedule` tie (stated for a `sll/sra` sign-extension pair,; evidence: byte-probed; from `func_8018D654`)* + +**§167-34 — A NAMED TEMP IS A *LOAD-ORDER* DIAL: it moves the load out of the expression and ABOVE the operand it was subtracted from — and NO statement permutation substitutes.** *(sharpens §49 (L3527, the LUID dial — a temp shifts an insn's expand-stream position; stated for a `sll/sra` pair moving EARLIER) and §165-14 (L14021, "source order IS the schedule" for three mutually-independent STATEMENTS); **bounds §164-43** (L12771), whose law is that splitting an expression into named temps changes the pre-reload stream "with **zero** change to the final instruction stream" — the COUNT is unchanged, the ORDER is not.)* + +**Target shape** — a block whose leading load group is a pure PERMUTATION of yours, same registers, same count: + + 8018d6e8 lh $a0,0x6($s1) <- self->f6 + 8018d6ec lui $v0,%hi(D_80126B62) + 8018d6f0 lhu $v0,%lo(D_80126B62)($v0) + 8018d6f4 lui $v1,%hi(D_80126B5E) + 8018d6f8 lh $v1,%lo(D_80126B5E)($v1) + +**THE LAW.** The loads are mutually independent, so `rank_for_schedule` ties on priority and on every dependence class and falls through to `INSN_LUID (tmp) - INSN_LUID (tmp2)` (`sched.c:2425-2428` — the fallback §165-14 cites). LUID is expand-stream position, and **a global's load is expanded at its FIRST source mention.** Hoisting the global into a temp of its own moves that mention from inside the expression to a statement ABOVE it, carrying the load above the operand it was going to be subtracted from: + + bx = *(s16*)&D_80126B5E; i = self->f6 - bx; + RTL lh B5E, lh f6, lhu B62 -> sched1 B5E, B62, f6 WRONG + + i = self->f6 - *(s16*)&D_80126B5E; v[0] = *(s16*)&D_80126B5E; + RTL lh f6, lh B5E, lhu B62 -> sched1 f6, B62, B5E TARGET + +Both emit the same 135 instructions in the same registers. Inlining also merges the two reads into the single `lh` the target has — cse commons them. + +**THE NEGATIVE CONTROL IS THE POINT.** With the temp present, **every** legal statement order was gated: 240 permutations of the six statements (`.run/wave6/func_8018D654/perm/res.txt`), floor **closeness 3** (`BI1320`, `BI3120`), never 0; 144 more on a second source shape (`perm2/`, all 137 ins) and 72 on a third (`perm3/`). Statement order cannot reach it, because the temp's assignment must precede its use — **the dependence you are trying to break is the one the temp created.** + +**DIAGNOSTIC TELL.** The residual is confined to the first N instructions of one block, is a pure PERMUTATION (same opcodes, same registers, `nins` exact), and does not move under statement reordering ⇒ **delete the temp and inline the global into the expression** (or, in the other direction, hoist it into a temp to move its load UP). Do not open the permuter, a pin, an `asm` barrier or §47/§158's allocno sliders: those steer a schedule TIE, and this is upstream of the tie, in the RTL you handed sched1. + +**BYTE EVIDENCE.** `func_8018D654` (ov_SC06_018, 135 ins, banked MATCH; `src/ov_SC06_018/ov_SC06_018_jr_80187AEC.c:5411`). Ladder 116 → 11 → 5 → 3 → 2 → 0; the 3 → 0 step is exactly the temp deletion. + +*(SHARPENS — sharpens §49, §165-14; bounds §164-43's "zero change to the final instruction stream"; evidence: byte-probed, 456 gated variants + the banked MATCH, one function; from `func_8018D654`.)* +*⚠ Wording corrected at vet time: the source note says "the post-reload schedule preserves the group order it receives", but **sched1 itself permutes** — RTL `f6,B5E,B62` comes out `f6,B62,B5E`. The sound statement is that the RTL order SEEDS the LUID tie-break, not that a pass preserves it. The `-fno-schedule-insns2` dumps quoted verify sched1's output only.* + +*(SHARPENS — sharpens §165-04 (L14622-14638), §165-02, §165z; evidence: single-instance; from `func_8017C294`)* + +**§167-35 — §165-04 ('a zero-emission =r/0 re-tie on a memory-loaded s16 buys +8 bytes of frame per site') does NOT replic** + +**⚠ §165-04 CONTESTED BY TWO LATER PROBES ON ITS OWN EXEMPLAR (P30 S48, waves 5 and 6).** §165-04's entire byte record is one A/B on `func_8017C294`'s pointer-walk body (`vars 240 → 256` for two re-tie sites on memory-loaded `s16`s). Two independent agents re-ran the instrument on that same function later the same day and measured the **opposite sign**: wave 5 got `vars 240 → 232 → 224` *with instructions added* (`.run/wave4/func_8017C294/var/r_a*.c`, `var/z3_retie.c`), and wave 6 independently reproduced the failure across **16** self-cast sites (`.run/wave6/func_8017C294/var/`). Neither ran §165-04's own prescribed 4-line reproducer, so the entry is **contested, not refuted** — but do not cite '+8 bytes of frame per re-tie site' as a dial, and do not spend a wave steering with it. The two laws it leans on are unaffected and both still stand: §165-02's memory precondition, and §165-YY's *post-`life_analysis` deletion* condition — which predicts that whether the re-tie buys a slot depends on **which pass** ends up deleting the widening, and therefore that its sign is body-dependent. **Retest is still the 4-line reproducer, ~2 minutes; do that before either banking or retiring it.** + +*(SHARPENS — sharpens §165-36 rule 1 (L14639-14646), §160g, §71; evidence: asserted; from `func_8017C294`)* + +**§167-36 — WAVE STEP 0, PART 3: INVENTORY YOUR OWN OUTPUT DIR *BEFORE* THE FIRST WRITE.** *(§165-36 rule 1 sends you to `find .run` for drafts in OTHER waves' trees; the collision that actually destroys work is inside the directory the harness just handed you. PROCESS, not a compiler law.)* + +`.run/waveN//` is not empty just because this wave is new — the harness re-uses the per-function path, so an earlier wave's entire variant tree, **including its best drafts**, can already be sitting in it. On `func_8017C294` the wave-5 agent opened by overwriting the top-level `func_8017C294.c` with the wave-3 draft and only found the pre-existing wave-4 tree's **three 2-diff drafts in `fin/`** by accident at the end of the run; the `prior_notes.json` it was handed described the wave-3 **11-diff** state and never mentioned them. + +**THE RULE.** Before writing anything: `ls -la` the assigned dir, and `match_one`-batch **every** pre-existing `.c` in it (they are cheap and local — §12/§19). Never overwrite the top-level `.c`; write your first draft under a new name. **The handed-down notes are a summary of ONE earlier wave, not an inventory — the directory is the inventory.** + +**Companion.** §165-36 rule 2 still applies to whatever you find: carry forward its reasoning, never its coordinates. + +*(SHARPENS — sharpens §21 wave-distilled idioms, L1872-1881 — 'store a call result AND test/reuse it in one expression → combined assignment *(T*)(p+k) = local = f();' (evidence func_80142DC4,; evidence: single-instance; from `func_80180450`)* + +**§167-37 — WHEN A CALL RESULT CROSSES *ZERO* CALLS AND ITS ONLY NON-STORE USE IS THE NEXT CALL'S SOLE ARGUMENT, NAME NOTHING: STORE IT TO THE FIELD AND RE-READ THE FIELD TWICE.** *(BOUNDS §21's `*(T*)(p+k) = local = f();` bullet (L1872-1881) — its own exemplar `func_80142DC4` has the result surviving a LATER call, the precondition it never states, and following it here costs +1. Adds a FOURTH row to §164-20's delay-slot discriminator table, which routes "copy missing from a `jal` delay slot" to §43 alone. Runs §164-82's re-read lever in the OPPOSITE SIGN. The MECHANISM is not new — it is §48-A4's ARG-register copy preference plus §50-D; cite those, do not re-derive.)* + +Target shape — one call's result stored, null-guarded, and handed straight to the next call: + + jal func_8012C1B8 + addu $s0,$a0,$zero + bnez $v0,.L288 + sw $v0,0x20($s0) <- the STORE is in the GUARD's delay slot + … + .L288: + lui $a1,%hi(D_SYM) ; addiu $a1,$a1,%lo(D_SYM) + jal func_8001C214 + addu $a0,$v0,$zero <- the ARG copy is in the CALL's OWN delay slot + +**THE LAW.** `$v0` reaching the `sw`, the `bnez` and the second call's `$a0` with no intervening copy means **the source named nothing.** Write it with no local at all — + + *(s32 *)(p + 0x20) = f(); + if (*(s32 *)(p + 0x20) == 0) { g(p); } + else { h(*(s32 *)(p + 0x20), (s32)D_SYM); … } + +— cse commons both re-reads onto the call's own pseudo, so each site is a bare register use, and dbr gets both slots (the store into the `bnez`'s, the argument copy into the `jal`'s). Bind the result to a local instead and that pseudo becomes a **cross-block global allocno whose one copy use is `(set (reg $a0) P)`**: `find_reg` checks copy preferences FIRST (§50-D, `global.c:1000-1030`), and because the pseudo is defined *after* one call and dead *before* the next it crosses **zero** calls, so `allocno_calls_crossed > 0` never fires to strip the caller-saved prefs (§48-A4, `global.c:906`). It takes `$a0` at its definition: the copy collapses into a `move $a0,$v0` **above** the guard, the guard tests `$a0`, and the `jal` loses its slot occupant. Measured **+1 (63 vs 62)**. + +**⚠ THIS IS THE PRECONDITION §21's BULLET NEVER STATES.** §21 (L1872) prescribes `*(T*)(p+k) = local = f();` for this *same* asm shape, and its evidence `func_80142DC4` is the *same callee* (`func_8012C1B8`) at the *same offset* (`0x20`) — but there the success arm reuses the value **across later calls**, which forces it callee-saved (`addu $sX,$v0,$zero`) and is exactly why naming it is free. **Read the arm before choosing: a later `jal` between the guard and the last use ⇒ §21's named combined assignment. No later `jal` ⇒ this entry, name nothing.** + +**AND NOTE THE SIGN INVERSION vs §164-20 / §164-82.** There the re-read *buys* an instruction (a copy that survives into a **branch** delay slot, curing a `-1`) and the guarded value is a plain field **LOAD**. Here the same source edit *deletes* one (curing a `+1`) and the guarded value is a **CALL RESULT already stored to memory**. Same three-line shape, opposite direction — the discriminator is what defines the value. + +**BYTE EVIDENCE.** `func_80180450` (ov_SC06_008, **62 ins**, banked, `src/ov_SC06_008/ov_SC06_008_jr_8017C294.c:4576-4621`; family ×7). Target read off the byte-identical sibling `asm/ov_SC06_018/nonmatchings/ov_SC06_018_jr_8017C24C/func_8018025C.s` — 62 ins, `sw $v0,0x20($s0)` in the `bnez` slot (`80180270`/`80180274`), `addu $a0,$v0,$zero` in the `jal func_8001C214` slot (`80180290`/`80180294`). Named-local spelling: **63 ins**, with `move $a0,$v0` above the guard. + +**THE DIAGNOSTIC TELL — the fourth row of §164-20's table.** An argument copy **missing from a `jal`'s own delay slot**, `LENGTH-DRIFT +1`, and that same `move $aN,$v0` sitting **above the guard branch** with the guard testing `$aN` instead of `$v0` ⇒ **you named the call result.** Not §43 (that row is a narrow prototyped `s16` param and shows no `+1`); not §164-20 (that row is a **branch** slot and a `-1`); not a pin (§164-22 — there is no copy insn to place). Delete the local: store and re-read. + +*Scope, honestly: **n = 1, one A/B, and the decisive control is missing.** The note reports the losing build only as "the named-local form was 63 ins" and never says whether that was the plain two-statement `v = f(); *(p+K) = v;` or §21's combined `*(p+K) = v = f();` — and only the latter decides the bound on §21. The losing build was not re-measured at vet time. **One probe settles it:** compile `*(s32 *)(p+0x20) = v = f(); if (v == 0) … else h(v, (s32)D_SYM);` on this body and count. Until then treat the §21 bound as a PREDICTION and "name nothing" as the byte-proven half. Related caution: §165z byte-REFUTED a similarly-worded "unname the call result" claim on `func_80181E18`; that refutation does not reach this shape (no null-guard, no argument-register consumer), but do not widen this entry past its four preconditions — call result, stored to a field, null-guarded, sole remaining use = the very next call's only argument.* + +*(SHARPENS — bounds §21 L1872, §164-20, §164-82; mechanism cited from §48-A4 / §50-D; evidence: single-instance; from `func_80180450`)* + +*(SHARPENS — sharpens §163c (L11793-11800), §162a1 (L11030-11056), §162a2/§161a (L10968-10989), §165-28 (L14450); evidence: single-instance; from `func_801810C8`)* + +**§167-38 — AN *INTERIOR* JTBL SLOT EQUAL TO THE DEFAULT/JOIN LABEL IS AN ABSENT CASE, NOT PROOF OF A WRITTEN `case k:`.** *(bounds §163c, whose "`jtbl[k] == the default label` is the fingerprint of a case label sharing the default body, and that empty label is LOAD-BEARING" reads as unconditional; and supplies the row §162a1's three-line edge table does not have — it rules on entry[0], on entry[N-1] and on "a real body", never on an interior slot.)* + +**Target shape** — a dense-looking 6-word table with one interior word pointing at the same address the range check's `beqz` targets: + + 801810E4 sltiu $v0,$v1,0x6 + 801810E8 beqz $v0,.L801812DC <- OUT-OF-RANGE goes to 801812DC + ... + jtbl_801B2BD4 = [80181108, 80181114, 8018113C, 80181168, **801812DC**, 801812A0] + ^ slot 4 == the default target + +**THE LAW.** `expand_end_case` emits one word for EVERY value in `[minval, maxval]` and fills any value that has no case node with `default_label`. An interior slot pointing at the default is therefore *the default*, and carries **no information** about whether the source wrote a label there. The label §163c calls load-bearing is load-bearing only through its two side effects on the table's SHAPE: the live-case **COUNT** (vs `case_values_threshold` = 5 — §163c, §55a, §165-28) and **minval/maxval** (§162a1/§162a2). When the count already clears the threshold and the gap is strictly interior, the source can simply omit the case — and here it does. + +**BYTE EVIDENCE.** `func_801810C8` (ov_SC02_016, 209 ins, **first-pass MATCH 209/209, closeness 0, zero iterations**; draft `.run/wave6/func_801810C8/func_801810C8.c`, sha1 `faaf37c3…`; target `asm/ov_SC02_016/nonmatchings/ov_SC02_016_jr_8017DC70/func_801810C8.s:11-18`; table `asm/ov_SC02_016/data/tail18.data.s:21-28`). The banked switch writes case nodes **{0,1,2,3,5} and no `case 4:`** — 5 nodes, minval 0, maxval 5 ⇒ `sltiu 6` + a 6-word table, first try. Slot 4 holds `.L801812DC`, which is both the join every arm `j`s to and the target of the range check's own `beqz`. + +**DIAGNOSTIC TELL.** Before adding an empty `case k:` on §163c's advice, ask two questions about k. **Interior gap AND you already have ≥5 live case nodes ⇒ write nothing** — the slot fills itself. Add the label only when (a) the node count is ≤4, where per §163c the extra label is the only thing that reaches a tablejump at all, or (b) k is an EDGE, where per §162a1/§162a2 the label moves `minval`/`maxval` and therefore the `sltiu` bound and the word count. + +**⚠ Bounds.** The converse A/B was not run: `case 4: break;` was never compiled against this target, so this entry proves the OMISSION matches — not that the label is byte-inert. (Mechanically both spellings should land slot 4 on the same join, but that is inference, not a measurement.) One function. + +*(SHARPENS — sharpens §1-I3 (L54-58), §1-I4 (L60-65), §28a div-by-constant bullet (L2349), §164-04 / §1-I5 (L11934-11942); evidence: single-instance; from `func_801810C8`)* + +**§167-39 — THE *SIGNED* DIV-BY-CONSTANT SHAPE IS `mult ; sra 31 ; mfhi ; sra k ; subu`, AND A BRANCH AGAINST THE RECONSTRUCTED PRODUCT IS `x == x / K * K`, NOT `x % K == 0`.** *(extends §1-I3, which gives only the UNSIGNED magic — `lui 0xcccc ; ori 0xcccd ; multu ; mfhi ; srl 3`; and promotes §28a's imported decomp.wiki bullet "0x66666667→/10" from a "worth trying" list to a byte-witnessed instruction shape with a C spelling.)* + +**Target shape** — six instructions of division, then the source's own multiply back, then an ordinary `bne`: + + 80181168 lui $v0,0x6666 ; ┐ the SIGNED magic for /10 + 8018116C lw $a0,0x1C($s0) ; │ + 80181170 ori $v0,$v0,0x6667 ; │ 0x66666667 + 80181174 mult $a0,$v0 ; │ `mult`, not `multu` + 80181178 sra $v0,$a0,31 ; │ the DIVIDEND's sign word + 8018117C mfhi $t0 ; │ + 80181180 sra $v1,$t0,2 ; │ shift 2, not §1-I3's 3 + 80181184 subu $v1,$v1,$v0 ; ┘ q = (hi>>2) - (x>>31) + 80181188 sll $v0,$v1,2 ; ┐ synth_mult's ×10 — the SOURCE's, not the division's + 8018118C addu $v0,$v0,$v1 ; │ ((q<<2)+q)<<1 + 80181190 sll $v0,$v0,1 ; ┘ + 80181194 bne $a0,$v0,.L801811B4 <- the ORIGINAL dividend against the product + +**THE LAW.** Two halves. (1) For a **signed** operand gcc picks `mult` with the signed magic and corrects by SUBTRACTING the dividend's sign word (`sra $x,31`); that trailing `subu` of an `sra 31` is the one-glance fingerprint separating the signed form from §1-I3's `multu`/`srl` unsigned form, and the `mfhi` shift differs too (2 here vs §1-I3's 3 for the same divisor). (2) Everything AFTER the `subu` is the source's own arithmetic. A `sll/addu/sll` chain reconstructing `q*K` followed by a branch comparing that product against the **dividend** is a literal transcription of `x == x / K * K`. Write it that way. + +**BYTE EVIDENCE.** `func_801810C8` (ov_SC02_016, 209 ins, **first-pass MATCH 209/209, zero iterations**; `asm/ov_SC02_016/nonmatchings/ov_SC02_016_jr_8017DC70/func_801810C8.s:45-56`; banked C `if (*(s32 *)(a0 + 0x1C) == *(s32 *)(a0 + 0x1C) / 10 * 10)`, draft `.run/wave6/func_801810C8/func_801810C8.c`, sha1 `faaf37c3…`). + +**DIAGNOSTIC TELL.** Magic multiply ⇒ reconstructed product ⇒ branch comparing the product to the **dividend** ⇒ write `x == x / K * K` and do not "clean up" Ghidra's spelling. A `%` would compute `x − q*K` and branch on zero: one extra `subu` and a compare against `$zero` instead of against `$a0`. Read the compare's operands before choosing the operator. + +**⚠ Bounds.** The `%` half is a **prediction** — no `x % 10 == 0` variant was compiled against this target, and the whole function landed first-pass, so nothing here was ablated. The signed-shape half is byte-witnessed once. §1-I4's `divu`/`mflo`/`mfhi` entry is a different regime (runtime divisor) and does not apply. + +*(SHARPENS — sharpens §164-04 / §1-I5 (L11934-11942), §1-I3 (L54); evidence: single-instance; from `func_801810C8`)* + +**§167-40 — THE SIGNED `/2^k` BIAS LEAVES `addiu`'s IMMEDIATE RANGE AT k = 16: LOOK FOR `ori (2^k−1) ; addu`, AND EXPECT THE `sra k` TWICE.** *(bounds §164-04/§1-I5, whose tell — "a `bgez` whose only job is to skip an `addiu` of `2^k − 1` immediately before an `sra k`" — structurally cannot fire for k ≥ 16, so an agent applying it as written mis-reads the block as a mask.)* + +**Target shape** — six instructions, no `addiu`, and the shift on both paths: + + 80181258 bgez $v1,.L8018126C + 8018125C sra $v0,$v1,16 <- the UNBIASED shift, in the delay slot + 80181260 ori $v0,$zero,0xFFFF <- the bias 2^16 − 1, MATERIALISED + 80181264 addu $v1,$v1,$v0 + 80181268 sra $v0,$v1,16 <- the BIASED shift, second copy + .L8018126C: + 8018126C negu $v0,$v0 <- the source's unary minus + +**THE LAW.** gcc's round-toward-zero correction for signed `x / 2^k` adds `2^k − 1` on the negative path. `addiu` carries a **signed** 16-bit immediate, so that constant fits inside the instruction only for **k ≤ 15** (2^15 − 1 = 32767). At **k = 16** it is 65535, out of range, and must be materialised — `ori $rD,$zero,0xFFFF ; addu` — after which dbr fills the `bgez`'s now-empty delay slot with the **unbiased** `sra k`, so the shift appears once in the slot and once on the biased fall-through. The block grows from §1-I5's 3 instructions to 5, and the `addiu` §1-I5 tells you to look for is absent. **The C does not change: write the division literally.** + +**BYTE EVIDENCE.** `func_801810C8` (ov_SC02_016, 209 ins, **first-pass MATCH 209/209, zero iterations**; `asm/ov_SC02_016/nonmatchings/ov_SC02_016_jr_8017DC70/func_801810C8.s:108-114`). Banked C: `*(s16 *)(a0 + 0xFE) = -(D_801B42D0 / 0x10000);` (`.run/wave6/func_801810C8/func_801810C8.c`, sha1 `faaf37c3…`), with `extern s32 D_801B42D0;`. + +**DIAGNOSTIC TELL.** A `bgez` whose fall-through is `ori $rD,$zero,M ; addu` with **M = 2^k − 1**, plus an `sra k` on BOTH paths (one of them in the `bgez`'s own delay slot) ⇒ a signed `/ 2^k` with k ≥ 16. Do not read the `ori 0xFFFF` as a mask, do not respell the division as `>> k` (which loses the entire block — §1-I5), and do not reach for an unsigned divide (a bare `srl`, no branch). A leading `negu` on the result is the source's unary minus, not part of the division. + +**⚠ Bounds.** k = 16 measured, on one function, first-pass (no ablation). k ≥ 17 is unmeasured — the bias would need `lui`+`ori`, so expect 6 instructions, not 5. The k ≤ 15 half is §1-I5's own byte evidence and is unchanged. + +*(SHARPENS — sharpens §165-43 (L14773), §88c (L6695), §162e (record_jump_equiv for a loop index), §164-52 (equivalence lifetime); evidence: single-instance; from `func_80181948`)* + +**§167-41 — cse's TAKEN-EQUALITY CLASS REACHES A STORE'S *SOURCE*, NOT ONLY A COMPARE: `sw $a0,%lo(sym)($at)` inside an arm guarded by `beq $a0,<1>` is `sym = 1;`.** *(sharpens §165-43 (L14773), which states the same `record_jump_equiv` class only for a register-to-register `bne` and whose tell — "an `==`/`!=` compare against a **callee-saved** register inside a guarded arm" — cannot fire on a store; and adds the ARGUMENT-register case, where the false reading is the incoming PARAMETER rather than a local.)* + +**Target shape** — a case arm reached by an equality test on the switch value: + + beq $a0,$v0,.Lcase1 <- $v0 = 1; on this edge cse learns $a0 == 1 + … + .Lcase1: … + sw $a0,%lo(D_801270C8)($at) <- reads as "store the parameter"; it is `= 1` + +**THE LAW.** `record_jump_cond` (`tools/reference/gcc-2.7.2/cse.c:5839`, via `record_jump_equiv` `:5791`) merges the two operands of a taken `EQ` into ONE equivalence class for the rest of the path, and cse then emits the CHEAPEST member of that class at **every** use — a use being any operand, the compare §165-43 documents **or the SET_SRC of a store**. `D_801270C8 = 1;` therefore materialises no `li` at all and stores whichever register the dispatch already proved equal to 1. Per §164-52 the class survives until a label with >=2 jump references resets the table. + +**BYTE EVIDENCE.** `func_80181948` (ov_SC01_077, 132 ins, `match_one` MATCH; draft `.run/wave6/func_80181948/func_80181948.c:42`): `case 1:` opens `if (*(s32*)(a0 + 0x1C) == 0x14) { D_801270C8 = 1; }` and emits the target's `sw $a0`. *(One sighting, one-sided: no counter-spelling was compiled.)* + +**⚠ It is BYTE-NEUTRAL inside the arm** — `D_801270C8 = state;` emits the same store, because cse proved them equal. What the reading buys is not writing `D_801270C8 = param;`, which `$a0` invites and which does not match once the parameter is anywhere else. + +**THE DIAGNOSTIC TELL.** A store whose source register is an **argument** register that this arm's own guard compared against a constant. Before transcribing it as a variable, ask what the guard proved on this path: §88c says a materialised constant proves nothing and §165-43 says its absence proves nothing — that now covers store sources, so let the guard, not the register name, decide. + +*(SHARPENS — sharpens §160d (L10936) — THE ASYMMETRIC INDEX RELOAD; the only store-then-reload entry, but its observable is an `lbu`/`sll` COUNT and its consumer is a table index, never a comp; evidence: single-instance; from `func_80183834`)* + +**§167-42 — A SATURATING `u16` DECREMENT-AND-TEST READS THE FIELD BACK, AND THE BRANCH'S DELAY SLOT IS THE TELL.** *(sharpens §160d, the only store-then-reload entry, whose observable is an `lbu`/`sll` COUNT for a table index and never a compare; and §164-20 / §164-82, whose re-read tells are both LOAD-guard-then-reload, the mirror construct. The ordering half is §164-29's register WAR fence, NOT §163d/§162j.)* + +Target shape — a timer field decremented and tested inside one arm: + + lhu $v1,0x86($s0) + beqz $v1,.Lskip + … + addiu $v0,$v1,-1 + sh $v0,0x86($s0) + andi $v0,$v0,0xFFFF <- the "reload", cse'd down to the STORED register, zero-extended IN PLACE + bnez $v0,.Lret + nop <- the slot stays EMPTY + +**THE LAW.** Write the second test against MEMORY, not against the local: + + t = *(u16 *)(p + 0x86); + if (t != 0) { + *(u16 *)(p + 0x86) = t - 1; + if (*(u16 *)(p + 0x86) != 0) { return; } + } + +cse replaces the reload with the just-stored pseudo, and the HImode→SImode zero-extend the compare needs is done **in place on that same register**. The `andi` therefore re-SETs a register the `sh` READ, and `sched_analyze_1` (`sched.c:1714-1715`) emits a hard `REG_DEP_ANTI` for exactly that (§164-29) — so the `sh` can never be moved below the `andi`, which is precisely what `reorg` would have to do to put it in the `bnez`'s slot. Slot stays `nop`: **5 words.** The local-variable form `t = t - 1; *(u16 *)(p + 0x86) = t; if (t != 0)` decrements in place (`addiu $v1,$v1,-1`), masks into an INDEPENDENT register (`andi $v0,$v1,0xFFFF`), frees the `sh` of the anti-dependence, and `reorg` steals it into the delay slot: **4 words, `LENGTH-DRIFT −1`.** + +**THE DIAGNOSTIC TELL.** On the SECOND branch of a decrement-and-test: **a `nop` in the slot ⇒ the source re-read the field; a `sh` in the slot ⇒ the source tested a local.** Read the slot before touching anything else — like §164-82's class this presents as a whole-tail shift, not as one missing instruction. + +**Byte evidence, and its limit.** `func_80183834` (ov_SC01_077, 135 ins) banks with the read-back form (`src/ov_SC01_077/ov_SC01_077_jr_80183324.c:3393-3398`), as does its family twin `func_8018127C` (`src/ov_SC02_000/ov_SC02_000_jr_80180D6C.c:3044-3049`); a corpus grep finds **zero** instances of the local-variable form anywhere in `src/`. + +⚠ **The counter-arm has no surviving artifact.** Every preserved probe in `.run/wave6/func_80183834/` (`base.c`, `varA`–`varR`) carries the read-back spelling, so the −1 is a transient measurement from the crack's front-half iteration and is **not re-runnable**. The mechanism above is a reconstruction — the crack note's own wording ("the `sh` is a live-def of the compare register") is wrong: the `sh` is a USE, and the edge is ANTI, not true. **Re-run the A/B before leaning on the −1 number; the transcription rule and the slot tell are the durable halves.** + +*(SHARPENS — sharpens §165-15 (L14033-14057), §163a (L11764-11780), §8d (L487-510), §17a-1 (L1445-1470); evidence: single-instance; from `func_80183BC8`)* + +**§167-43 — THE BLOCK-SCOPE SOLVENT HAS A MOVABILITY PRECONDITION, AND A RETURN-TYPE-ONLY CALLEE CONFLICT WAS NEVER ITS JOB.** *(bounds §165-15's discriminator table (L14049-14057), whose `conflicting types for 'X'` row says "⇒ §163a's block-scope solvent is live (move BOTH the typedef and the extern into the block)" without stating the precondition; the cure that row should hand over is §17a-1 point 1, which already names this exact case.)* + +**Target situation.** Your body needs the callee's `$v0` — `v0 = f(a0); if (v0 != 0) …` — while the destination TU already carries `extern void f(s32 a0);` at FILE scope, ABOVE your insertion point, feeding a function that is ALREADY BANKED. + +**THE LAW.** §163a's solvent is a property of a PAIR of declarations: *both* at block scope warn, *either* at file scope is the hard-error cell (§8d byte-proved that direction: `FILE(void *) -> BLOCK(int) -> conflicting types for 'D_801812A4'` ERROR). **You can only reach the solvent when you own BOTH decls.** A file-scope callee extern that an already-banked function in the same TU compiles against is immovable — demoting it changes that function's declaration environment (§8d, the whole reason `scope_data_externs.py` exists) and buys an R22 blast radius for a one-function problem. **When the only term that differs is the RETURN TYPE, do not redeclare at ANY scope: leave the TU's decl exactly as it stands and cast at the use site** — `v0 = ((s32 (*)(s32))f)(a0);` — codegen-neutral, draft-only, no fleet edit. §17a-1 point 1 already states this as *"Call-site casts, NOT redeclaration (the #1 recurring miss)"* and names the form verbatim for a `void`-canonical callee whose `$v0` is used. + +**BYTE EVIDENCE.** `func_80183BC8` (ov_SC02_041, 108 ins), banked whole-binary at `src/ov_SC02_041/ov_SC02_041_jr_8017BEBC.c:6477`. The TU declares `extern void func_8012CBCC(s32 a0);` at file scope `:5732` — consumed two lines later by the banked `func_80182484` through `((s32 (*)(void))func_8012CBCC)()` — and repeats the identical spelling at `:6470`. The new body takes the return through `v0 = ((s32 (*)(s32))func_8012CBCC)(a0);` at `:6509`. The crack note's proposed fix (add `extern s32 func_8012CBCC(s32 a0);` at block scope inside the body, leaving `:5732` in place) was **never compiled** and lands in §163a's file-scope cell. Cross-TU confirmation of the true arity/return: `src/ov_SC02_011/ov_SC02_011_jr_8017AE2C.c:4933/4953`. + +**THE DIAGNOSTIC TELL.** `match_one` MATCHes standalone, the whole-binary build reports `conflicting types for 'func_X'`, and diffing the two spellings term by term (§73's two axes) shows **only the return type differs**. Route it: *return-only* ⇒ use-site cast, zero blast radius (this entry). *Params* ⇒ §73's param row, cast at each use. *Arity / `too many arguments`* ⇒ §165-15, no scope helps. *A decl you own at BOTH ends* ⇒ §163a, move the pair into the block. **Check who else compiles against the file-scope decl before you touch it — if the answer is a banked function, the solvent is off the table.** + +*(SHARPENS — sharpens §43 (L3195), §164-48 (L12866), §163b (L11781), §164-30 (L12507); evidence: single-instance; from `func_80184E4C`)* + +**§167-44 — THE `sll $aN,$sX,16 ; sra $aN,$aN,16` PAIR IN THE SLOTS BEFORE A `jal` CAN BE A *PARAMETER* WIDTH TELL — §43's TRIAGE HAS NO BUCKET FOR IT.** *(SHARPENS §43 (L3195), whose tell is binary — in-place on `$aN` ⇒ s16 param, into a `$v0/$v1` temp ⇒ cast-of-s32 — and §164-48 (L12866), which owns this exact shape but reads it for a LOCAL and forwards the PARAMETER reading to §163b, which is switch-dispatch only. The ANSI-vs-K&R half is already indexed (cookbook-index L29; §164-20's delay-slot discriminator, L12286); the SHAPE is not.)* + +Target shape — a narrow param whose live range crosses a `jal`, re-extended into the ARG register at the next call: + + move $s1,$a2 ; move $s3,$a3 <- entry: the raw WORD is stashed, UNextended + jal + beqz $s0, + sll $a2,$s1,0x10 <- the extension READS the callee-saved copy… + sll $a3,$s3,0x10 + move $a0,$s0 ; move $a1,$s4 + sra $a2,$a2,0x10 <- …and WRITES the arg register + jal + sra $a3,$a3,0x10 <- second half rides the delay slot + +**THE LAW.** This is neither of §43's two buckets. The parameter is HImode; its live range crosses a call, so the raw word goes to a callee-saved pseudo at entry; the widening is the tree-level conversion of that HImode value to the callee's `s32` formal (`convert_arguments` → `convert_for_assignment`, `c-typeck.c:1623/1737`). **Declare the enclosing function's parameter `s16` in a plain ANSI prototype and pass it STRAIGHT THROUGH** — `f((u8 *)obj, a1, a2, a3)` against `extern void f(u8 *, s32, s32, s32);` — and ordinary integer promotion emits the pair for free. No `(s16)` cast on top (and a cast on an `s32` param cannot reach this shape, §43). + +**WHY THE PARAM STAYS NARROW INSIDE THE BODY (source-cited; §43 states this only as behaviour).** `config/mips/mips.h:1153` defines `PROMOTE_PROTOTYPES` but the port defines **no `PROMOTE_FUNCTION_ARGS`**, so `assign_parms` leaves `promoted_mode = passed_mode` (`function.c:3324-3329`) and an incoming `short` lives in HImode for the whole body — every SImode use pays its own extension, wherever it falls. + +**BYTE EVIDENCE.** `func_80184E4C` (ov_SC03_014, 50 ins, banked whole-binary, S48 wave 6). `src/ov_SC03_014/ov_SC03_014_jr_801848E4.c:2797` = `s32 func_80184E4C(s32 a0, s32 a1, s16 a2, s16 a3, s32 a4)`; the shipped `build/src/ov_SC03_014/ov_SC03_014_jr_801848E4.o` `0x568-0x62c` disassembles to the shape above (`move s1,a2` @0x588, `move s3,a3` @0x590, `sll a2,s1,0x10` @0x5a8, `sra a3,a3,0x10` in the `jal func_8001CB6C` delay slot @0x5c0). Both middle params flow untouched into `func_8001CB6C`'s `s32` formals; there is no cast in the body. + +*Scope, honestly: **n=1 and NO A/B was run** — the draft MATCHed first try, so the `s32` twin was never compiled. The direction is not in doubt (dropping to `s32` deletes both pairs by plain C semantics: LENGTH-DRIFT −4), and §164-48's one-character probe byte-proves the same width→extend-at-use coupling for a LOCAL. What is new here is the READING of the shape, not the coupling.* + +**THE DIAGNOSTIC TELL.** LENGTH-DRIFT −2 per narrow argument, the missing instructions a `sll 16`/`sra 16` pair straddling the call's `move $aN,…` arg setup, on a function that copies raw `$aN` into `$sX` at entry. **Ask whose value `$sX` holds before reaching for §164-48:** a call result or a computed local ⇒ §164-48 (and expect an `$s0/$s1` permutation with it); an incoming argument copied at entry with nothing done to it ⇒ **this entry — widen nothing, narrow the DECLARATION.** If the pair sits between a bias subtract and an `sltiu`, it is §164-37/§163b instead. + +*(SHARPENS — sharpens §20 scalar-global-RMW bullet (L1907-1918), §164-52 (L12943), §153 (L10463), §165-26 (L14399); evidence: single-instance; from `func_80188F60`)* + +**§167-45 — A GLOBAL THAT IS STORED TO *AND* HAS ITS ADDRESS PASSED TO A CALL IN THE SAME BLOCK NEEDS A POINTER LOCAL — AND AN `addiu $rD,$rB,-K` OFF THAT BASE NAMES WHICH SYMBOL THE SOURCE ANCHORED ON.** + +Target shape — one global's address held in a register serving BOTH a zero-displacement store and the call's pointer argument, while its neighbours in the same record stay plain symbol stores: + + la $16,D_800A5E98 + sw $2,0($16) <- the store goes through the BASE + ... + sw $2,D_800A5E9C <- the NEIGHBOUR keeps the plain %hi/%lo fold + sb $2,D_800A5EA4 + move $5,$16 ; jal func_80028620 + +**LAW 1 — THE ANCHOR.** `D_800A5E98 = -0x12; … f(1, &D_800A5E98);` emits the §18 single-store fold (`lui $at,%hi ; sw $v,%lo(SYM)($at)`) **plus a second, independent `la` into `$a1`** — two wrong instructions. `s32 *rec = (s32 *)&D_800A5E98; *rec = -0x12; … f(1, rec);` makes the store's address RTL the *same pseudo* the argument needs, cse ties them, and one `la` serves both. This is §20's scalar-global-RMW bullet with the READ replaced by a call argument: **the trigger is not "read and written", it is "the address is needed in a register for something else in the same block."** Check §153/§165-26 first — if the address is an argument to >=2 calls and there is no store, those apply instead and the declaration is what you DELETE. + +**LAW 1a — BOUNDS §164-52.** §164-52 says a pointer-to-global local evaporates into `lw SYM+k` unless a >=2-jump-ref label sits between its init and its uses. Here `rec` is init'd and used on ONE straight path with no join at all, and it survives — because §164-52's fold is `fold_rtx` rewriting `p[k]`, a **dereference**, into an absolute address. **A use of the pointer VALUE — a call argument, a compare, a store of the pointer itself — has no absolute form to fold into and keeps the base alive on any path.** Read §164-52's tell as scoped to all-dereference locals. + +**LAW 2 — THE READING HALF.** `la $rB,SYM ; addu/addiu $rD,$rB,-K`, with `$rD` a pointer argument and `$rB` a store base, means the source anchored the pointer local on the **LATER** symbol and subtracted: + + s32 *rec2 = &D_800A5E9C; *rec2 = 0x18; … f(1, rec2 - 1); + => la $3,D_800A5E9C ; sw $2,0($3) ; addu $5,$3,-4 + +Anchoring on the earlier symbol and writing `rec2[1]` instead gives `sw 0x4($base)` and a bare `la $a1` — 6 mismatched. **The STORE's displacement is the discriminator: `0(base)` means the anchor IS the stored symbol; `K(base)` means you anchored too low.** + +**PER-SITE, NOT PER-BLOCK.** Only the word the target routes through the base gets a pointer local; its neighbours (`D_800A5E9C` and the three `sb` bytes here) keep the plain symbol spelling. Converting the whole record to one struct-typed base is the failure mode — the same discipline as §162q1 and §165-27. + +**BYTE EVIDENCE.** `func_80188F60` (ov_SC02_000 + ov_SC02_003, `jr_8018173C`, MATCH, banked `commit:1713` at `src/ov_SC02_000/ov_SC02_000_jr_8018173C.c:3890`); asm at `.run/wave4/func_80188F60/t.s:90-94, 113-118, 308-309`. *Honest scope: the preserved rungs (`v1.c` 149-off -> `v2.c` 98-off) change FOUR things at once — this anchor, the branch polarity, the `rec2 - 1` anchoring and a counter merge — so the direction is established by the banked body and the losing spellings are prose, not preserved probes. Re-measure single-axis before quoting a magnitude.* + +*(SHARPENS — sharpens §164-80 Law 1 (L13566-13578), §20 cross-BB combine law (L2489), §165-06 (L13827), §164-24/§16x (L12405); evidence: single-instance; from `func_80188F60`)* + +**§167-46 — §164-80 LAW 1 EXTENDS TO `>>`/`+` ON A PLAIN REGISTER LOCAL, AND ITS SYMPTOM THERE IS A *LENGTH* DRIFT, NOT A REGISTER DRIFT.** + +§164-80 Law 1: `t = (t & M) | x;` evaluates the inner operation into an anonymous temp that takes a scratch, while `t &= M; t |= x;` expand in place on `t`'s own pseudo. Its evidence is one `and`/`or` pair on a memory-fed packed word, and its scope note says *"Re-measure before generalising Law 1 beyond `and`/`or`."* + +**It generalises, and it costs something different.** `i = (i >> 5) + 0x10;` on a register local mints a pseudo for `i >> 5` that lands in `$v0`. Two consequences, both LENGTH: + +1. the following `addiu` now carries a WAR dependence on `$v0`, so it can no longer sink into the branch delay slot; and +2. `reorg` fills the *preceding* branch's slot by DUPLICATING the `sra` — the same insn then appears both in the delay slot and at the join. **Net +1.** + +`i >>= 5; i += 0x10;` — both destructive on `i` — restores the target's stream. What matters is that each statement's DESTINATION is the variable itself; the `op=` token is the shortest spelling of that, not the mechanism (§164-24). + +**THE DIAGNOSTIC TELL (the new half).** **An instruction that appears BOTH in a branch delay slot and again at that branch's join, on a LENGTH-DRIFT +1, means your intermediate value is living in a scratch register.** Make the statement destructive on the variable before touching scheduling pins, allocno sliders or the permuter. **Discriminate against the one other cause of that exact signature:** §20's cross-BB combine law (L2489) produces a duplicated `sll` at a loop preheader *plus* the loop-back delay slot when a narrow value crosses a BACKEDGE — that one is a fold-defeat you WANT, and its duplicate straddles a loop, not an if-join. + +**BYTE EVIDENCE.** `func_80188F60` (ov_SC02_000, MATCH, banked `commit:1713`): `.run/wave4/func_80188F60/v2.c:67` (compound, 98 off) vs `v4_MATCH.c:68-69` (destructive, MATCH). *Honest scope: that rung also splits a decrement/test and block-scopes a pointer local, so the +1 attribution is the agent's reading of an intermediate diff, not an isolated A/B. The underlying fact — an expression temp is a distinct pseudo at expand — is §164-80's and is source-read.* + +### §167z — REFUTED IN WAVES 5/6: do NOT re-derive +* **`func_8017FE38`** — GENERALIZATION: when a long-latency load heads a dependency chain whose consumer is far away, the SOURCE STATEMENT BOUNDARY decides whether sched1 hoists it — an inline s + *Why:* Byte-refuted as worded. I built the banked body with `d = *(u16 *)&D_800B9A02;` as its OWN statement placed immediately above its consumer (`ot = (u32)&D_800A6610[d << 14];`): 12 mismatched at 239 ins — and the `.text` is BYTE-IDENTICAL to the fully-inline spelling (objdump diff is only the embedded filename). A statem +* **`func_8017FE38`** — THE FORCED COPY MUST BE CONSUMED IN PLACE OR THE STORE SCHEDULES AWAY FROM IT — rewriting `x -= (tp & 0xF) << 6;` as `tp = tp & 0xF; x -= tp << 6;` shortens the forced co + *Why:* The EFFECT is real — I measured the in-place vs inline A/B on the banked body at 0 vs 10 mismatched (239 ins both), with the `sh $rX,0x16($a3)` moving four slots and a register cascade behind it. But the stated LAW is a coupling claim about the forced copy, and the coupling is byte-refuted: with the `+ zr` copy REMOVED +* **`func_8018C3F8`** — BOUND ON §164-02: when the constant base reaches the addu via a SPILL RELOAD (`lw 0x30($sp)`) rather than a live `la` pseudo, "a round trip through memory kills the const + *Why:* The OBSERVATION is real and byte-confirmed by me in the shipped object (0x8018c554 `addu s1,t2,s1`, base first, source order won). The stated MECHANISM is impossible and must not be banked in this form. + +(1) PASS ORDER makes the spill story unreachable. §164-02's swap is a cse-time decision (`fold_rtx`, cse.c:5278-5304 +* **`func_8018C3F8`** — Corollary: "this ALSO explains why sibling func_8018F694 matched with the opposite spelling: there `ot` stays in a live register." + *Why:* BYTE-REFUTED from the shipped, matching binary. `func_8018F694` reloads its `ot` from the frame exactly like `func_8018C3F8` does: `8018f7d8 lw a3,0x38($sp)` / `8018f7dc sll s0,s0,0x2` / `8018f7f0 addu s0,s0,a3`. Its `ot` does NOT stay in a live register — it is stack-resident at the add, the same as the claimed "spill +* **`func_8017F17C`** — Lever 3 — FRAME: MEASURE THE BASELINE BEFORE ADDING DEAD LOCALS. The 8-byte hole at sp+0x18..0x1F reads exactly like §162's unreferenced-local oracle but is not one: gcc + *Why:* THE THREE ABLATION NUMBERS ARE RIGHT; THE MECHANISM IS BYTE-REFUTED AND THE PRESCRIPTION IS ALREADY BANKED THREE TIMES OVER. It must not enter the file in this form. + +(1) MECHANISM REFUTED. "gcc already rounds the align-1 struct's slot" is false. I compiled 7 controlled spellings on the pinned cc1: `struct { u8 c[8]; } +* **`func_801805D4`** — L8 — the accumulation must be ONE expression `f(x) + (f(y + K) >> 4)`, not two statements: gcc-2.7.2 evaluates the left operand first, parks it in $s0 across the second c + *Why:* The prescriptive half — 'ONE expression, not two statements' — is BYTE-REFUTED at vet time, twice. Splitting into two statements with per-site block-scope temps (`{ s32 x = f(..); s32 y = f(..); u = x + (y>>4); }`) -> MATCH 212. Splitting with a multi-set accumulator (`u = f(..); u = u + (f(..)>>4);`) -> MATCH 212. §16 +* **`func_80187F2C`** — The two field stores and the `func_8012E688(a0,0x98F,0)` call must be written call-LAST (`store; store; call;`). Interleaving (`store; call; store`) let gcc-2.7.2's delay + *Why:* THE TARGET DESCRIPTION IS FACTUALLY FALSE AND I BYTE-REFUTED IT. Disassembling the banked SHA1-green object at 0x80187FE0: `jal 8012e688` / `sh v0,254(a0)` — the 0xFE store IS the call's delay slot, not a `nop`. The whole claim is built on a target reading that does not exist, and the same TU's unbanked sibling `func_8 +* **`func_80187130`** — When EVERY arm of a 2-way branch ends in `return` (no single main-body arm), gcc-2.7.2's RTL-order-is-source-order places the arm written LAST adjacent to the shared epil + *Why:* BYTE-REFUTED — and refuted by the agent's own two builds, which are still on disk. I re-ran the pinned triple at vetting time (cpp / tools/bin/gcc-2.7.2-psx/cc1 -O2 -G0 -mips1 -mcpu=3000 / maspsx / as) on both surviving probes plus three new controls, and read cc1's own `.s`, not just the `.o`. + +(1) THE DIRECTION IS EX +* **`func_80183BC8`** — The accumulate must be spelled as the full raw expression twice — `*(u16*)ptr = *(u16*)ptr + r;` — and NOT as a `+=` compound assignment. + *Why:* No A/B was run on the `+=` token; it is bundled into the same edit as the per-branch duplication and the fresh local, so it credits an untested variable — the precise failure mode §164z refuted twice in this campaign (func_80184944: 'the A/B moved TWO variables at once and credited the wrong one'; func_801823E8: 'the A +* **`func_80183BC8`** — INTEGRATION PRESCRIPTION: because the destination TU already declares `extern void func_8012CBCC(s32 a0);` at file scope, bank the body by declaring `extern s32 func_8012 + *Why:* Three independent problems. (1) MOOT AND CONTRADICTED BY THE SHIPPED BANK: func_80183BC8 is already banked at src/ov_SC02_041/ov_SC02_041_jr_8017BEBC.c:6477 and the resolution actually used is the §17a-1 use-site cast — the file-scope decl is left as `extern void func_8012CBCC(s32);` (:6470, an exact-text duplicate of +* **`func_8017F3C8`** — FLAGGED CLAIM — 'a cross-block ternary-result allocno's global-alloc conflict is BLOCK-granularity, not instruction-precise': gcc-2.7.2 global-alloc records a conflict ag + *Why:* BYTE-REFUTED FROM THE PINNED SOURCE, and the cited evidence cannot support the claim it is offered for. + +(1) THE GRANULARITY HALF IS FALSE. `global_conflicts` (tools/reference/gcc-2.7.2/global.c:625-772) is an explicit per-INSN birth/death walk, not a block-granularity scan. Per block it (a) seeds `allocnos_live`/`hard +* **`func_801832F8`** — FLAGGED: the §164-82 / §46-L2 'test-then-re-read' idiom transfers from memory LOADs to constant-division results — writing `pan + out1.x/20` again in every arm of an if/e + *Why:* BYTE-REFUTED on the pinned triple (cpp / cc1-2.7.2 -O2 -G0 -mips1 -mcpu=3000 / maspsx 2.56 / as). I built the minimal repro the claim describes — `register s16 pan __asm__("$17")`, the field `out1.x` loaded through a real call, the clamp result captured into `s32 panv`, `panv` consumed by the trailing call — in two spe +* **`func_801832F8`** — Pinning `base` to ANY explicit hard register costs +2 instructions elsewhere in the SAME function — specifically it blocks fill_simple_delay_slots from duplicating the `a + *Why:* Byte-refuted, and already narrowed by the successor agent. In the shipped draft the pinned build emits the duplicated `addiu $a0,$sp,0x20` in BOTH positions (indices 62 and 64 of the mine-side stream, matching the target's pair at 801833F0/801833F8) at 111 ins — i.e. the pin does not suppress the duplication at all. Th +* **`func_801832F8`** — A second, more novel use of the §5a barrier: as a pure SCHEDULING fence with no return statement attached — placed after `if (v<0) goto y_neg;` to block the `bltz`'s dela + *Why:* Byte-refuted in its own function. Deleting exactly that barrier from the shipped draft and re-running `match_one` gives a BYTE-IDENTICAL result — 111 ins, 40 mismatched, and the two diff reports are `diff`-clean against each other (`.../vet832F8/w6_base.diff` vs `w6_nofence.diff`). The barrier is doing nothing at the s +* **`func_801832F8`** — (derived during this vet, from residual B) The target's pan-clamp `addu $v0,$s1,$v0 ; addu $s1,$v0,$zero` — sum into a scratch, then a copy into the callee-saved home — c + *Why:* Byte-refuted, and the cure is one deleted word. That copy is the deferred-truncation store an `s16` accumulator emits when it is ASSIGNED and then TESTED — and the `register __asm__("$17")` pin on `pan` is what deletes it. Minimal repro, same call skeleton, pinned triple: + pA `register s16 pan __asm__("$17"); pan = p +* **`func_80184E4C`** — MECHANISM: 'MIPS o32 leaves the upper bits of narrow-typed register params unspecified, so gcc re-extends on first actual use, even though the parameter was register-copi + *Why:* The OBSERVABLE half — raw whole-word stash to a callee-saved pseudo at entry, extension deferred to the use — is already §43 word for word ('with the raw values stashed to callee-saved pseudos first (s3←a1 …) and re-extended per use after calls'), so it is COVERED, not new. The ABI ATTRIBUTION is wrong and must not be +* **`func_8017EEEC`** — func_8017EEEC is a SECOND byte-verified instance of §163c's specific mechanism — an EMPTY `case 4:` glued onto `default:` is the load-bearing label that crossed the thres + *Why:* BYTE-REFUTED, and it matters: the draft is wrong and must not be banked. §163c's own diagnostic ('jtbl[k] == the default label is the fingerprint of a case label sharing the default body') falsifies the claim on sight. The target table is jtbl_801AA880 = [0x8017EFB0, 0x8017F074, 0x8017EFF0, 0x8017F074, 0x8017EF3C] (asm +* **`func_8017F5D4`** — (wave-5) FAMILY REACH: the ×2 sibling should remap mechanically — the family template is the (Morph_8017DC1C*, SVECTOR2*, SVECTOR2*, t) quadruple list; only the 11 per-ov + *Why:* Byte-refuted, first by the wave-6 agent and then independently by me: the only other func_8017F5D4.s in the tree (ov_SC01_008) is an 18-instruction body with frame 0x18 and two jals — nothing to template. The wave-5 note asserted a mechanical ×2 remap without ever opening the sibling's .s, which is precisely the failur +* **`func_80187D0C`** — Sub-claim of the same note: the oracle generalises to `<=` — "slti+bnez-to-label almost always means the source used >=/<= against the constant, with the label holding th + *Why:* BYTE-REFUTED, and it is the one part of the note that is not already on disk. The relational families split by polarity, not by strictness: the GREATER family (`>=`, `>`, and the operand-reversed `K <= x`) produces `slti`+`bnez`-to-else; the LESS family (`<`, `<=`) produces `slti`+`beqz`-to-else. I compiled all four on +* **`func_8017F2D4`** — ROOT CAUSE of the seven prior gate refusals was the wrong destination-TU citation, not a codegen residual — 'a banker splicing into jr_8017C340.c gets a no-op or a bad sp + *Why:* Already byte-refuted inside the very entry this campaign banked from these notes. §166a's scope correction states it was measured: `corpus.stubs()` derives each stub's TU from the actual INCLUDE_ASM site and `gate_stage` splices via corpus, so the harness was always editing the right file — only the PROSE was wrong. Af +* **`func_8017F2D4`** — NEW-1: gcc-2.7.2 cross-jumping keeps the FIRST copy while the target keeps the LAST, which is why the 14-arm shared tail must be hand-written ONCE at the last arm with `g + *Why:* The MECHANISM is on §165z's refuted list under this exact function, in these exact words. jump.c keeps the LAST on every path — `do_cross_jump` deletes the stream preceding `insn` and `jump_chain` is built push-front, so the earliest jump pairs with the latest copy (§162g, read out of jump.c:2537/219-221). The real rea +* **`func_8017F2D4`** — NEW-5: the surviving redundant compare at 0x8017F460 needs TWO levers TOGETHER — (a) `__asm__ ("" : "=r"(t) : "0"(t))` tied to SELF (tying to a fresh temp costs a real `m + *Why:* Split verdict, and both riders the agent re-asserts are already byte-refuted. Lever (a) itself is SOUND and banked as §165-03 (cse's record_jump_cond carries EQUALITY across a branch, cse.c:5944-5990; barrier deleted -> 3 mismatched, BRANCH-POLARITY) — COVERED. But the rider 'tying to a FRESH temp costs a real move' is +* **`func_8017F2D4`** — `s32 lt = mode < 0xA;` must be an explicit local because gcc-2.7.2 has no gcse and the single `slti` must precede the branch both arms need it after. + *Why:* On §165z's refuted list under this function, byte-refuted at vet time: deleting the local and inlining the compare at both use sites — `if ((mode < 0xA) || (D_8011515A == 0x100))` and `if (!(mode < 0xA))` — gives MATCH 279. The C form is byte-INERT here, so it is not a lever and must not be taught as one; the 'no gcse'