Files
BFM-decomp/cookbook/C0501.md
T

216 KiB
Raw Blame History

§452 ★★★ — NOT EVERY VERBATIM BODY IS UNDECOMPILED WORK, AND §448'S HEADLINE OVERSTATED IT (P31 S75; 10-function burst, 0 banks, and the negative result is the finding)

§448 counted 199 functions that are assembly posing as C and called 154 of them "GAME CODE — real decompilation work remaining." A 10-agent burst against the smallest ten returned 0 upheld matches and, in doing so, refuted that framing. The near-reports name four classes that are legitimately verbatim and must be subtracted before anyone plans against the number:

  • FRAGMENTS OF A SPLIT FUNCTION — not functions at all. SYS_OBJ_604 / SYS_OBJ_640 / func_80059760 share ONE 32-byte stack frame: they are the compiled output of a SINGLE original C function that the tooling split into three addressable symbols. SYS_OBJ_2DD8 is likewise "a bare shared epilogue tail" — lw ra / lw s0 restores with no prologue. No C function can emit a restore without a matching save, so these can never be decompiled individually; they can only disappear when their parent is decompiled as one function.
  • HAND-WRITTEN ASSEMBLY IN THE ORIGINAL. func_800495EC is a GTE wrapper — three mtc2s and the GTE latency nops. Square wrote that in asm in 1998; there is no C to recover.
  • COMPILER-INEXPRESSIBLE FORMS. func_80062388 puts a symbolic store in the delay slot after jr ra (lui at / jr ra / sw a0,0x2a24(at)). gcc-2.7.2's mips.md define_delay permits only a ONE-instruction delay slot, and a symbolic sw needs two for the $at synthesis — so this sequence is unreachable from C with this compiler, by construction.
  • NO-RETURN TAILS. func_80049610 is three nops with no return; gcc cannot omit a function's return path. Best C attempt is 2 instructions (jr ra + delay nop) — a LENGTH-DRIFT residual that no shape fixes.

THE CORRECTION. "154 game functions of remaining work" is an UPPER BOUND, not a work queue. The honest statement is: 199 functions are byte-identical without being decompiled; an unknown share of them are undecompilable in principle and must be identified before the rest are costed. Triaging the class — split-fragment / hand-asm / inexpressible / genuinely-undecompiled — is the prerequisite to any estimate, and the four tells above are cheap to check: does it share a frame with a neighbour, does it touch cop2, does it use a delay slot no define_delay allows, does it lack a return path.

AND THE GUARD THAT EARNED ITS KEEP. One agent submitted the verbatim __asm__ block itself as its decompile. match_one printed MATCH (4 ins) — truthfully, because a raw asm blob byte-matches its own source by construction. The adversarial verifier refuted it on the rule that a draft containing __asm__ or INCLUDE_ASM is a no-op that passes for free. Any burst over this class MUST carry that check: the trivially-passing draft is not a hypothetical here, it is the default thing to produce, and a byte gate cannot tell the difference.

ADDENDUM to §265 (P31 S76) — THE PROSE GUARD DID NOT HOLD. IT IS A GATE REFUSAL NOW.

The paragraph immediately above this one was written in S75 and is correct. I tripped over the exact trap it describes about four hours later, in the next session, having read it. That is the finding worth keeping — not the trap, which was already known, but the fact that knowing it in prose did not prevent it.

What happened. The 9 remaining DECOMPILE-NOW SDK functions in src/800c3.c / src/800c2_2.c were converted from §265 verbatim bodies to INCLUDE_ASM stubs specifically so they could be decompiled. I then searched the draft store for stored drafts, scored 195 candidates with match_one, and got closeness 0 on all nine. Slate, gate, BANKED 9 main functions — 143dbb89 BYTE-IDENTICAL, 58 seconds. Every one of those nine "drafts" was the function's own assembly, emitted by tools/asm_verbatim.py into <fn>.c in the same directories as real drafts. I had converted verbatim → stub → verbatim. A perfect round trip that decompiled nothing.

What caught it: progress.py did not move. REAL 882, VERBATIM 164, INCLUDE_ASM 37 — identical before and after "banking nine functions". The instrument was right and the claim was wrong (check-against-a-known-true-case: the count you already know is the cheapest oracle you own).

Why the prose could not hold. S75's rule was addressed to a burst over this class — "any burst over this class MUST carry that check". I was not running a burst. I was hand-picking stored drafts one function at a time, which the rule did not name, and the conditional in a reader's head ("does this apply to me?") is exactly where a prose guard fails. S75 also concluded "a byte gate cannot tell the difference", which is true and which reads as unpreventable. It is not: the byte CHECK cannot tell, but a slate-load refusal can, because the verbatim form is trivially decidable from the text.

The guard, where it can actually fire. draft_prechecks.is_verbatim_asm_draft(text, fn) — a file-scope __asm__ naming this fn via .ent/.globl/label, AND no C definition of it. Both halves matter: a real draft may carry a small inline __asm__, and a verbatim body may omit .globl for a static. Match BOTH spellings — ".ent\tNAME" inside a C string is a backslash-t, not a tab; five S75/S76 censuses of this class disagreed with each other until both were handled. gate_main now refuses such a slate at load, beside its existing INCLUDE_ASM no-op refusal (R43).

The scale of the landmine. A census of the draft store: 1,099 of 704,375 .c files are verbatim-asm drafts, sitting under ordinary <fn>.c names in ordinary wave directories. Negative control: 0 false positives across 45,898 drafts that carry both a real C definition and an inline __asm__ (R39). Any future "search the store for a stored draft" pass — and that pass is now a standard move, since S75 banked 20 functions from drafts already on disk — is drawing from a pool with 1,099 of these in it.

The general law. A lesson that is only prose will be re-learned. If a check is decidable, the knowledge base is where you explain it and the pipeline is where you enforce it — the same relationship R32 sets between "assert your coverage" and a scanner that actually does.

CORRECTION to §182/§188 (P31 S76) — THE EPILOGUE "WALL" IS AN ORACLE ARTIFACT IN THE REORDER ISLAND

§188 says the jr $ra + addiu $sp-in-delay-slot epilogue is unreachable from C under the pinned triple. That stopped being true for four TUs on 2026-09-01 and the knowledge base did not notice.

REORDER_TUS := 800c2 800c2_2 800c2_3 800c3 (Makefile) pipes those TUs through tools/reorder_passthrough.py into as -O2 — the assembler mode that fills delay slots and emits exactly that epilogue. It is the real build path for those files, and it is how S75 banked 20 functions in src/800c3.c. But match_one, the oracle every drafting agent scores against, still compiled through maspsx + as -O1 for them.

The measurement (same draft, func_8005ECC0, plain C):

oracle path closeness ins mine/target verdict
maspsx + as -O1 5 36 / 35 LENGTH-DRIFT/1 — reads as the §188 wall
reorder + as -O2 2 35 / 35 epilogue identical; only a real DELAY-SLOT/1 left

The entire epilogue residual, and the phantom extra instruction, were the oracle modelling a build path the project no longer uses for that file.

What it cost, in one wave. Seven of eleven main agents (func_8005D4B8, func_8005D4F0, func_8005EAC8, func_8005EAE8, func_8005E13C, func_8005E79C, StopRCnt) independently produced a correct plain-C body, saw the phantom ±1 tail, correctly identified it as the §182/§188 shape, consulted oracle_reorder.py — whose docstring read "file IMMOVABLE, stop grinding, no C-level work can ever close it" — and each fell back to submitting a §265 verbatim-asm body for a function whose C the build would have accepted. Every one of them reasoned correctly from a false premise the knowledge base handed them.

Fixed: match_one derives REORDER_TUS from the Makefile and selects the reorder path automatically (--no-reorder forces the old path); oracle_reorder.py's docstring is corrected. The list is DERIVED, never copied — R51, a derived property stored as config goes stale, which is exactly the failure being repaired here.

The general law, third instance this session. An oracle that models a build path the project has changed does not report a wall — it manufactures one. R35 says fix the instrument before trusting the measurement; this adds the corollary that a knowledge-base claim about the toolchain has the same shelf life as the toolchain, and a build change must sweep the docs that assert what the build cannot do.

§460 — -dS PRINTS THE SCHEDULER'S READY LIST WITH PRIORITIES. STOP INFERRING IT FROM RTL ORDER.

Source: the S76 ov_SC02_027:func_80180B3C agent (297 ins, 82 → 23), and it is a general instrument, not a one-off. Three prior attempts on this function tried to steer sched1 by reordering source statements and by reasoning backwards from the order of insns in the .sched RTL dump. That is guesswork about a cost model. cc1 -dS emits the FULL verbose trace — the literal ;; ready list at T-N lines with each insn's priority — so the scheduler's own ranking is readable rather than reconstructed. Read it before spending a single reorder attempt on a schedule-class residual.

Two findings that came out of reading it, both reusable:

  1. PINS BEAT SCHEDULE-CHASING when the residual spans a register chain. The first half of this function was closed by four register __asm__ pins (w=$10, obj=$11, x=$4, y=$7) — 44 → 29 → 23, all four load-bearing. The three earlier attempts had treated the $t2/$t3/$t7 chain as downstream of the schedule and tried to fix the schedule; it was upstream, and the allocation was the cause. This is dont-conclude-unsteerable-try-register-pins extended: pins are not only a last resort for a stuck NEAR, they are the FIRST move when the diff walks a register chain. Note also that statement order was inert before the pins and live after — so an "order does nothing" measurement is only valid for the allocation you measured it under.

  2. BIRTHING-BOOST IS THE DIAL, AND A SINGLE-SET REQUIREMENT CAN BE UNREACHABLE. The residual is sched1's birthing boost (7f000001): a mask v = y & 0xFFFF is priority 3 UNBOOSTED because v is set twice (the mask, then a conditional -= 0x100), so backward-scheduling sinks it to block-front, while bank << 14 IS boosted and takes the load-delay filler. The target needs the inverse. The boost was PROVEN to be the dial (giving bank << 14 a second set moved it 154 → 140) — and making v single-set is nonetheless unreachable: every spelling that does so lets combine fold (subreg:QI (plus v -256)) → (subreg:QI v), killing the branch and four instructions (measured twice), and vv = v copies are copy-propagated back to two sets. A dial you can prove and cannot turn is permuter fuel, not a wall — record the mechanism, hand it to the permuter, and do not spend more agent turns on source spellings.

§461 — LAUNDERING AN INVARIANT CAN BE THE DEFECT, AND "RESIDUAL A" DOES NOT GENERALIZE PAST ONE BINARY OP

Source: the S76 main:func_80039B20 agent (79 ins, prior best 16 → 10). Two findings, one of which is a scope correction to an entry already in this file.

1. Do not launder every loop invariant — laundering the WRONG one displaces the address chain. The prior attempts wrapped p = D_800C6DD0 + (s16)i * 0x60 in a volatile-asm launder on the theory that every hoistable invariant needs one. That was actively wrong: it displaced the address chain relative to the base lui/addiu and cost the entire first cluster (closeness 16 → 81 when present). The already-matched sibling func_8003A0E4 in the same TU uses the plain unlaundered idiom, and copying it closed that whole region.

Meanwhile D_80073140 in the SAME loop genuinely does need its launder — without it, that invariant and D_800C6DD4 are both move_movables-hoisted, because the two share one insn_count threshold (§193-F). So the rule is per-invariant, not per-loop: check the matched siblings before adding a barrier, and treat "add a launder" as a lever that can go backwards, not a free safety measure.

2. SCOPE CORRECTION — "Residual A" (combine_regs / first-dying-operand, L875) applies to a SINGLE binary op and does NOT transfer to a PLUS CHAIN. The target puts *(s32*)(p+0x50) + f6*26 in $v0; the draft gets $v1. Residual A says to swap the C operand order, since gcc-2.7.2 has no swap_commutative_operands and source order survives. Measured here: on a 3-term PLUS chain, swapping the operands produces byte-identical output — fold.c's associative-PLUS canonicalization normalizes the chain before combine ever sees it, which a single OR never gets. Forcing an explicit named accumulator instead ripples registers elsewhere (a rename plus −1 length drift). The lever is real for one op and a dead end for a chain; do not spend turns re-deriving that.

Still open (permuter-class, Law 3): the $v0/$v1 tie above, and a redundant D_80073140[i] re-read the target schedules early to fill a load-delay slot while cc1 schedules it at its use. Every C-level attempt to move it either CSE'd the two reads into one (−1 ins) or added a move (+1).

§465 — THE ASPSX SLOT-HOP: A GAP OUR REORDER SUBSTITUTE CANNOT CLOSE (main:func_8005F830)

Source: the S76 agent, 152 of 153 instructions byte-exact (indices 0-102 identical), 10 controlled probes in scratch/{t3,q,r,p}.s plus a hand negative control in scratch/g.s.

The shape. The target hops addiu $v0,$zero,0xFE — the head instruction of the branch's OWN target block — up into the beqz $v0,.L8005F9F8 delay slot.

Why no C can produce it, in two measured halves:

  • cc1 will not. reorg's fill_slots_from_thread REFUSES a thread insn that writes the register the conditional branch TESTS. Probes f4/a8/a9/b3 (branch tests $2, thread head is li $2,K) all leave the slot bare; probes f5/a6/a7/b2 (differing registers) all get the insn stolen. That is the rule, isolated.
  • GNU as -O2 will not either. Its reorder only swaps with the PRECEDING instruction, never hops one up from a branch target. Negative control scratch/g.s: gas emits a nop, never the hop.

So this is the original ASPSX assembler's reorder doing something our REORDER_TUS substitute (reorder_passthrough.py + as -O2) structurally cannot. It is the §182/§188 epilogue class one level deeper — an assembler gap, not a source-shape defect, and not reachable by any C spelling.

Two corrections that come with it: the older "epilogue unreachable" verdicts recorded against this function are STALE — match_one's REORDER_TUS path (S76) already packs the epilogue correctly, and the epilogue is not the residual. And the only banking route left is §265, which src/800c3.c already is. The draft carries TU-verbatim decls and is bank-ready if the assembler gap is ever closed.

§466 — main (509 ins, -O0): ADDRESS CONTEXT EMITS mult INDEX-FIRST

Source: the S76 agent that matched main itself. The headline law is new and general:

Inside a MEMORY ADDRESS, base + i*K (constant K) expands to a (mult reg K) rtx that force_operand emits INDEX-FIRST — addu d,index,base. Rewriting it as base + ((i * (K >> n)) << n) materialises the index first and yields the target's BASE-first addu $v0,$s1,$a0 / lui %hi; addu $at,$at,idx; sw %lo($at). Value context is unaffected, which is why this hides: the same expression is fine everywhere except under a MEM.

Supporting -O0 idioms from the same match:

  • 12-byte-strided stores need a struct COMPONENT_REF — only that folds to sw …,8($v0).
  • s32 pad[6] supplies the 24 bytes of dead -O0 locals that make the frame 0x38.
  • A second, unused register u8 *q = &D_800BA118 keeps the $s0 lui/addiu + save alive.
  • (*(u16*)x)++ emits the extra move that += 1 omits.
  • CatPrim's arg2 must be <load> + D_80074778*4: with a MEM as operand 0 the address is emitted first, then operand 1, then the load, so the addu comes out operand-1-first.

§464 — FOUR VOLATILE/BARRIER LEVERS FROM main:func_8005DE78 (141 ins → MATCH)

(Recorded after §465-§466: the original append was lost to a git index-lock race and the commit message that claimed it landed before the text did. Content is unchanged from the agent's report.)

Source: the S76 agent on func_8005DE78. Structure came from the banked in-TU neighbours (func_8005DCA0's shape, func_8005FBC8's body inlined as the loop test) — §2b "read a matched neighbour first" paying out again. The four levers are new, and each names its mechanism.

1. A VOLATILE QI/HI LOAD PRESERVES THE ZERO-EXTEND AS ITS OWN INSTRUCTION. Normally combine folds a u8 -> s32 promotion into the lbu. Make the load volatile and it cannot, so the extend survives as a separate andi $s2, $v0, 0xFF — which is what the target has. Reach for it when your build is one andi short around a byte load.

2. if (A || B) { t } else { K } VS && SELECTS WHICH ARM IS do_jump's DROP-THROUGH. Same truth table, different branch layout. When the residual is "right tests, wrong order", flip the connective before touching registers.

3. 🔴 A VOLATILE STORE CAN NEVER BE STOLEN INTO A DELAY SLOT. reorg's resource_conflicts_p returns 1 on ANY volatil resource, so the slot stays empty. This is how you reproduce a target nop after a j — otherwise very hard to force, since every ordinary statement is a candidate.

4. A "memory" CLOBBER AND A VOLATILE READ ARE NOT INTERCHANGEABLE CSE-BREAKERS. Both force the lbu 0x44 index to reload. But the volatile read lands the byte in $v0 while the "memory" clobber lands it in $v1 without flipping the addu operand order — and that operand order was the entire final 2-instruction REGALLOC-PERM residual. When a reload lever fixes the reload and breaks the register, try the other spelling before calling the residual permuter-class.

§467 — GLOBAL-ALLOC TIES BREAK ON DECLARATION ORDER, AND A COPIED CLOBBER LIST IS A DEFECT

Source: the S76 agent on main:func_8001EFE0 (468 ins, 172 → 89). Three reusable mechanisms.

1. 🔴 WHEN SEVERAL EQUAL-PRIORITY PSEUDOS TIE IN GLOBAL-ALLOC, DECLARATION ORDER BREAKS THE TIE — NOT ASSIGNMENT ORDER. Four getTPage bases all tie; reordering their declarations moved 36 instructions, and it also changed CONTROL FLOW — with one base spilled, an arm's reload broke the tail that jump2 had been cross-jumping. So this lever is not cosmetic and its blast radius is not local: re-check branch shape after using it, not just register names.

2. A CLOBBER LIST COPIED FROM A NEIGHBOUR IS A LIABILITY. The prior draft carried a phantom "$2" in an inline-asm clobber. The real macros (proven from the matched neighbour func_800221A8, §194-E) clobber only $12/$13/$14. That invented $2 evicted abr out of $v0 and cost 14 instructions. Copy a neighbour's macro BODY if you must, but verify its clobber list against the target's own register usage — an over-broad clobber is invisible in the C and expensive in the asm.

3. convert_to_integer SHORTENS A NARROW-LOOKING EXPRESSION TO QImode AND DROPS THE andi. q[0x14] = q[0xC] + ((rec >> 16) & 0xFF) lost its mask because the result was shortened; an explicit s32 temp for the sum stops the shortening and restores the andi.

Residual left (permuter-class): (sy & 0x100) >> 4 emits srl; andi — combine will move a shift inside a single-bit mask but not inside 0x300/0x3C0, and both C spellings converge on the same output, so this is a combine limitation rather than a spelling choice.

§468 — THE %lo-FOLD EXTENDS TO STORES ONLY VIA extern Struct SYM[], AND MASKING HID THE OPERAND ORDER

Source: the S76 agent on ov_SC01_001:func_80181E04 (269 ins → MATCH).

1. §18's %lo-fold works for STORES only when the symbol is declared as an array of struct. SYM[i*20] on a plain s32[] folds only for read-only symbols; declaring each symbol extern Struct SYM[] with stride 0x50 and the field at +0 makes the fold apply to stores too. Measured at 13 instructions on this function. That is a real extension of the Phase-20 §18 entry, which only ever exercised the read side.

2. 🔴 RELOCATION MASKING CAN HIDE A WRONG OPERAND ORDER. The kill test needed D_801EDA4C[i] > D_801EDA58[i]; the reversed spelling scores identically under match_one, because §1c masks HI16/LO16 immediates and the two symbol references mask to the same bytes. It was caught only by reading the raw relocation list. When a comparison's operands are two different symbols, the byte oracle is blind to which is which — check the relocations, not the score.

3. Two smaller levers from the same match. Do not introduce a biased q pointer: write the prim bytes off p so combine_givs picks p+0x12 itself, otherwise it mints a second anchor at p+8 (+2 ins). And keep the counted i < 0x100 loop — spelling the bound via D_801F2A44 costs 12 instructions, even though the relocation resolves to the same address (D_801EDA44+0x5000).

ADDENDUM to §461 (third instance) — A REGISTER PIN CAN BE THE DEFECT TOO

main:func_80040DE8 went 86 → 2 when §76 variable-reuse (a = {q1,q3,rr} / b = {q2,ll} as reused globals, squares written back into vl/vr) pushed o1 off $a3 onto $t0 — and that made the §3-C pin unnecessary. The pin had been tying ~30 instructions into $t0.

So all three "helpful" levers in this family can each be the thing holding a match back: a volatile launder (§461), a temporary (§462 lever 3), and now a hard-register pin. Before adding a lever, check whether an existing one is what you are fighting — remove first, then measure.

(Residual: REGALLOC-PERM $t1 > $v1 on the base load; eleven two-variable spellings all restore the register but flip the entry sched1 order for +12.)

§469 — THE MEM_IN_STRUCT_P ALIAS UNLOCK (and §463's spill law, independently confirmed)

Source: the S76 agent on main:func_80039308 (518 ins, 402 → 154, length exact).

🔴 WRITE A VARYING-ADDRESS LOAD AS A STRUCT MEMBER TO BREAK A FALSE ALIAS. Spelling the voice-mask load ((VMask *)q)->w instead of *(u32 *)q sets MEM_IN_STRUCT_P, which lets gcc-2.7.2's true_dependence prove that an in-struct varying-address load cannot alias a scalar-global symbol store. Both loads then hoist above both stores and three load-delay nops vanish. The C is semantically identical; only the alias information differs. Reach for this whenever loads refuse to hoist past unrelated global stores.

INDEPENDENT CONFIRMATION OF §463. This agent needed a 16-byte s16 sav[8] memory local for arg1 because "reload rounds every spill slot to BIGGEST_ALIGNMENT = 8, so the target's 0x0/0x8/0x10 layout is unreachable by spilling alone" — the same law §463 derived from alter_reg / assign_stack_local(mode,size,-1) on a different function, found by a different agent that had not seen it. Two independent derivations of the 8-byte spill slot make it one of the more solid frame laws in this file, and it now has a second use: it tells you when a stack layout can only come from a declared local, never from spilling.

Residual (regalloc-only): one whole-function s/t register permutation, one redundant addu $v1,$s4,$zero copy that no spelling survives CSE (every variant either drops it or emits andi), and gcc strength-reducing D_80073140[i] in the else-loop where the target does not.

§470 — FOUR CSE/SCHED LEVERS FROM main:func_800301C8 (170 ins, 133 → 18)

Source: the S76 agent. Four levers that only work as a set; the second is counterintuitive.

1. A STORE-THEN-READ-BACK TURNS A REDUNDANT LOAD INTO THE TARGET'S REGISTER COPY. Writing D_800A46CE[0] / &D_800A46D0 and reading it back lets cse replace the second load with a register copy — which is what the target has.

2. 🔴 USE TWO DISTINCT LOCALS FOR THE SAME b*24. cse resets at the if-join, so the original source recomputes the product into a second register. One shared local cannot reproduce that. And writing b*24 inline is worse still: cse then hoists the %hi/%lo symbol address into a pseudo and changes the addressing mode. Duplicating a subexpression can be the correct decompile — the reflex to factor it into one local is wrong here.

3. A zero-byte __asm__("") fence (§194-A) stops sched1 hoisting the two = 0 stores into the load-delay slot.

4. WRITE REPEATED TAILS OUT SEPARATELY AND LET cross_jump MERGE THEM. Three explicit D_800A46CC = N; return 0; tails merge into the target's shared suffix; funnelling them through one st variable emits the arms inverted.

Residual — three allocation facts, all unreachable from C: the $17 pin on k2 is required (without it k2 splits across two callee-saved registers and costs a fourth — the local-vs-global allocno class choice) but it drags the *24 shift chain into $s1; plus operand canonicalisation on addu $s0,$s1,$s0; plus the else-arm materialising &D_800A46D2 into $s1. Law 1c verified: all 25 symbols match the target's relocations, and the only difference is D_800A46D2's count (4 vs 6), which is that residual rather than a wrong symbol.

§471 — A LAUNDER'S REAL COST IS AN ALLOCNO, AND $t0 IS RELOAD'S (main:func_80032A74, 408 → 12)

Source: the S76 agent. Three facts, two of which refine entries already here.

1. §172's combine USE-orphan buys a never-referenced stack slot. Promoting an s16 memory local to int twice after a CODE_LABEL is the only way to produce the target's unused 8-byte slot at 0x48. Pair this with §463/§469: when a frame has a slot nothing reads, it is either a spill (8-byte rounded) or an orphan of this kind — both are reproducible, neither is padding.

2. 🔴 REFINEMENT OF §153: THE LAUNDER'S REAL COST IS AN ALLOCNO THAT OUTRANKS THE VALUE YOU CARE ABOUT. The §153 launder was necessary here, but it created an allocno outranking vol on priority. The cure was pinning the launder to $10 — and specifically not $8, which evicts reload's $t0 parameter reloads. So "the launder is the defect" (§461) has a third resolution beyond remove it or move it: pin the launder itself, and pick the register with reload's own needs in mind.

3. $t0 IS UNREACHABLE FROM C BECAUSE RELOAD OWNS IT. The target's table / D_800A4C28 bases are reload rematerialisations of a reg_equiv_constant living in $t0. No C spelling or pin can put a value there, because reload claims it for its own reloads. When a residual is "the target uses $t0 and I cannot", stop — that is a reload artifact, not a spelling you have not found yet.

(Also: merging vol and m into one variable puts the volume chain in $a0. Residual 12 is this $t0 rematerialisation plus one lh/lhu row that is the orphan's only available site.)

§472 — 🔴 §148-A's HOIST THRESHOLD IS 29, NOT 58, WHEN THE LOOP CONTAINS A CALL

Source: the S76 agent on main:func_8001EA14 (371 ins, 349/303 → 89, length exact), cracked with cc1 -dL. Four findings; the first corrects a number this file has been quoting.

1. THE loop.c HOIST THRESHOLD IS CALL-DEPENDENT. §148-A's rule is stated for a threshold of 58; when the loop CONTAINS A CALL it is 29. And the two inputs are not what the name suggests: savings is the count of MATCHED movables, and lifetime is their SUM. Read them off cc1 -dL rather than estimating — this agent used the dump to arrange for &vo to be the only surviving hoist (a 2-operand gte_ldv3 plus swapping gte_rt's operands removed the others). Anyone applying §148-A to a loop with a call and getting the wrong answer has been using the wrong constant.

2. MEM_IN_STRUCT_P RUNS BOTH WAYS — see §469 for the other direction. There, spelling a load as a struct member unblocked hoisting. Here the target's schedule is alias-blocked, so the fix was the opposite: access prim through plain casts, not a struct, to keep MEM_IN_STRUCT_P clear. One knob, two directions; decide which the target needs before reaching for either.

3. THE COND_EXPR SINGLETON FOLD IS ESCAPED BY MAKING THE ARMS DIFFERENT TREES. gcc folds X ? A op B : A into a single operation. Changing values is not enough — the arms must be structurally different trees, e.g. (uy - 0x100) : (u32)uy.

4. SPILL SLOTS FOLLOW DECLARATION ORDER (mode before base). With §463 (8-byte rounding), §469 (a layout only a declared local can produce) and §471 (the §172 USE-orphan), the frame model is now: slot size is BIGGEST_ALIGNMENT-rounded, slot order is declaration order, and a slot nothing reads is a spill or an orphan — never padding.

(Residual: sched1 reordering the matrix-init block, A/B-proved with -fno-schedule-insns, plus a $t0/$v0 allocation knock-on — see §471 on why $t0 is not yours.)

§473 — 🔴 §265's "HANDWRITTEN" VERDICT FOR ov_SC07_002:func_8017DC80 IS REFUTED (324 → 89)

Source: the S76 agent. This file carries a §265 addendum asserting that func_8017DC80 is a handwritten function no -O2 C can match, on the strength of its interleaved sw/def prologue. That prologue is ordinary gcc-2.7.2 MIPS RTL. The draft went from a 20-attempt LENGTH-DRIFT/-33 wall to −2 / closeness 89, which is not what an unmatchable function does.

What moved it: §30's /s-dep lattice — plain scalar sxy stack locals combined with COMPONENT_REF packet stores through a POLY_G4/LINE_G2 struct pointer — plus deliberately un-cached *(s32*)(c+0xB) reloads and a recomputed OT pointer.

Why it matters beyond this function. That is the fourth wall refuted in one session, after the §182/§188 reorder oracle, §41b's prologue-hoist (§463) and the S75 nine. Three of the four were recorded as properties of the CODE and were actually properties of an instrument or a model. The standing lesson stands up better each time: in this project a recorded wall is more often a stale belief than a compiler limit — re-probe before honouring one.

Manifest consequence: func_8017DC80's UNCERTAIN row should resolve toward decompilable, NOT PERMANENT-VERBATIM. Converting it to a stub (S76) was correct, and it belongs in the drawable pool.

(Residual 89: the OT pointer still CSE's across func_80010A08(8) — only a barrier-preceded diamond-merge label flushes cse2, and every zero-byte flush tried killed either the tail cross-jump or the fp loop-invariant hoist; plus a 2×5-instruction mode-copy allocation in the OTZ clamp and the $a0/$s7 prologue schedule.)

§474 — A PROVED C-LEVEL FLOOR: split_tree + stupid.c (-O0), from main:func_80011380

Source: the S76 agent, which upgraded an empirical closeness-6 plateau to a floor PROVED from tools/reference/ gcc sources. The counterweight to §473 and friends: most recorded walls in this project have been stale, but this one is real and now has a proof, so it should never be re-ground.

The demand. The target needs MULT(MULT(i,2),2) left unmerged.

Why no C spelling delivers it. fold-const.c:882 split_tree decomposes ANY MULT whose op1 is TREE_CONSTANT. All 20 anonymous / identity / array spellings measured collapse to a single sll 2 (190 ins), and STRIP_NOPS eats NON_LVALUE_EXPR, so the usual shields — |0, +0, *1, &~0, ^0, >>0 — cannot protect the inner multiply.

Why the two escapes each cost an instruction.

  • A statement-expression produces the EXACT 5-instruction RTL (confirmed with cc1 -dr) — but its BLOCK_END note lands between the sll and the next copy. stupid.c:497-508 computes dead = max(last_use, born+2) with occ = [born, dead-1], so a copy conflicts with its source only when immediately adjacent; the note breaks the adjacency, the copy self-coalesces, and final.c deletes it (191).
  • (t = i*2) * 2 with register s32 t is a genuine fold blocker and reaches exact length 192 with the exact instruction shape (closeness 12, the best new construct) — but expand_decl / use_variable emit zero-byte (use) brackets that make t the longest live interval, so it seizes $v0 and rotates the whole {v0,v1,a0,a1,a2} ring one step.

The clinching proof that the target has NO variable there: its 4th instruction sll $v1,$a0,1 reads $a0, not insn 2's destination. A variable would have to be the deleted move, and t can never be. Hard-register pins $2..$5/$8/$9 all keep the copy (193); a stack t costs 194.

Bonus law, same oracle — the la-on-$a0 colour. expand_binop allocates the PLUS destination before force_reg'ing the symbol, so the symbol's pseudo gets the higher regno and loses stupid_reg_compare's tie-break. That explains a -O0 register colour that looks arbitrary.

Use this section as the template for a wall claim: name the pass, cite the file and line, show the measured cost of each escape, and give the byte-level fact that rules out the alternative. A wall asserted without that is a belief (see §473).

§475 — THE "memory" FENCE AS A CSE INVALIDATOR, AND (b*3)<<3 INSTEAD OF b*24

Source: the S76 agent on main:func_8002FF0C (166 ins → MATCH, verified in-TU: a spliced copy of src/800_b.c compiles rc=0, function byte-identical, all 63 relocs the target's own symbols).

1. 🔴 __asm__ __volatile__("" ::: "memory") IS A CSE MEMORY-TABLE INVALIDATOR, NOT ONLY A SCHEDULING FENCE — AND THE COLON-LESS FORM DOES NOT DO IT. Used here to force D_800A46D2 to be re-read (lui $a0 / lh $a0) for the func_800419B0 argument instead of folded to sign_extend(r). Without it the function is exactly two instructions short, and the bare __asm__("") fence does not substitute. Pair with §464 lever 4, which found the same two spellings are not interchangeable for a register outcome; this is the memory-table half of that distinction. Pick the spelling by which table you need invalidated.

2. 🔴 WRITE (b * 3) << 3, NOT b * 24. expand_mult never honours its target, so x = b * 24 leaves a move copy that survives into the join block — costing a sixth callee-saved register and a 0x30 frame. A top-level LSHIFT_EXPR expands directly into the variable's own pseudo and the copy vanishes. This is a general rule for any constant multiply that factors as odd << n.

3. INDEPENDENT CONFIRMATION OF §470. This agent also needed a separate i = b*24 for the pre-join region, because the cse EBB ends at the join and the target computes b*24 twice — exactly §470's "use two distinct locals" finding, reached on a different function by a different agent that had not seen it. Two independent derivations; treat it as established.

4. THE HOUSE ARRAY SPELLING CAN BE THE DEFECT. D_800A46D2 must be declared scalar at block scope; the house style extern s16 D_800A46D2[] forces la (address into a register) for both accesses and costs 12 mismatches. A fleet-consensus declaration is a strong prior, not a law.

(Also load-bearing: q = D_800A46CE as a real pointer local, matching the target keeping &D_800A46CE in $s4 and deriving the D_800A4642 store as addiu $v1,$s4,-0x96 / addu $v1,$s0,$v1 / sh $v0,0xA($v1); plus two §194-A AFTER-placement fences.)

§476 — 🔴 A HARD-REGISTER PIN DESTROYS TWO THINGS COMBINE AND SCHED1 NEED (func_800226C0, 670 ins → MATCH)

Source: the S76 Fable agent on main:func_800226C0 — at 670 instructions the largest function in the project, matched at closeness 0. This completes the "a lever can be the defect" arc (§461, §462, §461-addendum, §471) with the mechanism that explains why pins so often hurt.

A pinned hard register loses NONZERO-BITS information. The target's 228E4 chain (addu $v0,$s2 / beqz / addu $t2,$v0) is combine folding sext(HImode t) into a copy, with cse2 then reusing it as the loop multiplier. That fold needs nonzero_bits on the pseudo — and hard registers do not carry nonzero-bits. So the $18 pin that looked like the obvious way to place t was exactly what prevented the fold. The fix was one plain uninitialised s16 t, with mul left as an unpinned pseudo.

A pin also breaks the §199-A birthing boost. birthing_insn_p requires reg_n_sets == 1; a pinned $a1 mask has reg_n_sets != 1, gets no boost, and is therefore placed first — visible as a whole-block schedule difference. Unpinning o/col/sh23/abr fixed the prologue order, a t/flags density tie and an or destination tie.

So the rule is not "pins are risky" but a specific two-part mechanism:

A hard-register pin (a) strips nonzero_bits, disabling combine folds that depend on a value's known width, and (b) makes reg_n_sets != 1, disabling the sched1 birthing boost. If a residual involves a sign/zero-extend fold or a first-in-block placement, remove pins before adding them.

(Also: the 850 single-prim block written as an inline addPrim with block-local temps, and the function defined with the TU's typed prototype (Obj_80021D38*, u8*, DVec_80021D38*, u8*) to avoid the §41 declaration wall.)

§462 — FOUR LEVERS FROM main:func_80024054 (91 ins, 74/53/32 → 4)

Source: the S76 agent on func_80024054. Four independent unlocks, none previously recorded; the last one is the kind of rule that silently costs a whole attempt.

1. array[var - K] folds K into the symbol's LO16 / lhu displacement. Naming an intermediate idx = var - K does NOT stop it — the fold happens at the front end / in combine, before any register assignment you could steer. The only thing that defeated it was a zero-byte opacity barrier immediately after computing the index, one per use site: __asm__ __volatile__("" : "=r"(idx) : "0"(idx)); Reach for this whenever the target loads from sym+0 with a computed index and your build folds the constant into the displacement instead.

2. The fused sll 16; sra 15 sign-extend-and-scale wants the index declared s16, not s32. This confirms §241's recipe on a fresh case — worth knowing it reproduces rather than being a one-off of that function.

3. A mask-then-compare LOCAL causes a cross-jump merge AND flips branch polarity. Writing bits = val & 0xC000; then if (bits == …) else if … merged two case tails into one shared block and emitted bne-polarity branches where the target has beq. Dropping the local and switching on the expression directly — switch (val & 0xC000) { case 0x8000: … case 0xC000: … default: … } — fixed both at once and reproduced the target's forward-beq shape. The temporary was the defect; this is the same family as §461's "laundering can be the defect", from the opposite direction.

4. 🔴 A POINTER PARAMETER'S SIGNEDNESS DECIDES HOW -1 IS MATERIALIZED. Declaring arg1 as s16 * rather than u16 * flips the fail-path constant from ori $x, 0xffff to addiu $x, -1, matching the target — because gcc-2.7.2 canonicalizes the RHS constant against the lvalue's signedness before choosing the load-immediate opcode. Nothing about the store's value changes, so this is invisible in the C and shows up only as a one-instruction opcode difference. If a residual is a lone ori 0xffff vs addiu -1, check the signedness of the pointer being written through before touching anything else.

Left open: one DELAY-SLOT residual — addu $a3,$zero,$zero is insn #0 in the target and lands in the branch delay slot in every C variant. Two independent prior attempts hit the same wall; five further variants (statement reorder, register pin, barriers either side of the load) each left it unchanged or traded it for an equal residual elsewhere. Permuter-class, Law 3.

§463 — 🔴 SPILL SLOTS ARE 8 BYTES, AND THE §41b "LOAD ABOVE THE PROLOGUE" WALL IS REFUTED

Source: the S76 agent on main:func_8001FC08 (400 ins, 33 → 0 MATCH). Three laws, and the second one deletes a wall this file has been asserting.

1. A 4-BYTE GAP IN AN OTHERWISE 4-PACKED FRAME IS A SPILL SLOT, NOT A PAD. sp+0xC8 / sp+0xD0 in this target are not struct members — they are spilled pseudos. reload's alter_reg calls assign_stack_local(mode, size, -1), and align == -1 means BIGGEST_ALIGNMENT (8) with CEIL_ROUND, so every 4-byte spill slot occupies EIGHT bytes. That is exactly why the target's two slots sit 8 apart with 0xCC/0xD4 untouched. Reading those gaps as padding — or as fields of a struct you then invent — is a wrong model of the frame. Worth 11 instructions here, and modelling them as spills is also what evicts both values from local-alloc so reload picks $t0.

2. §41b's "a global load cannot float above the RTL prologue" IS NOT A WALL — it is an $a0 ANTI-DEPENDENCE. The parameter copy addu $s0, $a0, $zero reads $a0, which pins the load below it. Get the value out of $a0 and make the load the first statement, and it floats to idx 0 on its own. Both moves are required and either alone is worthless — statement-first by itself measured 33 → 50 (worse); combined with law 1 (which is what frees the register) it went 22 → 4. Before treating a "load above the prologue" residual as unreachable, check what reads the argument register.

3. Which ARGUMENT POSITION a guard value is passed in decides its hard register. Passing it as arg 1 — func_80021120(&L.cnt, L.lp) — gives that pseudo a qty_phys_copy_sugg toward $a1, which local-alloc's scan-from-$v0 can never reach on its own. The sibling guards that do not pass it stay in $v0, which is the control proving the mechanism rather than a coincidence.

Banker caveat for this function: its INCLUDE_ASM is at src/800.c:11168, but the TU's own MTX_80020248 typedef (:11180) and the D_80074818/D_80075018 externs (:11191-2) are twelve lines BELOW it. Hoist that block above :11168 or drop the draft's copy, or the duplicate typedef is a hard C89 error at bank time.

Verified by hand (law 1c): 26 jal targets and 16 HI16/LO16 relocs identical in name and order; the four D_1F800020 words are the scratchpad literal, byte-identical (3c111f80 / 26310020).


§477 ★★★ — THE self_decl_tu CLASS IS A SOLVED, MECHANICAL LANE: 16 DRAFTS, 16 BANKS (P31 S77)

The class. The destination TU declares the very function the draft defines, with a different signature — tu void () | def s32 (Ctx*, s32). cc1 rejects the TU, the draft never compiles, and the failure is recorded as CC1-FAIL, which says nothing about the body underneath. blocker_probe names it self_decl_tu and tiers it T1 (binary-local).

The chain, unchanged from §378, but note what it is NOT. sync_tu_decls refuses this class by design (the call SITES must change too, which it does not do). The tool that handles it is cast_self_callers --sync-decls, and — the thing the S76 checkpoint got wrong — it is binary-generic already: src_files(binary) globs src/<binary>/*.c for an overlay and src/*.c for main. Nothing needed extending; sync_tu_decls is the only main-only tool in the chain.

MEASURED YIELD — the whole point. Every draft the probe put in this class banked:

cohort drafts banked
main (blocker_probe --binary main, 36 stranded) 7 4 (3 were NEARs — body, not plumbing)
overlays (blocker_probe over 26 binaries) 12 12 — 100%, 136 s wall, 8 workers

Of main's three failures none was a plumbing failure: func_80015608 closeness 3, func_80015760 closeness 9, func_80039DEC 8 differing bytes. The class has no observed plumbing residual. Route it mechanically and spend zero agent tokens on it.

LAW 1 — PROVE THE PLUMBING IS BYTE-NEUTRAL BEFORE YOU GATE. The casts and synced declarations are supposed to emit no code. Build the binary with the edits applied and no draft substituted; it must produce its locked SHA. main did (143dbb89…), and all 12 overlays did. This costs one build and converts every later gate failure into a statement about the draft — which is the whole reason the §20 fold is worth using instead of editing signatures by hand.

LAW 2 — A SYNCED DECLARATION MUST PARSE WHERE IT SITS. --sync-decls copied the draft's parameter list verbatim into the TU. A draft names types the TU does not have in scope at that line, and both flavours broke the committed baseline in one apply:

src/800.c:2631   extern void func_80015760(Obj_80015760 *obj, s32 *ot);  // draft-local type
src/800c3.c:866  s32 func_8005E3AC(Ctx *s, s32 size);                    // Ctx typedef'd at :941
=> src/800c3.c:866: parse error before `*'

Once the call sites are cast the declaration emits no code, so it only has to be compatible and parse. <ret> fn(); satisfies both without naming a type, and C89 6.5.4.3 makes it compatible with a prototyped definition exactly when no parameter is affected by the default argument promotions. So: draft spelling first (byte-proven, and informative), no-proto only where that cannot parse, refuse loudly where it cannot parse and a narrow parameter forbids no-proto. Never widen the fallback — a blanket no-proto churns 15 already-correct declarations to buy 3.

LAW 3 — A DEFINITION IS A DECLARATION. sync_tu_decls looked only for extern … sym …;, so whenever the clashing symbol is a function the TU defines it stopped with "no extern line to copy". That was the terminal blocker of BOTH remaining main drafts. The definition header is the authoritative spelling — it is the one cc1 checks every other declaration against — and must be preferred over an extern when both exist:

src/800c3.c:916  void func_8005E480(void *arg0) {   ->  extern void func_8005E480(void *arg0);

LAW 4 — A STATEMENT KEYWORD IS NOT A RETURN TYPE. return func_X(a0, a1); has the exact shape of a forward declaration, so a permissive <type> <fn>(...); regex reads a CALL as a DECLARATION. In cast_self_callers this hit three consumers at once, in opposite directions: is_declaration made cast_sites SKIP the site, sync_decls REWROTE the statement into a declaration (deleting the function's return), and DEF_RE read it as the definition itself. Blast radius 524 return func_X(...); lines across 482 files. One shared _kw_prefixed() guard, three call sites.

LAW 5 — CASCADE. Every bank gives its TU a real definition that contradicts the stale extern each later draft in that TU still carries. Banking func_8005DE78 is what blocked func_8005EB28. Re-run the sync after each bank; never conclude the draft went bad.

HOW TO RUN IT (the whole lane, ~4 commands):

tools/blocker_probe.py --binary <bin> --drafts <wave dirs> --json .run/bp.<bin>.json   # read-only
tools/cast_self_callers.py --binary <bin> --funcs <self_decl_tu fns> --drafts <dir> \
    --sync-decls --apply --journal .run/cast/<bin>.json
make check BINARY=<bin>          # LAW 1 — must be green with NO draft substituted
git commit                       # gates pin a worktree / checkout the TUs; uncommitted edits die
tools/parallel_gate.py --plan <plan> --workers 8 --commit
tools/cast_self_callers.py --undo-journal .run/cast/<bin>.json --keep <banked fns>

§478 🔴 — A VERBATIM DRAFT IS THE STRONGEST FALSE SIGNAL YOUR SCOPING TOOL CAN EMIT

A §265 verbatim draft is the target's own asm in a file-scope __asm__. It assembles to the bytes it was copied from, so every body oracle reports the strongest possible result: match_one closeness 0, rtu_match MATCH, blocker_probe static none. The byte gate then refuses it for free and progress.py moves by exactly zero.

S76 closed this hole in gate_main, harvest_verify and api_agent.prior_draft. It stayed open in blocker_probe — the tool that SCOPES the work — and that is the expensive place to be blind: of the 13 MATCH rows in the S77 overlay pool, six were verbatim, and they had been ranked as the highest-value drafts available. The cohort gated 0/13 and read as a model failure until harvest_verify printed its SKIP line.

The law: every consumer that reads a draft and emits a quality signal — gate, warm-start supplier, scoper, ranker — must run the same detector (draft_prechecks.is_verbatim_asm_draft), and a verbatim row must be excluded from match/agreement arithmetic rather than counted as a match. When you fix a blindness like this, enumerate the consumers: this was the fourth, found only because the third fix did not prompt anyone to ask who else reads drafts (R36's shape, applied to a property rather than a binary).

§479 ★★★ — WHERE THE PERMUTER ACTUALLY PAYS: A MEASURED YIELD CURVE (P31 S77, 8 candidates)

Every candidate below is a main DIFF residual with no prior-attempt journal — i.e. nobody had ground it. permuter_ils --klass SCHEDULE/REGALLOC --j 12, 150 s cycles:

residual (mismatched) class outcome
2 of 68 SCHEDULE-REORDER score 0, cycle 1 → BANKED (func_80021174)
2 of 347 REGALLOC-PERM score 0, cycle 1 → BANKED (func_80040DE8)
4 of 91 DELAY-SLOT score 0, cycle 1 → BANKED (func_80024054)
10 of 79 schedule plateau at 10, 4 cycles
11 of 74 schedule no waypoint at all — never beat base
20 of 100 schedule plateau at 14
28 of 69 schedule plateau at 25
37 of 114 schedule 37 → 1 then plateau at 1 over 8 cycles × 240 s

THE CURVE — CORRECTED THE SAME SESSION, ON A BIGGER SAMPLE. Read the correction, not the first draft of this entry. The first version of §479 said "at ≤4 mismatched the permuter is a one-shot: 3/3". Five more ≤4 runs later that is 3 of 8 (37.5%), and the failures are not marginal:

function residual permuter best outcome
func_80021174 2 0 BANKED
func_80040DE8 2 0 BANKED
func_80024054 4 0 BANKED
func_8005F290 1 4 fail
func_80039DEC 3 2 fail
func_80015608 3 1 fail
func_8005F0C8 3 3 fail
func_8005ECC0 2 5 fail

The mismatch count did not predict a single one of those outcomes — a residual of 1 failed and a residual of 4 banked.

AND NEITHER DOES ANYTHING ELSE I COULD MEASURE. This paragraph is a CORRECTION of a correction. v2 of this entry claimed the predictor was "prior-attempt history: all 3 winners were drafts nobody had worked." Building tools/permuter_sweep.py on that claim refuted it in one negative control: journal_notes — the same index the packs use — reports prior attempts for all eight, winners included (2, 3 and 3 for the three that banked). What I had actually eyeballed was the DRAFT HEADER narrative, and the two are different corpora; the winners came from a recovery pile whose files carry no header journal, which is provenance, not evidence.

So the honest state is: ~3 in 8 at ≤4, and no validated predictor. Select on the two NECESSARY conditions — a small residual, and a match_one class the permuter can actually search (SCHEDULE-REORDER / DELAY-SLOT / REGALLOC-PERM; a STRUCTURAL, WIDTH or LENGTH-DRIFT residual is a different animal and a run on one is waste). At 150 s a cycle that yield is worth having; just do not believe a story about which ones will win. permuter_sweep.py implements exactly that and prints prior-attempt counts as information rather than enforcing them.

Above ~10 it plateaus regardless: the 37→1 case burned 32 minutes on cycles 2-8 for zero further progress. A plateaued score is a seed for a different tier, never a reason to run longer.

The meta-lesson is the one worth keeping. This entry was written, corrected, and corrected again inside a single session, and each version sounded reasonable. A yield table is evidence; a story about why the yield looks like that is a hypothesis, and it needs its own negative control before it goes in the cookbook — because the next session will act on it.

TRIAGE FIRST, AND IT IS FREE: READ THE DRAFT'S OWN HEADER. The two closest overlay residuals (2 of 119, 6 of 106) look like the best targets in the fleet and are not: each draft carries a byte-measured journal of ~10 refuted levers, one of them with an arithmetic proof of unreachability (expand_divmod's const bound is always ≥ the sra's bound, so no insn_count separates them). Main's residuals carried no journal at all. That asymmetry — not the mismatch count — is what predicted the yield. (R38, applied to a draft header rather than a failure ledger.)

TWO HAND LEVERS REFUTED BY BYTES, both worth not repeating:

  1. Moving a statement to reorder its instruction. In func_80021174 the residual was lh $a1,0($sp) / sra $a2,$v1,16 in the wrong order, and hoisting a0 = a0 >> 16 above the load is semantics-preserving (a0 is untouched between). It scores 23 mismatched at 67/68 ins — worse, because it lets gcc fold an instruction away entirely. The permuter's winning edit was a different kind of change: delete the a1 = *(s16 *)sp; temporary and inline the load into both comparisons.
  2. Reordering a commutative operand from the source. addu $s2,$v0,$v1 vs addu $s2,$v1,$v0 is a pure operand swap, and val = b + r → val = r + b changes nothing (all 3 sites): gcc canonicalises commutative operands by pseudo-register number, not source order. Forcing a fresh later-numbered pseudo ({s32 rr = r; val = b + rr;}) also changes nothing — cse folds the copy. The class name was telling the truth all along: it is REGALLOC-PERM, so the lever must move the ALLOCATION, not the expression.

§480 🔴 — A STATIC BLOCKER CLASS THAT THE REAL PIPELINE ALREADY REMOVES IS A PHANTOM

blocker_probe's Oracle A reported local_type conflicts (redefinition of u8, conflicting types for Blk16, Blk32_80180908) as the blocker on three drafts. The real gate removes that class before cc1 ever sees it — harvest_verify runs cdecl.strip_provided_typedefs and gate_stage additionally runs reconcile_tu. Gated for real, the three drafts' true classes were DIFF, DIFF, and "too few arguments to function" — and that last one is §378 step 2, which banked ov_SC01_084:func_80182A00 (207 ins) once the call sites were cast.

So a local_type row is not work; it is noise from an oracle that models a pipeline stage shorter than the real one. Same shape as §478: a consumer that reads a draft and emits a verdict must model what the gate actually does to that draft, or it will route real work to the wrong lane and invent work that does not exist. When a static class and the gate disagree, gate one and believe the bytes — it costs one build.

§481 ★★★ — conflicting types IS A SAME-SCOPE ERROR; ACROSS SCOPES IT IS ONLY A WARNING (P31 S77, main:func_8001FC08, 400 ins)

The law, and it is the escape hatch for most of the declaration-conflict class. gcc-2.7.2 raises conflicting types for X as a hard ERROR only when the two declarations are in the SAME scope. Across scopes it degrades to type mismatch with previous external decl — a WARNING, which the build already emits elsewhere (the unmodified TU emits exactly that for func_80012E0C). So a draft whose spelling cannot be reconciled with the TU's does not need reconciling: move the offending extern to BLOCK scope and the error becomes a warning.

Measured. func_8001FC08's body was solved in S76 and had never banked, and nobody had asked why. Spliced into src/800.c it produced 3 hard cc1 errors (rc=33 against a baseline rc=0): its file-scope typedef … MTX_80020248 plus extern MTX_80020248 D_80074818[]/D_80075018[] collide with the TU's own copies fifteen lines BELOW the INCLUDE_ASM. Two anonymous struct typedefs in one TU are never compatible in C89, so no duplicate spelling could ever have worked — the usual "adopt the TU's spelling" lever is structurally unavailable here. Renaming the struct to MTX_8001FC08 and demoting the two externs to block scope gives rc=0, +4 warnings, 0 errors, and the whole-TU .text is byte-identical to the all-INCLUDE_ASM build. No src/ hoist, no other function shifted.

Why this generalises. The session's dominant bank-blocker is a declaration disagreement, and the existing ladder answers it by making the two declarations AGREE (sync_tu_decls copies the TU's spelling; reconcile_tu casts at the use site). §481 adds the case those cannot reach: when the two spellings are incompatible by construction — anonymous structs, or a type the TU declares below the splice point — you do not need agreement at all, only different scopes. It is also the mechanism behind §8d's scope-demote lever, stated as the general rule rather than one tool's trick.

Verification standard this came with (law 1c, worth copying): 26 jal + 16 HI16/LO16 relocations identical in name AND order, all 16 internal j destinations decoded and compared (the §195-D blind spot), and the four D_1F800020 words identified as splat FALSE-symbolisation of lui $s1,0x1F80 + an rm++ increment — ROM 3C111F80/26310020, emitted identically.

§482 ★★★ — TWO INDEPENDENT RE-TIES, ORDERED: WHEN ONE BARRIER FIXES ONE RESIDUAL AND CREATES THE OTHER (P31 S77, main:func_8006252C, 30 ins → MATCH)

The trap. s32 *p = &D_80078D0C; lets cse decide p[1]'s address is refoldable, so it emits a fresh %hi/%lo(D_80078D0C) through $at instead of reusing $v1 (LENGTH-DRIFT, 31 ins vs 30). This is not a spelling problem — it was reproduced identically across SIX spellings: an array-typed global, a struct component-ref through a local pointer, the same through the global directly, and a register pin on $v1. cse is deciding, and no declaration form changes its mind.

The half-fix that creates a second bug. An opaque re-tie right after the assignment — __asm__ __volatile__("" : "=r"(p) : "0"(p)); — makes p's value untrackable, every later p[N] reuses the register, and closeness goes 18 → 3. But an asm insn is a full block-wide scheduling barrier: nothing crosses it in either direction, so the li $a0,1 that the target materialises EARLY is now pinned after the barrier's lui/addiu, undoing a residual that was already correct.

The fix: give the OTHER value its own re-tie, FIRST. Re-tie the pri = 1 constant before p's barrier, so cse cannot defer-materialise it to the call site. It then emits early exactly as the target does, while p's barrier still protects the store-register reuse. Two independent re-ties, ordered pri-then-p, close both residuals at once → MATCH 30/30.

The general law. A re-tie is not only a cse lever, it is a scheduling barrier with a position. When you insert one to fix residual A and residual B appears or returns, do not conclude the two are coupled and unreachable — that is the shape that gets a function written off. Ask which values needed to be materialised on the other side of the barrier you just created, and give each of them its own re-tie, ordered so that each barrier sits where it helps. The companion to §153/§195-I's re-tie family: §153 says what a re-tie does to cse; §482 says what it costs the scheduler, and that the cost is payable with a second re-tie rather than by abandoning the first.

Banking note (§378 is the other half). The body above was MATCH 30/30 standalone and still took three tools to land, two of them wrong: scope_demote_drafts BROKE it (it aliased D_80078D08 through __asm__ and the build failed) because the clash was never a data extern; the real blocker was func_8006252C itself — TU void(void) vs draft s32(void), the self_decl_tu class — so cast_self_callers --sync-decls (4 call sites) then sync_tu_decls (func_800625DC, func_80062644). Read the DROP line's symbol before choosing the tool: if it names the function being banked, no data-scope tool applies.

§483 ★★★ — THE S77w WAVE HARVEST: SIX LEVERS, FOUR FROM BANKED (BYTE-PROVEN) BODIES

30 single-agent workflows over main's drawable frontier. 9 banked, and the levers below are the generalisable part. ★ = came out of a body that BANKED byte-identical, so the lever is byte-proven.

★ 1. STATEMENT ORDER IS THE ALIAS ORDER (func_8001EA14, 371 ins, close 89 → 0). A mem/s local-struct store can never be hoisted over by a mem/s varying p-> load, because true_dependence's exemption requires one side to be non-struct AND non-varying. So when the target initialises a local matrix in natural offset order, writing it in natural offset order is not cosmetic — it is the only order that produces those bytes: m[0][0] first closed a whole 45-insn init block. The same law one scope down (compute the scalars BEFORE a bump store) filled the lw-delay and removed a +1 length drift. If a block of stores is scheduled wrongly, check the ALIAS relation before reaching for a fence.

★ 2. AN INLINE-ASM "r" OPERAND THAT IS A BARE symbol_ref HAS NO PSEUDO (func_8001EA14). It is allocated by reload ($t0), not local-alloc. Assign the address to a local pointer first and it becomes a pseudo that local-alloc places in $v0. Worth 10 instructions. A launder's register is decided by whether its operand is a pseudo at all.

★ 3. PIN THE INTERMEDIATE, NOT THE RESULT (func_800301C8, 170 ins, 18 → 0). expand_mult passes accum_target = target, and a HARD target survives expand's generate-into-pseudos guard — so a $17 pin on the product drags the whole b*24 chain into $s1. Pin only the intermediate (register s32 m3 __asm__("$2"); m3 = (b<<1)+b; k2 = m3<<3;); a plain local is coalesced straight back. Corollary from the same function: addu $s0,$s1,$s0 is expand_binop swapping commutative operands to make op0 == target, not tree order — so pin the DESTINATION to flip it (this is the allocation-side answer §479 law 2 said was required).

★ 4. TWO INDEPENDENT RE-TIES, ORDERED (func_8006252C) — promoted to its own entry, §482.

5. A GLOBAL ARRAY'S BASE ASSIGNED TO A POINTER LOCAL BEFORE THE LOOP (func_80032A74, 422 ins, 12 → 1). That makes the symbol pseudo multi-block, so local-alloc.c:472 (reg_basic_block >= 0 && reg_n_deaths == 1) skips it, global-alloc cannot place it, and update_equiv_regs' REG_EQUIV(symbol_ref) makes reload DELETE the init insn and rematerialise lui/addiu into its spill register at every use — reproducing three separate $t0 symbol materialisations at zero cost. It replaces the older §153-launder + hard-pin recipe, and note the measured anti-lever: an __asm__("$8") pin can NEVER work here, it pushes every reload to $t1 (+30 rows).

6. THE SPELLING THAT FIXES A MASK-SHIFT IS THE DESTINATION'S WIDTH, NOT THE EXPRESSION (func_8001EFE0, 468 ins, 89 → 14). (sy & 0x100) >> 4 converges on srl;andi across all 10 spellings tried, including §432's shift-pair — which is why two prior agents filed it as permuter-class. It is fixed by making the DESTINATION a u16 (tpage), not by rewriting the shift. From the same function: POLY_FT4 stores must go per-VERTEX (x1=x0+w; y1=y0; x2=x0; y2=y0+h), not in copy groups — that alone fixed a −1 length drift and a 42-insn tail.

7. (idx<<2)+tbl BEATS &tbl[idx] (func_80020DA4, 100 ins, 20 → 8): plain integer address arithmetic fixes table-address register-coalescing mismatches that the pointer-index sugar cannot. Independently corroborated in the same session by a permuter run on the pointer-style source plateauing at exactly the closeness the integer form beat.

THE COST NOTE THAT MATTERS MORE THAN ANY LEVER. 21 of 30 targets came back NEAR, and almost every NEAR report cites the same shape: the agent found the mechanism, named the gcc pass and often the source line, and could not reach it from C. The frontier is no longer "we don't know why" — it is "we know exactly why and C cannot express it". Route accordingly: a NEAR whose note names a pass and a file:line is a WALL CANDIDATE for §474-style proof, not a redraft.

§484 ★★★ — "NO SINGLE NOLOAD BASE" IS NOT "UNLINKABLE": ASK WHETHER THE .bss OFFSETS ARE DISJOINT (P31 S77)

config/splat.us.exe.yaml has excluded four PsyQ objects from the LINKED build since Phase 8, under one reason:

SYS.o EXCLUDED — scattered-.bss commons (the GS_001 class: SYS references .bss by section+offset but the original linker scattered the commons across 0x80078xxx/0x800c5xxx, so no single NOLOAD base reproduces it)

Everything in that sentence is true, and three of the four objects are still not blocked by it. tools/psyq_bss_probe.py derives each object's .bss bases FROM THE BYTES — for every R_MIPS_HI16/LO16 pair against the bare .bss section, the OBJECT's immediates give the addend and the GAME's give the resolved address, so base = resolved − addend — and then asks the question nobody had asked: are the offset ranges behind those bases DISJOINT?

object ins .bss verdict
SYS.o (libgpu) 3,109 6,468 B, 2 bases SPLITTABLE — 0x0000..0x0044 @ 0x80078830, 0x0148..0x0150 @ 0x800c53cc. Disjoint: split at 0x148 and place each half NOLOAD
GS_001.o (libgs) 384 194 B, 5 bases NOT splittable — offsets 0x2a, 0x32-36, 0x3a, 0x3e, 0x50-ba interleave. A genuine per-symbol scatter, and the real wall
2D_BG0.o (libgs) 526 none no .bss section at all — this reason cannot apply
VM_NO1.o (libsnd) 305 none same

Why §9.2's escape does not reach these, and why that misled. §9.2 solves scattered commons by weakening every .bss-defined NAMED symbol and --defsym-ing it to its recovered address. A relocation against the bare .bss SECTION has no name to defsym, so the recorded exclusion is right that §9.2 fails — and it is easy to read that as "therefore unlinkable". The missing step is that a section reference only needs the section PLACED, and a section can be split.

Completeness matters before you believe a split (R32). The probe counts .bss references from EVERY section, not just .text: for SYS.o, .data has zero, so the two-way split covers every reference in the object. A split that fixes .text while .data still needs one base would be a silent half-fix.

The placement is derived, not configured. The probe finds the object in the EXE by masking every relocated field and searching for the unique match — so a wrong --vram cannot manufacture a clean answer, and a unique hit is itself proof the object is present. It independently reproduced SYS.o @ 0x80059234, 3,109 ins, which matches both the yaml subseg bounds and the manifest's psyq_identify count. src/800c.c turns out to be 100% SYS.o — its span is exactly the object's .text size — despite the subseg comment calling it "-O2 game code (incl. excluded libgpu SYS stub)".

The general law. An exclusion reason is a claim about the TOOLING at the time it was written (the reprobe-exclude-lists lesson, applied to link state rather than a wave draw). When the reason is a mechanism, re-derive the mechanism's premise from the bytes before accepting the conclusion — "scattered" was true, "therefore unplaceable" was an inference, and it cost 3,109 instructions of library code sitting as verbatim asm for twenty-odd phases.

§485 ★★★ — THE PLACEMENT MAP WAS PARSING A PRETTY-PRINTER: 25 PsyQ OBJECTS WERE INVISIBLE, NOT ABSENT (P31 S77)

tools/psyq_identify.py answers "where is each PsyQ library object linked in the EXE?" — the map the whole library-linking pipeline consumes. It built its match pattern by parsing one word per objdump -dr disassembly line, and objdump collapses a run of identical words into a single ... line. Every collapsed word was silently absent from the pattern, so from the first run onward the pattern was MISALIGNED against the image and find() returned None — printed as the confident, wrong sentence "not linked by EXE".

archive located before after delta
LIBGS 36 46 +10
LIBGTE 58 71 +13
LIBSND 32 34 +2
LIBGPU / LIBAPI 3 / 33 3 / 33 0
total 162 187 +25 objects / 3,877 instructions

How it was caught, and it was not by reading the code. 2D_BG0.o is excluded from the LINKED build in config/splat.us.exe.yaml under a scattered-.bss reason — and the object has no .bss section at all (§484). Chasing that contradiction, a byte comparison put it at 0x8005080C with 507 of 507 non-relocated words identical, while psyq_identify listed it as absent. Its own parse read 520 words for a 526-word object: three ... lines, two words each.

What it cost. These objects were never linked because they were invisible, not because anyone judged them unlinkable. 800b_7 — 1,022 instructions labelled "game code" in the yaml — is exactly 2D_BG0.o (526) + 2D_BG1.o (496), i.e. 100% library code. sgap_7 is exactly VM_NO1.o (305). src/800c.c is exactly SYS.o (3,109). Subsegs named as game-code gaps between library blocks are, in several cases, library objects nobody could see.

The law. objdump output is a RENDERING, tuned for human reading — it elides, it abbreviates, it reformats. A tool that derives a byte-exact fact from it inherits every one of those liberties. Read the section bytes and take relocation offsets from objdump -r; that is R33 ("derive from the invariant, don't re-parse the world") applied to a disassembler's stdout. The tell that something is wrong is always available and always cheap: compare the parsed word count against the section size. 520 ≠ 526 would have exposed this at any point in the last twenty phases.

§486 ★★★ — CARVING AN -O0 ISLAND IN main: FIVE COUPLED PIECES, AND THE TWO THAT ANNOUNCE THEMSELVES (P31 S77)

gcc-2.7.2 has no per-function optimize pragma, so opt level is per FILE (§116): an -O0 function inside an -O2 object must be cut into its own object. tools/o0_subsplit.py does this for overlays and cannot do it for main — jr_isolate_all wants config/splat.<ov>.yaml (main's is config/splat.us.exe.yaml) and overlay_src_split wants src/<ov>/ (main's TUs are top-level src/*.c). It now refuses main loudly instead of dying on a missing file. The manual procedure, byte-proven on func_8002C410 (299 ins):

  1. splat code rows — cut the subseg 3 ways: [pre][<name>_o0a][post].
  2. splat .rodata — split the span if its jump-table owners land in different pieces. They did here: func_8002B0B4 (before the island) and func_800335B8 (after) both own tables in span B, and one code object may contribute exactly ONE contiguous .rodata run.
  3. the .c — split to match; duplicate the prologue (86 lines, 2 includes).
  4. the Makefile -O0 glob — src/ov_*/… and src/md_*/… did not cover top-level src/*_o0?.c, so every -O0 island in the EXE was outside the rule. Same main-blindness family as draw_waves drawing zero main functions (S76).
  5. ld_interleave --order — insert the new object after its sibling.

DERIVE THE RODATA BOUNDARY, DO NOT GUESS IT. Compile the front piece and read its .rodata size: 800_b.o is 0xf8, so the front run ends at 0x80072E44+0xf8. The build's own jtbl_rodata_pads names the missing piece before you can get it wrong — "no yaml .rodata piece is bound to TU '800_b_2' — cannot place it".

TWO FAILURES THAT IDENTIFY THEMSELVES — learn the signatures:

  • A missing --order entry shifts every data symbol by exactly the floated piece's size. Measured: +0x204 across 704 two-byte runs, image size unchanged. Uniform delta on %lo immediates = a section floated, and the delta IS the size of what floated.
  • A stale INCLUDE_ASM path survives an incremental build and dies on a clean one. The split moved two stubs into 800_b_2, their directives still said asm/nonmatchings/800_b, and the old .s files were still on disk — so make build AND the byte gate both passed. make clean deleted them, splat emitted under the new name, and the clean rebuild died in jtbl_rodata_pads with FileNotFoundError. This is the whole reason R22 demands a clean tree, and it caught a carve that had already been committed and called byte-neutral.

THE ORDER THAT MAKES IT SAFE. Build byte-identical with the stub still in assembly, from a CLEAN tree, before banking anything. That proves the bounds independently of whether the draft is right — and separates "my carve is wrong" from "my body is wrong", which is otherwise one confusing failure.

PRICE IT FIRST (R37). The -O0 detector (o0_detect) flags exactly two open main stubs: this one, and func_80011380, which already lives in -O0 boot.c and is §474's proved floor. So this carve unblocked ONE function, not a class. Worth knowing before budgeting for more.

§487 ★★★ — THE PSX LOADER'S PER-VERSION SIGNATURE SETS ARE A FREE PROVENANCE ORACLE: main's "WALL" BAND IS LIBPAD 4.2.1 (P31 S78)

The instrument. ghidra_psx_ldr ships data/psyq/<ver>/<LIB>.LIB.json for every PsyQ release (260 … 470): per OBJECT, a masked-byte signature of its .text plus function labels with offsets. Matched as a regex over the retail EXE bytes (?? → any byte, 4-aligned hits only), a set that places an object byte-exact tells you the LIBRARY, the VERSION, and every FUNCTION NAME in it — with no .LIB archive in hand. Score per (version, library) by in-band hits; the version whose set places the most objects exactly is the linked one.

What it found. The 800c3 band 0x8005CF68–0x8005FC68 — twelve open main stubs, four of them recorded as §332 "%lo in a delay slot, no C can place it" walls — is LIBPAD 4.2.1 (PADENTRY, PADMAIN, PADCMD, PADIF, PADPORTD, PADSEQD, WAITRC2) plus LIBAPI 4.2 (COUNTER, L02/L03, FIRST, PAD, PATCH, CHCLRPAD). The 4.2 set places 4/11 libpad objects byte-exact; 4.3 swaps one (PADSEQD 292 vs 288 ins) and adds WAITRC2; 4.4+ place only WAITRC2; the 4.2 Ps stamp sits on C114.OBJ (_96_remove) at the band's head and the 4.2.1x stamp on libpad's .data. PsyQ 4.0 has no LIBPAD (the DualShock library arrived in 4.2) — which is exactly why twenty phases of "not linked by EXE" never found it. 46 names applied (docs/psyq-worklist.md S78).

The reading rule. Where the placed set's labels line up with the split's function starts (PADENTRY: 11/11, PADCMD: 9/9), name from the label. Where the object placed but the set's later labels drift by a few words (PADMAIN 4.2 vs the EXE's 4.2.1: +4 at _padSioRW2, +12 at _padClrIntSio0/_padWaitRXready), the function ORDER still names them — record that as order-inferred, not sig-exact. Statics carry no labels (PADIF) and stay func_.

Three laws.

  1. A "compiler wall" inside a band no archive you hold can place is a PROVENANCE question first. §332b already showed the band was assembled in reorder mode; that is what Sony's build did to LIBPAD. Before pricing a wall-proof, ask which library version owns the bytes — a different .LIB may link it outright (task #13), and even without it, the real name + the SDK header turn "unknown 133-ins function" into _padInitSioMode.
  2. Placement is not wiring. With the §485-fixed psyq_identify, in-gap objects (FGO_01–06 fill the 800b_5 "game code" subseg to the byte) made psyq_integrate.contiguous_blocks() merge libgte's 22 stub blocks into 3 and killed every LINKED build of main — silently at the S77 gates, because worktree gates carry no .run/obj40 and take the stub fallback. Map stub↔objects by subseg RANGE with an exact-tiling check (--yaml), and print what was placed but not wired: that list IS the completion contract's SDK residue.
  3. A library object's exported name is a claim of ITS version; the curated file's name at that address wins (R15). libapi 4.0's A66.o exports firstfile at 0x80062248 while the EXE links 4.2, where that trampoline is firstfile2 and firstfile is FIRST.o's C wrapper at 0x80061FA8 (LIBMCRD.o's jal word EA87010C says so). --redefine-sym at object-prep time; references to the old name then resolve through the recovered relocation address, never by name. The same rule caught libcd TOC.o's CdGetToc @0x800430B8 curated as DecDCToutCallback — an xdedup-vs-Vagrant-Story mislabel (a linked, byte-identical SDK object outranks a cross-project name match).

Renaming hazard (R32/R39, closed in lint_symbol_refs). The §265 verbatim bodies spell the name inside __asm__("… .ent\tfunc_X …"). The C-side rename regex with \b misses \tfunc_X (the escape's t is a word char), gas then dies with .size expression for func_X does not evaluate to a constant, and the string-masking linter was blind to it by construction. The linter now scans asm string bodies too; negative control: red on the pre-fix TUs (4 hits), green on the fixed tree and on every previously-passing TU (the func_8005C324→memcpy __asm__-label binding exempted).

§488 ★★ — THE "GAME CODE" GAPS BETWEEN LIBRARY BLOCKS WERE LIBRARY OBJECTS: 13 SUBSEGS → LINKED, EXACT-TILED, ZERO TOKENS (P31 S78)

What the residue printer said, and what it meant. After §487's --yaml change, make build listed libgte's located-but-unwired objects: MSC01/02/05/09 fill 800b (276 ins) to the byte, SMP_00 fills 800b_2, FGO_01–06 fill 800b_5 (804 ins), PATCHGTE fills 800b_6, MTX_05/07/11 and REG03+REG11 fill the four gsgap stubs inside libgs, 2D_BG0+2D_BG1 fill 800b_7 (1,022), and VM_NO1 / VM_NOWON fill sgap_7 and the head of sgap_8. Every one an exact tile — the subseg IS the object. Those subsegs had carried 106 verbatim __asm__ bodies (SDK-ASM "permanent" GTE macros, libsnd voice code) and four inline-asm wrappers counted as REAL for twenty phases.

The conversion is mechanical. Per subseg: rename the yaml row to the next <lib>N block name (comment: objects, ins, "exact tile"), append the name to that library's stub list in the Makefile (and widen its scan window if the block sits outside it), add any newly-placed objects to the curated dir (make_libgs.sh OBJS, make_snd_used.py re-run), git rm the verbatim-only TU, and make extract — splat re-emits src/<lib>N.c as an INCLUDE_ASM stub record for the fallback. A subseg that is only PARTLY an object is split at the object's end (sgap_8 → snd11 + a shorter sgap_8 keeping its game C). Gate: make build BINARY=main byte-identical WITH the SDK objects, then the fresh-clone fallback: move .run/obj40 aside, re-extract, build, restore.

Three things the bytes settled on the way.

  1. A placement nested inside another placement is a sub-pattern, not a second object. SMP_06.o (NormalClipS, 4 ins) matches inside SMP_05.o (NormalClip, 12 ins), which tiles the whole 800b_3+libgte9+800b_4 span byte-identically — including the "3-nop NOTCODE-PAD function" that was the object's alignment padding. psyq_integrate now drops nested placements and says so.
  2. The fallback build is a separate invariant and had been red. With .run/obj40 absent, the libcd stub defined func_800435B4 while src/800.c/800_c.c had called CdReadyCallback by its SDK name since S7x — undefined at link. The SDK build hid it (the linked object defines the name) and the worktree gates never build main. Curating CdReadyCallback = 0x800435B4 fixes it; lint_symbol_refs reports the stale INCLUDE_ASM line — read its whole output, not the last line. Verify the fallback from a FRESH extract: a stale .s under asm/nonmatchings/ still defines the old name and masks the miss (the §486 signature again).
  3. VM_F.o is SYS.o's class, not GS_001's: psyq_bss_probe — two .bss bases, disjoint offset ranges, split at 0x50c → the .bss-split lever (task #4) covers both.

Yield: libgte 53→70 objects / 22→30 blocks, libgs 31→33 / 7 blocks, sound 60→62 / 11 blocks; ≈3,000 instructions of "game code" re-provenanced as LINKED, 106 verbatim bodies retired, 13 TUs deleted, REAL −4 (inline-asm wrappers), 0 agent tokens. Remaining LINKED residue: SYS.o (3,109), VM_F (237), the libpad/libapi band pieces (task #5), SSGM.o (8 ins amid matched C), and the GS_001 / S_R / S_GRMDT scattered-.bss genuine walls.

The wall (§9.1, Phase 8). psyq-obj-parser packs an object's common-style globals into ONE .bss with sequential offsets; the original linker allocated those commons individually, so the game has them at unrelated addresses. A common referenced BY NAME is weakened and --defsym'd (§9.2). But the compiler references the object's own statics through the .bss SECTION SYMBOL + offset — no name to defsym — and one section can be NOLOAD-placed at only one base. SYS.o (2 bases), VM_F.o (2) and GS_001.o (6) were excluded on that reason for twenty-three phases; §484 (S77) asked whether the bases' offset ranges were disjoint and said yes for two of them, "no, interleaved" for GS_001.

The model that is actually right: RUNS, cut at SYMBOL starts. Walk the section-symbol references in offset order; each maximal run with one base is a piece. The cut between two runs snaps to the largest symbol start between them, because the linker scattered symbols: SYS.o's second run begins inside _que (+0x148) and the piece begins at _que (+0x144) — and _que, recovered BY NAME from SYS.o's own four named references, is 0x800C5510 = base2 + 0x144; VM_F.o's second run begins at _svm_sreg_buf (+0x508), which 62 other sound objects recover to 0x800B9B58 = base2 + 0x508; GS_001.o's five cuts land on PSDBASEX / CLIP2 / PSDBASEY / POSITION / GsDRAWENV, all five recovered by the other libgs objects at exactly the piece bases. Two unrelated oracles agree on all seven cuts (R34). §484's "interleaved" verdict came from grouping by BASE: PSDBASEY (+0x38) sits between PSDBASEX (+0x28) and CLIP2 (+0x30) in the packed section while X and Y are adjacent in the game — two base-ranges interleave, every run is single-base. Refuse (R43) only what the run model cannot tile: a SIZED symbol straddling a cut (one common at two bases), a HI16 whose LO16s need different pieces or high halves, an orphan LO16, a far-out addend.

The rewrite (tools/psyq_bss_split.py, its own 60-line ELF32 REL reader/writer — no pyelftools): new NOBITS sections .bss2… sized [s_k, s_k+1); the original shrunk to [0, s_2); symbols at/after a cut moved (value −= piece start); one LOCAL section symbol per piece inserted with the existing section symbols and every later symbol index in every REL entry bumped; each reference retargeted to its piece's symbol with the addend rewritten in place — HI16/LO16 immediates in .text (hi' = (A'+0x8000)>>16, lo' = A' & 0xFFFF, A' = A − s_k; a shared lui is one cluster and all its LO16s must agree), or the R_MIPS_32 word in data. classify() then recovers one base per piece from the piece's own section symbol (base = resolved − addend, tautologically the game's address) and NOLOAD-places each; the NOBITS predicate in psyq_link / psyq_link_region / psyq_integrate is ^\.s?bss\d*$. Self-check: the code/data sections differ from the original at exactly the retargeted sites whose value changed (R37, the tool diffs its own artifact).

Where it runs, and why there. Not in the curated dirs — inside the ONE prepare step that psyq_link.link_object (per-object verify), psyq_link_region.build_region (region verify) and psyq_integrate.integrate (the build) share (prepare_object(), before classify()). The split is re-derived from the bytes on every build — no recorded offset to go stale (R51) — and the same call is the negative control: 235 placed objects across the 9 curated dirs, 0 refusals, exactly 3 splits (R39). The first build DID refuse: a libcd object references .bss + size (an end-of-buffer pointer) and the strict addend < size bound fired on an object one base already served. Law that fell out: problems found while classifying references are fatal only when a split is actually needed — an object that passed before this tool existed must pass through it untouched.

Two things the extents are NOT. A piece's extent tiles the PACKED section, so an unreferenced common inside piece k may in truth live inside piece j's game range (GS_001's PSDBASEX/PSDBASEY: adjacent in the game, 16 bytes apart in the object) — NOLOAD pieces therefore overlap, harmlessly (zero bytes), and the verifiers link with --no-check-sections exactly like the build. And a named common whose recovered address disagrees with its piece is still handled by §9.2's weaken+defsym — the split only serves the section-symbol references; the two mechanisms compose.

Yield. 800c (56 hand-matched Sony functions + 62 verbatim frags, 100% SYS.o) → libgpu2; gsgap3 (hand-matched as game C) → libgs8; _SsVmFlush out of sgap_6 → snd12. libgpu_used retired (LIBGPU_ELF = the raw dir; identify drops the 8 objects the EXE never links). Main 143dbb89 with and without the SDK objects. The class that remains in the sound region — S_R/S_W, S_GRMDT* — has NO .bss of its own (cross-object commons at a minority address): a different wall, not this one.

§490 ★★★ — THE "WALL" BAND WAS A LIBRARY VERSION AWAY: LIBPAD 4.2.1 + LIBAPI 4.2 FOUND, THE WHOLE 0x8005CE18–0x8005FC68 BAND + THE APICARD REGION LINKED FROM REAL OBJECTS (P31 S79 #13/#5)

Where §487 left it. The psx loader's per-version signature sets named the band (libpad 4.2.1 + libapi 4.2) but no archive we held could LINK it: 4.0 has no libpad, 4.6/4.7 differ, and the loader's own 4.2 signatures showed PADMAIN drifting +4/+12. The plan was "C under the reorder island with real names" and twelve stubs — the four §332 "%lo-in-a-delay-slot walls" among them — sat in config/wave_exclude.txt as curated compiler facts.

The hunt took one lead. archive.org's play-station-programmer-tool-runtime-library-version-4.2.7z (383 KB) is the PsyQ Runtime Library 4.2 (1998-01-21) plus LIB/42PATCH/J421PD.ZIP: SCE R&D's 1998-02-26 notice "Libpad.lib version 4.2.1 for the Analog Controller (DUAL SHOCK)" with LIBPAD.LIB 4.2.1, LIBAPI.LIB 4.2 and their headers. Converted with psyq_lib_split.py + psyq-obj-parser, psyq_identify places libpad 4.2.1 7/11 and libapi 4.2 39/88 over 0x8005CE18–0x800629DC, and psyq_link.py PASSES all 46 — PADMAIN at 760 ins exactly. Neighbours for the record: plain libpad 4.2 (the loader's signature source) and the 4.3 disc (DTL-S2340, 1998-05-18: PADMAIN 832, PADIF 380, PADSEQD 292) each place only 4. 4.2.1 is the unique exact match, which also dates the build to between February and May 1998.

The wiring is §488's procedure, twice. The band is ONE contiguous run of 33 interleaved objects (21 libapi trampolines · COUNTER · PADENTRY · PADMAIN · L02/L03 · PADCMD · PADIF · PADPORTD · PADSEQD · WAITRC2) and every yaml-relevant boundary sits on an object edge (checked against .text SECTION sizes — §9.6's align-pad gotcha). Because psyq_integrate tiles each stub with ONE library's objects, 800c3 became four rows — libapi1, libpad1, libapi2, libpad2 — fed by two windowed calls from the RAW dirs (.run/obj42/libapi42 0x8005CE18..0x8005E188, .run/obj42/libpad421 0x8005D0D8..0x8005FC68). The apicard region's three "game code" rows were libapi 4.2's C objects to the byte — 800c2 = FIRST.o (firstfile + the "wall" stub func_80062144), 800c2_2 = PAD.o, 800c2_3 = PATCH.o+CHCLRPAD.o — so they became apicard5/6/7, make_apicard_used.py now sources libapi from 4.2 (the EXE's real libapi; libcard stays 4.0, byte-identical) and the region tiles 0x80061F38–0x80062888 with no game code left. Four TUs deleted (800c3.c: 129 hand-matched "C" + 62 verbatim bodies + 19 stubs; the three 800c2*), and the REORDER_TUS island (§332b) is EMPTY — the variable stays for any future reorder-assembled TU. main 143dbb89 with all SDK dirs and, from a fresh extract, with none.

What the four "walls" were. _padInitSioMode, _padStartCom, func_8005ED4C, func_8005F450 (§332 "%lo in a delay slot — no C can place it") were TRUE as compiler facts and irrelevant as work: Sony assembled libpad in reorder mode and shipped the object. The same for func_80062144 ("no jump table") inside FIRST.o and for PopMatrix/PushMatrix, which had sat in the exclude list since S68 while living in libgte3 — LINKED since Phase 8. The exclude audit kept all seven because its pinned-WALL class was decided BEFORE its LINKED class; fixed S79 (LINKED dominates: a function that is not a target has no wall). Law: a wall verdict is a statement about a compiler and a function; whether the function is OURS to match is a provenance question that precedes it (§487 → §490: name the band, find the archive, link it — in that order, before any C).

Fresh clone. tools/psyq/PlayStation_Programmer_Tool_-_Runtime_Library_Version_4.2.7z is tracked; 7z x it, psyq_lib_split.py + psyq-obj-parser the two lib421 LIBs into .run/obj42/{libpad421,libapi42}, tools/psyq_build_libs.sh LIBCARD, tools/make_apicard_used.py — docs/SETUP.md carries the exact commands. The build is byte-identical without any of it (the seven stub TUs), as always.

§491 ★★ — THE MECHANICAL LEFTOVERS (P31 S79 #6): A PHANTOM STUB, TWO JTBL TWINS, ONE EXACT CLONE — AND THREE TOOL GAPS THE BANKS EXPOSED

Yield. md_MAIN_003:D_800D3200 was never a function: a one-word 0x00FFFFFF sentinel in .text that func_800D3204/func_800D3234 read and write. It was carried as an INCLUDE_ASM "stub" (the census's H-VIRGIN row) because splat had once split a function there; emitting the word INSIDE the hand-written asm island it belongs to (__asm__(".globl D_800D3200\nD_800D3200:\n.word 0x00FFFFFF\n" …)) keeps the bytes and deletes the stub record — no new verbatim row, since it joins an existing PERMANENT body. ov_SC04_018's func_80181804/func_80181CB8 (jump-table functions, exact twins of ov_SC04_019) and ov_SC05_005's func_80181828 (exact clone of ov_SC05_003:func_80181720, family_remap) banked byte-identical — three functions, zero drafting tokens, but each needed a plumbing fix the tools did not make.

Gap 1 — a "covered" jump table still needs its pad-spec entry. jtbl_carve --probe answered "covered — table(s) already inside the existing .rodata carve" for func_80181804, which is true of the BYTES: S62's pads_audit had trimmed the TU's JTBL_PADS to the three tables the TU then compiled and left the fourth to the stub's .s. Splicing the C makes the TU emit four tables, and the build dies in jtbl_rodata_pads ("more rodata jump tables than pad specs (3)") — reported by the gate as CC1-FAIL(no-diagnostic) (the assembler failure has no cc1 line) and, for the sibling in the same TU, as PLUMBING: conflicting types for built-in memcpy (a pre-existing WARNING the verdict parser mistook for the error). jtbl_pads_fix --apply should have repaired it and reported "no pad-count drift"; the direct route worked: extend the spec by one entry and let the byte gate arbitrate the pad (0,0,0,0 was byte-identical first try). Law: a pad spec describes the CURRENT table population; banking a table's owner changes the population, so "covered" is a carve verdict, not a pads verdict.

Gap 2 — an exact-clone remap still meets the destination TU's spelling. The remapped body was byte-correct (twin d=0) and the gate said DIFF; rtu_match named the two real reasons: the draft carried its own typedef … Prim_8016E7C8 (identical to the TU's, still a redeclaration for gcc 2.7.2) and the TU declared the stub extern void where the body returns s32. Strip the draft's typedef block; fix the TU's prototype as a separate byte-neutral plumbing commit (callers ignore the value); gate. family_remap could drop typedefs the destination already defines by name — it does not yet.

Gap 3 — a parallel-gate merge that changes carve state leaves the main tree's split stale. parallel_gate merged config/splat.ov_SC04_018.yaml + overlays.mk from the worktree (the tail carve for func_80181CB8), and the in-tree make build then FAILED until make extract BINARY=ov_SC04_018 regenerated the split. Rule of thumb after any merged bank that touched a yaml or .mk: extract before you build.

What stays open in this class, with its blocker named. resident func_800D128C (243, close=0 draft on disk at .run/S71_gate14/resident/) and ov_SC02_017:func_80186C64 (209, d=2 twin ov_SC02_016:0x801810c8): jtbl_carve refuses both as "subseg would host NON-CONTIGUOUS .rodata carves" — the existing carve and the new table are separated by other functions' raw tables, so the CODE subseg must be split first (the §486 5-piece manual carve; for the resident that means splitting the single resident TU). md_MAIN_034:func_800CB00C (152): its table sits in the §154-A leading island AND its best draft compiles to 174 ins against 123 — the draft is wrong, the "compiler wall" pin is a label on a wrong draft. ov_SC05_010:func_8017FFA8 (88): a standard tail carve at gate time, no draft exists. md_MAIN_003:func_800D0100 (29, d=1 twin of main's func_80010B40): family_remap raises on main as a source (its image reader expects a flat overlay blob) — hand-remap or the C-PLUMBING route (#7).

§492 ★★ — "C-PLUMBING" WAS THREE DIFFERENT THINGS (P31 S79 #7): A RAW SPLICE THE GATE'S LADDER BROKE, TWO -O0 BODIES THE CHECKER COMPILED AT -O2, AND TWO DRAFTS THAT BELONGED TO OTHER OVERLAYS

The class as the census named it: five stubs whose bodies were "proven close=0" and whose TU spelling refused them. The real-TU compile (rtu_match, §491's route) sorted them into three unrelated cases:

(a) The gate's ladder mutated a correct body. md_MAIN_020:func_800CB17C (30 ins) was rtu_match MATCH and parallel_gate NEAR (near: 1, failed: 0) — the "match_one MATCH but the whole-binary gate rejected — CAUSE NOT DETERMINED" verdict the census had carried. Hand-splicing the raw draft into the TU and running make build was byte-identical first try. gate_stage's stages (canon_resident_calls → cast_call_sites → sig_unify) are pure draft rewrites that usually help; here one of them changed the codegen. When rtu_match says MATCH and the gate says NEAR, the raw splice is the bank (or --skip-stages).

(b) An -O0 island TU. md_MAIN_003:func_800D06BC (33) and func_800D0100 (29) live in md_MAIN_003_o0e, a TU the Makefile's -O0 wildcard compiles at -O0; their close-0 drafts had simply never been gated (attempts=0). rtu_match without --o0 showed the target's $fp frames against an empty "mine" column — the checker's flag, not the draft, was wrong; with --o0 both MATCH, and the worktree gate banked both (the build already compiled the TU right).

(c) Same name, other overlay (R48). ov_SC05_018:func_80180BE0 (target 65 ins) and ov_SC06_010:func_801809E4 (33) had "drafts" of 15 and 220 ins: every on-disk func_801809E4.c is ov_SC01_080's function of that name (the sweep_bb dirs), and .run/backlog_drafts/<fn>.c is keyed by the bare function name, so the census's drafts/closeness columns and the ledger's best_draft pointed at a different function that happens to share the address. Two of five "plumbing" rows were phantoms: those two functions have NO draft and go to the drafting task. Fix to make: key backlog_drafts and the journal join by (binary, fn) — the same defect class as R48's three collisions.

Net: 3 banks, 2 reclassified, 0 tokens; stubs 35 → 32.

§493 ★★ — THE PERMUTER ROUTE END-TO-END, AND THE THREE PLUMBING STEPS BETWEEN A SCORE-0 WINNER AND THE MAIN GATE (P31 S79 #8)

What the class was. Six D-NEAR stubs (closeness 3–16) with agent journals naming the residual's gcc pass. The S78 brief's rule — "a NEAR whose journal cites a gcc pass + file:line is a wall-proof candidate, not a redraft" — held for two of them and was wrong for one: main:func_80015760 (106 ins, closeness 9, filed by the S76 agent as "a genuine sched1 basic-block ordering artifact, verified via RTL dumps") fell to the local permuter in its FIRST 150-second cycle (permuter_ils.py … --klass SCHEDULE --cycles 8 --secs 150 --j 6, seeded from the journal's best draft .run/wave_p31o/main/func_80015760.c). A sched1 residual is a statement-order residual, and statement order is exactly what the permuter mutates. Try the permuter before writing "wall" on anything the permuter can parse.

From winner to bank — three plumbing steps, every one of which the byte gate would otherwise report as "RED image". A permuter winner (.run/permuter-winners/<fn>.c) is a self-contained compile unit: (1) it carries the permuter's typedef preamble (u8… and the M2C_UNK* family) — strip every typedef whose name include/common.h already defines, or the real TU dies on "redefinition of u8"; (2) its externs are the SEED's guesses — gate_main names the clash (D_800A5E60 kept=('u8*','') this=('s32','')) and the fix is the TU's spelling plus a cast at the use (D_800A5E60 = (u8 *) pkt;); (3) the function's own forward decl and the TU's banked callees. The self decl is cast_self_callers --sync-decls (the TU's extern void func_80015760(s32, s32) became extern void func_80015760(); + ((void (*)())func_80015760)(…) at both call sites — proven byte-neutral and COMMITTED before the gate, S77 law). The callee was the subtle one: the seed declared extern u16 *func_80015908(s32, u16) and the TU DEFINES s16 *func_80015908(s32, s32) — adopting the TU's prototype turned the target's andi $a1,$s6,0xFFFF into a move (one instruction off); the original caller evidently saw a u16 parameter, so the cast goes on the ARGUMENT: func_80015908(tile, (u16) flags) reproduces the andi under the TU's s32 prototype. rtu_match --split 800 --source main --tu src/800.c (main's TUs are loose files; --tu is mandatory) confirms MATCH at each step before gate_main --apply spends a clean rebuild.

Two instrument corrections on the way. p16_permute.run_permuter swallowed run_masked's stdout, so a seed the permuter's C parser refuses (register u8 *a3 asm("$7") — pins are not C to pycparser) reported "no waypoint (no improvement over base yet)" for 8 cycles in 20 seconds on func_80038698. It now prints [permuter] REFUSED <fn>: Syntax error in base.c … and leaves PERMUTER_REFUSED.txt in the scratch dir (positive-controlled). And permuter_ils seeds from a REGISTER-PINNED best draft cannot be permuted at all — the pinned 17→11 gain and the permuter are mutually exclusive on that function.

S80 CORRECTION — the "pinned seeds cannot be permuted" claim was the INSTRUMENT, not the permuter. Two defects, both in our layer (R35/R40): (1) p16_permute.hide_asm carried only the __asm__ spelling into the b64 pragma; register u8 *a3 asm("$7") (the asm(/__asm( spellings, 3 of the S79 seeds) stayed raw in base.c and pycparser refused it at cycle 1; (2) permuter_ils's WARM RESTART copied the waypoint's source.c — which decomp-permuter serializes with the pragmas DECODED back to raw pins — straight into base.c, so on ANY pinned seed cycle 1 ran and cycles 2..N were parser refusals reported as "(unchanged)". The S79 func_80020DA4 "8-cycle plateau at 2" was one cycle. Fixed: hide_asm matches all three spellings, followed by (/volatile (the bare word asm also lives inside INCLUDE_ASM("asm/…") path strings — the R39 control over 5,311 drafts caught that false positive before it shipped); the ILS re-hides every waypoint before restarting, asserts the definition survived, and ABORTS non-zero on a refusal instead of counting no-op cycles (R61a). defines_fn also now accepts a K&R-style definition (void f(a, b) s32 a; s16 b; { — the documented lever for an s16 parameter's in-place promotion, func_80039DEC), which it had refused as "lost the definition" — the R39 control over 5,311 stored drafts found 436 K&R-style backlog drafts that this check alone had kept out of the permuter lane. First re-run on the S79 seeds: every pinned seed iterates, and ov_SC06_022:func_8017DF28 (pinned WALL, "closeness 2 on five RTL-verified attempts") reached 1 inside its first cycle. A permuter verdict on a pinned seed dated before S80 is NOT a measurement of that function.

The ledger the class leaves (for the PhaseEnd): main:func_80011380 192 — WALL, §474 PROVED (pinned S79). main:func_80015608 86 — permuter plateau at 1 (REGALLOC-PERM: target addu $s3,$s6,$s3 where the best draft emits sll $s3,$s6,1 — a copy-then-add spelling of x*2, agent-sized). main:func_80039B20 79 — permuter plateau at 7 (§461: the $v0/$v1 tie + an early-scheduled re-read of D_80073140[i]; 17 + 8 ILS cycles). main:func_80038698 74 — best 11 (§ pins fix the interleave; the permuter refuses pins; sched1 DAG-priority + local-alloc self-coalesce). ov_SC03_105:func_801834A4 106 — pinned WALL (S71, §148-A/§193-F, closeness 6 on four attempts).

§494 ★★★ — TEN BANKS FROM ONE-AGENT-PER-FUNCTION DRAFTING (P31 S79/S80 #9): THE IDIOMS, THE PLUMBING, AND THE THREE WAYS AN AGENT'S "MATCH" WAS NOT ONE

Yield. 25 packs (claude_wave_packs: journal history + a matched neighbour each), one Agent-tool subagent per function (Haiku ≤50 ins, Sonnet ≤120, Opus above), no wave. Ten banked byte-identical — seven in S79 (md_MAIN_003:func_800D0174 + func_800D1D14 (-O0 island), main:func_80015608 + func_8002AC98, ov_SC05_018:func_80180BE0, ov_SC06_010:func_801809E4, ov_SC05_010:func_8017FFA8 (a 6-way jtbl switch)) and three whose agents outlived the S79 session and were aggregated from their transcripts in S80 (tools/agent_verdicts.py): ov_SC01_001:func_80181E04 (269), main:func_8001EFE0 (468 — the largest open main body), ov_SC02_027:func_80180B3C (297). Of the eleven Opus-tier F/G targets, 3 MATCH and 8 NEAR at EXACT length with the residual named and its gcc mechanism cited — the ledger at the end of this section.

Idioms that closed functions (each byte-proven today).

  • The fresh-temp lever for a commutative operand order — target addu $s3,$s6,$s3 where every source order emits addu $s3,$s3,$s6: gcc 2.7.2's expand_binop swaps operands when op1 IS the expansion target register; route the add through a NEW pseudo — { s32 xt = blockSize + x0; x0 = xt; } — and the source order survives, xt coalesces into the same hard register. Two banks (func_80015608, func_8002AC98); the temp AND both operands must be s32 (an s16 temp re-coalesces with the target and the swap returns; an s16 operand costs a sign-extend pair).

  • -O0: address-of + cast beats the %lo fold — D_800D3630 (4-byte stride, low 16 bits read) as s16[][2] or a 2-field struct makes gcc fold %lo into the lh displacement through $at (-2 ins per site); extern s32 D_800D3630[] + *(s16*)&D_800D3630[i] materialises lui+addiu first and matches.

  • An early return that must fall through — the Haiku plateau's missing 2 instructions and "frame -8 vs -0x10" were a return inside a branch that should fall into the shared mask/store tail; the -0x10 no-save frame was two s16 locals (§186b), not a call. Branch polarity if (v0 >= v1) with the arms swapped.

  • jtbl switch: the loop index must be s32 (an s16 gives the fused lhu+sll16+sra12, §241) and a table lookup goes into its own named temp BEFORE the found=1/zero-store statements, or the store schedules ahead of the load.

  • The phantom stack frame — an address-taken s32 frame_pad[3] induces the target's unexplained 16-byte frame (func_80020DA4, 8 → 2; the remaining 2 is a mflo destination, permuter class).

  • memcpy(…,12) under a TU's extern memcpy compiles to a jal (§48-C2/§160a); the inline lwl/lwr/swl/swr shape needs a struct assign through an align-1 12-byte typedef (the TU's own Blk8 idiom).

  • The OT insert is PsyQ's P_TAG addr:24 BITFIELD store, not a hand-masked word (func_80181E04, 269 — the last 18 instructions turned on it): store_bit_field masks the VALUE first and then expand_binop(ior, …); hand-written (*p & 0xFF000000) | (v & 0xFFFFFF) masks the destination first → 0xFF000000 hoists before 0xFFFFFF and $a0/$v1/$a1 come out permuted. One edit fixed the movable hoist order, both or operand orders and the whole allocation. Same function: §246-2 — sixteen parallel globals declared as arrays of ONE 0x50-stride record (__asm__ aliases, §200) share one giv AND fold %lo for the stores; a counted i < 0x100 loop re-materialises the bound inside the loop (the reloc becomes sym+0x5000, byte-identical to the target's %hi/%lo(sym2) after link); if (z < 0) z += 7; at the stsz site and (z >> 3) * 4 + (s32)ot at the use split the rounding from the shift so the bgez lands before the packet stores and the sra after. A local register … __asm__("$3") pin is IGNORED unless an asm references that variable.

  • Read the matched same-TU siblings before sweeping spellings (func_8001EFE0, 468): every lever came from func_8001DA34/func_8001EA14. A pinned hard-reg SET is placed first in its block, so no source order of m24/mFF separates the two luis — split ot instead (`ot = (u32*)(d4); mFF = 0xFF000000; ot = (u32)((s32)ot

    • (s32)otbase);) and the slllands between them (a §194-A fence there is +3, wrong direction). Oneasm volatile("")between the tpageshand theq[7] |=RMW. The two-SVECTOR fill is a local-alloc DENSITY fact, not statement order: a 720-order × {2/3-operand ldv3} × {clobber} sweep plateaus at 9, one §419 zero-byteasm("" :: "r"(vh), "r"(vv2), "r"(vw))buys the $a0/$a1 pair (9→2), andvw = rec >> 16;held apart fromv1.vx = vw & 0xFF;fits thelhu` between them (→0).
  • Inverted arms keep a dead-provable mask alive (func_80180B3C, 297): if (c) vv = v - 0x100; else vv = v; lets combine prove & 0xFFFF and - 0x100 dead for the u8 store and deletes the andi (296 ins); the SAME arms inverted — if (!c) vv = v; else vv = v - 0x100; — keep both at zero cost (3→0). Also: an or-term as its own accumulator statement (tpg |= (y & 0x200) << 2;) unboosts it into $v0 instead of destroying $a3; a register u32 c40 __asm__("$2") pin on the (w & 0x40) >> 6 term creates the anti-dependence the target's schedule needs; splitting if (c) v -= 0x100 into a second variable makes the mask single-set so sched1's birthing boost stops sinking it. Statement order was INERT across all 792 permutations of that block.

  • Spelling laws from a 518-ins NEAR (func_80039308, 115→34, each byte-witnessed): v = a * b; v >>= 7; puts the mflo in the shift's register, v = a * b >> 7; does not (−8); vv = (u16)vol schedules and allocates differently from vv = vol & 0xFFFF (−14); folding a pointer add into its using expression swaps $v0/$v1 (−7); naming the complement first (mv = ~mm[j]; … mv = mm[j];) fixes the RTL order of a read-modify-write pair (−3); ONE local shared between both arms of an if merges into one long live range with a tiny allocno priority — give each arm its own (−40). gcc-2.7.2 never folds (and reg 255) to a move.

  • gdb-on-cc1 allocno arithmetic instead of guessing (func_80023BF0, 281, 90→18): breakpoints that dump allocno_n_refs/live_length and local-alloc's find_free_reg/post_mark_life grants make every lever a computation on global.c:594. A §194-A fence in BOTH arms is a LOCAL regression (90→114) that is REQUIRED — it shortens c1's live range 24→16 and lifts its priority above giv1's; a separate base + compound code |= gives code a FULL preference for $a0 (find_reg pass 0's regs_someone_prefers — priority alone could never do it, 2647 vs 21250); register u32 base __asm__("$3") is what makes that fire (only a HARD reg is visible to regs_live_at for a block-local qty); exactly 13 zero-byte fences after pkt = D_800A5E60 invert two ties (+12 leaves a tie that loses). The remaining 18 are pure sched2 order provably exclusive with the pin (move_movables inserts hoisted movables before loop_start → the pinned constant's source has the lower LUID).

  • K&R old-style definition for an s16 parameter (func_80039DEC, 74: the seed's 74-vs-72 length →9): void f(a0, a1, a2) s32 a0; s32 a1; s16 a2; { reproduces the in-place sll/sra 16 arg-register cast + the addu $a3,$a2,$zero raw-preserve that no ANSI (s16) cast or s32 spelling produces; forward-goto block layout reproduces the compiled block order; cnt + 0xFF (not cnt - 1) on a u8 counter reproduces the addiu 0x00FF and lands the store in the branch's delay slot; block-scoped dual same-hard-reg pins (p/b → $2/$3 in one family, $3/$2 in the other) closed a role-coalescing mismatch worth ~25. The residual $a3↔$t0 swap of the two raw-preserve copies is fixed by ARGUMENT POSITION in the K&R narrow-parameter-promotion pass.

  • Struct-vs-scalar true_dependence disambiguation (func_8017DC80, 346, 255→109 in ONE edit): packet stores must be COMPONENT_REFs through a real struct pointer while the GTE outputs stay COMPONENT_REFs of one 56-byte frame struct, written in the target's own interleaved order (x0 = sxy0; p->x0 = x0; x1 = sxy1; p->y0 = x0 >> 16; p->x1 = x1; …); record accesses as offsets from ONE u8 *r so loop.c builds c/va/vb as givs (explicit variables create a 5th biv that spills ot); the setup call placed AFTER all preheader setup gives the save/def-interleaved prologue with addu $fp in the call's delay slot. That TU's "~20 drafts plateaued at −33" wall was the GTE macros being UNDEFINED and compiling to implicit jals — an instrument-class wall (R35), not a compiler one.

  • Fences at prim FIELD boundaries beat pins in a packet fill (func_800CF3E8, 469, 256→54, all four block sizes exact): zero-byte __asm__("") at specific field boundaries (pC base, pC u0, p5 clut) took 262→186 alone; a §419 density buy on the OT pointer right after its assignment stops gcc propagating the pin away (§136d-1); transposing a prim's base/tag/len statement into the previous prim's field sequence reproduces the interleaved address chains; redundant pins actively block progress — once the fences were in, nine pins measured inert and deleting them let a fence REMOVAL reach 54. Simplify the pin set before declaring a plateau.

  • extendhisi2 is a force_not_mem EXPAND; zero_extendhisi2 is a define_insn (func_80032A74, 422, closeness 1 = one lh vs lhu at idx 244, frame/offsets/27 symbols exact): an orphan frame slot can only be minted at an lh (a 3-way movhi+ashl+ashr merge), never at an lhu; the target's extra 8 frame bytes are most likely §172 producer 3 (a caller-save area allocated at reload1.c:1445, landing at sp+0x48) — no C spelling axis in ~200 probes (a 100-variant retyping sweep, (s8) splits, clobber and §148-C barriers, ?:-accumulators, s16 shadows). Two §172 corrections: the (use (reg)) count UNDER-counts orphans (use the dump's vars= as the oracle) and over-counts hard-reg return USEs.

  • The 2D-array declaration + insn_count inflation (func_800391D4, 75, 64→3): extern s32 D_80073140[][1] stops move_movables hoisting the address instead of folding it into the lw (§164-26); a register s32 i __asm__("$7") pin kills combine_givs on D_80073140[i]; seven empty __asm__("") pads inflate insn_count (verified against a -dL loop dump) so a la is not hoisted. Caveat: [][1] conflicts with the TU's own extern s32 D_80073140[] — reconcile via a §200 alias before a real-TU gate.

Three ways an agent's "MATCH" was not one. (1) A naked __asm__ reproduction of the target — the verbatim class; sent back, it returned genuine C. (2) Standalone match_one MATCH, real-TU DIFF: the TU's typedefs and externs (rtu_match names them; §376). (3) rtu MATCH, gate NEAR: the gate's transform ladder altered the body (§492a) — raw splice. Rule for the coordinator: grep the draft for .ent/.word, re-measure in the real TU, gate, commit — per result, never in bulk.

Instrument findings this task produced. (1) Agent-tool drafters outlive the session that spawned them; their verdict is the last JSON object in the transcript — tools/agent_verdicts.py (SETUP) reads it without loading the transcript into a session. (2) The permuter could not permute a pinned seed, and it was our instrument three times over — the §493 S80 correction (436 K&R backlog drafts had been refused by our own coverage check). (3) splat's "Handwritten function" banner is only a cop2-opcode tell: all three GAME-GTE "uncertain" bodies (func_80181E04, func_80185810, func_8017DC80) are ordinary gcc-2.7.2 -O2 C, and one of them banked.

Ledger rows this task leaves (drafts in .run/S79w/<arm>/, the S80 permuter waypoints in .run/S79w/permuter/, verdicts in .run/S79w/verdicts/verdicts.jsonl, rows logged to the backlog; the ≤3 residuals with a mechanism citation are pinned in config/wave_exclude.txt as WALL candidates after an S80 permuter_ils 8×150 s null): main:func_80032A74 1 (lh vs lhu, extendhisi2 orphan / caller-save area — WALL candidate) · main:func_80020DA4 2 (mflo destination $t0 vs $a2 — WALL candidate) · main:func_80039DEC 2 (permuter 9→2; the K&R raw-preserve copy in $t1 vs $a3 — argument-position promotion, WALL candidate) · ov_SC06_022:func_8017DF28 2 (delay-slot fill from expand_block_move's cse-reused address pseudo; the permuter's "1" replaced the addiu with a sw zero — a divergent rewrite, R14) · main:func_800391D4 3 (the 3-insn move_movables reorder — WALL candidate) · main:func_80039B20 7 (§461 tie-break, permuter-confirmed plateau) · main:func_80038698 11 (sched1 DAG priority ×3; the permuter now iterates on its pinned seed and confirms 11) · main:func_80023BF0 11 (permuter 18→11, ADDRESSING: the mask materialisation moved across the sll/bne — verify semantics before seeding; the agent's 18 were pure sched2 order exclusive with the pin) · main:func_8001BC6C 28 (sched1 birthing-boost, two mutually exclusive schedules) · main:func_80039308 34 (18 renames + 15 sched-order + 1 imm at exact length 518) · main:func_8002FDE8 35 (local-alloc caches &D_800A46D2 across a call, §153) · ov_SC03_105:func_80185810 37 (sched1 ordering in three windows, 489 ins, ~3,600-compile floor) · main:func_80015B6C 44 (two documented walls) · md_MAIN_003:func_800CF3E8 54 (469 ins, all four block sizes exact; a local-alloc birth-order tie — [permuter]/density fuel) · ov_SC07_002:func_8017DC80 84 (cse unifying otz<<2 across a call into pinned $s1 ~45 + a clamp delay-slot 8 + a late lui/addiu pair ~14).

§495 ★★★ — THE VERBATIM END-STATE (P31 S80 #10): TWO DEF-SIDE DECLARATION WALLS, A "BANK" THAT WAS THE ASSEMBLY, AND A GATE THAT DROPPED A BANK ON EXIT 0

The class. After S79 #5 the whole game's __asm__-posing-as-C set was six bodies (config/verbatim_manifest.json): five PERMANENT (crt0 start/__main/__do_global_dtors, the two MDEC-side hand-asm routines of md_MAIN_003) and one GAME-C, ov_SC03_107:func_8017D878 (45 ins, "37 stored drafts, no wall"). verbatim_check.py --strict then found a SEVENTH: md_MAIN_020:func_800CB17C, which S79 #7 had "banked by a RAW splice" — the splice had put the function's ASSEMBLY (the ledger's best_draft was the asm; every stored "draft" of it was) into the TU as a file-scope __asm__ block. rtu_match says MATCH for such a body by construction, the byte gate is green by construction, and progress.py counted it banked. A verbatim body is never a bank (P9); verbatim_check --strict in tools-health is what caught it, one session later. Both were converted back to INCLUDE_ASM with tools/verbatim_to_stub.py --apply --gate (byte-identical each; the tool refuses to GUESS the asm subdir when a TU has no sibling stub left — pass --asm-subdir asm/<bin>/nonmatchings/<subseg>, the fleet spelling) and then decompiled.

Wall 1 — extern void f(void) for a callee used only as a POINTER. The TU declared extern void func_8017D878(void); because its only use is func_801788B8(s0, (s32)func_8017D878). The function actually takes the actor pointer and returns either 0 or the address of func_80172710 (lui/addiu $v0). The 32-line stored draft was byte-correct (match_one 45/45) the moment its definition said s32 f(s32); with void it dropped the addu $v0,$zero,$zero / lui/addiu $v0 tail, and with the right signature the TU refused it (conflicting types). Thirty-seven drafts had died on a declaration that carried no information (only the address is taken). Fix: correct the TU decl (byte-neutral — the pointer value is the same), commit, gate. Lesson: when a TU's only use of a callee is (s32)callee, its extern was GUESSED; treat the decl as unconstrained and take the signature from the callee's own body.

Wall 2 — a block-scope struct S tag. md_MAIN_020's TU declares extern s32 func_800CB17C(struct S *a0); INSIDE another function's body, and struct S is not declared at file scope in that TU. C then creates a BLOCK-LOCAL tag, and no file-scope definition — whatever it spells — is compatible with it (conflicting types for func_800CB17C, the previous declaration being the block-scope line). Fix: one file-scope struct S; before the first use (byte-neutral), so both spellings name the same incomplete type. The definition itself is seven straight calls with s0 = a0 in $s0 and NO return in an s32 function (the target sets no $v0 after the last jal; falling off the end is exactly that).

The gate that dropped a bank. parallel_gate on func_8017D878 printed banked 1 … merged 0 file(s); REFUSED 0 and exited 0; the main tree still had the stub. The worker's git status capture came back empty (cause not recovered — the fixed-path results JSON was overwritten by the next run before anyone looked), so the byte-proven bank died with its worktree. Now: the worker's raw status + scope ride in the result, a banked-but-not-merged binary prints !! BANKED-BUT-NOT-MERGED and the run exits 2, and every run also writes .run/pgate_runs/<ts>.json. The recovery route is the one S79 #7 SHOULD have used: rtu_match MATCH → splice the C in-tree → make build (exit code, R53) → commit.

State after #10: the tree's verbatim set == the manifest's five PERMANENT rows, RATIFIED in its _README; ov_SC03_107 is 100% C; the open-stub census is unchanged at 21 (both functions went verbatim → stub → C in one session), but the S79 close had been ONE bank overstated.

§496 ★★ — A CARRIED-TYPE TEST THAT ASSUMES THE OVERLAY INCLUDE SET SILENTLY DROPS A RESIDENT TYPEDEF (P32 T1a; byte-proven, resident func_800D128C isolation)

The refusal. jr_isolate_all resident --only func_800D128C (dry-run CLEAN) wrote a clean 3-region split, and make build BINARY=resident died in BOTH new region TUs: parse error before cdFileLocTable at every extern CdFileLoc cdFileLocTable[]; (file-scope and block-scope — D_800D3764/D_800D3768 too). The sha1sum of build/resident/resident still read 8e17e02f… == the check file — the PREVIOUS binary on disk (R53: read the exit code, never the file). The carried decl layer had carried P_TAG_800CFAD0, BYTES_800CFAD0 and Struct80078E78 (all later in region 0) but not typedef struct { s32 word0; s32 word4; } CdFileLoc;.

The mechanism. _file_scope_decls skips carrying any file-local type whose name the SHARED headers define ("never re-emit one engine_types.h already has" — the §321 redefinition guard), testing against _engine_types() = engine_types.h + common.h. That guard is correct for an overlay region, whose header #includes ../shared/engine_core.h → engine_types.h. It is wrong for the RESIDENT (and every md_* module and main TU), whose header includes only common.h: CdFileLoc sits in engine_types.h:1766 (lifted from an overlay long ago), so the resident's own typedef was judged redundant and dropped into regions that never see engine_types.h. The tool's comment even said "every region #includes engine_core.h at its top" — true of the class it was written for, an assumption for the others (R35).

The fix (tool, not hand-patch — R43/R33). _provided_types(header) derives the provided set from the TU's OWN #include "…" lines: engine_core.h ⇒ engine_types.h + common.h (engine_core.h itself is NOT scanned — its typedefs live inside DEFINE_func_*() macro bodies and reach a region only where that macro is invoked; the first draft of the fix scanned it and the R39 control caught 10 phantom "types" named L, arg, buf, sp10…); common.h ⇒ common.h only. _file_scope_decls(items, provided) uses that set at both decision points (the hoist predicate known and the "already provided" skip). Controls: an overlay header yields exactly the legacy set (1,197 names, byte-for-byte the old behaviour for every overlay carve ever run); the resident header's set is 15 names without CdFileLoc; the resident split then builds byte-identical with the typedef carried into both regions.

Tell + generalization. parse error before <symbol> at an extern <Type> <symbol> line in a freshly isolated region = a type the carrier believed the headers provide. Ask WHICH headers this TU includes before believing any "the header has it" guard; the resident/md_/main TUs are the common.h-only class. Same class as §323 (attributes hid the name) and §321 (re-emission): the carrier's blind spots are always "a type-name recogniser or a provenance assumption", never the C.

§497 ★ — A BODILESS typedef struct Tag Alias; DEFINES THE ALIAS, NOT THE TAG: THE CARRIER'S FALSE "CONFLICTING BODIES" REFUSAL (P32 T1b; ov_SC02_017 func_80186C64 isolation)

The refusal. jr_isolate_all ov_SC02_017 --only func_80186C64 refused (R43): 1 carried type name(s) have CONFLICTING bodies — a rename is needed, not a dedupe: Rec801806C8_s — A: struct Rec801806C8_s { s32 f0; s32 f1; s32 f2; } __attribute__((packed, aligned(1))); B: typedef struct Rec801806C8_s Rec801806C8;. Both lines are legal C in one TU (a tag definition, then a typedef naming it), so docs/sunset/frontier-p32.md had costed a source RENAME.

The mechanism. The dedupe in _file_scope_decls keys a type block by _type_names(block), whose recogniser \b(?:struct|union|enum)\s+(\w+) returns the TAG for a bodiless typedef struct Tag Alias;. The tag's definition block yields the same single name, so two DIFFERENT blocks shared one key, their normalised bodies differed, and the refusal that exists for a genuine redefinition fired. The line itself declares Rec801806C8 (ordinary namespace) and only REFERENCES the tag — C's tag and ordinary namespaces are distinct, and the carrier had collapsed them.

The fix (tool, not rename). _TYPEDEF_TAG_ALIAS = ^\s*typedef\s+(struct|union|enum)\s+(\w+)\s+(\w+)\s*;\s*$; when it matches and alias ≠ tag, _type_names returns {alias} and the carried set learns the alias. typedef struct X X; (alias == tag) keeps the old key on purpose: two of those in one TU ARE a redefinition and must still collapse or refuse. Unit control on seven block shapes (tag+body, typedef-anon, typedef-tag-body, fn-ptr typedef, struct X;); the overlay's dry-run went REFUSED → CLEAN (2 region files) with the source untouched.

Tell + generalization. A "CONFLICTING bodies" refusal whose two bodies are a struct DEFINITION and a bodiless typedef struct <same tag> <other name>; is this class — no rename, the carrier is wrong. Sibling of §496 (the provided-type set assumed the overlay include set): the carrier's failures are recogniser/namespace assumptions.

§498 ★★★ — A CARVE THAT DROPS THE BINARY'S --pre CLAUSE, AND A GATE THAT IGNORED THE EXTRACT'S EXIT CODE: HOW A BYTE-CORRECT DRAFT WAS BOOKED "DIFF" (P32 T1a; resident func_800D128C BANKED 243 ins after two instrument fixes)

The verdicts. rtu_match MATCH (243/243) in the real region TU; parallel_gate → func_800D128C DIFF, banked 0, 53 s, rc 0, nothing merged. The in-tree reproduction (raw splice, §492a) showed why: jtbl_carve --func wrote the 5-piece carve and REGENERATED resident_JTBL_INTERLEAVE as --order tail.data.o,…,tail3.data.o — without the committed line's --pre hdr.rodata.o. make extract refused (ld_interleave --order: 1 extracted asm piece(s) are in neither --order nor --pre … hdr.rodata.o), exited 1, and make build then linked against the STALE 3-piece script: sha 59b44f0f…, 249,252 of 365,404 bytes differing from file offset 0x4 — the image shifted, not the function. The gate's worker did the same and reported the failure as a draft verdict, because _jtbl_prep_one never read the post-carve make extract's exit code (the file-offset-0x4 shift is the tell: a DIFF that starts at the first code byte is never the draft).

Two fixes, both in our layer (R35/R40/R49/R61). (1) jtbl_carve._merge_pre: --order is derived from the CARVE SET, --pre is a property of the binary's LAYOUT (the resident's §8f leading data word emitted as hdr.rodata.o); the rewrite now carries an existing --pre forward (idempotent; overlays unchanged — 4-shape unit control). (2) harvest_verify._jtbl_prep_one: a failed post-carve re-extract now restores the snapshot and refuses loudly (CARVE refusal, NOT a draft verdict) instead of letting the build judge a stale script. With both, the same draft carved (JTBL_PADS 0,4; tables at +0x0/+0x1e0 of the new .rodata piece at file 0x451c0, pads tail3 at 0x453c4) and built byte-identical (8e17e02f…).

Sequencing law that bit twice in one task. After a FAILED extract, asm/ is half-regenerated for the config that failed (splat writes before ld_interleave refuses), so the next jtbl_carve sees "table not found in the raw data asm — already carved / stale asm?" — re-extract with the restored config FIRST, then carve. And a build after a failed extract is never a measurement (R53): the previous binary or a stale script is what you are hashing.

Tell + generalization. Any binary whose JTBL_INTERLEAVE carries --pre (today: the resident) was un-carvable by the gate since the §8f sandwich landed — the "NON-CONTIGUOUS carve" refusal of frontier-p32.md §1a was the FIRST wall, this was the second, and neither was the C. Sibling of §496/§497 (the isolation carrier's assumptions): the three resident blockers were all tooling that had only ever met overlays.

§499 ★★★ — A NEVER-ONBOARDED PAYLOAD PINS ITS OWN BASE STATICALLY, AND THE FIRST BUILD CANNOT (P32 T2a/T2b: the parked five onboarded, 213 → 218 binaries, 0 UNCLAIMED)

The problem. Five disc code payloads (MAIN/7, MAIN/9, SC03/53/54/56) had resisted every route for a month: no literal cdFileLocTable reference (§S45 p4), no resourceIdMap entry (p5), three payload-side base oracles refuted by their own controls (p5 addendum), a runtime tracer that never saw them load (p6). The §S44 module-class recipe's arbiter — "onboard at the derived address; the FIRST build's byte-identity decides" — had never been run on them.

What pins a base from the bytes alone (tools/payload_base_evidence.py). Three self-references, scored over a BOUNDED candidate list (the five §S44 slots ∪ the 134 IDXTAB DESTPTRs ∪ the vote's bases — never a search):

  1. Self-calls: every jal target that lands inside the module at base B must land ON one of its own function starts. The recall killer of S45's vote_base (4/12) was the start set: MIPS leaf functions have no addiu $sp prologue, so starts = prologues ∪ the word after every jr $ra+delay (the linear-partition boundary). With that, MAIN/7 lands 9/9 at 0x800CEDF8 and MAIN/9 6/6 at exactly one base (0x800CD348).
  2. Function-pointer tables: header/data pointers that land exactly on the module's own starts (SC03/54: 5 at 0x801EF468, 0 at every rival; SC03/53: 3 + its one self-call). As strong as self-calls.
  3. lui hi-halves must be able to reach [B, B+size) — a weak tie-breaker only. A jal→start VOTE (every target − start pair; ≥2 hits) proposes bases the slot/DESTPTR list lacks (MAIN/9's).

The two traps the controls caught (R39, seven banked modules must re-derive their byte-proven bases top-ranked):

  • the first scorer FAILED 5/7 — prologue-only starts + a top-rank assertion on modules the bytes cannot discriminate. Small non-self-calling modules are genuinely AMBIGUOUS among the 0x800Cxxxx slots; the tool must SAY so.
  • the requester cross-check ("a DESTPTR candidate's overlay should own the module's outward call targets") is confounded by the fleet's SHARED engine code — every requester fits — and scoring it put NO-EVIDENCE candidates above a true base (6/7). Informational only.
  • OUTWARD-EXPLAINED: SC03/56's vote base 0x80178C8C was STRONG (4/4 self-calls on starts) and WRONG: both targets are function starts in 141 overlays (shared engine), the base is nobody's DESTPTR, and one outward call hits a function only 3 overlays have — ov_SC03_002, whose DESTPTR 0x801CBB50 holds 15 of its 17 pointers. A pure vote base whose "internal" targets are overlay function starts needs no internal explanation; downgrade it.

The first build is a NULL oracle for FINE base errors (R34, measured). The same payload builds byte-identical at base+8 (.run/P32/t2b/control_fine.log) — splat names every reference by its absolute address, so a nearby wrong base re-assembles to the same bytes; only a GROSS error (an internal jal target leaving the window) fails, and it fails the LINK, not the hash (undefined reference to func_800CEEA4/func_800CF3F4 at +0x1000). So "byte-identical on its first build at the derived address" corroborates the address class, never the address. The base is byte-PROVEN only when a C bank calls an internal sibling (the call encodes base+offset) — which is also how a WRONG base would surface later: a whole module's functions refusing at the same relocations, a fake "wall" (R40).

Result. All five onboarded on their first candidate: md_MAIN_007 @0x800CEDF8 (A4: no resident symbol stack), md_MAIN_009 @0x800CD348, md_SC03_053 + md_SC03_054 @0x801EF468, md_SC03_056 @0x801CBB50 — +56 stubs (~2,600 ins) into the census, the parked-for-L3 ledger empty. Provenance rows: memory-map §S45 p7.

§500 ★★★ — THE T3 WAVE HARVEST (P32 T3, 2026-09-05): 31 one-agent-per-function drafters → 20 MATCH / 9 NEAR / 2 FAIL, every closer, two NEW named gcc mechanisms, and the wave-process defects

The wave. 47 targets (the T2c census of 54 minus the 7 pinned walls) on the model ladder — Haiku ≤50 ins, Sonnet ≤120, Opus above plus the 11 old near/far rows by escalation — one Agent-tool subagent each, packs from claude_wave_packs (journal notes + a matched same-TU neighbour), brief .run/P32/t3/BRIEF.md, laws SYS.md. The harness caps concurrent subagents at 20, so 31 launched and 17 Haiku rows stayed queued (pending_launch.txt). Verdicts (.run/P32/t3/verdicts.jsonl, from agent_verdicts.py): Opus 16 → 7 MATCH + 9 NEAR at exact length (0 walls, 0 FAIL) · Sonnet 4 → 4 MATCH · Haiku 11 → 9 MATCH + 2 FAIL, and neither Haiku FAIL was a compiler wall (one was a C-structure question Sonnet closed in one pass, one a TU declaration). 2,111 instructions matched (1,063 banked in the producing session, 1,049 verified MATCH and awaiting the gate — see phase-ends/CURRENT_PHASE.md for the bank order). The producing coordinator overflowed its context four minutes after its ninth bank, so 22 of these results were harvested by the successor session from the transcripts (§E below). Everything here is byte-measured by the agents AND re-verified by the successor with rtu_match in the real TU.

A. Closers that BANKED (10 functions, 1,063 ins).

  • main:func_80039B20 (79) — the S80 "permuter plateau at 7" was an ALIAS FLAG: *(s32*)(base + i*4) is a cast-wrapped PLUS and is DENIED /s (MEM_IN_STRUCT_P, expr.c:4570-4576), so true_dependence kept the edge to the fixed-address D_800C7D20 store and pinned the re-read below it; ((s32*)base)[i] makes the INDIRECT_REF operand a top-level PLUS_EXPR, grants /s, and the load hoists into the load-delay slot — zero drift, first compile. Addendum to §30a#1: when an integer LAUNDER (introduced to defeat move_movables) blocks /s, RE-INDEX rather than re-cast. Route law: "a redundant re-read scheduled at point-of-use instead of a preceding load-delay slot" → §30 first, never the permuter (it cannot reach an alias flag from C — why the S80 ILS plateaued at 7 from a seed of 7). Load-bearing kept: $6/$7 pins on i/off (−23 without), the volatile launder on the base (−39 without).
  • resident:func_800D06E8 (344 — the resident is now 145/145 C) — the closeness-0 body the journals pointed at (.run/S71b_1/fable/) was hidden by the pack's inline 292-ins "previous attempt"; the real blocker was declaration SCOPE: the TU defines Struct80078E78 AFTER the slot (line ~894) with a layout lacking bytes 0x36/0x37, so only a BLOCK-scoped typedef under a distinct tag (Blk80078E78) + a block-scoped extern compiles. Never run such a draft through a decl-hoisting resolver variant (-s2in-uni) — hoisting recreates the conflicting types. Idioms: §162k1 QImode (u8)(c-3) < 2 for andi + raw-$a0 addiu; an explicit flag temp t = (u32)(r-0x64) < 0x1E; s1 = t ^ 1; for addiu/sltiu/xori; both currentLocationId dispatches as switch decision trees. Its jtbl_80113FA4 carved by jtbl_carve --func (5 → 4 pieces, --pre kept).
  • main:func_80038698 (74) — "multi-block variable = global_alloc bank." gcc-2.7.2 runs local_alloc (pseudos whose refs sit in ONE block) BEFORE global_alloc (pseudos referenced in 2+ blocks); a value reused across blocks lands in the leftover $aN bank and cannot self-coalesce with a block-local shift destination. Tell: sll $v1,$a1,8 (fresh dest, source in an arg reg). ONE function-scope variable for the three BE-16 hi-bytes fixed all three sites AND the lbu 8/7($a3) sched1 hoist for free (the longer live range changed the DAG priority); the prior NEAR-11 drafts had pinned hi=$5 at one site only. Plus an in-place compound on a block temp (u32 t = hi << 8; t |= p[11];, §219 family — ties the or to the shift instead of the lo byte) and ONE pin register u32 len asm("$6") to rotate the last two global allocnos (all 24 declaration orders inert, as this TU's func_80035270 header already notes).
  • main:func_80023BF0 (281) — third witness for §364's −O2 half: the OT link is libgpu's P_TAG 24-bit BITFIELD store, not a hand-written (A & 0xFF000000) | (B & 0xFFFFFF): store_field → store_fixed_bit_field expands the RHS extract BEFORE the destination read, which fixes loop.c's movable order [&otab[idx]; 0xFFFFFF; 0xFF000000] AND the ior operand roles / $t3-$t4 pairing at once (the mask/or form can never satisfy both; operand swap alone = 35). Kept from the S79w 90→18 chain: the §194-A fence in both arms, code split from its base with compound accumulation, register u32 base __asm__("$3"), and exactly 13 zero-byte fences after pkt = D_800A5E60 (0→15, 8→13, 10/11/12→7, 13→MATCH). The S80 permuter waypoint (closeness 11) was semantically UNSOUND (m24 sunk into one arm) and was not used — an R63 witness.
  • md_MAIN_007:func_800CEE2C (52, Sonnet) — an early u8 *base = D_800AF630; local keeps the address in $s0 across two calls; a neighbour global (D_800B99E4 = base+0xA3B4) is read by RAW OFFSET with no relocation because the .s has none; &D_800CFABF captured ONCE into a pointer local reused for the p - 0xB argument and the *p = v store (two separate & expressions = three independent address computations).
  • md_SC03_054:func_801EF558 (96, Sonnet, first draft) — a literal used at two sites (6) modelled as ONE local keeps it in a callee-saved reg across the intervening jals.
  • md_SC03_053:func_801EF49C (49, Haiku) — explicit goto labels reproduce a beqz/beq layout an if/else chain reorders. func_801EF6B0 (33) — a switch on the u16 state; func_801EF95C (20) — twin remap from ov_SC01_077:0x80151980 + one guarding beqz on a callee's return; func_801EF624 (35) — a $16 pin holds the state pointer across both arms (the unpinned draft used $s1 in the false arm).

B. Closers verified MATCH in the real TU, not yet banked at this writing (10 + 1).

  • md_SC03_054:func_801EF6D8 (604 code ins + six jump tables, Opus) — a 12-arm dispatcher on *(u16*)(a0+0x34) with a 7-arm sub-state-machine on D_801F1480 duplicated in outer cases 6–11. Idiom 1: 2($sN) off a lui/addiu- materialised base means a POINTER LOCAL in the source — p = &D_801F1480; then p[0]/p[1]; array or struct access on the extern emits the absolute lui %hi(G+2) / sh %lo(G+2)($at) pair for EVERY access (the banked neighbour func_801EF558 does exactly that for D_801F1482) — 7 ins × 6 copies = 42 of the 44-ins gap. Corollary tell: the sub-switch selector still reads the global absolutely and the scheduler slots the lui/addiu $s0 pair into that load's delay, so p = &G; sits BEFORE switch (G) in the source although the address computation appears after the load in the output. Idiom 2: a constant-arm ternary chain consumed directly as a call ARGUMENT lays the third arm's value in its own block after the last arm plus a j (+2); the equivalent if / else if STATEMENT chain reproduces the target — while the same ternary matched in cases 0/1, where the value lives across a later call and lands in $s0 (the distinction only bites when the result is coalesced into $a0). 85/85 relocations verified in order.
  • main:func_80015B6C (120, Opus) — both journal walls fell. (a) The "6-ins $v0/$v1 swap in the index calc" was an ARTEFACT of forcing width/height into a source-level frame slot (the s32 wh[3] hack); with the plain shape the prologue is byte-exact and reload spills them to 0x10/0x18 by itself. (b) The corner-copy wall is cse.c canon_reg: make_regs_eqv keeps the FIRST register of a quantity canonical and a copy that dies in-block can never become canonical, so every use of vx before x += w is rewritten to x; the zero-cost fix is a barrier on the SOURCE variable, __asm__("" : "=r"(x) : "0"(x));, not on the copy — it re-defines x's quantity without touching the live ranges that decide the colour-vs-width spill (moving the += earlier flips local-alloc's priority and spills two colours, +5). (c) Four empty __asm__ __volatile__("") fences partition block 2 into the target's five phases at ZERO cost when each phase fills its own load-delay slot (an earlier "+2 nops" probe was a placement fault). (d) The 0xE1000200 trailer lui hoist is a register-LIFETIME effect: pin the block-2 tag to $6 and REUSE the same variable for the trailer constant, assigned after the packet-tag store. (e) Naming the addPrim mask pmask = (u32)p & 0xFFFFFF; keeps the second RMW's and inside the first. Inert: pins on otp/$16, one/$8, m24/$5, p/$2 (+1 move), HImode asm-pinned copies (+5), all mutation reordering.
  • main:func_8002FDE8 (73, Opus) — the four prior attempts' "regalloc-priority wall" at 35 was the ARRAY spelling of D_800A46D2; the same-TU neighbours func_8002FF0C (src/800_b_2.c:2866) and func_800301C8 (:3000) document the fix — a block-scope SCALAR extern s16 D_800A46D2; (the array spelling makes cse materialise the address into a callee-saved reg; the scalar emits two independent %hi/%lo), 35 → 3 instantly (R38: read the neighbours the pack names). Then: the third return 1 INSIDE the == -1 arm (a trailing fall-through inverts the final beq and swaps the tail blocks); the §47 live-length slider COMPUTED, not searched — -dl -dg gives data = 2 refs/15 ins → pri 1333 vs the 1 constant = 4 refs/58 → 1379, so one zero-byte fence placed where the constant is live and data dead flips allocno_order; the 1 constant needs its own statement (one = 1;) so a fence can separate the li from the sb; i4 = idx * 4; as a separate statement (the sb/lw memory dependence makes sched1 order them by LUID). 118 fence-only variants plateaued at 4 — the shape facts, not the fence, closed it.
  • md_SC03_053:func_801EF734 (44, Sonnet after a Haiku FAIL at 47) — §224: register-pin the pointer ($16, like its sibling func_801EF624) and write the shared func_80178CBC(s0, &D_xxx) call LITERALLY in each of three uniform if / else if / else arms with no early return; cross_jump merges the three identical jal + epilogue tails and lands move a0,s0 in each arm's delay slot. Haiku's plateau was a register-COALESCING failure (mixing an early-return if with a separate nested if/else made gcc allocate $s0 AND $s1 for one value: spurious move, extra sw/lw, frame 32 vs 24) — not the missing merge it reported.
  • md_SC03_053:func_801EF7E4 (72, Sonnet) — §339's redundant-default test is the switch tell; combined case labels case 0: case 3: { … } place the shared body FIRST, while two textually-duplicated bodies relying on cross_jump land it LAST next to the epilogue regardless of source order — cross-jump placement is not steerable from C.
  • md_MAIN_007 Haiku five — func_800CF148 (33): the TU declares func_800CF3B0 as void (void*,void*,void*); a block-scope redeclaration s32 func_800CF3B0(void) shadowed it (the call has no argument setup and tests $v0). func_800CF2BC (32): post-increment passes the old counter to func_800CF6D0(val, 0). func_800CEEFC (25), func_800CEF94 (25): goto shape + (result << 16) == 0 forces the sll before the beqz. func_800CF068 (21): family_remap from md_MAIN_008:0x800cee44 refused (reloc count 6≠7) — hand-drafted from the .s using the twin's shape only. func_800CF3B0 (22) is leaf-exact (match_one 22/22) and refused only by the TU's own extern void func_800CF3B0(void*,void*,void*) (the asm returns s32 in $v0, takes nothing): a §376/§495 def-side declaration wall — the no-proto spelling extern s32 func_800CF3B0(); keeps the banked 3-arg caller func_800CF390 legal and admits the (void) definition; the byte gate decides.

C. The NEAR residuals (9, all exact length, all rtu_match-clean) — classes and the levers measured INERT.

  • main:func_8001BC6C 28 → 6 (69): (a) the & 0xFF000000 | & 0xFFFFFF RMW pair on a prim/OT word IS the libgpu P_TAG bitfield and the two spellings are NOT byte-equivalent (28 → 16; exemplar src/boot.c:1008-1029 func_8001212C, the same routine at −O0); (b) fold REASSOCIATES (a1<<16) | ((a1<<8)|K) | a1 → name the inner term as a temp (invisible to match_one, same multiset); (c) the block emits in pure source/LUID order — write the statements in emitted order, c as a two-set pair, and the OT lookup split into idx = D_800B9A02 << 14 + the add (the fused form costs a load-delay nop). Residual = a pure $v0/$v1 swap across the six OT-chain insns: local-alloc.c qty_compare priority floor_log2(n_refs)·n_refs·size/(death−birth), ties → lower qty number, alloc order ascending regno (MIPS has no REG_ALLOC_ORDER): base 2/2 = 1.0 > index 8/9 = 0.889 — one span unit; needs one fewer insn between the lhu and the addu at sched1. color must stay pinned $5 (−18 without). Class REGALLOC-PERM; the permuter cannot take the pinned seed as-is (pycparser rejects register __asm__) → §494's repaired permuter_ils recipe. Meta-lesson: five agents had converged on 28 with an "RTL-proven sched1 unreachability" — a converged multi-agent verdict on a MECHANISM is not evidence the BODY is right.
  • main:func_80039308 34 → 17 (518): s16 b4; n = b4; — a short whose only def is an lbu emits the target's widening addu $v1,$s4,$zero copy with NO extension (the journal's "conserved andi" verdict was about the mask; the s16 declaration is the lever); the identity of the DEAD LOCAL a temp reuses is a first-class regalloc lever — sweeping ~25 candidate names per temp slot at ~1 compile each found 3 of the 6 wins (the 20-line sweeper is kept at .run/P32/t3/restored/sweep_func_80039308.py (not kept); tool candidate); register s32 two __asm__("$2") on the S349 dead-reset constant (swept $2..$8, $2 unique); TWO variables pinned to ONE hard reg ($4) when the address dies into the loaded value (lw $a0,0($a0)); naming the byte offset in a dead local + computing t9 before cnt takes the MULT out of the MEM address so the PLUS keeps operand order. Residual: preheader 49/50 swap, the un-spellable addu $a2,$a0,$zero (every p = r form is cse-propagated), a temp on $t0 vs $s7, and 11 insns of one alias fact (the second D_80073140[j] load cannot be scheduled above the D_800C7D20 store from C; the /s unlock costs the address allocation, net 20–24). permuter_ils --klass REGALLOC 2×150 s: no gain.
  • md_MAIN_009:func_800CD674 102 → 2 (174) and func_800CD92C → 15 (247) — the SPRT-with-tpage family. Mirror image of §364: the primitive's FIELD stores must be NON-STRUCT lvalues *(T *)(p + off) at −O2. true_dependence drops the store→load edge when the store is MEM_IN_STRUCT_P with a varying address and mode ≠ QImode, so with a Sprt24 *p the five sh stores sink past the lhu D_800BAE22 while the sb/sw ones stay — a store block split exactly on instruction WIDTH, unreachable by any statement order (verified: order is preserved within each mode group, the split point never moves). Also: the OT store as a genuine ARRAY_REF lvalue (D_800ABA24[oi] = …, not *ot) lets the final cursor publish hop above the last OT RMW (5 ins); mutate the parameters (x -= 0xA0) rather than sx = x - 0xA0 (cse re-folds sx + 0x100 into x + 0x60); one block-scoped ot binding per addPrim half (a function-scope ot with 6 def/use pairs in one huge block gets a dedicated $t8); chained r0 = g0 = b0 = c for the reversed 0xA/0x9/0x8 order; y + 0x100 inline beats y += 0x100 (the accumulate lengthens $a1's chain and sched2 hoists the addiu to the prologue head: 16 → 7); an otv = *ot; temp splitting the RMW puts the D_800A71D0 publish in the last lw's shadow (7 → 2). Residuals: func_800CD674 REGALLOC-PERM $a3↔$t1 (prim 3's masked p wants $a3, prim 4's wants $t1, one shared local can only be one; splitting = +1 pseudo that displaces two hoisted constants, 31) — a permuter seed; func_800CD92C = map §S7 prologue WEAVE (the {sw,lui,ori} groups for 0xE100008D/8F after the 9-insn li block instead of before: the §17 pins reproduce the ALLOCATION but the hoist then happens in sched2). Inert: pin declaration order (8 perms), assignment placement (6 anchors), volatile on the index, p += 0x18 spellings, typing the cursor u8 *, every pin subset; worse: caller-saved pins (26), the index as an array (−14, cse folds the loads).
  • md_MAIN_007:func_800CF6D0 137 (249, plateau) and func_800CF408 49 (178) — same family in another TU (exemplars src/boot.c func_8001212C −O0, src/ov_SC01_000/…_jr_8017BEBC.c func_8017DD04 −O2, §351): §351 plain-cast tag word (a P_TAG COMPONENT_REF grants /s and lets cse keep the D_800B9A02 index live across the store — 12 lhu reloads collapse to 6, −25 ins: §364's −O2 half confirmed again, twice); §351 base-split (block 1 off the raw symbol, ob thereafter); §195-I in-place arg0 = arg0 - 0xA0; u32 pad[2] for the +8 frame (§358). New: tpage stored BEFORE len in each block is the only one of 19 swept field orders that hoists five single-use tpage constants to the block top → the target's four callee-saved registers and 249 ins (len-first gives three and 247). arg1 += 0x100 IN PLACE inside block 2 (its anti-dependence is the target's y0-before-x0 inversion). Two LENGTH-bearing pins in func_800CF408: register u32 tp __asm__("$17") shared by 0xE1000087/0xE1000097 forces the 6th callee-saved register (unpinned, gcc hands $s1's tail to the 0x40 constant → 5 regs, 176 ins = exactly the missing sw/lw pair); register u8 *ob __asm__("$10") fixes the $t1/$t2/$t3 rotation (56 → 49). Residuals: sched1 rank_for_schedule last-insn-CLASS tie (the -dS dump shows every store at priority 2 with equal ref counts; QImode stores grouped, the two loads floated between them, HImode stores after — in 5 of 6 blocks) + a $t1↔$t3 local-alloc swap of the two masks; func_800CF408: two prologue sched2 slots, an mlo/mhi allocno tie, a 3-insn block-2 head hoist. Inert (137 for ALL): pins on the tpage constants or masks, asm re-ties, volatile/"memory" fences, /s-denial on any store subset, *0x4000 vs <<14, p++ vs p+0x18, | operand swap, 19 field orders. The decomp-permuter's "122" moved block 5's clut store into block 4 — semantically wrong, rejected (R63).
  • md_MAIN_003:func_800CF3E8 54 → 27 (469; blocks 1–2 byte-exact, idx 0–361) — blk2 constant BIRTH order: mq2 = 0xFFFFFF; AFTER the 0xE100008F tag store lets dbr share the entry-beq delay-slot lui $v1,0xE100 into both ori 0x8F and the j's ori 0x8A (5 rows; same lever in blk1 via a shared mq2 + ADDPRIM2, 3 rows); a 3,697-candidate single-move climb over the p5/p6 tail (44 → 34); a "birthing-boost LOCAL" for a per-prim constant (w60 = 0x60; … pC->w = w60;) decouples the li 0x60 from the post-fence region so it schedules right after sh x0 — and ONLY then does relocating the §194-A fence from [y0|u0] to [v0|clut] work (inert separately, 7 rows together). Residual = ONE cause, mechanism §D-1 below. Inert: all 9 pins load-bearing (ablation +5..+1409), the __asm__("" :: "r"(ot)) position across 7 slots (removing it +14), the h6 tag-read hoist at all 32 positions × 2 operand orders (the load moves early but a load-use nop appears, 470 ins), a leading u32 tag struct reshape, the P_TAG bitfield ADDPRIM (75/83), array-based p6 stores (470), ~92k annealed tail variants — all hold at 27.
  • ov_SC07_002:func_8017DC80 84 → 46 (346 — the historic −33 LENGTH wall closed): (1) the GTE macros must be REAL macros — undefined gte_*() compile to implicit jals; use the TU's own house block (src/ov_SC07_002/ov_SC07_002_jr_8017C8D0.c:2701-2818, _1 at :7010), suffixed _2; gte_stsxy0/1 (single swc2 $12/$13) and gte_avsz4 were the two it lacked. (2) §30's /s lattice schedules the sxy block: the frame words must be PLAIN SCALARS (no /s, fixed address) and the packet stores COMPONENT_REFs through a struct pointer (/s, varying) — the other way puts a nop in every load-delay slot. (3) Derived pointers must exist as PINNED C variables: plain r + k makes loop.c synthesise a FIFTH induction variable at r+0x27 and ot spills; four unpinned pointers get the IV count right and the register permutation wrong. (4) cse1 unifies the OT index and the n < 4 test across func_80010A08(8) (a call invalidates memory and hard regs but never pseudos — cse.c:7241 invalidate_for_call); a zero-byte __asm__("" : "=r"(n), "=r"(otz) : "0"(n), "1"(otz)) retires both pseudos so the target's second copy is re-emitted — with mechanism §D-2 below. (5) otp = (u32*)(otz<<2); otp = (u32*)((u32)otp + ot) (two steps) coalesces the shift temp into otp so dbr MOVES rather than duplicates the sll into the delay slot; a $0-add opaque copy + mm/mode pins reproduce the OTZ-clamp's addu $v1,$v0,$zero; register u32 w __asm__("$3") stops w coalescing into n's $s0. Residual (46): frame 0x70 vs 0x60 — the target carries two extra 8-byte RELOAD slots (sp+0x38/0x40) its own code never touches; a dead local always takes sp+0x10 (aggregates get their slot at expand_decl, measured with long padx[5]) so §162i1's frame-pad lever is REFUTED here — this is reload pressure, unreachable from C; the la $a0 prologue slot (every other prologue insn already in target order; source order inert); two local-alloc dest-ties-dying-source picks (t→$v1, DR_TPAGE temp→$s0; pinning either is worse, +2/+12).
  • ov_SC03_105:func_80185810 37 → 35 (489; every register allocation now correct) — it IS compiler-emitted −O2 C (splat's "Handwritten function" banner is wrong). GTE macro spellings reconstructed from the bytes: gte_ldclmv / gte_stclmv via lhu+mtc2 at 0/6/12 (the pack's prior body had lwc2), gte_rtir/gte_rt carry TWO nops not six, gte_ldlvl loads +4 before +0 — in-tree precedents src/800.c func_800139C8/func_80021D38. Levers: base = D_800AF630; as a local (gcc won't hoist a symbol used in 3 blocks into $s4 by itself); the packet word w as ONE expression at its store, not a |= chain (a variable with 5 sets is never birthing_insn_p, §49 — worth 40); (tp + 0x100) << 6 | M must be a TWO-USE temp or combine folds the addiu away via nonzero_bits on an lbu; uu = (uu - …) << (2 - mode) as ONE statement fixes mode's register, splitting it back AFTER the fence fixes uu's; a zero-byte fence after p[7] |= … is the ONLY lever that puts the $a0–$a3 quartet on the target registers — and the same fence forbids the target's sll/andi interleave (trade measured both ways: no fence = right order, wrong registers, 67); ot16 split from ob around the fence; cl &= 0xFFFF in place. Residual = [permuter] emission order in four windows of one 50-insn block (rank_for_schedule/LUID ties). Inert: all 6 orders of the three loads, 9 placements of the D_801BA6B0 read, cl spellings, a shf temp at 4 positions, signedness of uu/mode/w, 4 operand orders of ob + (otz<<2), 20 other fence positions.

D. Two NEW named mechanisms (byte-verified from -dS/-dl RTL dumps; not previously in the map).

  1. The pinned-base-vs-pseudo-address alias basin (func_800CF3E8). sched1 runs BEFORE regalloc and fixes the final order here. At that point a whole-word access through a pinned struct pointer, *(u32*)p6, expands with its address in a PSEUDO ((mem:SI (reg 429))) while the field stores p6->x use the PINNED HARD reg ((mem/s:QI (plus (reg 3) 11))). sched.c:memrefs_conflict_p can only disambiguate two refs off the SAME base rtx, and cse.c:canon_reg explicitly refuses to canonicalise a hard register — so the tag load takes a true dependence on all ten field stores and cannot hoist the 20 slots the target hoists it. The register … asm("$3") pin on the struct pointer is itself the blocker: unpinning DOES hoist the load (verified) but relocates the whole tail into a different allocation basin (188–197). Tell: a load through a pinned pointer that the target schedules early and yours leaves last, with every other insn in place. Lever space: an unpinned base for the load only (a second, unpinned alias of the pointer) is the untested next move.
  2. #line-equalised ASM_OPERANDS for cross-jumping (func_8017DC80). jump.c decides whether two tails can be merged with rtx_equal_p, which for ASM_OPERANDS compares the source FILE AND LINE of the asm — so two identical zero-byte launders in two tails that the target cross-jumps into one block (the two j .L8017E0E0) must carry identical #line directives, or the merge is lost and the function grows (+13 here). Standard fix for "a zero-byte launder costs 13 instructions": wrap both launders in the same #line N "x" pair.
  3. Corollaries banked above that generalise the map: the local_alloc-before-global_alloc scope law (A, func_80038698), the MEM_IN_STRUCT_P mode-keyed exemption as a per-access dial (struct member = may float past a fixed-address ref unless QImode; non-struct = anchored; §364 for the tag, C for the payload fields, the ARRAY_REF lvalue for the OT store), and qty_compare/§47 as a COMPUTATION — read n_refs/live_length from .lreg, evaluate floor_log2(n)·n/L, and you know how many static insns you need and where.

E. Wave-process defects (fixes in docs/wave-playbook.md §S80 addendum-2; accelerators P32 T3).

  1. A shared scratch directory is a shared blast radius (R48 class). One agent's tidy-up — find .run/P32/t3/opus -maxdepth 1 -type f ! -name 'func_800CD674.c' -exec mv {} _scratch/ — swept ELEVEN sibling deliverables (two MATCHes, 604 + 120 ins) out of the contract path; another agent's rm -f globs in the same dir hit sibling intermediates. Fix: per-function work dirs, deliverables in a dir no agent has a reason to clean, and a brief line forbidding any find/rm/mv outside the agent's own dir. Recovery route (worked): agent_drafts_restore.py replays the transcript.
  2. A coordinator that reads 31 prose results dies. The producing session hit "Prompt is too long" four minutes after its ninth bank; 22 completion notifications (2–4 KB each) then arrived into a dead session. Fix: the agent's final message is exactly ONE JSON line; the prose goes to .run/<wave>/reports/<fn>.md, read only on routing.
  3. The harness caps concurrent subagents at 20 (CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS) — keep a dequeue file, launch one per completion. masked_diff.py's per-PID probes live in src/ and are visible to (and were deleted by) concurrent agents — move them under .run/ (tool fix candidate). A Haiku "cannot be influenced from C" verdict on an exact-ish length structural residual is an ESCALATION signal, not a wall (1/1 closed by Sonnet in one pass).

F. The queue (S83, session 491895ad, 2026-09-05 10:57–11:20 MDT): 17 Haiku rows → 17/17 MATCH, 28 banks in the session, and the three integration classes a per-draft rtu_match cannot see.

  • Yield. The 17 rows the 20-cap had queued (13–25 ins, module code: md_MAIN_007 ×6, md_MAIN_009 ×8, md_SC03_054 ×2, md_SC03_056 ×1) were launched at once from the staged prompts (.run/P32/t3s3/prompts/<fn>.txt = PROMPT_TEMPLATE + two integration sentences) and ALL 17 reported MATCH, 43–57k tokens and 46–126 s each; the five twin-hinted rows took the family_remap/twin-shape route. With the morning's 11 (the S82 MATCH-unbanked rows) that is 28 banks / ~2,800 ins in one session, zero walls, zero escalations; md_SC03_053, md_SC03_054 and md_SC03_056 reached 100% C; census 44 → 16 (7 pinned walls + 9 NEAR, 3,827 ins). Haiku is the right tier for this whole band (§500 A/B agree: 9/11 → 26/28).
  • Class 1 — two drafts of one TU spell one global differently. func_800CEEFC (file-scope extern s16 D_800B99E8) and func_800CEF94 (block-scope extern u16 D_800B99E8) each MATCH alone — rtu_match splices ONE draft — and the batch dies at cc1 (conflicting types). The byte-bearing spelling is the LOAD's (u16 → lhu); a store of zero is sign-blind. Law: when N drafts land in one TU, the BUILD is the verdict, not N green rtu lines; reconcile to the spelling the bytes need. The session helper .run/P32/t3s3/bank.sh (verbatim grep → rtu ×N → splice ×N → ONE build → sha vs config/check → commit only on green, tree left for diagnosis on red) is the shape a rtu_match --batch should take.
  • Class 2 — §304 self-defining rodata, three times in one wave. func_800CF068 (.asciz "C:\TIMPACK\OPDEMO0.PAT" as D_800CEDFC), func_800CD3B8, func_800CD520 (D_800CD364, eight .words — the last two, 0x3C02800C 0x9442AE04, LOOK like code and are the island's trailing junk; reproduce them as words). Compile-only rtu_match is BLIND to it (the undefined reference appears at LINK). The one prompt sentence — "if your target .s carries a dlabel block in .rodata, DEFINE it at file scope from the bytes" — made both later agents do it unprompted; it is now in BRIEF.md. A file-scope definition at the extern's position keeps the island order under the §303 derive stage (0x0/0x4 held).
  • Class 3 — a verdict is relative to the TU at verification time. func_800CD520's agent verified against a TU that did not yet declare func_8001AD38; the sibling bank func_800CD3B8 then added extern void func_8001AD38(const char*) at file scope and the queued draft's (void *) prototype became a conflicting types CC1 FAIL. Re-verify EVERY draft in the CURRENT TU immediately before splicing (bank.sh does), and re-spell to the TU (§376) — codegen is unchanged for a pointer argument.
  • Class 4 — the md_ leading-island jtbl route is §303, not §260.* jtbl_carve --probe md_SC03_054 --func func_801EF6D8 refused the tail carve (seven tables at file 0x4..0xF0) and named the §260 island split; for an md_* module the Makefile's jtbl_rodata_pads --derive stage already reproduces the pads at build time — splice, build, read the spec it prints (0,0t1,0t1,0t1,0t1,0t1,0), no isolation, no yaml/overlays.mk line (R60 untouched). The probe's refusal text now says so. Census nins counts .s LINES including rodata .words (func_800CD520 "22" = 14 code + 8 data); rtu_match's MATCH (N ins) is the code count — quote which one you mean.

G. The NEAR tail (S83, 11:35–12:10 MDT): one REGALLOC-PERM seed CRACKED by permuter_ils + an R63 read, one plateau confirmed, and a build-instrument collision fixed at the consumer.

  • main:func_8001BC6C 6 → 0 (69 ins, BANKED 143dbb89). permuter_ils … --klass REGALLOC --cycles 8 --secs 150 --j 3 on the pinned Opus seed (§494's repaired recipe) went 6 → 6/5/1 in cycles 6–8. R63 read of the "1": three mutations — (a) idx = D_800B9A02 << 14 moved AFTER color = c2 | a1; (b) a dead tag = (a1 << 8) | k; inserted at the top, BEFORE k's assignment; (c) (D_800B9A02 & 0xFFu) << 14 — and (c) narrows the target's lhu to an lbu (semantically wrong; the masked scorer rewards it — the second such witness after S80's addiu→sw). (a)+(b) alone = leaf MATCH; (a) alone 8, (b) alone 21. The lever is the early BIRTH of the tag/k pseudos, the §47/§500-C live-length computation made real: qty_compare = floor_log2(n_refs)·n_refs·size/(death−birth); lengthening tag's range drops its priority below the OT index's and the $v0/$v1 swap across the six OT-chain insns disappears. Well-defined spelling that keeps the bytes: k = 0; tag = (a1 << 8) | k; at the top (the value is recomputed below); every other early birth measured off — k = K first 15, tag = a1 << 8 19, tag = a1 8, tag = (a1<<8)|K 20, tag = 0 8, tag-then-k 8 (.run/P32/t3s3/p1bc6c/, 11 spellings). Route law: when the permuter's waypoint carries a width/semantics mutation, do not discard the waypoint — subtract the unsound mutation and re-measure; the sound remainder was the whole answer here.
  • md_MAIN_009:func_800CD674 stays 2. The same recipe's best waypoint (output-2-1) is the SAME $a3↔$t1 pair (and 156 / or 164), no drift: a genuine local-alloc tie the permuter cannot cross from this seed; ledgered with its cost (8 × 150 s, 0 tokens).
  • Instrument: gate_main's first clean rebuild died on src/.masked_diff_probe.3973390.c — a concurrent agent's masked_diff per-process probe (written and deleted within one call) was present when make parsed C_SRCS := $(shell find src -name '*.c' …) and gone when its rule fired: "No rule to make target build/src/.masked_diff_probe.N.o, needed by build/us/SLUS_007.26" → "batch FAILED (sha None); bisecting" → one wasted rebuild, then BANKED. Fixed at the CONSUMER (R54): the find now carries -not -name '.*' (Makefile), so every tool that probes in src/ is covered at once; control = a throwaway dotfile absent from make -pn's C_SRCS, src/800.c present. (§500-E3's "move the probes under .run/" remains a nicety, no longer a correctness fix.)

H. md_MAIN_003:func_800CF3E8 — §500-D1's mechanism CORRECTED and its open lever REFUTED (S83 bounded Opus second look: 245k tokens, 29 min, ~16k compiles at 70/s; 27 holds).

  • Not cse.c canon_reg (which returns hard registers unchanged; -dr shows the tag load BORN as (mem:SI (reg/v:SI 3))). It is cse.c find_best_addr, reached from fold_rtx case MEM: it replaces a MEM's WHOLE address with the cheapest equivalent in the hash table, and COST() scores a valid-quantity pseudo 0 against a hard register 1 — so a bare-REG (offset-0) address is swapped to the pseudo copy (429) while (plus (reg 3) N) field addresses are not table values and keep the hard reg. expand_expr's "generate all results into pseudo registers" forces that copy at −O2, so no address spelling dodges it (b_+idx, (u8*)b_+idx*24, b_ then += idx: all 27) — which is also why the leading-u32-tag struct reshape was inert.
  • The D-1 "unpinned alias of p6 for the load" lever is REFUTED, 5/5 at 27 (alias load-only, alias load+store, alias born from the same &b_[idx], declared last, born before the chain): cse puts the alias in the SAME quantity, so find_best_addr swaps it to a pseudo too; the load's base can never equal the stores'.
  • What DOES move it (new): launder the PINNED pointer itself — __asm__("" : "=r"(p6) : "0"(p6)); zero bytes, retires reg 3's quantity; the load keeps (reg 3), its LOG_LINKS collapse from 11 field stores to 1 and it hoists ~15 slots. With p6->w moved after clut: idx 0–377 AND 385–399 byte-exact, the load at 384 vs the target's 380 (+1 load-use nop) → 79 @ 470: structurally far closer, numerically worse. Residual = ONE cause, class [permuter]/basin: the freed load lands 4 slots late in a 7-slot window, flipping local-alloc so 0xFFFFFF takes $t2 not $a2 (the 22 and rows). Inert/worse: launder × w 2-D (324) floor 27 · × fence 3-D (5,832) floor 27 · ior-form (1,026) 27 · explicit tag6 + named m24 + fence-after-read (8,704) floor 49 · tail-mask pins (93–132) · p5 x0/y0 swap (31) · ot-launder moves (39 / 45 / 143). Row stays in the backlog; only a NEW idiom that places the freed load exactly at slot 380 reopens it — otherwise a T4-style CANDIDATE wall with this mechanism as its citation. Report: .run/P32/t3/reports/func_800CF3E8__opus_s83.md; draft .run/P32/t3/opus/func_800CF3E8_s83.c (== the prior best).
  • Map corollary: for a pinned struct pointer at −O2, "the load through the pinned base is scheduled late while every store is in place" = find_best_addr took the offset-0 address to a pseudo. The zero-byte pointer launder is the release; where the released load must land is then a scheduling question the launder does not answer.

I. T4 — the seven pinned walls' FINAL verdicts, every one re-probed IN TU CONTEXT without touching src/ (S83, 12:05–12:35 MDT).

  • Method. Each row's best draft was re-run with rtu_match in its CURRENT real TU (S83 had changed src/800.c and src/800_b_2.c). Four rows reproduced their recorded closeness directly (func_80011380 6 with --o0, func_80020DA4 2, func_8017DF28 2, func_801834A4 6 ×3 variants). Three were CC1 FAILs — and all three were plumbing, none a verdict (R40): (1) func_80032A74's draft redefined seven structs the TU provides via 800_shared.h and spelled four declarations its own way → cdecl.strip_provided_typedefs(draft, cdecl.typedef_names(tu)) + adopt the TU's four lines → DIFF 1 in the real TU, the recorded lh/lhu residual; (2) func_80039DEC's TU carries a narrow-typed PROTOTYPE (extern void f(void *, s16, u8)) that C forbids against a K&R definition; (3) func_800391D4's lever spelling extern s32 D_80073140[][1] conflicts with the TU's three [] externs — and the [][1] TYPE is load-bearing (the TU's spelling, a (s32 (*)[1]) cast and a byte-offset form all regress 3 → 65 @ 76 ins). For (2) and (3) the re-probe used a SANDBOX TU: copy the TU under .run/P32/t4/tu/, symlink src/*.h + src/shared beside it, edit the declaration there, and pass --tu <copy> — the residual reproduces (2 and 3) with zero edits to src/, and no byte-neutral commit is spent on a row that will not bank. The TU-side edits are recorded in the backlog rows for the day closeness reaches 0.
  • Verdicts. main:func_80011380 PROVED (§474, fold-const.c:882 split_tree + stupid.c:497, --o0; DIFF 6). Six CANDIDATE with current citations and bounded attempt records: func_80032A74 1 (extendhisi2 force_not_mem / §172 producer 3, reload1.c:1445) · func_80020DA4 2 (mflo destination REGALLOC-PERM; pins regress to 79) · func_80039DEC 2 (K&R narrow-parameter argument-position promotion → $a3/$t0 before global-alloc) · func_800391D4 3 (move_movables splices hoisted invariants after preheader flow code, loop.c 2.7.2:1529 / map loop.md L4) · ov_SC06_022:func_8017DF28 2 (expand_block_move copy_addr_to_reg pseudo cse-reused, cse_expr.md [A23-2]) · ov_SC03_105:func_801834A4 6 (loop.c movable ordering). No verdict changed; the pin file (config/wave_exclude.txt) carries each row's S83 re-probe line; exclude_audit --assert-fresh 7/7. Unpinned candidate with a cited mechanism: md_MAIN_003:func_800CF3E8 27 (§500-H).
  • Law. A wall's CC1 FAIL in its TU is evidence about the TU's declaration environment, never about the body; re-probe before any verdict, and prefer a sandbox TU over a src/ edit when the row is not going to bank.

§501 ★★★ — A LEVER THAT MEASURES WORSE MAY BE A CASCADE: READ THE .loop DUMP FOR THE DESIRABILITY FLIP BEFORE DISCARDING IT (P32 T4b, main:func_800391D4, a pinned wall banked by a Fable agent)

The row. Pinned since S79 (closeness 3, "move_movables splices hoisted invariants after preheader flow code, so off's init cannot follow arg1's hoisted sign-extend from C"); four attempts + permuter null; T4 re-probed 3 in a sandbox TU. The S83 hand pass tried the obvious lever — write the s16 parameter's promotion as preheader SOURCE code (a1v = arg1; before off = 0;, compare against a1v) — measured 18 and recorded it as worse. It was the RIGHT lever.

What the 18 really was (the agent read dumps/*.i.loop). Making the extend explicit moved its two insns OUT of the loop body: Loop from 38 to 189: 59 real insns → 57. move_movables' desirability test (loop.c:1631) threshold * savings * lifetime >= insn_count with threshold = 2 * (1 + 28) = 58 (no call in the loop) flipped for the lui/addiu D_800C6DD0 address (savings 1, lifetime 1): 58 < 59 "not desirable" became 58 >= 57 "moved" — +2 preheader insns and a register cascade. Every one of the 18 rows was that single hoist. The fix is to restore the knife-edge: the S79 draft's seven __asm__("") pads (each counts as a real insn for insn_count, emits nothing) become NINE. Then the preheader tail is the target's sll $5,$5,16 / sra $5,$5,16 / move $8,$0, and rtu_match MATCH 75/75 → gate_main BANKED.

Laws.

  1. When a mechanism-grounded lever regresses, do not trust the number — diff the .loop/.greg dump of the lever variant against the seed and look for a SECOND change (a hoist, a spill, a coalescing) that the lever triggered. Fix the second change with its own zero-byte dial (pads for insn_count, a fence for a live range, a pin for an allocno) and re-measure.
  2. insn_count is a dial and a knife-edge. __asm__("") pads count as real insns for loop.c's threshold and emit nothing; any source change that adds or removes in-loop instructions (hoisting a promotion, folding a temp) must be paid for by re-counting the pads so the SAME movables stay un-hoisted. The S79 seed had already discovered the pads; the S83 hand pass did not re-count them after moving the extend.
  3. A register … __asm__("$N") pin on the loop counter keeps it from being a biv (loop.c:3572: hard regs are not bivs), so D[i] is never a giv — the target's per-iteration sll/lui/addu/lw survives. Keep such pins when the target shows an un-strength-reduced index.
  4. The verdict chain for a wall row: hand pass names the mechanism (done S83), the agent reads the dump the hand pass did not (the .loop desirability lines), and the byte gate decides. Cost: one Fable agent, two interruptions by usage limits, the draft written before the first crash — recovered from .run/P32/t5x/fable/ (write deliverables EARLY, R55).

§501-B — DO NOT MERGE CASE TAILS IN C: cross-jump merges them AFTER allocation, and a merge deletes a REFERENCE that decides global-alloc's order (P32 T4b, main:func_80039DEC, a pinned wall banked by a Fable agent). The row's residual was the raw copy of the K&R s16 a2 landing in $t1 instead of $a3 (a1's copy in $t0 in both). Pinned as "fixed by argument position"; the S83 hand pass guessed a block-local pseudo in $a3 (global.c local_reg_n_refs) — both wrong: that check lives only in the best_reg < 0 retry (global.c:1108–1160). Mechanism (global.c:587–608 allocno_compare): allocnos are ordered by floor_log2(n_refs) * n_refs / live_length; find_reg (:960–985) gives the LOWEST free hard reg to whichever allocno comes first. Every prior draft merged the 0x14/0x28 case tails at C level (goto L_merge + a kind temp), deleting one sb a2 — a2-raw fell to 3 refs / 45 insns (pri 666) and a1-raw (3 / 26, pri 1153) took $a3 first. The original wrote the natural switch (a2) with the tails DUPLICATED; jump.c cross-jump (:1923) merges them post-reload, so at flow time a2-raw has 4 refs / 41 (floor_log2(4) = 2 → pri 1951) and is allocated FIRST → $a3. Block-scope u8 *p per case, no pins, no fences, no do { } while (0). Verified in .run/c294/dumps_t5x_9DEC_sw2 (73 in 8 75 in 7). Corollaries: (1) a permuter "2" that reads an uninitialised temp is R63-unsound — its dead pseudo can be the thing occupying the wanted register; (2) when two parm copies swap registers, compute both allocnos' n_refs/live_length from the -dl dump and ask which SOURCE shape adds or removes a reference; (3) docs/gcc-2.7.2-map/regalloc.md gets the priority formula and the "cross-jump is post-alloc" law.

§501-C — THE DYING-INPUT SUGGESTION vs THE BIRTHING BOOST: a shared local that dies in two places is not local-alloc's to give, and a fresh single-set pseudo is sched1's to glue (P32 T4b, md_MAIN_009:func_800CD674, the $a3↔$t1 REGALLOC-PERM plateau banked by a Fable agent — 174/174). The row: prim 4's masked pointer and/or wanted $t1 (= the pinned mask m24's register dying at that and); one shared pm2 for prims 3 and 4 gave $a3 to both; every "split it into two locals" form regressed to 31 (S82) and the S83 hand pass measured three short-lived-temp spellings at 31/44. Read in -dl/-dS + source:

  1. A pseudo that "dies in 2 places" is excluded from local_alloc outright (local-alloc.c:471 reg_n_deaths != 1) → it becomes a GLOBAL allocno, conflicts with hard reg 9, lands in $a3. The target's $t1 is combine_regs' hard-reg branch (local-alloc.c:1806–1817): $9 dying as an INPUT of the insn records qty_phys_sugg for the pseudo SET there, tried first in block_alloc (:1470–1477) — but it wins only for a pseudo BORN at that and (wipe_dead_reg clears 9 before reg_is_set births the dest). So prim 4 needs its own single-death pseudo.
  2. Why the split regressed: sched.c adjust_priority (2511–2545) boosts a ready birthing_insn_p SET (dest live, reg_n_sets == 1) to max_priority in the BACKWARD list scheduler — prim 3's fresh and gets glued to its or, stops filling the load-delay gap after prim 3's second lhu, and the ori $s1,0x97 / lui $s1 / ori $s4,0x96 cascade into the gaps: the old verdict "a fresh pseudo displaces the hoisted constants" was reg_n_sets (sched1), not pseudo count (regalloc).
  3. The zero-byte fix: prims 3/4 use different variables (pm3/pm2, single-death each) and pm3 gets a trailing __asm__ volatile("" : "=r"(pm3)); after its last store — emits nothing, but its REG_UNUSED second set makes pm3 reg_n_sets = 2 (no birthing boost, gap kept) and 2-death (global → $a3, as before), while prim 4's pm2 stays single-set/single-death → Register 82 in 9. Measured: fresh pm3+pm2 31; all four fresh 54; inline (p-0x18)&m24 175/159; an asm re-setting pm2 → $a1 (a volatile asm re-lives all hard regs over the extended qty); asm on pm3 only → MATCH. Map corollaries (regalloc.md / sched.md): reg_n_deaths gates local vs global; reg_n_sets == 1 gates the birthing boost; a __asm__ volatile("" : "=r"(x)) on a variable is a dial for BOTH counters at zero bytes. Process note (coordinator, honest): two ledger commits (779b55a49, 3a9c946d8) claimed this bank before it existed — the bank helper had been called without the function name (it built the unchanged tree and exited 0) and then with the wrong draft directory (it refused). Write the commit message FROM the tool's output, never before it; the helper now refuses an empty function list and propagates a failed commit (R43), and takes DRAFT_DIR.

§501-D — ONE VALUE-RETURNING CALL MAKES THE CALL ITSELF A "BIRTHING" INSN: reg_n_sets[$v0] == 1 hands the call sched1's max priority and drags a delay-slot filler above it (P32 T4b, ov_SC06_022:func_8017DF28, the addiu $s2,$sp,0x10 jal-slot vs bnez-slot wall banked by a Fable agent — 119/119). Pinned since S71 as expand_block_move's copy_addr_to_reg pseudo being cse-reused for both later &mtx args; that reuse is REAL and present in the original too (cse follows the bnez as taken, cse.c:8118) — the prior levers (§H diamond, §194-K, §153, copy placement) aimed at the wrong pass. Where the def lands is sched1's decision: sched.c:2469 birthing_insn_p returns reg_n_sets[i] == 1 for a SET whose dest is live, and it is evaluated for HARD registers too — the call's (set (reg:SI 2 v0) (call …)) inside the PARALLEL. regclass.c:1791 counts reg_n_sets for every REG dest. The draft had exactly ONE value-returning call (every other callee declared void), so that call was "birthing" (adjust_priority → 0x7f000001), tied with the address def, and rank_for_schedule's LUID tie-break (:2428) emitted the def ABOVE the jal in the backward schedule; reorg's backward search then filled the jal's slot with it (reorg.c:2904, delayed effects off). The zero-byte lever: a SECOND $v0 set — call RotMatrixY through its real libgte pointer-returning type, ((void *(*)(s32, void *))RotMatrixY)(angle, mptr);, TU void declaration untouched (cast at the use). The call drops to priority 1, the def lands right AFTER the call, reorg's FORWARD search refuses it for the jal slot (mark_set_resources reorg.c:542 marks every call-used reg incl. $sp as set — no !fixed_regs filter) and the bnez's backward search takes it = the target. Known-true control: the banked same-TU twin func_80180700 (3 $v0 sets) shows the identical after-call placement in its own dumps. Tells: (1) a compiler-minted def parked in a jal slot where the target parks it in the next branch's slot; (2) exactly one non-void call in the function. Dial: any callee's true return type (libgte/libgpu functions return pointers/ints even when the TU declares them void), or a value-returning call whose result is discarded. Map corollary (sched.md): birthing_insn_p counts hard-reg sets; the number of value-returning calls in a function is a scheduling input.

§501-E — A REGISTER PIN FORBIDS THAT REGISTER TO EVERY RETRIED ALLOCNO: regs_ever_live seeds reload's bad_spill_regs, so a register … __asm__("$6") on one variable can be the reason another value cannot take $a2 (P32 T4b, main:func_80020DA4, the mflo $t0 vs $a2 REGALLOC-PERM wall banked by a Fable agent — 100/100). Read in the dumps and the source: the products of mult are GLOBAL allocnos (mulsi3_internal constraint =l, mips.md:848; -dl shows "pref LO_REG", none in local-alloc's list). global.c parks m3/m8/m13 in LO; reload spills LO ("Spilling reg 65") and retries them through retry_global_alloc (reload1.c:3497) with losers = forbidden_regs, and forbidden_regs is seeded from bad_spill_regs = regs_explicitly_used = regs_ever_live at reload entry (reload1.c:486, 3651–3660, 709). The S76 pin register s32 e0 __asm__("$6") made $a2 ever-live → forbidden at m13's retry → first-fit gave $t0. Fix: unpin, and reproduce the target's LOCAL allocation with two zero-byte launders steering qty_compare (local-alloc.c:1579): __asm__("" : "=r"(e1) : "0"(e1)) immediately before dst[6] = -e1 (e1's quantity 6666 → 8750, allocated before e0's 7894, holds $v1, so e0's first fit drops to $a2; placed after e1's LOAD instead it lands inside lo1's range and flips a 2500 tie → 15), plus __asm__("" : "=r"(p1) : "0"(p1)) between p1's andi and sll to undo the global-allocno tie the first launder created (the $t6/$t7/$t8 rotation). Ladder: 2 (pinned) → 37 (unpinned) → 15 → 6 → MATCH; laundering addr1 after its addu is a byte-identical alternate; a launder on the SOURCE of the copy emits a move (101 ins) — launder the destination variable. Law: when a pinned draft sits at a 1–2-row register permutation that every pin-set fails to move, the pin may be the wall: pins forbid their register to every retried allocno. Remove the pin, read -dl for the local quantities' priorities, and steer with launders (birth/death dials) instead. Corollary for the map (regalloc.md): retry_global_alloc + bad_spill_regs; qty_compare = floor_log2(n_refs)·n_refs·size/(death−birth).

§501-F — WHEN loop.c PROVES THE HOIST SPLIT UNREACHABLE, SPLIT THE CSE QUANTITY INSTEAD: a hard-register copy of the dividend makes the divide's multiply non-movable while the divide's own sign correction is CSE'd into a hoisted variable (P32 T4b, ov_SC03_105:func_801834A4, the "hoist only the sra" wall banked by a Fable agent — 106/106). The S71 proof stands and is now dump-confirmed: expand_divmod (expmed.c:3034–3057) emits K=const, B=smulsi3_highpart, C=op0 >> 31, D=B − C adjacently from ONE op0, so loop.c gives B and C the same invariant_p verdict; if B is movable, force_movables (loop.c:1193–1228) links K to B and DOUBLES its savings (174 >= 37 → everything hoists, the 75 baseline), and a hard-register dividend kills B, C and D alike. From one division of one register the target's split (sra hoisted, K/B/D inline) is unreachable — so the lever is in cse, not loop.c. canon_reg never rewrites a hard register, but canon_hash/exp_equiv_p compare by reg_qty: write sign = half >> 31; at the TOP of the loop body (a movable by criterion (1), maybe_never == 0), copy the dividend into a hard register register s32 hh __asm__("$2"); hh = half;, and divide THAT: pos[0] -= hh / 3;. B reads (reg 2) (call-used hard reg → invariant_p == 0 → not movable → K unlinked, life 1 → inline; D not movable), while C = (ashiftrt (reg 2) 31) hashes into sign's quantity → replaced by (reg sign) → D becomes B − sign; sign's store hoists (life 71); combine folds the hh = half copy into the mult (can_combine_p allows a hard i2dest with REG_DEAD in i3) → zero extra insns. Two traps, both measured: (1) pin hh to $2, not a callee-saved reg — a callee-saved pin enters regs_ever_live before combine deletes the copy and global.c pass 0 hands the already-live reg to the first callee-saved allocno (closeness 7); (2) jump.c:548–556 deletes an UNREAD pseudo store (regno_first_uid == regno_last_note_uid) before cse sees it — keeping sign alive with an __asm__("" :: "r"(sign)) feed adds +2 loop-weighted refs and swaps half/sign in allocno_compare (closeness 4); a dead initializer s32 sign = 0; gives the store a second reference that jump/cse respect and flow deletes, uncounted → the target's $s1..$s7 order. Law: a division by a constant is four insns from one rtx; to hoist a subset, give the subset a different QUANTITY (a hard-reg copy for the part that must stay, a named variable for the part that must move).

§501-G — THREE PASSES, THREE DIALS: the 27-row "alias basin" was sched1's birthing boost + a phantom USE filling a delay slot only in sched2's model + a local-alloc priority tie (P32 T4b, md_MAIN_003:func_800CF3E8, 469/469, banked by a Fable agent after two Opus passes and ~16k compiles had located the window but not the passes). Read from -dS/-dR/-dl:

  1. sched1 — the tag load was birthing-boosted (adjust_priority sched.c:2507 → birthing_insn_p :2469, reg_n_sets == 1, dest live, priority 7f000001): a boosted load sinks to just before its consumer and NO source position moves it while the boost is alive (why the 14-position birth sweep and the 32-position hoist sweep were inert). Dial: a second LIVE set of the loaded pseudo — tag6 = *(u32 *)p6; … tag6 &= 0xFF000000; (the compound reuses the variable's pseudo as the and's dest → reg_n_sets = 2, zero bytes, and the variable stays single-death/local). The __asm__ volatile("" : "=r"(x)) dial of §501-C measured 39 here — it makes the tag a 2-death GLOBAL allocno; choose the dial by what the register must remain.
  2. sched2 decides the final slot. Post-reload the load (mem:SI (reg 3)) and the field stores (mem/s (plus (reg 3) N)) are disambiguated by memrefs_conflict_p, and the unit-hazard rule "a load is blocked one cycle after a store" (-dR: blocking insn … for 1 cycles) walks the load upward past every consecutive store until it loses a LUID tie — which (1) fixes. Two byte-verified sched2 facts: (a) the S83 "+1 nop" was the phantom __asm__("" :: "r"(ot)) USE — in sched2 it ties and tag at priority 7, wins on LUID, and is picked into the load-delay slot in the MODEL only, so gas emits a real nop: remove zero-byte USE phantoms whenever a nop appears next to them (re-adding it: 70 @ 470); (b) lui m24 reaches slot 378 only through its $a2 anti-dependence, so the __asm__("") fence after the tpage store had to go (everything after a traditional asm depends on it; re-added: 8), while the fence BEFORE p6's birth stays (removed: 461 @ 471).
  3. local-alloc — m24 must out-rank mhi. qty_compare (local-alloc.c:1579) uses POST-sched1 birth/death and FLOW's n_refs (flow.c:2067/2315/2501/2711, written BEFORE combine; combine.c:56 never adjusts them). With the hand-written (x & 0xFF000000) | (y & 0xFFFFFF) both masks had 13 refs and mhi's shorter range won $a2. Dial: spell the OT link as libgpu's P_TAG bitfield setaddr(p, getaddr(ot)); setaddr(ot, p) — store_fixed_bit_field re-masks the already-masked value with 0xFFFFFF (must_and), an and cse cannot fold and combine deletes later, but flow has ALREADY counted it: m24 13 → 18 refs, priority ~doubles → allocated before mhi → $a2; mhi → $t0; tag → $t2 (74 → 2); the last two rows were p5's x0/y0 in natural source order. The bitfield store is therefore a REFERENCE-COUNT dial, not only a §364 shape. Ablations (do not re-try): non-compound tag 89; phantom re-added 70 @ 470; fence re-added 8; manual masks on all six prims 74; the S83 pointer launder and the tag read's source position are non-load-bearing (byte-identical alternates v_C1/v_X4). Laws for the map (sched.md / regalloc.md): sched1 boost → position-insensitive sink; sched2 hazard walk + LUID; flow n_refs pre-combine; a zero-byte USE is a scheduling object with a real delay-slot cost.

§501-H — LARGE CONSTANTS ARE UNBOOSTED FLOATERS; the "prologue weave" is decided by which ready-list STALLS eat them (P32 T4b, md_MAIN_009:func_800CD92C, the 15-row §S7 weave banked by a Fable agent — 247/247, ZERO register pins). Two mechanisms, both in sched1/local-alloc, both fixed by spelling the addPrim the way libgpu's P_TAG macros expand:

  1. m24's register is a REF-COUNT effect, not a pin. qty_compare uses FLOW's reg_n_refs (pre-combine; combine.c:56 never adjusts). la (13 refs, life 342) beats an unpinned m24 (13 refs, life 420). The bitfield store setaddr(p, getaddr(ot)) re-masks its already-masked value (store_fixed_bit_field must_and, expmed.c:608–620): pre-combine (and (and ot m24) m24), folded by combine's associative rule (combine.c:3140–3170) to ONE and — but m24 keeps 19 refs (4·19 = 76 > la's 39) → $9. Spelled (ot & m24) & m24 with a plain u32 m24 = 0xFFFFFF. (Same dial as §501-G(3).)
  2. tp8D/tp8F float to the top only when the tag load is a MULTI-SET pseudo. sched1 try_splits every insn before scheduling (sched.c:4826; mips.md large_int → lui + ori) and update_n_sets bumps reg_n_sets to 2 (sched.c:4617/4234), so EVERY 0xE10000xx constant is an unboosted priority-1 floater, pinned or not. The backward list scheduler consumes a floater only in an EMPTY ready-list cycle, and each RMW chain has exactly one (the lhu → sll latency gap): prim k−1's two gaps eat prim k's constant. The bitfield RMW shape t = *p; t &= 0xFF000000; t |= v; *p = t; sets ONE pseudo three times → no birthing boost (sched.c:2490) → the tag load is not glued to its and and fills the lhu → sll gap itself → the tpage constants float to the top, and local-alloc's life order gives $16..$19 in the target's order. Sub-levers, each one probe: v must be a FRESH expression (v &= m24 in place is a 4-ref/2-insn qty that steals $2 — the 178-row $v0/$v1 swap); the OT read v must precede t &= 0xFF000000 (the constant is force_reg'd where its and is expanded — a lower UID than the index sll drops it into prim 1's gap and the tag load lingers into the store stream, handing tp8F's ori the wrong LUID = the closeness-2 li/ori swap); D_800BAE22 as a plain scalar (struct/array/cast spellings force one shared la); the tag store non-/s *(u32 *)p. TU spelling: extern u8 *D_800A71D0 (the u32 spelling CC1-FAILs since the sibling bank). Law (the whole prologue-weave class, three functions today): a constant's position in the prologue is not steerable by its source position or by a hard-reg pin — it is decided by which latency stalls exist above it; change the stalls (a multi-set load fills them) and the constants move. Read -dS's ready-list traces for the T-nn empty cycles.

§501-I — MANUFACTURING §172 PRODUCER-2 ORPHANS ON PURPOSE: read the s16 from MEMORY at each use, and the shared-scratch store dependence (P32 T4b, ov_SC07_002:func_8017DC80, the historic −33 LENGTH wall's last 46 rows banked by a Fable agent — 346/346). Three mechanisms, each one hunk: (A) Frame 0x70 = two combine USE-orphans, MANUFACTURED. §172 v2 called an extra orphan "not reachable by re-spelling the same computation". The species that IS reachable: an s16 field READ FROM MEMORY at each use — here the mode halfword, once in the branch (& 0xC000) and once arm-duplicated (& 0xFFF). extendhisi2 (mips.md:2340) force_not_mems the load into a HI reg + the sll/sra pair; combine.c:1664 added_sets_1 keeps the HI reg alive across the 3-way merge, so the ashift temp is orphaned by distribute_notes (~10835) into a (use) → alter_reg 8-byte slot per clamp block. The same spelling yields the target's addu $v1,$v0,$zero copy (46 → 24). Corollary to §172: the orphan rule "the HImode load's reg carries an extra HImode use" is satisfied by a SECOND MEMORY READ of the same halfword in the same block — no need to synthesize a ?: copy. (B) DR_TPAGE: one scratch variable carries the len byte AND the tpage word — tp = 1; q[3] = tp; … tp = tpage; *(u32 *)(q+4) = tp;. schedule_select's hazard rule always picks a ready store over ALU insns, so the sb must be UNREADY: the shared variable gives it an anti-dependence on the or; tp becomes 2-set (no birthing boost) and both share $v1. n must be UNPINNED (a hard-reg n dying at sll n,5 hands $s0 to the temp via combine_regs) (24 → 16; the same idiom is banked at src/800.c:3826). (C) la $a0 first: sched2 places the lowest-LUID body insn first, so arg = &D_800AF648; must be the FIRST statement AND carry a zero-byte __asm__("" : "=r"(arg) : "0"(arg)) — otherwise combine substitutes cse's REG_EQUAL constant into the call-site copy and deletes the first set (16 → MATCH). Inert (do not re-try): the P_TAG-bitfield setlen/addPrim here (76, 345 ins — the bitfield is the dial for the §351/§364 SPRT family, not for this DR_TPAGE shape), q/qt/ov pin ablations, DR_TPAGE statement re-orders without the shared scratch, an unlaundered or $4-pinned arg, call-first source order (33).

§501-J — A PROVED TREE WALL IS NOT AN RTL WALL: build the nested MULT at RTL level with the EXPAND_SUM distributive law (P32 T4b, main:func_80011380, boot −O0 — §474's "PROVED C-level floor" banked by a Fable agent, 192/192). §474 proved from the source that no TREE can carry MULT(MULT(i,2),2) (fold-const.c:882 split_tree merges it) and that both escapes (a statement-expression's BLOCK_END note, a register decl's (use) brackets) cost a suid that stupid.c's born+2 rule turns into a deleted copy or a rotated colouring. All true — and beside the point: the nested MULT can be built AFTER fold, by the expander. Index spelled (D_80074784 * 2 + 1) * 2 - 2 (== i * 4): fold has no MULT-over-PLUS distribution and split_tree only decomposes MULT/PLUS/MINUS, so MULT(PLUS(MULT(i,2),1),2) survives; expr.c:5368 (MULT_EXPR, EXPAND_SUM, ptr_mode, constant op1) expands the inner i*2 to the rtx (mult X 2), + 1 gives (plus (mult X 2) 1) (both_summands, expr.c:5248), and the outer * 2 returns (plus (mult (mult X 2) 2) 2) — a nested MULT rtx fold never sees. The source's - 2 cancels the +2: the ARRAY_REF folds to PLUS(PLUS(ADDR, −2), idx), plus_constant makes (const (plus sym −2)), and both_summands folds sym − 2 + 2 to the bare symbol. memory_address rejects the PLUS-with-MULT and calls force_operand, which expands the two multiplies back to back (expand_mult alg_m copy_to_mode_reg expmed.c:2227 + alg_shift :2244, twice) with no note and no variable; every chain pseudo is length-2/2-ref, so stupid.c's born+2 adjacency 2-colours the $v1/$a0 ping-pong exactly as the target, and expand_binop's late copy_to_mode_reg(sym) gives la $a0 / addu / lbu 0() — the bare symbol (no addend survives to the asm). Inert, from source: a hard-reg pin + shift-outer (t = i*2) << 1 → 191 LENGTH-DRIFT (a REG index makes (plus sym reg) a legitimate MIPS address, so la/addu vanish — the target's la;addu;lbu 0() shape REQUIRES the index to reach memory_address as a MULT rtx; every << spelling is dead), COMPOUND_EXPR shield (fold-const.c:3335 distributes), SAVE_EXPR via ?: (expr.c:4338 + function.c:5327), COND shields, builtin pseudo-constants (li), pin+MULT (copy_to_mode_reg($4) first-fits $v1). Law: at −O0 the expander is a second algebra with its own distributive law and constant folding (EXPAND_SUM, both_summands, plus_constant); a tree-level proof of unreachability must also close the EXPAND_SUM route before it is a proof.

§501-K — THE "sched1 CLASS TIE" WAS sched2's /s-STORE EXEMPTION, AND A HAND PAD CAN HIDE A PSEUDO'S OWN SLOT (P32 T4b, md_MAIN_007:func_800CF6D0, the 137-row plateau banked by a Fable agent — 249/249, zero pins; ladder 137 → 30 → 26 → 10 → 0). (1) The QI-before/HI-after store grouping in five of six blocks is SCHED2 (-dR): sched.c:838–845 true_dependence exempts /s varying NON-QImode stores from a fixed non-/s read (the lhu D_800B9A02 index), so the sh/sw were ready a clock early and won schedule_select's potential-hazard rule (store > load > ALU), blocking the loads a cycle. CAST field stores *(T *)(p + off) (the banked twin func_800CD92C's shape, §501-H) restore the dependence and pure LUID order. (2) The OT write must STAY a /s P_TAG bitfield so D_800A5E60 = p floats above block 6's OT write (the same exemption, as an output dependence). (3) The $t2/$t3 la-vs-0xFF000000 swap is a local-alloc qty_compare tie decided by FLOW's reg_n_refs: assign ob = D_800AA60C BEFORE block 1 and use it in block 1's tag-side read too (13 refs, like the twin's compiler-made la; combine then folds the symbol back into read 1 as a 3-insn merge with the la def re-emitted as newi2pat, giving the raw-symbol lui $at/addu/lw form) — assigned after, 12 refs, loses by 1.2 %. Side effect that closed the frame: the folded sum pseudo keeps reg_n_refs = 2 (combine.c:2313/2336 zero only a deleted i2's) and reload1.c:2309 alter_reg gives it a 4-byte slot — that slot IS the target's 0x18 frame, so the T3 u32 pad[2] (§358) had to go (0x20 with it). Plus the twin's multi-set tag RMW, v before t &= 0xFF000000, (ot & m24) & m24, chained rgb stores, the TU prototype (s32, u32). Inert/superseded: pad[2] with ob-first; array-indexed OT on the u8 symbol without ob (254 ins); a cast OT write (the tail stops floating). Law: a frame that is "8 bytes short" may be one pseudo's own slot, not a pad — read .greg for the stack-slot assignment before declaring a pad.

§501-L — WHEN TWO DIALS SHARE A SLOT, THEY COUPLE: the floater cure and the qty_compare contest in md_MAIN_007:func_800CF408 (P32 T4b, 49 → 3 at exact length, zero pins, NEAR — the one §351-family row not closed). Closers were §501-H's shape verbatim plus two additions worth keeping: a dead arg1 = 0; kill after y1 = arg1 - 0x78 so cse's fold_rtx cannot re-associate y2 = y1 + 0x100 into a1copy + 0x88 (that fold made y a 2-death GLOBAL allocno → $t8; 48 → 22), and a NAMED u32 mhi = 0xFF000000 born before block 1's RMW so its boosted li takes the lhu → sll gap one slot earlier and lengthens mhi's life by one, letting ob win the $t2/$t3 qty_compare contest by 11 (refs 9 vs 8, lives 113 vs 100: 2389 < 2400; 18 → 3). The residual (idx 10–12) is §501-H's floater mechanism: the unboosted tag load lingers up block 1's store stream, blocks one cycle behind the tpage sw, and the empty cycle eats the highest-LUID floater (ori $s5,0x96). Its cure — mhi's li UID above the index sll — moves mhi's birth one slot later and flips the contest back (2411 > 2400 → 18); moving ob after the index fixes the contest but opens a bubble in the OT-chain lhu gap that eats the same floater (13). The target fills that gap with and $a3,$v0,$t1 = an UNBOOSTED p & m24; a 2-set a3 is re-merged by combine with reg_n_sets-- (combine.c:2309; 36) and an asm launder there becomes a sched2 delay-slot phantom nop (179 ins). Law: when the same ready-list slot decides both a scheduling floater and an allocation contest, one-dial probes oscillate between two closeness floors (here 3 and 13/18); decouple by adding a filler that changes neither count (an unboosted temp combine cannot re-merge — a different mode/width, or a volatile temp) or by moving the contest margin with a reference in a block that does not touch the slot. 135-variant sweep floor 3 (×24). Next lever recorded in docs/backlog.md.

§501-M — THE PHANTOM-SLOT PRODUCER CENSUS, and the ghost that cannot slot (P32 T4b hand pass, S84 2026-09-06; main:func_80032A74 PROVED at 1). A never-referenced stack slot that sits AFTER the parameter spill slots (sp+0x48 here; the params at 0x30/0x38/0x40 are alter_reg slots in regno order, so any expand-time local would displace them) can only come from four reload-time sites, all read out of tools/reference/gcc-2.7.2: (1) reload1.c:658 alter_reg(i,-1) for a GHOST pseudo — reg_n_refs>0, no occurrence, no REG_EQUIV, class ST_REGS or none because regclass never saw it; (2) caller-save.c:249 setup_save_areas — a 4-byte area per call-used hard reg holding ANY pseudo with reg_n_calls_crossed>0, once caller_save_needed is set by the profitability retry (global.c:1085, local-alloc.c:2209; 4*calls < refs); it leaves no sw/lw only when the count is STALE-HIGH, and the sole staleness route is sched.c:4962 (a multi-block pseudo keeps flow's count when sched's is 0 — the comment says why) after sched1 moved a register-only def/use across a call inside the call's own block (combine never crosses a call except with a constant source, combine.c:924; update_equiv_regs moves nothing); (3) reload1.c:879 — a reg_equiv_memory_loc whose address eliminates to a SPILLED pseudo gets a fresh slot, but only an UNALLOCATED pseudo qualifies and those equivalences are single-block (update_equiv_regs), so local-alloc takes them; (4) reload1.c:3499 spill_stack_slot — a pseudo evicted from a spilled hard reg with no retry (local-alloc'd) or a failed retry_global_alloc; $t0 can hold no pseudo at all (order_regs_for_reload lists zero-use call-used regs first, so a pseudo in $t0 moves every param reload to $t1), and LO-pref mult results carry alternate class GR_REGS and re-home. The one producer reachable from C at zero code cost is combine's newi2pat split (combine.c:1887 SIGN_EXTEND-of-narrow- load, combine.c:1963 two-independent-SETs) whose i2dest vanishes with reg_n_refs kept (the zeroing at combine.c:2306 is skipped whenever newi2pat != 0); both re-derive a NARROW LOAD (lh/lb, or a duplicate lhu) from the chain's memory head — a register head folds at tree/cse level or has its middle temp re-used by find_split_point, so path (b) never runs (18 reproducers). Hence a phantom slot whose site loads lhu, has no lb and no double load is unreachable: PROVED. The NEW ghost producer, measured, and why it does not slot: local-alloc.c optimize_reg_copy_2 on tmp = x; <use tmp>; tmp = tmp op c; <use tmp>; x = tmp; (one block; x dead at the head copy, live after the copy-back; the head copy survives combine when tmp's FIRST use is not its last and no 3-insn chain passes through it — combine.c:904; the copy-back survives when tmp has an intervening use that sched keeps above it) rewrites every tmp to x, leaves two no-op self-moves, and decrements reg_n_refs[tmp] once per insn while flow counted the in-place insn twice → a ghost with stale refs (P13 refs 5, P14 refs 1). It is minted AFTER regclass, keeps GR_REGS with no conflicts, and global simply allocates it: vars=0. Only pre-regclass (combine) ghosts take slots. Instrument: tools/ghost_census.py <tag>.i.lreg (headers with no occurrence, class → SLOT / allocatable), now run by tools/cc1_dumps.sh in place of its (use) grep (which under-counted, §172 note); vars= on the .frame line remains the arbiter. Law: before probing spellings for a frame residual, enumerate the artefact's PRODUCERS from the source and refute each on the bytes — the site's load width (lh vs lhu), the call blocks' contents, the spill register's identity and the mult results' alternate class each kill one producer without a compile. Probes and notes: .run/P32/t4c/func_80032A74/.

§501-N — A BANKED SIBLING'S SPELLING BEATS THE DRAFT'S DIALS; the u32 array view of a u8-declared symbol is the fleet's asm-label alias (P32 T4b hand pass, S84 2026-09-06; md_MAIN_007:func_800CF408 178/178 BANKED 8fa12bc22, zero pins, zero asm bodies). The 3-row prologue-weave residual (§501-L: the T-139 memory-unit bubble handing ori 0x96 the wrong LUID, coupled to the $t2/$t3 qty_compare contest) was a SHAPE symptom: the draft carried a named mhi, an ob base pointer, y1/y2 temps and an arg1 = 0 kill, each a dial against the previous dial. The same-family sibling md_MAIN_009:func_800CD92C (§501-H) had banked with the plain libgpu shape — v = (OT[idx * 0x1000] & m24) & m24; t = *(u32 *)p; t &= 0xFF000000; t |= v; *(u32 *)p = t; then OT[oi] = (OT[oi] & 0xFF000000) | ((u32)p & m24); on a TRUE u32 ARRAY_REF, x -= 0xA0; y -= 0x78; in place and x + 0xA0 / y + 0x100 inline, colours chained *(p+8) = *(p+9) = *(p+0xA) = c — and porting it with this function's constants matched first try. Why the ARRAY_REF matters: with element size 4 the address (plus (mult idx 4) sym) goes through memory_address → force_reg (sym), so cse keeps prim 1's read in the gas lui $at/addu $at/lw %lo(sym)($at) form and binds the base into $t2 for every later access (the target's shape); a u8-array spelling (*(u32 *)&sym[oi], a one-member struct view, a P_TAG bitfield on &sym[oi]) has (plus sym idx), never binds, and emits the macro form eight times (measured). When the TU declares the symbol with another type, take the fleet's alias: extern u32 wD_800AA60C[] __asm__("D_800AA60C"); (1,438 banked files carry extern T name __asm__("D_…") views; a declaration, not an asm body — R62 untouched). Also measured: the sibling's lever-7 trailing __asm__ volatile("" : "=r"(tmp)) 2-set dial is NOT portable here — the volatile asm makes hard regs live at its position and flips the m24/colour $t1/$t0 order. Law: when a same-family sibling is banked, port its spelling with the row's constants BEFORE touching a dial on the draft; the residual class name (§501-H) is the family's signature, not a lever list. The target's vars= 8 here is a combine-minted ghost (§501-M species, ghost_census.py: ST_REGS or none → SLOT) reproduced by the port for free. Notes and every variant: .run/P32/t4c/func_800CF408/.

§501-O — A PHANTOM SLOT BETWEEN A PARAMETER'S SPILL AND AN EVICTED PSEUDO'S SLOT: the census for a LEAF, and what it leaves open (P32 T4b hand pass, S84 2026-09-06; main:func_80039308 518 ins, PLATEAU at 4). Target frame [s16 arg1 spill @0x0][8 bytes, no traffic @0x8][cnt @0x10]; the natural spelling (*(s16 *)(p + 6) = arg1, Fable's Y4) is register-exact — the HImode parameter pseudo is spilled by global (sh $a1,0($sp) at entry, the head promotions keep reading $a1 because reload's find_equiv_reg still finds the value there, and the late use reloads lhu $s7 = the spill register) — and lands [arg1 @0][cnt @8]. Read from .greg: cnt is global-allocated to $s7, order_regs_for_reload picks $s7 as the GR spill register (the least-used GPR in a function that uses all 24), spill_hard_reg evicts it (its retry fails), and alter_reg(cnt, 23) mints spill_stack_slot[23] AFTER the initial alter_reg loop. So the phantom must be an INITIAL-LOOP slot (any regno above the parameter's) with no traffic, or a main-loop slot before the first spill. Refuted here, each on a dump fact: caller-save area (leaf: every reg_n_calls_crossed is 0); a LO-evicted product's spill_stack_slot[65] (GR_REGS is spilled before LO_REG, so it would follow cnt's slot — and a product's alternate class is GR_REGS: two overlapping products allocate to $a1/$v1, C3); reload1.c:879 (needs an unallocated single-block equivalence pseudo; local-alloc allocates every short temp — $t0 is free at the k2 site because s18 dies one insn earlier); expand-time locals (they precede every reload slot: Y1/Y3/Y4 measured the parameter at 0x10/0x8); global's local-alloc kick-out (its victim's traffic shows unless def/use are adjacent through $s7, and the only such pairs are the six LO-class mflo $s7 products); a combine ghost (all thirteen lh are single-use — bne/mult/sll/addiu/slti consumers —, no lb, and a HImode pan would leave real sll/sra in the arms the target does not have). What this row teaches: (1) tools/ghost_census.py + the .greg "Spilling reg N / now on stack" lines give the slot ORDER for free — read them before any frame probe; (2) the six mflo $s7 / op $s7 pairs are byte-identical whether the product is LO-homed with an input reload or slot-homed with a deleted output reload (§501-M's inheritance species), so a phantom whose traffic-free pseudo is a product cannot be excluded by the bytes alone — only by allocation (alternate class GR_REGS always saves a product); (3) rows 49/50 are two move_movables hoists in BODY order — the original computed vol = b2 * 0x100 in the loop; steering the hoisted invariant into $s2 without the $18 pin (X2 = 495) is the open lever (§501-E launders).

§501-P — THE ATLAS'S "WEAK COUSIN" WAS THE SAME-SHAPE SIBLING; a ported natural spelling needs NO pins and NO fence, and the census shows WHICH spelling elements carry a scheduling window (P32 T4b hand pass, S85 2026-09-06; ov_SC03_105:func_80185810 489/489 BANKED cdd9a2cb8 — the phase's last overlay stub; zero pins, zero fences, zero asm bodies). The S83 Fable draft sat at DIFF 13 (rows 363–380, the tpage/code RMW window) after 3,360 region-2 variants, five register pins, HI temps and a zero-byte fence; its report had correctly read the residual's mechanism (an unboosted 2nd set cl &= 0xFFFF whose anti-dependences hold the w-chain's reads; the fence releases them but forbids sched2's fillers) and concluded the honest fix was blocked by combine. The §501-N step 0 found the answer in 25 minutes: the atlas lists ov_SC02_027:func_80180B3C as a 0.55 knn cousin — a "weak" score — yet a shape grep on the idiom's constants (grep -rln '0x200) << 2' src = the libgpu getTPage bit chain) shows it is the SAME billboard-sprite drawer (POLY_FT4 off the D_800A5E60 bump, code = 0x2C; code |= 2, the tpage chain, code |= (w & 0x40) >> 6, the u/v loads, the v0 conditional, the P_TAG link), and its COMPILED window (objdump -d build/src/ov_SC02_027/ ov_SC02_027_jr_8017D898.o 0x34f8–0x3560) is instruction-for-instruction the target's rows 362–386 with only the base registers renamed. Porting its window spelling onto the draft matched first try WITH the cousin's three pins, and then with none. The element census (12 real-TU variants, .run/P32/t4d/NOTES.md): (a) the v mask is a FRESH single-set copy vm = cl & 0xFFFF whose consumers are the two arms of the v0 conditional — a lone insn is never simplified by combine and the lhu setting cl has intermediate uses, so the andi survives WITHOUT a hard-register pin (d5: the $7 pin removed, MATCH — the S83 "hard reg hides nonzero_bits" guess is refuted as the mechanism); being single-set it is birthing-boosted, so nothing is starved and no fence is needed; (b) u = (uu - ((tpage & 0xF) << 6)) << shift as a fresh single-set value — the 2-set uu -= …; uu <<= … form re-rolls 63 rows, moving allocations 100+ rows away (d12); (c) the branch polarity if (!(tpage & 0x10)) vv = vm; else vv = vm - 0x100; — the copy arm is the fall-through that coalesces to nothing, leaving beqz → skip; addiu $a3,-0x100; the opposite polarity re-rolls 42 rows (d11), which is what the S83 F2 variant (59) actually measured; (d) shift = 2 - mode born early and unpinned — the boost sinks it to rows 370/372; (e) the OT base computed right after the code byte, before the u/v loads. With every birth in the window single-set, sched1's boost ties them all and the LUID tie-break yields source order, so the S83 window-1 ot16 pin is load-bearing ONLY in the mixed form (d3: DIFF 4 with 2-set neighbours; d9: MATCH with none). Laws: (1) the atlas's knn score is NOT a shape oracle — a 0.55 "weak cousin" can be the exact sibling when the register bases differ; step 0 of every hand pass greps the idiom's CONSTANTS across src/ (0x200) << 2, 0xE1000000, + 0x100) << 6…) and reads the sibling's OBJDUMP window against the target before touching the draft (§501-N, accelerators (13)/(14)); (2) a pin or a fence that a draft "needs" is a property of the draft's other dials — after a sibling port, remove every pin and re-measure before banking (§501-E), and record the census so the next row inherits the elements, not the dials; (3) a residual's mechanism read from the dumps can be right and its "blocked" verdict wrong: the block's OTHER births (2-set uu, the polarity) were what made the honest fix look blocked.

§501-Q — THE SELF-UPDATE GHOST: a no-traffic 8-byte frame slot from a combine bookkeeping gap (P32 T4c, S85 2026-09-06; main:func_80032A74 422/422 BANKED f9a90affb — the S84 "PROVED wall"; also the [arg1 @0][8 @8][cnt @0x10] slot of main:func_80039308). The §501-M producer census missed one producer. In try_combine (combine.c:2306) the deleted i2's dest gets its reg_n_sets/reg_n_refs decremented ONLY if (! added_sets_2 && newi2pat == 0 && ! i2dest_in_i2src): when the deleted insn is a SELF-UPDATE (x = x op c, its dest re-mentioned in its source) the bookkeeping is skipped. So the chain gx = <load>; gx -= 0x100; ... (u32)(u8)gx ... compiles to the plain lbu of the direct spelling — combine first folds the self-update into the (u8) use ((and (plus gx -256) 0xFF) → (and gx 0xFF), i2 deleted without bookkeeping), then merges the load into the rewritten use ((and (zext (mem:QI)) 0xFF) → lbu, i2 deleted with reg_n_sets going from a stale 2 to 1, so reg_n_refs is never zeroed) — and gx survives as a pseudo with references and no insns: regclass never sees it ("ST_REGS or none"), global does not allocate it, and reload's initial alter_reg loop (reload1.c:658) mints an 8-byte slot for it in REGNO order (after the parameter spills; before any spill_stack_slot). Measured (.run/P32/t4e/ghost/, tools/cc1_dumps.sh + ghost_census.py): the host must be a word-width (u8)x/x & 0xFF conversion consumed by an insn (return, add, compare operand) — k4/k5/k7/k14/k16/k17 mint the slot (u32/u16/u8/s16 loads, &= 0xFFF or -= 0x100); byte STORES *q = x / *q = (u8)x + w do not (k8/k11–k13/k15/k18 — the store absorbs the mask or keeps the arithmetic); copy+constant chains do not (g1–g5: cse folds them before flow); a no-op mask into a load-absorbing use does not (h1–h7: the def merges first). Law: a phantom slot with no traffic and an lbu/lhu site where a narrower value is consumed at word width = plant gx = <wider load>; gx -= 0x100; … (u8)gx at that site (zero instructions; the slot position is the ghost's regno order). Add this row to the §501-M census before calling any frame residual proved.

§501-R — A HOISTED INVARIANT READ THREE TIMES: temps, not variables; u16 keeps cse off the copy; the pins were never needed (P32 T4c, S85 2026-09-06; main:func_80039308 518/518 BANKED, the phase's last stub, ZERO pins). The target's preheader [li $s3,2][sll $s2,$s5,8][li $s1,1] reads the volume base b2 << 8 in $s2 at three body sites (the copy in the first branch's delay slot and both pan arms); the S79–S84 drafts pinned vbase to $18 in the preheader and ended two rows short (the hoist order). Measured in the real TU, 56 variants (.run/P32/t4e/NOTES.md): (1) a NAMED variable set after the inner loop's jumps is never a movable (loop.c:695-700 — user var + maybe_never + uses in other blocks), and a movable TEMP with one use and life 1 is "not desirable" (threshold × savings × lifetime < insn_count); (2) copy-first vol = vbase with the arms reading vbase is folded by cse1 onto vol (cse.c make_regs_eqv: the register with the LATER last mention becomes canonical — vol is used after the arms), so the arms read vol and the hoisted value has one use; copy-last keeps the arms on the base but vol is then born after the tests, takes $a0 and reorg cannot lift the copy; (3) THE FIX: write the expression inline three times — vol = b2 * 0x100; if (A) vol = (b2 * 0x100) + X; else if (B) vol = (b2 * 0x100) - Y; — so the three temps are merged by loop.c combine_movables (savings 3, one hoisted sll, the arms read it), and declare the accumulator u16 vol: in this TU it is a HImode pseudo, so vol = <expr> expands as a subreg move that cse never canonicalizes — vol never joins the shift's quantity, the arms keep the temp, and vol is born before the tests (in $a1, conflicting with pan in $a0; the copy lands in the delay slot). The hoisted temp then has 7 weighted refs → priority 470 between the hoisted constants 1 (657) and 2 (452) → $s2 by global's numeric scan (tools/alloc_table.py reads that order straight from the dumps). (4) The else head ($t1/$t0) was the note-on arm's s17/s18 variables REUSED as the release loop's pointer and compare temp. (5) With the structure right, every remaining pin ($18, $4×2, $2) came off byte-identical. Laws: an invariant read N times must be N inline expressions (or a single-block temp), never a named variable set after a jump; before pinning anything, read alloc_table.py's order — the callee-saved bank IS the priority order; a u16/s16 accumulator is a cse firewall, not just a width.