Files
BFM-decomp/cookbook/C0299.md
T

20 KiB

§278 — ADDENDA HARVESTED FROM WAVE cf (P31 S60): 34 candidates, 13 already covered, 8 sharpenings, 4 new laws

Provenance. Wave cf — the first wave drafted after the soft-429 fix, and the largest single-wave harvest yet reviewed. One Sonnet reviewer read all 34 notes against a corpus that had just grown by §274-§277 from the 18-wave batch earlier the same day, checking each ADDENDUM/NEW claim against the banked C in src/ and the target .s before proposing it.

Verification (R14/G3). Two claims did NOT survive byte-checking and are recorded as refuted rather than laundered in: func_800CAE74's note described a "parameter reassignment" idiom that is not in the banked C at all (the real law is §279), and func_800CBCD4's note named the wrong register ($v1 where the bytes use $v0). At merge I re-verified §279 independently: the target carries one slt and zero sltu, and the banked C drives the bound test with a plain s32 cursor beside a separate s32 * — the direction the law predicts.

Three harness suspicions, not idioms (recorded, not acted on): func_800CAFE4, func_8017F39C and func_800CBC14 each describe a byte-correct body that would not bank. All three resolve to documented mechanisms (§42b stale-object gate trap, §238 overlay-homonym trap), but func_8017F39C suspects a binary-registration mixup with a same-named function in ov_SC01_009 — worth a maintenance look at the GATE side, not the drafting side.

ADDENDUM to §74 (func_800D0E30, resident)

DELIBERATELY CO-PINNING TWO NAMED LOCALS TO ONE HARD REGISTER, NOT JUST TOLERATING INCIDENTAL SCRATCH REUSE

§74's "SCRATCH REUSE" bullet documents an incidental case: gcc uses a pinned register as a temporary for an unrelated value before the pinned variable's own value lands there, and calls it "usually benign." func_800D0E30 (resident, 37/37 MATCH) shows the deliberate form of the same fact, usable as an active decoding lever rather than something to merely tolerate: when the target's disassembly shows ONE physical register carrying two unrelated short-lived value-chains at two different points in the function with a genuine gap between them (no instruction anywhere reads chain A's value after chain B's first write), declare TWO separate register T x __asm__("$N") locals — same $N, different C names — one per chain, rather than forcing a single C variable to span the whole gap (which would wrongly extend its live range and risk cse/global-alloc treating it as one value) or leaving the register unpinned (which gives the allocator no reason to reuse the same physical register for both chains at all).

Byte evidence (src/resident/resident.c:1369-1398 vs asm/resident/nonmatchings/resident/func_800D0E30.s): register u32 t3 __asm__("$3") and register u32 h __asm__("$3") are two different C locals sharing register $3. t3 is defined and consumed only inside the first if block (t3=(b+0x3C)&0xFF; b=t2|t3;), dead afterward; h is defined right after that block (h=t2>>8;) and consumed in the second if's condition, entirely inside t3's dead gap; t3 is then redefined fresh in the second if body. The two pinned locals' live ranges are provably disjoint.

Precondition: verify the gap is real (no read of the earlier chain survives past the later chain's first write) before co-pinning — this is not a general "reuse pins freely" law, it is a same-hard-reg case of §72's "preference, not reservation."

Note also: the PARAM-CARRY half of this candidate's note (assign a packed value back into the parameter variable when the target redefines an argument register in place) is already §167-21 verbatim — covered, not repeated here. Its TWO-USE SPLIT claim (hoisting a shift into its own statement to block a fold) describes an untried counterfactual spelling and was not promoted.



ADDENDUM to §172b-4 (func_80185054, ov_SC03_097)

A hand-written (x + (x>>31)) >> 1 on an explicitly-unsigned expression is sometimes REQUIRED, not just "not required" — §172b-4's own boundary, sharpened.

§172b-4's existing addendum shows a plain (s16)w / 2 (SIGNED context) auto-producing srl 31/addu/srl 1 from gcc's own constant-division expansion — no hand-rolling needed. §1-I5 separately states an unsigned divide-by-2 emits a bare srl with no bias at all. Both are true and do not contradict this entry: func_80185054 (ov_SC03_097, MATCH 66/66) needs the FULL 3-instruction bias tail on a value the C code casts to u32, in a context where gcc's automatic lowering (per §1-I5) would never emit it, because unsigned division by a power of two needs no correction.

Why. gcc cannot prove a u32 expression is always < 0x80000000, so a literally hand-written round-toward-zero bias — ((u32)(t*3) + ((u32)(t*3) >> 31)) >> 1 — is not simplified away even though it is a mathematical no-op for every value the expression can realistically take (t*3 is far below 2^31 here). This is dead-in-practice-but-live-to-the-compiler code, faithfully lowered instruction-for-instruction: sll/addu (the *3), srl $r,$r,31 (extract bit 31), addu (apply bias), srl $r,$r,1 (final shift — LOGICAL because the C type of the shifted expression is unsigned).

The discriminator vs the automatic case (§172b-4/§1-I5). If the target's bias tail sits on a value that is genuinely SIGNED in the source's own arithmetic (no (u32) needed to explain it) — the plain cast-division (s16)x / K already produces it; do not hand-roll. If the target's tail sits on a value that is unsigned, OR signed-but-explicitly-cast-to-unsigned for that one sub-expression, and a bare srl (§1-I5's automatic-unsigned form) is too short by 2 instructions — the bias must be spelled out literally, (u32) cast included even when the underlying variable is s32/s16.

Evidence. func_80185054: block 1 (u16 dividend, w declared u32) — w*3 then srl 31/addu/srl 1, all logical, final op srl. Block 2 (dividend is t = (s16)func_8004787C(...), itself signed) — the identical 3-op tail is reached ONLY by writing ((u32)(t*3) + ((u32)(t*3)>>31))>>1; omitting the cast routes through gcc's own signed-division machinery, giving a shorter sra-terminated shape that fails to match. Source: src/ov_SC03_097/ov_SC03_097_jr_80182B9C.c:4084-4085,4093-4094. Target: asm/ov_SC03_097/nonmatchings/ov_SC03_097_jr_80182B9C/func_80185054.s, idx 8018508C-80185098 and 801850FC-80185108.



ADDENDUM to §265 (func_8017DC80, ov_SC07_002)

A card's embedded seed transcription can itself be a stale sibling copy — re-open the real target .s and diff line-by-line before finalizing, even inside this same verbatim-asm lane.

Symptom. A §265-lane card ships a pre-transcribed __asm__ block as its seed (a prior session's attempt, or a copy adapted from a structurally similar sibling function). The seed looks right — same instruction count in the ballpark, same handwritten-asm tells — and it is tempting to trust it since §265's own workflow is "transcribe verbatim, model nothing," which reads as if there is nothing left to verify once a transcription exists. Byte-proven case: func_8017DC80 (ov_SC07_002, 346 ins, banked at src/ov_SC07_002/ov_SC07_002_jr_8017C8D0.c:3298). Its precedent function in the same handler family, func_80185810, is structurally close enough that a seed built by adapting one for the other reads as plausible but is not byte-identical.

Mechanism. §265 already establishes that a raw __asm__ body must be transcribed verbatim with no modelling — but that rule presumes the thing being transcribed is a correct read of this function's .s. Two concrete deltas surfaced only on a fresh read of the true target, both invisible from the seed alone: (1) one conditional arm's tail jumps to a shared label (j .L8017E0D4) instead of duplicating the inline tail the seed had written out (+2 ins if missed); (2) one arm has no nop between its last lbu …,0x29($v0) and the following j, while a structurally similar twin arm (which has two more sb stores after the same lbu 0x29) does carry a nop there — the seed had copied the twin's slot layout onto the wrong arm (+1 ins). Both are exactly the class of instruction §265 says to "leave alone, transcribe it" — but that only holds if the source of the transcription is the real .s, not a look-alike seed.

The law. Before finalizing a §265-lane transcription, read the actual target .s fresh and diff it against the seed's __asm__ block instruction-by-instruction — do not trust a seed just because it already looks like a correct verbatim-asm transcription. This is the same "wrong-body" risk §238 documents for ordinary C drafts, but it survives inside the §265 lane specifically because that lane's own instruction ("transcribe verbatim, model nothing") reads as license to stop checking once a transcription exists.

Verification. Confirmed in the current tree: func_8017DC80 is banked verbatim-asm at the cited line; asm/ov_SC07_002/nonmatchings/ov_SC07_002_jr_8017C8D0/func_8017DC80.s contains both the lbu $v0, 0x29($v0) / nop pair and, later, the j .L8017E0D4 the note describes.



ADDENDUM to §17 (func_800CB794, md_MAIN_036)

THE ASYMMETRIC LOOP-BOUND LEVER, RESOLVING THE §17-NOTED "IV-FINAL-VALUE ADDRESSING" STUB

§17's own residual notes (the func_8012C2D0 passage) flagged a case where a loop's two bound tests address logically the same value with two DIFFERENT instruction shapes — one folded (base + N materialized as a single lui/addiu off the array symbol), one a wholly separate symbol (lui/addiu of its OWN name) — and concluded "no clean C form (pointer-var, struct-array index) reproduces the separate base materialization. Still a genuine residual — stub."

The resolved lever, byte-proven on func_800CB794 (md_MAIN_036, MATCH 0x120/0x120): write the two tests as two DIFFERENT C-level bound expressions, not one reused bound:

u8 *p = D_801202A0;
if (p < D_801202A0 + 0x6480) {          /* entry guard: folds to lui/addiu(D_801202A0) + addiu 0x6480 */
    do {
        ...
        p += 0x10C;
    } while (p < D_80126720);           /* back-edge: a SEPARATE, independently-materialized symbol */
}

The entry test's bound is spelled as <array-base> + <constant> (gcc folds this into one lui/addiu off the base plus a single addiu for the offset). The back-edge test's bound is spelled as a distinct extern symbol name (D_80126720) even though it is numerically the same address as D_801202A0 + 0x6480 — this materializes its OWN independent lui %hi/addiu %lo pair rather than reusing/folding through the base. There is no single bound variable or expression that produces both shapes; the two tests must be written with two different spellings of "the same" address.

Evidence. asm/md_MAIN_036/nonmatchings/md_MAIN_036/func_800CB794.s: entry guard at 800CB7A4-800CB7B0 (lui $s0,%hi(D_801202A0); addiu $s0,%lo; addiu $v0,$s0,0x6480; sltu); back-edge at 800CB884-800CB88C (lui $v0,%hi(D_80126720); addiu $v0,%lo(D_80126720); sltu) — two independent %hi/%lo materializations for numerically-identical addresses. Source: src/md_MAIN_036/md_MAIN_036.c:287-317.

When it applies. A loop whose entry guard and back-edge test bound the SAME logical range but whose target asm shows one bound folded (addiu off an already-loaded base) and the other as its own lui/addiu pair — do not chase a single C bound expression; declare a second extern for the numerically-identical end address and use it only at the back edge.

(The candidate's induction-variable-fusion half — "write NO second pointer, let gcc's strength reduction fuse the givs" — is fully covered already, §30-2/§52-2/§135-17/§145/§36; not repeated here.)



ADDENDUM to §162p (func_8017C120, ov_MAIN_012)

RUNG 1 ALSO COVERS A POINTER-PLUS-OFFSET ADDRESS, NOT ONLY A BARE COPY OR A GLOBAL'S ADDRESS

§162p rung 1's ladder ("use inside a loop, and the loop already hoists something → write the copy IN THE LOOP BODY... move_movables carries it to the preheader, in BODY order") was stated for a plain register-register copy (§48-B extended to dim = flag;). func_8017C120 (ov_MAIN_012, jr_801789AC, MATCH 34/34) shows the same lever fires for a computed address, s16 *q = (s16 *)p + 20;, when declared as a loop-body statement inside a loop that already hoists another invariant (D_80115130[i] array-load chain).

Placement is body-order, exactly as rung 1 predicts. Declaring q's initializer OUTSIDE the loop (as a preheader statement written by hand) schedules it EAGERLY — before the array-load chain that also gets hoisted — landing it in the wrong preheader slot. Declaring it as a loop-body statement lets LICM hoist it, and its position among the OTHER hoisted movables is set by where it sits in the body relative to them, landing it in the target's slot 47 (immediately after the array-load chain's preheader materialization).

Precondition already covered elsewhere (§167-33), noted for completeness: the base pointer must be routed through an actual pointer local (p = &D_80115130; *p = 0;) rather than a bare D_80115130 = 0;, so both the zero-store and the loop's q = p + 20 derive from the SAME register — this is §167-33's own law, not new here.

Evidence. src/ov_MAIN_012/ov_MAIN_012_jr_801789AC.c:5843-5850; target asm/ov_MAIN_012/nonmatchings/ov_MAIN_012_jr_801789AC/func_8017C120.s.

When it applies. A loop-invariant ADDRESS (not just a value) that must land in a specific preheader slot relative to another invariant the loop already hoists — try the loop-body-initializer placement before reaching for statement reordering outside the loop or a register pin.



ADDENDUM to §20 (~L1947) (func_80186AD0, ov_SC06_032)

THE "FILL A LOAD-DELAY SLOT IN A TAIL-STORE SEQUENCE" LEVER ALSO FIRES FROM PLAIN SOURCE-ORDER PLACEMENT, NO MANUAL TEMP NEEDED, AND REACHES PAST A NON-CONSTANT WORD OP

§20's existing entry for this lever requires pulling the deferred field into an explicit named temp one statement early (c = out.c; <constant stores>; *(short*)(p+0xe)=c;) and is worded only for a tail of CONSTANT stores. func_80186AD0 (ov_SC06_032, 40/40 MATCH) shows a narrower-effort variant: simply writing the halfword copy statement *(u16*)(dst+K) = *(u16*)(src+K2); in SOURCE POSITION before an unrelated statement — here a non-constant *(s32*)(dst+4) |= 0x50000000; word RMW, not another constant store — is enough on its own; no manual temp hoist is required. sched1 still hoists the halfword LOAD early (to fill the load-delay slot ahead of the OR's own load) while independently sinking the STORE to the end of the block, past the OR and three later constant stores.

Read as: the entry's law is broader than its own example — the "tail" can contain a non-constant RMW, and the C lever is plain statement placement, not necessarily a named early temp.

Byte evidence. src/ov_SC06_032/ov_SC06_032_jr_80182890.c:5347-5364 vs asm/ov_SC06_032/nonmatchings/ov_SC06_032_jr_80182890/func_80186AD0.s:5E9C4-5E9E8 (lhu $a0,0xE($s1) at the OR's load site; sh $a0,0xC($s0) as the LAST store in the block, after three unrelated constant stores).

(The candidate's other claim — that a register __asm__("$4") pin on the copy source would delete the copy instruction entirely — describes a spelling the final banked code does not use (it uses a plain local s32 a; a = s0;); this could not be checked against the banked bytes and was not promoted. See "could not verify.")



ADDENDUM to §164-56 (func_8017F644, ov_SC04_005)

switch NEVER PRODUCES range_test'S SHAPE; A MIXED TARGET (SOME CASES FOLDED, SOME NOT) MEANS RESPELL AS SEPARATE || GROUPS BY K,K+1 ADJACENCY

Symptom. A target dispatching on a small set of case constants shows range_test's addiu(-K)/sltiu(2) shape (§164-56) for SOME of the values but a plain andi/compare chain for the rest — inside what a switch-shaped seed reads as one statement. switch{case 0x31,0x32,0x33,0x27B: ...} compiles via expand_end_case's dense-compare tree (stmt.c), which never calls fold_truthop/range_test at all: no spelling of a switch can ever reproduce a mixed range_test/plain-compare target, regardless of case values.

Mechanism, byte-confirmed on func_8017F644 (ov_SC04_005). The case set {0x31,0x32,0x33,0x27B} splits into exactly two if-groups by §164-56's own K,K+1-consecutive rule: {0x31,0x32} IS a K,K+1 pair and folds to addiu $v0,$v1,-0x31 ; sltiu $v0,$v0,2 ; bnez — no andi 0xffff mask, because the operand is a fresh lhu result (already zero-extended, no prior arithmetic to lose that guarantee). {0x33,0x27B} is NOT K,K+1-adjacent, so range_test's consecutive-constant test never fires and it stays a plain beq/bne pair — preceded by a defensive andi $v1,$v1,0xFFFF, because the earlier addiu/sltiu on the same register loses the compiler's zero-extension tracking on it. Source: two SEPARATE if statements, each its own || pair — if (x==0x31||x==0x32) return 1; if (x==0x33||x==0x27B) return 1; — banked at src/ov_SC04_005/ov_SC04_005_jr_8017BEBC.c:5057-5075; asm at asm/ov_SC04_005/nonmatchings/ov_SC04_005_jr_8017BEBC/func_8017F644.s.

Rule. When a target's case-dispatch mixes range_test and plain-compare shapes: (1) never write it as a switch — go straight to if/||; (2) group the case constants into pairs strictly by K,K+1 adjacency, one if per pair/singleton, matching whichever pairs the target folded — don't chain all cases into one ||. §164-56's own "count andi 0xffff between addiu and sltiu" tell can point backwards here: the FOLDED pair may show no mask (raw unsigned load operand) while the UNFOLDED pair still shows one (zero-extension tracking lost after the folded pair's arithmetic). Trust the K,K+1- adjacency test over the mask-counting tell when the two disagree.



ADDENDUM to §237 (func_8017F7FC, ov_SC03_092)

A BLOCK-SCOPE UNPROTOTYPED EXTERN DOES NOT ALWAYS SHADOW A HOSTILE FILE-SCOPE PROTOTYPE

Symptom. §236 point 5's fix ("block-scope the () decl inside your function so it shadows the prototype") is not universal. func_8017F7FC (ov_SC03_092): a block-scope extern void func_80178D18(); was written to call the callee WITH an argument, but the TU's dominant declaration for that identifier is the FULL prototype extern void func_80178D18(void); — present at file scope (src/ov_SC03_092/ov_SC03_092_jr_8017AE2C.c:2542) and repeated at nearly every call site in this TU, including line 5358 immediately above the target's own definition. Under C89's composite-type rule, a declaration of an identifier with external linkage composites with every other visible declaration of the SAME external object — the inner unprototyped () form does not erase the outer (void) prototype, it merges with it. The call with an argument then fails "too many arguments" at the CALL SITE, not a warning at either declaration, which reads as an unrelated breakage because nothing points at the declaration itself.

Fix. Do not contest the arity via a competing declaration. Adopt the TU's own prototype AND its cast-at-call-site idiom (§237 "IT FIXES" bullet 2): extern void func_80178D18(void); ... ((void (*)(s32))func_80178D18)(arg0); — banked verbatim at src/ov_SC03_092/ov_SC03_092_jr_8017AE2C.c:5358-5364, and the identical pattern recurs for this exact callee at nine-plus other sites across this overlay family (lines 4872, 5107, 5301, and ov_SC03_092_jr_80178D40.c / jr_8017A4AC.c / jr_80181694.c).

Boundary for §236-5. When a callee's dominant TU declaration is a full prototype, redeclaring it unprototyped at block scope is not a reliable escape from its arity — cast at the use site instead.