37 KiB
§283 — ADDENDA HARVESTED FROM THE 36-WAVE BATCH (P31 S60): 1,216 candidates, 859 already covered, 46 sharpenings, 9 new laws
Provenance. Waves cg..dt — the ore that accumulated while the campaign ran uncollapsed waves. Five reviewers on disjoint wave groups (D1 dd/de/df · D2 dg–dm · D3 cg–cm · D4 cn–cw · D5 cx–dt), each checking every ADDENDUM/NEW claim against the banked C in src/ and the target .s before proposing it. 859 COVERED is the healthy result: §193–§282 are earlier rounds of this same cycle.
What verification caught, both directions. Six reviewer claims were REFUTED at merge rather than laundered in — three 'the harvester leaked an unbanked NEAR into a banked-only harvest' reports (all three functions are genuinely banked; the notes were written before their gate landed and read as harvester bugs hours later), a 'match_one resolves targets by bare symbol name' claim (it resolves an explicit path, and the drafting path always passes it), and two idiom claims whose mechanism was absent from the banked code. §284's fence was described as NON-volatile; the banked code uses __asm__ volatile and the text is corrected here, because that is a detail readers copy verbatim.
The most-rediscovered gap. Eight cards across four waves independently re-derived that the whole-object gate needs every sibling matched — implicit in the corpus, never stated as its own law. Recorded here as the strongest signal of a missing section; it wants writing up properly rather than being buried in an addendum.
A process note for whoever runs the next batch. Forked sub-reviewers exceeded their brief in three of five groups: one re-derived five waves it was not assigned and self-merged over the shared output path, one silently dropped 11 rows including a whole function, one produced nothing. Each parent caught its own fork. Tell forks explicitly not to spawn further forks — they infer it from the parent's own earlier actions in inherited context — and give each one a private output path.
ADDENDUM to §164-75 — the fold-reassociation law also fires at a variable's INITIALIZER, not only a later expression
Confirms and extends §164-75 ("FOLD RE-ASSOCIATES A CONSTANT OUT OF AN INLINED +/-; ONLY A STATEMENT BOUNDARY PINS IT"), whose own scope note flags it as "one function, one direction... measure it together with §163z's independent claim." func_801916BC (ov_SC06_033, wave dd, MATCH 24/24) supplies a second, independent instance of the same underlying fold/RTL-reassociation law, at a DIFFERENT syntactic position: a pointer's own INITIALIZER, not a mid-expression arithmetic term.
Byte evidence. src/ov_SC06_033/ov_SC06_033_jr_8018D98C.c:5087:
p = &D_801CB512;
q = &D_801CB512 - 1;
asm/ov_SC06_033/nonmatchings/ov_SC06_033_jr_8018D98C/func_801916BC.s idx 5-7:
addiu $v1, $v1, %lo(D_801CB512)
addiu $a1, $v1, -0x2
$a1's base operand is $v1 — the register that JUST received D_801CB512's own address one instruction earlier — not a fresh %hi/%lo materialization of D_801CB512 - 2. This is the same fingerprint §164-75 already names ("an addiu $rD,$rS,-K whose $rS is the WRONG-looking operand of a nearby address computation — same instruction count, one register wrong, no length drift") but observed in an INITIALIZER, where the alternative failure mode is not "the constant rides the wrong operand of a nearby addu" but "the compiler independently materializes the offset pointer from scratch instead of folding it onto the just-computed base register."
The law, extended. §164-75's fold/reassociation law is not confined to a later arithmetic expression — it also governs how gcc-2.7.2 builds a SECOND pointer's initializer immediately adjacent to a first pointer's own address materialization. Writing q = &SYM - K; directly (not through an intermediate decrement inside a loop, and not as q = p - K; off an already-named p) is what keeps the -K riding the SAME register the base symbol's address just landed in, reproducing a single addiu $rD,$baseReg,-K rather than a fresh lui/addiu pair.
TELL. A target's second pointer-typed local is initialized with a single addiu off the register that the PREVIOUS statement's symbol address just computed (not a fresh %hi/%lo pair) ⇒ write that second local's initializer as a literal &SYM ± K expression, evaluated immediately after the first pointer's own &SYM statement — do not decrement inside the loop and do not route it through the first pointer's own C variable.
ADDENDUM to §21 — the bltz+slti (or N-separate-compares) signed-range-split bullet is now CONFIRMED on three independent functions, and generalizes beyond lbu/u8
Two independent byte-proofs land on the same bullet this batch — merged into one entry.
Instance A — the exact bltz+slti mechanism, second byte-proof (wave dd, func_800CCBF0, md_MAIN_042).
The bullet at cookbook L1984 ("lbu value with SIGNED compares (bltz + slti, NOT sltu) → write each range bound as a SEPARATE if (…) goto statement, never a chained ||") is marked "⚠ CANDIDATE (unverified) — evidence func_801775E0, a NEAR-MISS (closeness 2, NOT byte-banked)." func_800CCBF0 (md_MAIN_042, wave dd, MATCH 92/92) is a second, independent, fully byte-banked instance of the identical mechanism.
Byte evidence. src/md_MAIN_042/md_MAIN_042.c:66-72:
s32 var = *(s32 *)(a0 + 0x23C);
if (var == 0) {
...
} else {
if (var >= 0) {
if (var < 7) {
...
asm/md_MAIN_042/nonmatchings/md_MAIN_042/func_800CCBF0.s idx 22-24:
bltz $v1, .L800CCC54
slti $v0, $v1, 0x7
beqz $v0, .L800CCC54
— exactly the target-side fingerprint the L1984 bullet predicts (a provably-non-negative-looking s32 field still carrying a real bltz, immediately followed by an independent slti, both as SEPARATE nested if statements rather than a chained if (var >= 0 && var < 7)). This is a different function, different overlay/binary, and a different reviewer/wave than the original func_801775E0 near-miss, closing exactly the gap the original bullet asked for ("A drafter MAY try them... but VERIFY before trusting").
The law, upgraded. Remove the "⚠ CANDIDATE (unverified)" qualifier from the L1984 bltz+slti bullet — it is now confirmed byte-banked on two independent functions (func_801775E0's original near-miss evidence, and func_800CCBF0, md_MAIN_042, MATCH 92/92). Keep the bullet's prescription unchanged: split a signed range/sign guard on a value that LOOKS unsigned (u8/small non-negative field) into separate nested if statements — never a chained &&/|| — whenever the target shows an independent bltz immediately followed by a separate slti/sltiu.
Instance B — the fold generalizes beyond lbu/u8 to a plain s32 local, and this is (at least) the THIRD confirmation overall (wave dd, func_8017FD34, ov_SC05_017).
Symptom/context. §21's bullet (originally appended from a NEAR-MISS, func_801775E0, and still marked "⚠ CANDIDATE (unverified) ... NOT confirmed by a byte-match") describes chained bound tests folding into a sltu-based unsigned range trick when the source reads a u8 global via lbu, prescribing a separate if (...) goto per bound instead of a chained ||. func_8017FD34 (ov_SC05_017, MATCH 33/33) shows the identical fold — and identical fix — firing on a PLAIN s32 local holding a function-call return value, never loaded with lbu at all.
Byte evidence. src/ov_SC05_017/ov_SC05_017_jr_8017AE2C.c:5892-5920; target asm/ov_SC05_017/nonmatchings/ov_SC05_017_jr_8017AE2C/func_8017FD34.s. The dispatch value code (an s32 local, set from two different call-result chains) is tested against 1 and 2 with THREE separate compares — beq $v1,$v0(1), slti $v0,$v1,2, bne $v1,$v0(2) — exactly the un-folded shape §21's bullet describes, reproduced by writing the checks as a ladder_done: if (code==1) goto do_call; if (code<2) goto epilogue; if (code!=2) goto epilogue; chain rather than a chained code==1||code==2 test.
The boundary, sharpened. The fold this bullet describes is a general property of gcc-2.7.2's range_test/fold_truthop machinery acting on any small-integer comparison chain against ADJACENT constants — it has nothing to do with lbu's zero-extension per se. The lbu+signed-compare framing in the existing bullet is one INSTANCE (where the un-collapsed shape happens to keep a provably-dead bltz), not the boundary of when the lever is needed. Apply the same separate-if-per-bound fix whenever a target keeps N separate compares against small adjacent constants instead of a folded sltu/addiu-range test, regardless of the tested value's declared width or load instruction.
Also worth recording. This is at least the THIRD independent byte-proof of this bullet's core mechanism (alongside func_800CCBF0/md_MAIN_042, cited at §274's ADDENDUM to §1-I5, and the original func_801775E0 context) — the "⚠ CANDIDATE (unverified)" banner immediately above the bullet (cookbook ~L1980-1983) is stale and should be removed at the next corpus edit pass.
ADDENDUM to §202 — the DEF-SIDE ALIAS also resolves a function-vs-DATA-symbol identifier clash, not only a function-vs-function prototype clash
§202 documents the definition-side __asm__ alias escape for a case where the destination TU declares the SAME function under a conflicting FUNCTION prototype (its own worked example: TU-wide void, real body s32). func_80181940 (ov_SC06_000, wave dd, MATCH 45/45) shows the identical fix required by a different TRIGGER: the TU declares the function's own symbol as a plain DATA object, used as a stored callback VALUE, not as a function prototype at all.
Byte evidence. src/ov_SC06_000/ov_SC06_000_jr_8017AE2C.c:7219-7230:
extern s32 func_80181940;
...
((s32 (*)(s32 *))func_8012A568)(&func_80181940); /* install-handler registration, elsewhere in the TU */
...
void aF80181940(void) __asm__("func_80181940");
void aF80181940(void) {
...
}
The TU-wide extern s32 func_80181940; is a data declaration (the symbol's address is registered as a callback pointer via func_8012A568); defining void func_80181940(void) { ... } directly under that identifier is a conflicting-types DEFINITION-side wall a standalone match_one compile cannot see (§25's over-prediction class), because the standalone draft never carries the TU's competing data decl.
The law, extended. §202's def-side alias escape is not limited to a function-vs-function prototype/return-type clash — it equally resolves a function-vs-DATA-symbol clash, where the destination TU has declared the INCLUDE_ASM function's own name as an extern <scalar-type> object (typically because some other function in the TU takes its address as a callback/handler value). Define the body under a private alias identifier (aF<addr>) bound with __asm__("func_<addr>"); the TU's data declaration and its callback-registration call site keep working untouched, and the linker resolves both spellings to one symbol. Symptom to index this under: "an INCLUDE_ASM function's own name already appears as extern s32/extern u8/etc. (not a function prototype) elsewhere in the destination TU."
ADDENDUM to §22 (volatile-qualified-global reload lever, cookbook ~L1922) — for a NON-constant, same-address double RMW, a plain memory clobber beats volatile, and volatile actively breaks a delay-slot fill
Symptom. §22 prescribes volatile to force a reload between two consecutive stores to the SAME global when gcc constant-folds/forwards a just-stored LITERAL. func_80182AB4 (ov_SC02_031, MATCH 44/44) shows the boundary case where volatile is the WRONG lever: two |= read-modify-writes to the SAME memory word (*(u32*)(p+4) |= 0x50000000; then *(u32*)(p+4) |= 0x80000000;), where the second RMW's value is not a compile-time-known constant merge (it depends on the first RMW's runtime result), and the target additionally needs the SECOND store to remain schedulable into a following jal's delay slot.
Byte evidence. src/ov_SC02_031/ov_SC02_031_jr_8017AE2C.c:7175-7179:
*(u32 *)(p + 4) = *(u32 *)(p + 4) | 0x50000000;
__asm__("" ::: "memory");
*(u32 *)(p + 4) = *(u32 *)(p + 4) | 0x80000000;
func_80128EA8((s32)p, t, (s32)&D_800D3888);
Measured on this function (per the drafting note, consistent with the committed source): qualifying the field volatile does defeat the unwanted store-forwarding (the second load becomes a genuine re-read instead of reusing the first RMW's known-result register) but ALSO pins the second store above the following call, leaving the jal's delay slot a bare nop (a 13-instruction residual, near-miss). The bare, non-volatile __asm__("" ::: "memory") clobber between the two RMW statements gets the same forwarding-defeat WITHOUT constraining scheduling — the second store still sinks into the jal's delay slot, matching the target exactly (MATCH 44/44).
The law, sharpened. §22's volatile lever is for the specific case of re-reading a just-stored COMPILE-TIME CONSTANT (gcc's constant-fold path, which only volatile defeats). For a same-address double RMW where the SECOND value is not a known constant (an OR/AND of a runtime-derived first result), a plain non-volatile memory-clobber barrier between the two statements already defeats the unwanted store-to-load forwarding — and, unlike volatile, it does not additionally forbid the store from being scheduled into a later instruction's delay slot. Reach for volatile only when the re-read value is provably a literal gcc can constant-fold; otherwise the plain clobber is both sufficient and strictly cheaper on scheduling freedom.
ADDENDUM to §195-E — a goto-ladder's STORES must sit AT the labels, after the gotos, not inline before them
Symptom. A flag-materialising join (§195-E's own shape: bnez/beqz branches each carrying a
constant in their own delay slot, converging on one consumer test) is attempted as a goto ladder —
already §195-E's documented technique for "naming across a join" beyond plain if/else — but the
draft still collapses to jump.c's store-flag sltiu form on one or more arms, or the target's per-arm
delay-slot constants don't materialise.
Mechanism, byte-verified (func_801E56A8, md_SC03_135, MATCH 71/71,
src/md_SC03_135/md_SC03_135_jr_801E5358.c:363-397;
asm/md_SC03_135/nonmatchings/md_SC03_135_jr_801E5358/func_801E56A8.s). The target keeps every arm's
flag store in its OWN branch's delay slot (addu $v0,$zero,$zero for the two zero arms,
addiu $v0,$zero,0x1 for the "u>=0x12C" arm, and a plain fall-through addiu $v0,$zero,0x1 for the
"one" arm), with all three true-arms landing at the same consumer beqz (.L801E570C). The banked
source writes the ladder as:
if ((u32)(u - 0xC8) >= 0x64) {
if (u >= 0x12C) {
if (func_80029178(0xFA) & 0xFF) {
goto zero;
}
} else {
goto one;
}
} else {
goto zero;
}
flag = 1;
goto eval;
zero:
flag = 0;
goto eval;
one:
flag = 1;
eval:
if (flag != 0) { ... } else { ... }
Every test reaches a bare goto with NO store attached to the test itself; the stores live entirely
at the two labels (zero:/one:), each holding exactly one assignment followed by goto eval;. An
earlier attempt on this same function that wrote store-then-goto INSIDE each arm (i.e. put the
flag = 0;/flag = 1; assignment before the goto, at the branch site rather than at the label)
emitted store-then-branch instead of branch-with-delay-slot-store, and did not reach MATCH.
The law. For a §195-E-shaped goto-ladder join, the constant assignments must sit AT the labelled
targets (after the gotos land), not inline before the gotos that jump to them. This gives reorg
multiple independent fall-through/branch candidates — one store per incoming edge — so each bnez/
beqz steals its own arm's constant into its own delay slot exactly as the target shows. Writing the
store before the goto at the branch site instead produces a store-then-unconditional-branch
sequence that reorg cannot redistribute into each test's own slot.
When it applies. Any §195-E join reached via a goto ladder (not plain if/else) where a
store-before-goto spelling leaves the target's per-arm delay-slot constants unreproduced, or where
jump.c's store-flag sltiu collapse persists on one or more arms despite the ladder shape.
ADDENDUM to §195-N — a GNU statement-expression slider must sit INSIDE the conditional arm's value position, not as a post-hoc barrier, to block the store-flag transform on a ternary chain
Symptom. A chained-ternary flag computation (flag = A ? 0 : (B ? 1 : (C ? 0 : <expr>));) that
should reach one arm via a genuinely-computed value collapses one leaf to jump.c's sltiu
store-flag transform, and a zero-byte __asm__ fence placed AFTER the assignment (§195-N's own
"outer if + __asm__ slider blocks the store-flag transform" aside, stated there only for a
call-bearing if (f()) return 1; chain) is inert against it.
Mechanism, byte-verified (func_801E559C, md_SC03_135, MATCH,
src/md_SC03_135/md_SC03_135_jr_801E5358.c:226-241;
asm/md_SC03_135/nonmatchings/md_SC03_135_jr_801E5358/func_801E559C.s). The banked source is:
flag = ((u32)(t - 200) < 100u) ? 0
: ((t < 300) ? 1
: (((func_80029178(250) & 0xFF) ? 0 : ({ __asm__(""); 1; }))));
if (flag) {
*(s32 *)(*(s32 *)&D_801E256C + 4) = (s32)D_801E68BC;
} else {
*(s32 *)(*(s32 *)&D_801E256C + 4) = (s32)D_801E68F0;
}
The GNU statement-expression ({ __asm__(""); 1; }) sits as the innermost ternary's ELSE-arm VALUE
itself — not between the ternary and the if, and not between an assignment and its later use. An
earlier draft (per the note, "draft #22") placed an equivalent slider AFTER the whole assignment,
between it and the consuming if; that placement is inert because jump.c's fold-based store-flag
pattern match on x=0; if(c)x=1;-shaped trees runs at TREE-EXPANSION time on the ternary's own
sub-expression, before a post-hoc statement barrier downstream of the completed assignment can have
any effect. Placing the barrier inside the arm's own value expression prevents the constant-folding
tree match from recognising that one leaf as a bare 0/1 in the first place, which is what lets
reorg later distribute each leaf's constant into its own branch's delay slot instead of jump.c
collapsing the tail to sltiu. The note also records that the if/else arms must each hold their
OWN store (not a shared post-join store through a selected pointer) to prevent jump-threading/arm
duplication — consistent with, and reusing, the tail-shape lever already sourced from the banked twin
func_8017F9D8 (ov_SC03_124).
The law. A zero-byte __asm__ "slider" that is meant to block jump.c's store-flag/constant-fold
transform must sit INSIDE the value position of the specific conditional arm whose 0/1 leaf must
survive as a real branch target — as a GNU statement-expression (({ __asm__(""); 1; })) supplying
that arm's value — not as a bare statement-level fence placed between the completed assignment and
its later use. §195-N's own aside documents the slider for a different C shape (a call-bearing
if-chain closed by return 0;); this extends the same "asm barrier defeats jump.c's constant-fold
recognition" mechanism to a chained-ternary flag computation, and pins down WHERE the barrier must be
written for it to work.
When it applies. A chained ternary (or similar nested conditional-value expression) where one leaf
is a compile-time constant and the target's branch to that leaf carries the constant in its own delay
slot (i.e. §195-E's "named across a join" shape), but a sltiu/store-flag collapse persists on that
leaf despite an attempted asm-fence placed after the assignment.
ADDENDUM to §176-B2 — in a micro-function with no long/short lifetime asymmetry, BOTH contending pseudos need their own hard-register pin
Symptom. §176-B2 prescribes pinning the short-lived INTERLOPER (not the long-lived contested value) when local-alloc's first-fit hands one hard register to two competing pseudos. In a very short leaf function, pinning only one of the two contenders does not close a REGALLOC-PERM residual, and the section's own precondition (a long-lived value vs. a short-lived single-use constant) does not obviously identify which pseudo is "the interloper."
Mechanism, byte-verified (func_80181EF0, ov_SC02_031, MATCH 12/12,
src/ov_SC02_031/ov_SC02_031_jr_8017AE2C.c:6852-6861;
asm/ov_SC02_031/nonmatchings/ov_SC02_031_jr_8017AE2C/func_80181EF0.s). The function is 12
instructions total — every pseudo in it is short-lived by construction, so there is no long-lived
value for a single interloper pin to protect. The banked source pins BOTH:
register s32 m __asm__("$4");
register s32 c __asm__("$3");
m = D_80078EBA;
c = 4;
if (m == c) {
return (u32)(D_80078EB1 - 7) < 5;
}
return 0;
(✔ byte-checked: the target's lbu $a0,%lo(D_80078EBA)($a0) and addiu $v1,$zero,0x4 land in $a0
($4) and $v1 ($3) exactly as declared.) An unpinned draft left both values on a REGALLOC-PERM
$v0>$v1>$a0 residual that no source-level reordering (named local, operand swap, ternary, zero-init
accumulator) moved — consistent with §137's "source-level levers are a dead end" for this class.
Pinning only m to $4 closed part of the gap (4→2); the residual only fully closed once c was
ALSO pinned to $3.
The law. §176-B2's "pin the interloper, not the contested value" presumes a long/short lifetime
asymmetry between the two competing pseudos, which lets you identify one of them as the thing to move
out of the way. A micro-function (single basic block, every value short-lived) has no such asymmetry —
first-fit can hand the wrong register to either pseudo, or both. When the target's own registers for
BOTH contenders are known (read them off the .s), pin BOTH pseudos directly to their target hard
registers rather than searching for a single interloper to displace.
When it applies. A REGALLOC-PERM residual of 1-4 instructions in a short, single-block leaf function where §176-B2's single-interloper-pin approach does not close the gap and no clear long-lived value exists to leave unpinned.
ADDENDUM to §199-G — A default: LABEL GROUPED ONTO THE LAST CASE REMOVES THE j default TAIL, EVEN THOUGH THE 2-NODE HEADER STAYS ALL-POSITIVE
Symptom. A 2-case-node switch whose target .s shows the expected all-positive equality-test header (§199-G's own tell), but with no trailing j <default> at all — the second test's fall-through lands directly inside a CASE body, not a separate default arm.
Mechanism, byte-verified (func_8017DFE4, ov_SC03_011, 42/42 MATCH). Source
(src/ov_SC03_011/ov_SC03_011_jr_8017C730.c:3547-3567):
void func_8017DFE4(s32 param_1)
{
s32 state;
state = func_801399F0(*(s32 *)(param_1 + 0x198));
if (state != 0) {
func_80139914(*(s32 *)(param_1 + 0x198));
*(s32 *)(param_1 + 0x198) = 0;
switch (state) {
default:
case 1:
*(s32 *)(param_1 + 0x198) = func_8013767C((s32)&D_80185F80);
func_80171990((u8 *)param_1);
break;
case 2:
*(s32 *)(param_1 + 0x198) = func_8013767C((s32)&D_80185FA8);
*(u8 *)(param_1 + 0x216) = 5;
break;
}
}
}
Target (asm/ov_SC03_011/nonmatchings/ov_SC03_011_jr_8017C730/func_8017DFE4.s):
beq $s1, 1, .L8017E034
sw $zero, 0x198($s0) ; delay slot
addiu $v0, $zero, 2
beq $s1, 2, .L8017E058
nop
.L8017E034: ; case 1 AND default fall straight in here
lui $a0, %hi(D_80185F80)
...
Both equality tests remain positive (beq→.L8017E034 for state==1, beq→.L8017E058 for
state==2), matching §199-G's header shape exactly — but there is no third branch and no j default: when neither test fires, control simply falls out of the header into .L8017E034,
which is also case 1's own body. This is only possible because default: and case 1: share
one label in the source — emit_case_nodes (stmt.c) has no separate default arm to jump to, so
expand_end_case's emit_jump_if_reachable(default_label) (stmt.c:4909, the very call that
produces the j default §199-G's rule 5 says is unconditional) has nothing to jump to and emits
nothing.
The law. §199-G rule 5 states that an explicit default: still emits the all-positive header
and a j default. That holds only when default: has its own body/label. When default:
is grouped onto (shares a label with) one of the case arms, the trailing j default disappears —
the header still reads all-positive per node, but the LAST node's fall-through is itself the
shared arm, not a jump. Tell: a 2-node all-positive header with NO trailing j at all, where
the fall-through address is a real case body rather than a distinct join/epilogue block ⇒ write
default: stacked directly above that case's label in the switch, not as a separate arm.
Boundary. Distinct from §199-G's own "empty arm" bound (rule 2, case 0: break; folding onto
the join and inverting the SECOND test) — here neither arm is empty; the grouped label is what
removes the jump, not label-collapse from a lack of body.
ADDENDUM to §220 — REFERENCING THE RAW PARAMETER (NO NAMED COPY, NOT EVEN A PIN) LETS THE CALLEE-SAVED SPILL LAND IN THE FIRST CALL'S OWN DELAY SLOT
Symptom. A function's prologue shows sw $sN (the callee-saved slot) BEFORE an unrelated
lui/addiu address materialization, with the actual addu $sN,$aM,$zero copy embedded inside a
following jal's OWN delay slot — not as a separate early move. Every spelling that introduces an
explicit copy variable for the parameter (plain named local, register __asm__ pin, or a
non-volatile fence around the copy) reorders this: the sw $sN slides to AFTER the lui/addiu
pair, or an extra move appears.
Mechanism, byte-verified (func_8017E8B0, ov_SC03_002, 79/79 MATCH,
src/ov_SC03_002/ov_SC03_002_jr_8017D604.c:3280-3299):
void func_8017E8B0(void *a0) {
s32 v1;
if (func_8012C354((s32)a0, (s32)&D_801892C8) == 0) {
return;
}
...
No s0 = a0; (or pinned/copy variant) appears anywhere — a0 is referenced directly at every
site through the rest of the function. Target
(asm/ov_SC03_002/nonmatchings/ov_SC03_002_jr_8017D604/func_8017E8B0.s):
addiu $sp, $sp, -0x18
sw $s0, 0x10($sp) ; save slot FIRST
lui $a1, %hi(D_801892C8) ; unrelated address materialization SECOND
addiu $a1, $a1, %lo(D_801892C8)
sw $ra, 0x14($sp)
jal func_8012C354
addu $s0, $a0, $zero ; the copy itself, in the CALL's own delay slot
$a0 is live only because it is used again after this first call (inside the later if-guarded
block); gcc has no reason to establish $s0 until the point where $a0 must actually survive
across a call — which is exactly the first jal. Since a delay slot needs filling there anyway,
the compiler's own callee-saved spill fills it for free, landing the register-save sw at
function entry (ordinary prologue placement, independent of when the value is copied) but the
VALUE copy itself only at its first genuine cross-call use.
The law. When the target embeds a parameter's callee-saved copy (addu $sN,$aM,$zero) inside
a call's own delay slot rather than as a standalone early instruction, and every explicit-copy
spelling (plain named local, register pin, or fence around the copy) instead produces either a
standalone move or shifts the sw $sN save to a different position — do not name a copy variable
at all. Reference the incoming parameter directly everywhere in the body; let the compiler's own
save-placement machinery decide where the spill lands, which will be at the first point the
parameter's live range must actually cross a call. This sharpens §220's "explicit named local is
the reliable way to pin a callee-saved copy" general rule with its own boundary case: a named copy
(pinned or not) is the WRONG tool specifically when the target's spill sits inside a call's delay
slot rather than at a fixed early position — that shape wants the raw, un-copied parameter instead.
ADDENDUM to §252 — a >=0/<0 split on an unconditionally-decremented value needs the POSTFIX operator INSIDE the branch condition, not a prior statement
Symptom. The target shows a decrement's STORE landing in the branch's OWN delay slot (§252's second reading rule: "the decrement is unconditional and the test follows") for a sign-test split (bgez/< 0 or the reverse) — but the branch's own compare must read the pre-decrement value while the delay-slot store writes the post-decrement value. §252's existing prescription for this reading ("write it before the if, not inside the taken arm") is ambiguous between two spellings that are NOT equivalent here: a prior statement (x--; if (x < 0) {...}) tests the value after the decrement, while the target needs the test on the value before it.
Mechanism, byte-verified (func_80186140, ov_SC03_006, MATCH 38/38): the only C spelling that reproduces both facts at once is the postfix decrement used directly as the controlling expression, relying on C's postfix semantics (yield the old value to the comparison, store the new value as a side effect):
if ((*(s16 *)(a0 + 0x2C))-- < 0) {
func_801292C8((u8 *)a0);
} else {
...
}
(✔ byte-checked: src/ov_SC03_006/ov_SC03_006_jr_8017AE2C.c:9938; asm/ov_SC03_006/nonmatchings/ov_SC03_006_jr_8017AE2C/func_80186140.s idx 10-13 —
lhu $v0,0x2C($s0) / addiu $v1,$v0,-0x1 / sll $v0,$v0,16 / bgez $v0,.L80186190 with sh $v1,0x2C($s0) in the delay slot: the sign test (sll+bgez) operates on the freshly-loaded, PRE-decrement $v0, while the delay slot unconditionally stores the already-computed POST-decrement $v1 — exactly the postfix-operator split.)
Splitting into two statements (s16 old = *(s16*)(a0+0x2C); *(s16*)(a0+0x2C) = old - 1; if (old < 0)) or writing the decrement after the store both fail to reproduce this: either the load-delay-slot content changes or the branch ends up testing the wrong value.
The law. When a target's guarded decrement shows the store in the branch's own delay slot (§252's "unconditional decrement" reading) AND the branch's own compare must read the value from BEFORE that decrement (a >=0/<0, or similarly value-dependent, split) — write the decrement as a bare postfix operator (x-- </>= K) directly inside the if's controlling expression. This is a narrower, single spelling within §252's "write it before the if" prescription, not a free choice among equivalent-looking alternatives: any two-statement decompose of "decrement, then test" tests the post-decrement value instead.
TELL. A sll(sign-extend)/bgez-or-bltz pair operating on the SAME register a following delay-slot sh/sw stores a -1'd copy of, with no intervening reload — the sign test's operand and the stored operand are two different pseudos derived from ONE load, exactly what a postfix --/++ used as a condition produces and a decompose-then-test spelling does not.
ADDENDUM to §215 — FIFTH SHAPE: reused mask constants across two call-free merge sites each get their own whole-function hard-register pin, and a shared sub-expression at the second site must be its own statement
Merge-time reclassification. The wave-df submitter proposed this as a bare NEW section; the wave-dd submitter independently reviewed the same function and called it COVERED by §215's existing pin-economy shapes 1-4 plus the L1816 non-volatile-fence law. On inspection, §215's shape 3 ("pin the CONSTANT interloper, not the pointer") is the closest existing match but only covers ONE contested constant at ONE site — it does not state the two-mask/two-site/whole-function-residency condition, or the shared-sub-expression-as-its-own-statement technique, that this function actually needs. Filing this as §215's fifth shape rather than a free-standing section, since it is a direct extension of an established list, not an unrelated mechanism. The wave-dd submitter's "fence must not be volatile" sub-claim IS already covered verbatim by the L1816 bullet (cited correctly by that note) and is not repeated here.
FIFTH SHAPE — REUSED MASKS ACROSS TWO SITES, PLUS THE SPLIT-EXPRESSION TECHNIQUE (P31, func_80187980, ov_SC06_032, byte-proven 55/55; independently re-derived in both wave dd and wave df)
The symptom. A function builds one 32-bit word from two 24/8-bit masked halves twice — once merging
a freshly-loaded word w with a passed-in value t ((w & mB) | (t & mA)), and a second time merging
that same passed-in slot's value again with the OBJECT'S OWN ADDRESS ((r2 & mB) | ((u32)p & mA)) — and
the target keeps mA/mB resident in the SAME two hard registers across both sites (no call
intervenes) while the two or-destinations land in two DIFFERENT registers ($a1 at the first site,
$v1 at the second). A plain-local draft lets cse/regalloc pick whichever registers it likes for the
masks and whichever destination coalescing prefers, and both diverge from the target.
Byte evidence. src/ov_SC06_032/ov_SC06_032_jr_80182890.c:5934-5960; target
asm/ov_SC06_032/nonmatchings/ov_SC06_032_jr_80182890/func_80187980.s.
register u32 w __asm__("$5");
register u32 mA __asm__("$4");
register u32 mB __asm__("$6");
register u32 r2 __asm__("$3");
u8 *p;
u32 t, ps;
p = (u8 *)func_80010A08(0x10);
p[3] = 3;
p[7] = 0x42;
w = *(u32 *)p;
__asm__ ("" : : "r" (w));
p[4] = a2[0];
p[5] = a2[1];
p[6] = a2[2];
*(u16 *)(p + 8) = a0[0];
*(u16 *)(p + 0xA) = a0[1];
*(u16 *)(p + 0xC) = a1[0];
*(u16 *)(p + 0xE) = a1[1];
mA = 0xFFFFFF;
mB = 0xFF000000;
t = *a3;
*(u32 *)p = (w & mB) | (t & mA);
r2 = *a3;
ps = (u32)p & mA; /* <-- the shared sub-expression, pulled into its own statement */
r2 = (r2 & mB) | ps;
*a3 = r2;
Target asm confirms every register: lui $a0,%hi(0xFFFFFF) / ori $a0,... materialises mA into $a0
(=$4); lui $a2,%hi(0xFF000000) materialises mB into $a2(? — note: mA/mB's C-level register
names are $4/$6; the compiled instructions use gcc's own $a0/$a2 mnemonics for those same
physical registers, i.e. $4=$a0, $6=$a2 under the o32 convention). First merge:
and $a1,$a1,$a2 (w&mB) / and $v1,$v1,$a0 (t&mA) / or $a1,$a1,$v1 / sw $a1,0x0($v0). Second
merge: and $v0,$v0,$a0 — ps = (u32)p & mA lands in $v0, p's OWN register, reused because p's
last read is this exact instruction; and $v1,$v1,$a2 (r2&mB, $v1=$3); or $v1,$v1,$v0
($v1=$3, matching the register u32 r2 __asm__("$3") pin exactly); sw $v1,0x0($s1).
The law. When a target reuses the SAME two mask constants across two independent masked-merge
sites with no intervening call, pin each mask to the hard register the target holds it in for the
WHOLE function (register u32 mA __asm__("$4"), etc.) rather than using a plain local — a plain local
gives cse/regalloc no reason to keep both masks resident in fixed registers across the gap between
sites. When the SECOND site's shared operand is a freshly-computed value (here, the malloc'd pointer's
own address masked) rather than a parameter already in a register, write it as its OWN preceding
statement (ps = (u32)p & mA;) instead of inlining it into the | expression — this is what lets the
about-to-die variable's register (p's $v0) be reused for the intermediate rather than forcing a
fresh allocation, and lets the merge's destination land in the mask-adjacent pinned register ($3) the
plain-inlined form does not reach.
Refuted sub-claim, worth recording explicitly (per the narrowest-form rule). The drafting note (both
the wave-df and wave-dd cards for this same function) additionally claims the FIRST site's load,
w = *(u32 *)p;, must be written textually AFTER p[4] = a2[0]; in source order for the target's
lw $a1,0x0($v0) to land in the first byte-load's delay slot, "without the fence cc1 sinks the load
next to its lui/ori." This does not hold against the committed source: line 5941 (w = *(u32*)p;)
sits BEFORE line 5946 (p[4] = a2[0];), the opposite of the claimed order, and the function still bytes
MATCH (55/55). The target's lw $a1,0x0($v0) (idx 17 of the asm) does land in the delay slot of the
FIRST byte-load (lbu $v1,0x0($s0) at idx 16, loading a2[0]) — but that is sched1 reordering the
compiled instruction stream, independent of which of the two adjacent statements was written first in
C; the non-volatile use-fence (__asm__("" : : "r"(w))) immediately after w's assignment is the part
that is load-bearing (it is present in the banked code and keeps w from being folded/sunk into its
own lui/ori constant-build window), not the claimed statement ordering. Do not carry the
"AFTER p[4]" ordering claim forward; only the fence + the mask/split/pin levers above are byte-verified.