2.6 KiB
§212 — THE WALKING CURSOR IS COUNTABLE: *wp++ emits one addiu PER STORE, wp[0..2] emits one (P31 S58)
THE TELL. Three stores at displacement 0 off a base that steps between them:
sh $v0, 0x0($a1)
...
addiu $a1, $a1, 0x2
sh $v0, 0x0($a1)
...
addiu $a1, $a1, 0x2
sh $v0, 0x0($a1)
(✔ byte-checked: asm/ov_SC04_011/.../func_80184B3C.s.)
THE LAW. Three sh 0(reg) with three separate addiu reg,2 is *wp++ written three times.
wp[0] = …; wp[1] = …; wp[2] = …; compiles to sh 0 / sh 2 / sh 4 plus ONE addiu reg,6 at the
loop bottom — a different instruction count and a different delay-slot fill. The two spellings are
semantically identical and byte-distinct; read the displacement column, not the C you would naturally
write.
The same discrimination runs the other way inside one function. func_801861D4 (ov_SC04_011,
wave af) has two loops with two different answers: loop 1 over D_801EFCA8 is a WALKING POINTER
(*ptr++, 16 entries) — the indexed form gave the same 52-ins count but lost the register allocation
— while loop 2 over D_801EFCF8 stays INDEXED (D[i], i+1 as the call argument) exactly as the
asm shows. Do not unify them.
AND THE THIRD FORM — two same-base walkers of DIFFERENT scale must both be INDEXED.
func_80187754 (ov_SC06_000, 17 ins) burned four drafts on this. A naive two-pointer walk gave
LENGTH-DRIFT; a walked-pointer p[5], p++ gave addiu $a0,0x14 where the target holds $v1 as an
addu COPY of $a0; stepping the halfword walker by elements gave the wrong byte step, and moving
the increment earlier let gcc fold it into the store address (losing the standalone addiu). The
crack: spell BOTH walkers as indexed expressions off the single base — ((s16*)a0)[i+2] and
a0[i+5]. gcc-2.7.2's IV elimination then creates the two givs in the target's order ($v1 = word
walker +4, $a0 = halfword walker +2 in the branch delay slot). The residual was giv-CREATION
order between two same-base index expressions of different scale, and it fell out of the
indexed-vs-walked spelling with no pin.
RELATED, and worth reading with §162e2/§164-03: func_80184738 (ov_SC02_011) — gcc hoists
v1 = p + 0x70 out of the loop, so every offset in the body is printed relative to p+0x70;
sh …,0x8C($v1) is really *(s16*)(p+0xFC). And on the same card, writing p += 0x10C in the loop
BODY emits addiu $v1 before addiu $a2; moving it into the for-increment clause
(i++, p += 0x10C) flips the order to match. Pure C-source placement, invisible in any pseudo-C
rendering of the asm.