Files
BFM-decomp/cookbook/C0340.md
T

3.7 KiB

§306 — A HAZARD nop IN FRONT OF A DIV-RESULT STORE IS A STATEMENT-ORDER DEFECT: THE INDEPENDENT TRAILING STATEMENT MUST BE WRITTEN BEFORE THE DIVISION-CONSUMING ONE (P31 S62 T4; byte-proven func_8017E7D0)

THE TELL. A LENGTH-DRIFT +1 whose extra instruction is a bare nop sitting immediately before the sw that stores a division result, while an adjacent, data-independent global store (lui $at,%hi(G) / sh $zero,%lo(G)($at)) appears AFTER that sw in your draft and BEFORE it in the target:

mine:    … mflo $a0 ; nop ; sw $a0,0x48($s0) ; lui $at ; sh $zero,%lo(G)($at)
target:  … mflo $a0 ; lui $at ; sh $zero,%lo(G)($at) ; sw $a0,0x48($s0)

Do not touch registers, pins, barriers or the permuter — the whole residual is one statement in the wrong place.

THE MECHANISM (two halves; the second is source-verified). (1) cc1 sees div as ONE insn carrying the long imuldiv latency, so its consumer store is not ready for many cycles and the scheduler fills the shadow — but only from insns already available in LUID (= source) order. An independent statement written BEFORE the division statement sinks into that shadow; one written AFTER it is not pulled back into it. (2) The visible nop is not gcc's — maspsx splices it: after the --expand-div expansion, _handle_nop_before_next_instruction (tools/maspsx/maspsx/__init__.py:642-675, called at :1137) emits nop # DEBUG: Reuse of '<rd>' whenever the instruction following the expanded mflo reads the quotient register. So an empty shadow costs exactly one instruction, and any insn that does not read the quotient kills it. Corollary (source-read, not A/B'd): that same rule exempts a next instruction which uses $at while nop_at_expansion is false — which our pinned --aspsx-version=2.56 gives (maspsx.py:99-104) — so a division result stored straight to a global through lui $at / %lo(SYM)($at) never pays this nop at all.

THE C SHAPE. Swap the two independent trailing statements so the one with NO dependence on the division comes first:

/* target order — shadow filled, no nop (54 ins) */
D_8019F70C = 0;
*(s32 *)(a0 + 0x48) = -D_80188A34 / ((s16)D_80188A2A[0] / 2);

/* wrong order — empty shadow, maspsx nop (55 ins) */
*(s32 *)(a0 + 0x48) = -D_80188A34 / ((s16)D_80188A2A[0] / 2);
D_8019F70C = 0;

BYTE EVIDENCE. func_8017E7D0 (ov_SC06_016, _jr_8017C8D0, 54 ins). Target tail asm/ov_SC06_016/nonmatchings/ov_SC06_016_jr_8017C8D0/func_8017E7D0.s:50-53 = mflo $a0 / lui $at,%hi(D_8019F70C) / sh $zero,%lo(D_8019F70C)($at) / sw $a0,0x48($s0), under the signed --expand-div form (break 7 / break 6, §228-3). v1 (division store first): mine=55, target=54, 10 mismatched, class LENGTH-DRIFT [structural], first mismatch idx45: 00000000 nop | 3c01801a lui $at,%hi(D_8019F70C). v2 = the same file with ONLY those two statements swapped: MATCH (54 ins); --json verify {"status":"match","closeness":0,"nins":54,"residual":[]}. One function, both directions measured.

(Extends §16Xb (L13142) — the mult→mflo window read at the div's shadow: §16Xb's tell is a ZERO-drift reorder and its lever is "move the statement that FOLLOWS the multiplying statement"; here the drift is +1 and the filler comes from the statement BEFORE, so §16Xb's prescription points the wrong way. Instance of §2-T2 (L78). Kin of the §253/§165-06 note (L25420: postfix (*p)++ keeps the store after the mfhi/bnez pair, bare ++ sinks it before the div) and of §50-E (L3616), which prices global-store REORDER the other way (an assembler-merged lui $at). §179-B rule 4 (L17393) owns this same maspsx nop splicer for hand-written asm; this entry is its C-side face.)