Files
BFM-decomp/cookbook/C0212.md
T

3.1 KiB

§190 — THREE PRESCRIPTIONS FROM THE SAME HARVEST (weaker evidence than §189, honestly labelled)

§190-A — THE PREHEADER HAS THREE FIXED STRATA, and an init in the wrong stratum is not schedulable. In stream order: (1) a source biv's own init, before NOTE_INSN_LOOP_BEG, never moved; (2) move_movables' hoisted invariants (loop.c:1652/1708, called from scan_loop:966); (3) strength_reduce's reduced-giv inits (emit_iv_add_mult(..., loop_start), loop.c:3879, called at :976). Strata 2 and 3 both insert immediately before loop_start, so the pass order 966 < 976 is what puts giv inits last. A register whose zero-init sits after a hoisted invariant cannot be a source biv — it is a reduced giv of a stride-1 counter the source never named, and maybe_eliminate_biv (loop.c:5952) then deletes that counter, which is why the target's exit test reads slti rGIV, N*K instead of slti rCOUNTER, N. Tell: count and content already match; the residual is a lone move rX,zero on the wrong side of the preheader's la/lui block, with zero missing or extra instructions. Do not open the permuter on it. Fix: n = 0; ... i = n * K; ... n++; — and i = n * K must be its own named statement, or every &SYM + n*K becomes its own address giv and floods the preheader. Proven twice: src/ov_SC04_011/ov_SC04_011_jr_8017D494.c:7740 (81/81) and :7882 (122/122). Resolves §2-T2's open case; scope-fenced to the strata BOUNDARY — order within stratum 2 is §162e2, within stratum 3 is §164-06.

§190-B — COMPILE THE PLAIN NATURAL ORDER FIRST: an interleave in the target is not evidence the source was interleaved. sched1 produces interleaving from natural order, so a hand-"pre-scheduled" draft is itself a defect class — the tell is a draft you deliberately reordered "to help gcc" sitting at a small stubborn residual filed as instruction-scheduling. Measured, and this is the new part: per-block rigidity is asymmetric within one function, so the winning set is a product-structured plateau, not a point. Permuting stores that feed later reloads is rigid (4 of 24 orders reach 0; 552/576 cells nonzero); permuting independent same-base writes whose values die locally is loose (8 of 24 byte-identical, natural among them). So "the natural order won" and "order is a live dial" are both true at once — do not generalise a loose block into §178's "statement-order sweeps are worthless", nor a rigid block into "you must hand-craft the order".

§190-C — A CALL-ARG CONSTANT WEDGED INTO A DEPENDENT LOAD'S DELAY SLOT IS A sched2 HOIST. Attribute it first: if -fno-schedule-insns2 alone reproduces the target order, sched1 was already right and post-reload sched2 is hoisting the load above the constant. The fence is both halves together — arg registers pinned above a __volatile__ memory clobber: plain s32 locals get folded into the call's own arg setup and sink below the fence (verified), and it is __volatile__ that makes the asm a barrier at all (sched.c:1957, if (code != ASM_OPERANDS || MEM_VOLATILE_P (x)), gating flush_pending_lists).