Files
BFM-decomp/cookbook/C0211.md
T

6.4 KiB
Raw Blame History

§189 — FIVE COMPILER LAWS MINED FROM THE WAVE R/S JOURNALS (P31 S53), each source-cited and re-derived by a second agent

🔴 THE INFERENCE DIRECTION IS BYTE-REFUTED — see §199-A/§199-E (next session's harvest). The SPLIT-TIMING half below survives (-fno-schedule-insns emits an unsplit li and no pair at all). What is FALSE is everything downstream of it: LUID adjacency does NOT imply emission adjacency, an interloper licenses NO conclusion about the source spelling, and the "no statement order and no pin" absolute is wrong. Byte-proof: the banked one-statement slice prim.col[1] = 0x101010; compiles with SEVEN insns between its lui and ori; the two-step spelling prescribed below is BYTE-IDENTICAL (the fix is inert); moving an unrelated statement moves a third constant in and out of the gap; and one separated pair is 0x88888889, gcc's own reciprocal magic for / 0x3C — a constant with no source spelling at all, so "the target wrote two steps" is unsatisfiable there. rank_for_schedule tests INSN_PRIORITY first (sched.c:2395) and reaches the LUID tie-break only at :2428; the separator is the BIRTHING BOOST (birthing_insn_p, gated on reg_n_sets == 1), which the split pair can never have because try_split gives its pseudo two sets. 8 separated pairs across 5 functions in 3 binaries.

§189-A — SPLIT-CONSTANT LUID ADJACENCY: an interloper between lui/ori proves the target wrote TWO source steps. mips.md:3208's large_int define_split fires in sched1's per-block pre-pass (sched.c:4830 try_split, reload_completed == 0) — before sched_analyze hands out LUIDs (sched.c:2175). The two halves are therefore chain-adjacent with consecutive LUIDs, and since rank_for_schedule's only live discriminator among equal-priority ALU constants is the INSN_LUID tie-break (sched.c:2428), no statement order and no pin can put a third constant between them. So if the target shows one there, the target did not write one constant. Fix (both edits needed; either alone still scores 2): split it in source with a §30#3 zero-byte re-tie between the halves (x = HI; __asm__("" : "=r"(x) : "0"(x)); x |= LO;) so the |= carries its own LUID, and route the interloper through an already-multi-set / pinned variable so birthing_insn_p (sched.c:2469) does not boost it below both halves. Four-form A/B on func_8001BBBC, identical 44-instruction multiset each time, only the li 4's slot moving: one constant → after both · two steps, no re-tie → before both · two steps + re-tie + anonymous temp → below both · two steps + re-tie + multi-set pinned temp → BETWEEN, MATCH 44/44 (src/800.c:5453, idiom at :5465-5476). In-function control: the same function's 0xE1000040 is a one-statement split and its lui/ori sit adjacent with nothing between — exactly as the law predicts. Narrative correction: the re-tie does not stop cse folding the halves back into one CONST_INT (A/B shows they stay separate without it, with only a REG_EQUAL note). It stops cse deleting and re-materialising the HI set later in the chain, which is what lifts the lui's LUID above the interloper. The comment in src/800.c:5430-5431 says the former and is wrong.

§189-B — SELF-ACCUMULATE OPERAND ORDER IS FIXED AT RTL EXPANSION, SO REORDERING THE C ADDENDS IS A GUARANTEED NO-OP. optabs.c:399-421 swaps op0/op1 whenever target == op1 by rtx pointer identity; for a local in a pseudo, expand_expr's VAR_DECL arm returns DECL_RTL itself (expr.c:4258) and store_expr passes that same rtx as target. So x = b + x is swapped back to (plus x b) and emits the identical addu $x,$x,$b as x = x + b. If the target shows addu $x,$b,$x, permuting the addends will look like a refutation and prove nothing. Levers, both of which break the identity: route through a temp ({ s32 xt = b + x; x = xt; }), or pin the destination to a hard reg. This BOUNDS §10 Residual A / Fix A1, §164-02 and §167-32's "write the operand you want in rs first" — all three are cse/front-end mechanisms with no target-identity constraint, and all three are inert on a self-accumulating statement.

§189-C — A §30#3 RE-TIE THAT MUST LIVE IN bb0 NEEDS A "memory" CLOBBER. The boost-kill itself is a sched1 effect, but the asm insn you added survives into sched2, where prologue saves, stack-arg loads and ALU fillers all tie at priority 1 and fall through to INSN_LUID (§167-13). Inside bb0 a bare __asm__("" : "=r"(x) : "0"(x)) floats through that interleave and rotates the filler chain one triple early. Escalating to __asm__("" : "=r"(x) : "0"(x) : "memory") makes expand_asm_operands emit (clobber (mem:BLK (scratch))) (stmt.c:1656-1663), which sched_analyze_insn routes through the write-memory path (sched.c:2035-2042 → :1736-1790) and pins it. volatile over-fences. Bounds §30#3, whose prescription ("place the re-tie in a LATER basic block") has no answer when the kill must happen in the entry block.

§189-D — A NARROW PARAMETER IS BORN INTO TWO PSEUDOS, so its target shape is a callee-saved COPY-OF-A-COPY with no extension anywhere. mips.h:1153 defines PROMOTE_PROTOTYPES (the caller widens) but the port defines no PROMOTE_FUNCTION_ARGS and no PROMOTE_MODE, so for a prototyped u16/s16 parameter nominal_mode(HI) != passed_mode(SI) and assign_parms takes function.c:3643 rather than the one-insn emit_move_insn at :3679 — emitting a tempreg (:3665/:3667) and a parmreg conversion (:3669-3673). Target tell: move $sA,$aN then move $sB,$sA, the two halves feeding disjoint use sites, with no sll/sra and no andi 0xffff at any of them. §167-44's extension-based tell cannot fire on this — there is nothing widened to see.

§189-E — THE COMPARE-CONSTANT ROW FLIP: naming the constant changes the comparison's SHAPE, not just its schedule. fold-const.c:4417/4430 rewrites X < CST → X <= CST-1 only when arg1 is an INTEGER_CST. A named local is a VAR_DECL, so the fold never fires and three things flip together: bare literal → li t,0xE0FF ; slt d,t,dzsq ; beqz d (constant−1, operands swapped, materialised at the compare) versus named local → li t,0xE100 ; slt d,dzsq,t ; bnez d (constant verbatim, natural order, materialised at its own def and therefore schedulable into an earlier delay slot). Sharpens §48-C4, which treated this as a scheduling lever only.