Files
BFM-decomp/cookbook/C0351.md
T

5.5 KiB
Raw Blame History

§315 — ALL-CONSTANT AGGREGATE FILL: THE EMISSION ORDER IS SHARED-LITERAL GROUPS × DESCENDING INDEX, WITH PAIRED SUB-FIELDS INTERLEAVED (P31 S66; ⚠ UNPROVEN — func_80185214, gate-REFUSED draft, closeness 32→5, never MATCHed)

⚠ STATUS FIRST. This section is filed on a differential measurement, not on bytes. func_80185214 never reached MATCH; its best draft bottoms at {"status":"near","closeness":5,...,"klass":"SCHEDULE-REORDER","bucket":"permuter"}. Treat the ordering rule below as a probe worth two minutes, not a law. No gcc-2.7.2 source cite ties it to a pass — the agent grepped sched.c for birthing_insn_p (§3-C) and tried §3-G block reordering, and both named mechanisms were flat.

THE TELL. A run of li/addiu constant loads into one or two scratch regs ($v0/$v1), each immediately followed by several sh/sw stores into a local struct/array (PS1 POLY_FT4-style vertex/uv prim is the archetype), where the store OFFSETS run counter to naive left-to-right source order — e.g. one constant group hitting 0x28, 0x18, 0x2A, 0x22 = array indices 3, 1, 3, 2 — and where two sub-fields of the same element (uv[i].u then uv[i].v) land as adjacent stores rather than being split apart by field.

THE PROBE (what the sweep selected). When a local aggregate is filled entirely with compile-time constants and several sibling fields share literals, write the source as: (1) one group per distinct literal (all the 0x40s, then all the -0x40s, …); (2) inside each group, walk the array index DOWN — [3], [2], [1], [0]; (3) for paired sub-fields, interleave per index instead of writing whole fields.

prim.v[3].vx = 0x40;  prim.v[2].vx = -0x40; prim.v[1].vx = 0x40;  prim.v[0].vx = -0x40;
prim.v[3].vy = 0x40;  prim.v[2].vy = 0x40;  prim.v[1].vy = -0x40; prim.v[0].vy = -0x40;
prim.clut = 0x79;
prim.v[3].vz = 0; prim.v[2].vz = 0; prim.v[1].vz = 0; prim.v[0].vz = 0;
prim.uv[0].u = 0xC00; prim.uv[0].v = 0x180; prim.uv[1].u = 0xC7F; prim.uv[1].v = 0x180;
prim.uv[2].u = 0xC00; prim.uv[2].v = 0x1FF; prim.uv[3].u = 0xC7F; prim.uv[3].v = 0x1FF;

EVIDENCE — a 40-variant combinatorial sweep, differential only. .run/.../sweep2.py (not kept), 5 v-shapes × 4 u-shapes × 2 tails, masked-diff against the real target .s via match_one. Baseline backlog draft (source order + single-expression tail) = 32 (sig":"WIDTH/sw!=lw"). A load-hoist variant = 34 (worse). Splitting the colour expression alone with unchanged write order = 38 (worse still). Then: grp+nat+B = 16, vdsc+nat+B = 10, fdsc+nat+B = 5 (best), against fdsc+base+A = 28, fdsc+natd+A = 28, vasc+nat+A = 28. The three levers are jointly required — swap any one out and it reverts to double digits.

WHAT THE FLOOR ACTUALLY IS, AND WHY IT MATTERS. The surviving 5-instruction window is not the fill body — it is the prologue schedule:

mine:   sw $v0,0x50($sp) ; li $v0,0x40 ; sw $ra,0x58($sp) ; lw $a3,0x1C($a0) ; li $v1,-0x40
target: sw $ra,0x58($sp) ; lw $a3,0x1C($a0) ; addiu $v1,$zero,-0x40 ; sw $v0,0x50($sp) ; addiu $v0,$zero,0x40

Per §234 the li 0x40 / addiu $zero,0x40 pair is the same encoding (bit-15-clear ⇒ the signedness dial is inert), so the residual is pure position: the target sinks the first frame store below the $ra save and the $a3 load, and births $v1 before $v0. So the fill-order lever appears to have done its job and the remainder is a separate prologue-window defect — which is the honest reading, and also why the section cannot be promoted: nothing byte-confirms it. From that floor, 484 statement placements of the flags store/load, 57 block permutations, §205 chained forms, §55a/§3-C re-tie fences, §220 param typings and empty-asm barriers all left the number at 5.

BOUNDS AND RELATIONS (read these before reaching for it).

  1. This is §213's declared blind spot, which is why it is its own section. §213's BOUNDARY says verbatim: "This is about independent stores — no aliasing, no shared value, no call between them." It then routes shared-value runs to §30/§193-D/§194-M — none of which cover an all-constant aggregate fill. §213's own three-step procedure (asm order verbatim → ascending offset → rotate-one-left) was not among the winners here, and its S58b addendum's "a non-monotonic offset sequence in the target is order-faithful" reading is exactly the trap this card fell into at closeness 32.
  2. Not §145(c). §145(c) is descending emission from one chained statement a = b = c = 0;. This is descending across separate statements, grouped by literal.
  3. Not §281. §281 groups by OPERATION KIND (copies then RMWs) in a copy-then-adjust block; here every store is a constant and there is no dependent RMW to group.
  4. The tail-split half of this card is a REDISCOVERY — do not file it. t = t | (t<<8) | (t<<16); col = 0x808080 - t; beating the single combined expression is §164-80 Law 1 (L13628: an inner op evaluates into an anonymous temp that takes a scratch; make each statement's destination the variable itself) as extended by §167-46 (L16204). Cite those, not this section.
  5. n=1, one shape. One function, one PS1 prim layout, sh-width stores, no call inside the run. Re-measure before generalising to word stores, to non-aggregate locals, or to any fill with a non-constant term.

Grep bait: all-constant struct fill, prim vertex store order, descending array index stores, constant group store order, uv interleave, POLY_FT4 fill order.