5.5 KiB
§315 — ALL-CONSTANT AGGREGATE FILL: THE EMISSION ORDER IS SHARED-LITERAL GROUPS × DESCENDING INDEX, WITH PAIRED SUB-FIELDS INTERLEAVED (P31 S66; ⚠ UNPROVEN — func_80185214, gate-REFUSED draft, closeness 32→5, never MATCHed)
⚠ STATUS FIRST. This section is filed on a differential measurement, not on bytes. func_80185214 never reached MATCH; its best draft bottoms at {"status":"near","closeness":5,...,"klass":"SCHEDULE-REORDER","bucket":"permuter"}. Treat the ordering rule below as a probe worth two minutes, not a law. No gcc-2.7.2 source cite ties it to a pass — the agent grepped sched.c for birthing_insn_p (§3-C) and tried §3-G block reordering, and both named mechanisms were flat.
THE TELL. A run of li/addiu constant loads into one or two scratch regs ($v0/$v1), each immediately followed by several sh/sw stores into a local struct/array (PS1 POLY_FT4-style vertex/uv prim is the archetype), where the store OFFSETS run counter to naive left-to-right source order — e.g. one constant group hitting 0x28, 0x18, 0x2A, 0x22 = array indices 3, 1, 3, 2 — and where two sub-fields of the same element (uv[i].u then uv[i].v) land as adjacent stores rather than being split apart by field.
THE PROBE (what the sweep selected). When a local aggregate is filled entirely with compile-time constants and several sibling fields share literals, write the source as: (1) one group per distinct literal (all the 0x40s, then all the -0x40s, …); (2) inside each group, walk the array index DOWN — [3], [2], [1], [0]; (3) for paired sub-fields, interleave per index instead of writing whole fields.
prim.v[3].vx = 0x40; prim.v[2].vx = -0x40; prim.v[1].vx = 0x40; prim.v[0].vx = -0x40;
prim.v[3].vy = 0x40; prim.v[2].vy = 0x40; prim.v[1].vy = -0x40; prim.v[0].vy = -0x40;
prim.clut = 0x79;
prim.v[3].vz = 0; prim.v[2].vz = 0; prim.v[1].vz = 0; prim.v[0].vz = 0;
prim.uv[0].u = 0xC00; prim.uv[0].v = 0x180; prim.uv[1].u = 0xC7F; prim.uv[1].v = 0x180;
prim.uv[2].u = 0xC00; prim.uv[2].v = 0x1FF; prim.uv[3].u = 0xC7F; prim.uv[3].v = 0x1FF;
EVIDENCE — a 40-variant combinatorial sweep, differential only. .run/.../sweep2.py (not kept), 5 v-shapes × 4 u-shapes × 2 tails, masked-diff against the real target .s via match_one. Baseline backlog draft (source order + single-expression tail) = 32 (sig":"WIDTH/sw!=lw"). A load-hoist variant = 34 (worse). Splitting the colour expression alone with unchanged write order = 38 (worse still). Then: grp+nat+B = 16, vdsc+nat+B = 10, fdsc+nat+B = 5 (best), against fdsc+base+A = 28, fdsc+natd+A = 28, vasc+nat+A = 28. The three levers are jointly required — swap any one out and it reverts to double digits.
WHAT THE FLOOR ACTUALLY IS, AND WHY IT MATTERS. The surviving 5-instruction window is not the fill body — it is the prologue schedule:
mine: sw $v0,0x50($sp) ; li $v0,0x40 ; sw $ra,0x58($sp) ; lw $a3,0x1C($a0) ; li $v1,-0x40
target: sw $ra,0x58($sp) ; lw $a3,0x1C($a0) ; addiu $v1,$zero,-0x40 ; sw $v0,0x50($sp) ; addiu $v0,$zero,0x40
Per §234 the li 0x40 / addiu $zero,0x40 pair is the same encoding (bit-15-clear ⇒ the signedness dial is inert), so the residual is pure position: the target sinks the first frame store below the $ra save and the $a3 load, and births $v1 before $v0. So the fill-order lever appears to have done its job and the remainder is a separate prologue-window defect — which is the honest reading, and also why the section cannot be promoted: nothing byte-confirms it. From that floor, 484 statement placements of the flags store/load, 57 block permutations, §205 chained forms, §55a/§3-C re-tie fences, §220 param typings and empty-asm barriers all left the number at 5.
BOUNDS AND RELATIONS (read these before reaching for it).
- This is §213's declared blind spot, which is why it is its own section. §213's BOUNDARY says verbatim: "This is about independent stores — no aliasing, no shared value, no call between them." It then routes shared-value runs to §30/§193-D/§194-M — none of which cover an all-constant aggregate fill. §213's own three-step procedure (asm order verbatim → ascending offset → rotate-one-left) was not among the winners here, and its S58b addendum's "a non-monotonic offset sequence in the target is order-faithful" reading is exactly the trap this card fell into at closeness 32.
- Not §145(c). §145(c) is descending emission from one chained statement
a = b = c = 0;. This is descending across separate statements, grouped by literal. - Not §281. §281 groups by OPERATION KIND (copies then RMWs) in a copy-then-adjust block; here every store is a constant and there is no dependent RMW to group.
- The tail-split half of this card is a REDISCOVERY — do not file it.
t = t | (t<<8) | (t<<16); col = 0x808080 - t;beating the single combined expression is §164-80 Law 1 (L13628: an inner op evaluates into an anonymous temp that takes a scratch; make each statement's destination the variable itself) as extended by §167-46 (L16204). Cite those, not this section. - n=1, one shape. One function, one PS1 prim layout,
sh-width stores, no call inside the run. Re-measure before generalising to word stores, to non-aggregate locals, or to any fill with a non-constant term.
Grep bait: all-constant struct fill, prim vertex store order, descending array index stores, constant group store order, uv interleave, POLY_FT4 fill order.