Files
BFM-decomp/.run/giants/func_80133CD4.fable.md
T
Drew T 6e99157bcf chore(phase-27 T3 addendum): widen the .run/giants allowlist — cookbook §45 cited untracked files
Found while reading the seeds for T1: cookbook §45 names
.run/giants/func_80133CD4.fable.c as its worked example and .run/giants/fable_cd4/
as the flagship's gdb oracle — BOTH were untracked. The docs cite artifacts that
were not in the repo.

- .gitignore: widen by FILE TYPE, not directory — .run/giants/*.{c,md,sh} +
  fable_cd4/*.{c,md,sh,gdb,txt}. +49 files / 460K.
- Now preserved: the flagship func_80133CD4 crack + its gdb oracle (§45's cited
  worked example); the byte-verified pf*.c regression ladder (the seeds' own
  Method/reproducibility section cites it: pf2 78, pf_c2 30, pf_d1 35, pf_h1 280);
  the dump.sh/mon*.sh RTL harnesses; the banked giants' drafts (80135480, 80163EC8,
  80166994).
- Still ignored (regenerable via dump.sh, R33): d_pf*.i.*, *.s, dumps_m*/, and the
  ILS/permuter .log files. Negative control re-verified: all 5 probes IGNORED, no
  db.*.gbf staged (R23).

Lesson (R31 candidate): a doc that cites a path is an untested claim about the repo.
The §45 citation and its file were 4 days out of sync; only reading the seed for an
unrelated reason caught it. Candidate lint: cookbook path citations must resolve to
tracked files.
2026-07-15 17:35:29 -06:00

7.4 KiB
Raw Blame History

func_80133CD4 (399 ins, flagship giant) — MATCH, pin-free, ×134-ready

Result: match_one MATCH (399 ins) + rtu_match --split ov_SC01_077_a MATCH (399 ins) (real-TU-faithful, ambient decls live). Draft: .run/giants/func_80133CD4.fable.c. Baseline was the pin-free structural seed at 378 mismatched (permuter-walled at masked-172). Whole-binary byte-gate = the human's step (G3/P9); no //@EDIT, no ec_edit, zero file-scope footprint, so it should splice + dedup_propagate/family_sweep cleanly.

Pin-free: YES. No register T x __asm__("$N") anywhere. Three asm statements total: the two pre-existing GTE lwc2/swc2/sqr volatile blocks (blessed, real opcodes) + ONE new generic-constraint in-out asm __asm__("lh %0, 2(%2)" : "=r"(h) : "0"(h), "r"(pb0) : "memory") — real opcode, no hard-reg names, same ×134-safety class as the GTE blocks (the §42e SIGABRTs come from fixed-reg pins, which this is not; rtu_match compiles the whole real TU fine).

The levers, in landing order (all byte-verified stepwise)

  1. Merged accumulator variables — the headline (378→147). The target holds $s0 = {call3-result, denom, s0-loop-accum} and $s1 = {call2-result, second-neg, s1-loop-accum} across disjoint regions. K8: global.c has no coalescing, so one hard reg across disjoint regions ⇒ ONE source pseudo ⇒ the original reused one variable per chain. Merging (s0var, s1var) makes both allocnos call-crossing (K4 global.c:917-922) with merged refs ≈19/18 → top density (K2 global.c:594) → they allocate FIRST → plain regno first-fit (K3) reproduces the entire 9-callee-saved permutation including arg0→$s7/copy→$fp (the §43 K&R s16 arg0 double-copy, kept from the seed) and the +addu $s0,$v0 399th instruction. Verified in .greg: dispositions 90→16, 91→17, 92→18, 76→18(arr shares $s2 with s2a), 77→19, 74→20, 75→21, 83→22, 73→23, 72→30.
  2. Block-scoped pointer splits (§44-3) (147→141). Target pc0-regions live in $a0/$a2/$v0 = three source pointers (pw store-group / pc0 main / loop-local pl); pb0 similarly split from the loop-local pb.
  3. The 1-death shared read-temp — the gdb-on-cc1 find (67→13). The accumulator-init reads pb0[0]/pb0[1] serialize through ONE reg with a byte-visible nop ⇒ a shared temp h (anti-dependence). But a plain 2-set h has 2 REG_DEADs → fails local-alloc.c:472's reg_n_deaths==1 gate → GLOBAL allocno → allocated after all locals → the upd2-chain local qty steals $v0 and six caller-saved identities permute (q1/q2/q3, pb0, chain, lw-t). Proof by oracle: gdb-patching reg_n_deaths[h]=1 at local_alloc entry (breakpoint *0x814575b, pointer at 0x82d3a4c) flipped all six to target in one shot (67→46 with everything else unchanged). No pure-C spelling yields 2 sets + 1 death: flow emits REG_DEAD per region (flow.c:2533), REG_UNUSED also counts (flow.c:2101), combine's failed 2-insn merges undo cleanly and its split path needs i1 (3-insn combos only, combine.c:1737), and cse/combine dissolve every 1-set spelling back into independent temps (byte-tested: HI-h, volatile-h(→lhu ✗), s32-h-split, mixed, mirror). The escape is flow.c:2511: no REG_DEAD when the reg is set in the same insn it last uses — expressible only as an in-out asm ("0"(h) ties read2's lh to use+set h in one insn) → 1 death → LOCAL → wins $v0 (priority-10000 tie broken by earlier qty birth) → chain→$v1, pb0→$a0, q1/q2→$v1, q3→$a1 all cascade by first-fit.
  4. upd2 below the reads + "memory" clobber (fills read2's delay slot, not read1's). q3v = s1var * pc0[2] / s0var; named before the reads (semantics: s1var still holds the neg), pb0[2] += q3v; after; the mem-clobber on the in-out asm stops the chain-lhu hoisting above it → nop stays at read1's slot, lhu fills read2's.
  5. Tail: branch polarity off the opcode (§32#2, both y >= -0xBCB and the *(u8*)cmd return); per-element serialized store groups (loads can't cross /s struct-member stores — ((H16*)&D_x)->h = v keeps the dep that a plain scalar-global store drops); single w local for the s3u[1] one-load-two-stores pair; early bp pointer local (fills the first delay slot); goto-shared-ret1: — giving the common return 1 its own (jumped-to) BB stops sched1 hoisting the li v0,1 into the last store pair cross-BB, which also frees $v0 for the last lhu temp; dbr still steals the li into the bnez delay slot.
  6. The offset-0 /s store asymmetry — last 5 diffs. p[0] = x expands non-/s (mem (reg)) while p[1]/p[2] are mem/s ⇒ the reload-born $t0 asm-operand load kept a true-dep on S0 only (visible in the sched2 dump: insn 763's dep-list contains 401; §30/sched.c:820 drop-clause needs /s+varying vs non-/s+fixed) and parked in the SECOND lh-delay gap. ((struct { s32 w; } *)pw)->w = s3[0]; forces /s at offset 0 → dep dropped → the lui/lw pair floats to the FIRST gap = target.

Method notes (§34 flywheel)

  • -dS/-dR give full sched1/sched2 traces (gccdump.sched, .sched2) — the sched2 dump's per-insn dependence lists on reload-born insns (763) handed over lever 6 directly; read those before hand-modeling.
  • gdb oracle pattern: when a hypothesis reduces to one compiler-internal quantity, patch it mid-compile and diff the output (break *local_alloc; set $nd[245]=1) — one run converts "plausible" into "proven root cause" and licenses spending on the C-form search.
  • The find_free_reg tracer (ffr2.gdb) works symbolically except qty_first_reg — use the pointer at 0x82c5404 (the info address symbol 0x82c496c is stale for this binary).
  • Dumps/traces/variants preserved in .run/giants/fable_cd4/ (dumps_seed/m1..m5, ffr traces, oracle.s, variant files).

§44 cookbook distillation (proposed)

Lever 6 — the merged-variable permutation-breaker + the 1-death local-alloc gate (func_80133CD4, 399 ins ×134). (a) When a target holds ONE $sN across disjoint value-regions (divisor→denominator→loop-accumulator), global-alloc's no-coalescing law (K8) means the original reused one C variable; merging makes the allocno call-crossing (K4) + top-density (K2) and the whole callee permutation falls out by first-fit — check for reused-variable chains before concluding "unsteerable whole-function permutation". (b) A serialized shared read-temp (lh/lh into one reg + nop) is a 2-set variable, which local-alloc REJECTS (reg_n_deaths==1, local-alloc.c:472) → it goes global and loses the low-scratch first-fit to any block-local qty → caller-saved permutation. No pure-C form gives 2 sets/1 death; the pin-free fix is the in-out asm read __asm__("lh %0, 2(%2)" : "=r"(h) : "0"(h), "r"(p) : "memory") — use+set in one insn suppresses the region-1 death (flow.c:2511 dead_or_set), making h a LOCAL qty that wins $v0 by qty-birth tie-break; the "memory" clobber doubles as the fence that keeps a following load in the later delay slot. (c) p[0]=x stores are non-/s while p[k≥1] are /s — a fixed-address load's dep survives only against the offset-0 store; ((struct{s32 w;}*)p)->w = x /s-ifies it to release the load into the first delay gap (store-side twin of §37's load-side /s lever). (d) A shared return 1 reached by goto ret1 gets its own BB — stopping the return-li from filling a last-element load-delay slot cross-BB (and dbr still steals it into the branch slot).