Files
BFM-decomp/cookbook/C0166.md
T

14 KiB
Raw Blame History

§155 — hi/lo literal scanning MUST track base registers (S45)

A "find who references address X" sweep that pairs any lui with any later lo16-bearing op in a window produces PHANTOM cross-references: the lo16 may ride a DIFFERENT base register (e.g. lui $s2,0x800B … lui $at,0x8019; sw $s1,-0x1798($at) — the window-pairer reads 0x800AE868, the truth is 0x8018E868). One such phantom steered an evening of MAIN/7 hunting (S45). Track the register: record lui rt → hi, match only ops whose BASE is that rt (addiu rs==rt / mem-op base==rt), invalidate on clobber. Register-blind results are candidates for triage only, never evidence (G3/R14). The corrected pattern lives in the S45 rescan (checkpoint p4 → tools).

§155a — the same failure class, one level up: SHAPE-blind table scanning (S45 p5)

§155's lesson generalizes past instructions. Hunting a data table by its shape alone ("-1-terminated s16 run whose values are all valid indices") produces the identical brand of phantom, and it will pass a coverage assertion while doing so. The S45-p5 scan re-found its known-good control table exactly (R32 green) and still returned 664 "tables" across 212 payloads whose parked-index hits were transparently (offset, count) pair data — [44, 2, 48, 7, 62, 6, 74, 7, ...] "contains 7". The discriminating power of a shape predicate collapses when the value you are hunting is small and common (a global index of 7 or 9 looks like every other small integer in the binary).

The law: a coverage assertion (R32) proves the scanner ran over everything; it says nothing about whether the predicate discriminates. Those are two different oracles (R34). Before trusting a shape scan, ask: would a random data region satisfy this predicate? If yes, the scan is a triage filter, never evidence — derive the table from the code that indexes it (register- tracked, §155) instead of from the values it holds.

Cheap test to apply first: compute the predicate's hit-rate on the corpus. 664 hits where the truth is ~1-per-overlay is itself the refutation — a discriminating predicate is rare.

§155b — check the TYPE your oracle returns before comparing against it (S45 p5)

A membership test against the wrong key type fails silently and always, and it looks exactly like a real finding. corpus.stubs(binary) returns a dict keyed by integer address (2148696600), not a set of names. Testing "func_80132018" in corpus.stubs(b) is therefore always False — and it produced, in one session, two confident and completely wrong conclusions: "none of the wave's matches are live stubs" and "the target pool was never filtered". The pool was in fact 160/160 and 166/166 correct, and 9 of the matches were genuinely bankable.

This is the R32/R35 family's blind spot: those rules make a tool assert its own coverage and its own correctness, but neither catches an interface mismatch at the call site. A silent always-False comparison has no coverage gap to detect and no instrument to repair — the tool is fine; the caller is wrong.

The law: before using any oracle's return value in a comparison, print one element of it. print(type(x), next(iter(x))) costs one line and would have caught this instantly. Corollary for this repo: corpus.stubs is address-keyed — convert with {f"func_{a:08X}" for a in corpus.stubs(b)} before comparing against names.

Smell test: a membership test that returns 0/N — exactly zero, across the whole corpus — is far more often a type error than a discovery. Real negatives are usually ragged. When a check comes back perfectly empty, verify the comparison before believing the conclusion (R14/R37).

§155c — the ZERO-REFERENCE trap: gcc splits a global-array address across the lui and the LOAD (S46)

§155a's law says a shape scan needs a discriminator, and names the obvious one: require a register-verified code reference to the candidate's address. Run exactly that against the two byte-proved IDXTABs and it returns zero references — and the naive reading ("nothing references these tables") is wrong in the most expensive way, because it looks like a discovery.

The tables are read by gcc's indexed global-array form:

lui  $at, 0x8019          ; hi half
addu $at, $at, $a0        ; + index   <- $at is WRITTEN here
lh   $v0, -0x2844($at)    ; lo half, in the LOAD  -> 0x8018D7BC

The address exists only as (lui imm, load offset) — the index add sits between them. Any tracker that invalidates a register when it is written (which §155 correctly demands!) kills $at at the addu and can never rejoin the halves. So the strictness that makes §155 sound also creates a blind spot for the single most common way a compiler reads a table.

The fix (in tools/find_addr_refs.py): carry the hi half through an index addu — still strictly register-tracked, never window-paired — and label what it feeds -indexed so the two shapes stay distinguishable. That one change took tools/idxtab_map.py from 0/2 controls to 2/2 and produced the fleet load map (docs/idxtab-map.md).

The law: "no code references X" is a claim about your DECODER, not about the binary, until you have shown the decoder recognises the addressing forms the compiler actually emits. Before believing a zero, hand-disassemble ONE known-good case and check your tracker sees it (§155b's smell test: exactly zero is more often an instrument gap than a fact).

§156 — an ORPHANED reconcile poisons the fleet: dedup_propagate's kept edit (S45 p6/p7)

ATTRIBUTION CORRECTED (R14). This section first blamed gate_stage's arity pre-pass (finding F1 of docs/concurrency-design.md). That was wrong. No arity journal from that session mentions func_80146A6C (checked: 74/26/4 entries), and the arity undo reported success in every log. F1 is real and still worth guarding — it just did not cause this. The commit message on dcc76228b carries the same wrong attribution; corrected forward here, history not rewritten.

The real mechanism. dedup_propagate --recover's Part B reconciles an overlay's conflicting caller extern and, when that buys the byte-match, deliberately leaves the edit on disk (dedup_propagate.py, "keep the reconcile on disk"). That is correct while the function survives. But a function can still be dropped by a later iteration against a different fail_ov, and when plan finally empties, the sys.exit("[error] all candidates dropped …") fired with no restore.

Observed live: reconciles kept for ov_SC07_001..009, then everything dropped, then exit — leaving no-proto'd caller externs for functions that were never propagated → ov_SC07_010: passing arg 2 of 'func_80146A6C' makes pointer from integer → 141 of 213 binaries failed check-all.

Fix (landed): a reconcile ledger — every kept reconcile is recorded against its function, undone the moment that function leaves plan, and all outstanding reconciles are restored before the failure exit. Proved by tools/test_reconcile_ledger.py: applying a real reconcile for the exact overlay+fn (35 edits across 18 files) then driving the ledger undo restores all 25 files byte-identical.

The generalizable law: a tool that deliberately leaves an edit on disk pending an outcome owes a ledger for it. "Keep it if this succeeds" is only half a transaction — the other half is undoing it on every path that can later invalidate the success, including the exit paths.

What it does NOT do: it cannot false-bank. The gate compares against config/check.<bin>.sha (the original retail bytes, written by no pipeline stage) and INCLUDE_ASM pastes the original assembly, so wrong C always diverges. The failure is loud and fail-closed — it costs time, never integrity.

The trap it sets: a broken tree makes EVERY subsequent gate report near. Two batches (4/4 and 20/20 "near") were read as verdicts about the drafts when they were verdicts about the tree. A gate result measured on a tree you have not just verified is not evidence (R35).

Standing practice:

  • Gate with GATE_NO_ARITY=1 unless you specifically want the arity lane; take the arity-needing drafts through a separate serial pass. Measured cost of the guard: 2 banks of 24 — cheap.
  • Assert the bracketing invariant after every gate batch: git status --porcelain src/shared config must be empty. This is the only cheap detector.
  • On breakage, do NOT surgically patch: git checkout -- src/ config/ and replay from the on-disk drafts. Replay is deterministic (7/9 and 22/39 reproduced exactly), so recovery costs build time only.

§157 — the cheap-tier size cliff, measured (S45 p6)

Two controlled Haiku waves, same prompt, same pool construction, same independent verification — the only variable was function size:

band close rate tokens/match
4–27 ins 43/50 = 86% ~44k
30–39 ins 9/17 = 53%
40–49 ins 5/8 = 62%
50–59 ins 2/9 = 22%
60–69 ins 1/8 = 12%
70–85 ins 2/8 = 25%
≥50 combined 5/25 = 20% ~177k (4× worse)

The documented "Haiku ≤~50 ins" band is optimistic. The cliff starts around 30 and collapses past 50. Route ≤30 → Haiku (unbeatable cost), 30–50 → Haiku only when targets are plentiful, ≥50 → Sonnet.

Agent honesty at the cheap tier is excellent and should be relied on as a FILTER (never as the gate): across 100 drafters, 63 MATCH claims, 63 confirmed by independent match_one re-runs, 0 false. The one apparent false claim was the verifier's own fault — an -O0-cluster function (func_8013C360) checked without --o0. Always retry a failed verification with --o0 before calling an agent wrong (§116).

§158 — The RANGE-EXTENDER: a fourth zero-emission asm lever completes the allocno toolkit (P30 S46 tier-3, func_8017CE58, 733 ins)

Fourth member of the zero-emission-asm family: §148-C moves the priority NUMERATOR (refs), §47 slides the DENOMINATOR window (+1 slot, no refs), §153 launders an address out of cse, and this one MOVES A DEATH — __asm__ __volatile__("" : : "r"(var)) placed AFTER var's natural last use extends var's live range to the asm, growing its reg_live_length by the whole gap (every insn crossed, +1 each) at the cost of +loop-depth-weighted refs. It reaches the one fork §47 declares one-directional: a chained order↔registers fork on a USER-VARIABLE pair.

Symptom

"Source order buys the ORDER or the REGISTERS but never both" on a same-source pair (here mny = wy; my = wy >> 16; — copy/shift of one loaded word straddling the next load), with the rest of the function matching. SCHEDULE-REORDER/2 by the classifier, but the resolution lives in global.c, not sched.c.

Why the fork is chained (gcc source, validated insn-by-insn against -dS/-dR dumps)

  • sched.c schedules each block BACKWARD; rank = priority → independence-from-last-scheduled (an insn ANTI-dependent on the just-placed one is class 2 and waits a tick — that is what slots the next lw BETWEEN the pair) → INSN_LUID = source order. adjust_priority's REG_DEAD cases are dead code (the ??? comment is accurate); only the birthing boost runs, and only for reg_n_sets==1 pseudos — single-set load temps get boosted, multi-set user vars never do. Net: the pair's final order follows source statement order, full stop.
  • global.c allocno_compare: pri = floor_log2(refs)·refs·10000·size / live_length; equal refs → shorter life allocated FIRST → LOWER free register; exact tie → lower allocno = declaration order. Live length is counted on SCHED1's output order (sched.c sometimes_live, +1 per insn per segment) — so the earlier-born pair member is always longer → always loses the low register. Order and allocation are chained to the same source order; inside the block the fork is unwinnable.

The method (dump-arithmetic first, then place — no probing)

  1. -dl the current best: read both pseudos' Register N used R times across L insns. Compute both priorities. Work out the needed inequality (here: extend my so pri(my) < pri(mny)).
  2. The extender adds ~+depth+1 refs and +gap length; solve for the required L before placing (here refs 46→49 ⇒ need L ≥ 111 from 102 — the first placement gave 110 and flipped back; ONE slot deeper was exact). Knife-edge is normal, the dump re-check costs seconds.
  3. Placement rules: after the pair's rival is DEAD (so only the extended one grows); adjacent to an existing volatile asm (§47's rule — no new barrier); NEVER before a call-crossing gap or a loop-entry (an upward-exposed use makes the var live-on-entry/call-crossing → it loses its caller-saved register entirely and the cascade is catastrophic).
  4. Expect §47-plateau collateral and compose the levers. The extender's +1 slot shifts EVERY pseudo spanning it; here it pushed the two loop-invariant &gsz address pseudos (refs 7, 560/559) off their int(140000/L)=250 plateau → $s5/$s6 swapped. A bare asm volatile("") immediately after the extender (+1 slot, ZERO refs) put them back on a tie (562/561 → 249==249 → allocno order). One lever per priority relation: refs-carrying extender for the pair, ref-free slider for the plateau.

Bonus facts worth keeping

  • USE insns are the mechanism, and gcc plants its own: loop.c leaves (use (reg)) insns in the stream that exist at sched1/alloc time and vanish before final — live length is measurably a function of insns final never emits. The keep-alive asm is just a plantable one.
  • The sched1-vs-sched2 uid-sequence diff (-dS vs -dR, ~14 windows on this function) is a fast map of WHERE zero-byte freedom exists; both my failed "swap the min/max arm statements" probes sat OUTSIDE any window and predictably broke bytes (arm order + the y0/y1 $a0/$a1 tie flip together — measured, 6 mismatches).
  • Declaration order = allocno tie-break is a real lever (s16 mny, my.. won a 103/103 tie in a probe) but it only fires on EXACT length ties; the extender makes the inequality strict instead.

Symptom lines for the index: "order or registers, never both" · "copy/shift pair swapped around a load, registers correct" · "first-born always gets the higher register" · "keep-alive flipped an unrelated $sN pair" (→ compose with §47's bare slider).