14 KiB
§155 — hi/lo literal scanning MUST track base registers (S45)
A "find who references address X" sweep that pairs any lui with any later lo16-bearing op in a
window produces PHANTOM cross-references: the lo16 may ride a DIFFERENT base register (e.g.
lui $s2,0x800B … lui $at,0x8019; sw $s1,-0x1798($at) — the window-pairer reads 0x800AE868, the
truth is 0x8018E868). One such phantom steered an evening of MAIN/7 hunting (S45). Track the
register: record lui rt → hi, match only ops whose BASE is that rt (addiu rs==rt / mem-op
base==rt), invalidate on clobber. Register-blind results are candidates for triage only, never
evidence (G3/R14). The corrected pattern lives in the S45 rescan (checkpoint p4 → tools).
§155a — the same failure class, one level up: SHAPE-blind table scanning (S45 p5)
§155's lesson generalizes past instructions. Hunting a data table by its shape alone
("-1-terminated s16 run whose values are all valid indices") produces the identical brand of
phantom, and it will pass a coverage assertion while doing so. The S45-p5 scan re-found its
known-good control table exactly (R32 green) and still returned 664 "tables" across 212 payloads
whose parked-index hits were transparently (offset, count) pair data — [44, 2, 48, 7, 62, 6, 74, 7, ...] "contains 7". The discriminating power of a shape predicate collapses when the value
you are hunting is small and common (a global index of 7 or 9 looks like every other small
integer in the binary).
The law: a coverage assertion (R32) proves the scanner ran over everything; it says nothing about whether the predicate discriminates. Those are two different oracles (R34). Before trusting a shape scan, ask: would a random data region satisfy this predicate? If yes, the scan is a triage filter, never evidence — derive the table from the code that indexes it (register- tracked, §155) instead of from the values it holds.
Cheap test to apply first: compute the predicate's hit-rate on the corpus. 664 hits where the truth is ~1-per-overlay is itself the refutation — a discriminating predicate is rare.
§155b — check the TYPE your oracle returns before comparing against it (S45 p5)
A membership test against the wrong key type fails silently and always, and it looks exactly
like a real finding. corpus.stubs(binary) returns a dict keyed by integer address
(2148696600), not a set of names. Testing "func_80132018" in corpus.stubs(b) is therefore
always False — and it produced, in one session, two confident and completely wrong conclusions:
"none of the wave's matches are live stubs" and "the target pool was never filtered". The pool
was in fact 160/160 and 166/166 correct, and 9 of the matches were genuinely bankable.
This is the R32/R35 family's blind spot: those rules make a tool assert its own coverage and its own correctness, but neither catches an interface mismatch at the call site. A silent always-False comparison has no coverage gap to detect and no instrument to repair — the tool is fine; the caller is wrong.
The law: before using any oracle's return value in a comparison, print one element of it.
print(type(x), next(iter(x))) costs one line and would have caught this instantly. Corollary
for this repo: corpus.stubs is address-keyed — convert with
{f"func_{a:08X}" for a in corpus.stubs(b)} before comparing against names.
Smell test: a membership test that returns 0/N — exactly zero, across the whole corpus — is far more often a type error than a discovery. Real negatives are usually ragged. When a check comes back perfectly empty, verify the comparison before believing the conclusion (R14/R37).
§155c — the ZERO-REFERENCE trap: gcc splits a global-array address across the lui and the LOAD (S46)
§155a's law says a shape scan needs a discriminator, and names the obvious one: require a register-verified code reference to the candidate's address. Run exactly that against the two byte-proved IDXTABs and it returns zero references — and the naive reading ("nothing references these tables") is wrong in the most expensive way, because it looks like a discovery.
The tables are read by gcc's indexed global-array form:
lui $at, 0x8019 ; hi half
addu $at, $at, $a0 ; + index <- $at is WRITTEN here
lh $v0, -0x2844($at) ; lo half, in the LOAD -> 0x8018D7BC
The address exists only as (lui imm, load offset) — the index add sits between them. Any tracker
that invalidates a register when it is written (which §155 correctly demands!) kills $at at the
addu and can never rejoin the halves. So the strictness that makes §155 sound also creates a blind
spot for the single most common way a compiler reads a table.
The fix (in tools/find_addr_refs.py): carry the hi half through an index addu — still strictly
register-tracked, never window-paired — and label what it feeds -indexed so the two shapes stay
distinguishable. That one change took tools/idxtab_map.py from 0/2 controls to 2/2 and produced the
fleet load map (docs/idxtab-map.md).
The law: "no code references X" is a claim about your DECODER, not about the binary, until you have shown the decoder recognises the addressing forms the compiler actually emits. Before believing a zero, hand-disassemble ONE known-good case and check your tracker sees it (§155b's smell test: exactly zero is more often an instrument gap than a fact).
§156 — an ORPHANED reconcile poisons the fleet: dedup_propagate's kept edit (S45 p6/p7)
ATTRIBUTION CORRECTED (R14). This section first blamed
gate_stage's arity pre-pass (finding F1 ofdocs/concurrency-design.md). That was wrong. No arity journal from that session mentionsfunc_80146A6C(checked: 74/26/4 entries), and the arity undo reported success in every log. F1 is real and still worth guarding — it just did not cause this. The commit message ondcc76228bcarries the same wrong attribution; corrected forward here, history not rewritten.
The real mechanism. dedup_propagate --recover's Part B reconciles an overlay's conflicting
caller extern and, when that buys the byte-match, deliberately leaves the edit on disk
(dedup_propagate.py, "keep the reconcile on disk"). That is correct while the function survives.
But a function can still be dropped by a later iteration against a different fail_ov, and when
plan finally empties, the sys.exit("[error] all candidates dropped …") fired with no restore.
Observed live: reconciles kept for ov_SC07_001..009, then everything dropped, then exit — leaving
no-proto'd caller externs for functions that were never propagated →
ov_SC07_010: passing arg 2 of 'func_80146A6C' makes pointer from integer → 141 of 213 binaries
failed check-all.
Fix (landed): a reconcile ledger — every kept reconcile is recorded against its function,
undone the moment that function leaves plan, and all outstanding reconciles are restored before the
failure exit. Proved by tools/test_reconcile_ledger.py: applying a real reconcile for the exact
overlay+fn (35 edits across 18 files) then driving the ledger undo restores all 25 files
byte-identical.
The generalizable law: a tool that deliberately leaves an edit on disk pending an outcome owes a ledger for it. "Keep it if this succeeds" is only half a transaction — the other half is undoing it on every path that can later invalidate the success, including the exit paths.
What it does NOT do: it cannot false-bank. The gate compares against config/check.<bin>.sha
(the original retail bytes, written by no pipeline stage) and INCLUDE_ASM pastes the original
assembly, so wrong C always diverges. The failure is loud and fail-closed — it costs time, never
integrity.
The trap it sets: a broken tree makes EVERY subsequent gate report near. Two batches
(4/4 and 20/20 "near") were read as verdicts about the drafts when they were verdicts about the
tree. A gate result measured on a tree you have not just verified is not evidence (R35).
Standing practice:
- Gate with
GATE_NO_ARITY=1unless you specifically want the arity lane; take the arity-needing drafts through a separate serial pass. Measured cost of the guard: 2 banks of 24 — cheap. - Assert the bracketing invariant after every gate batch:
git status --porcelain src/shared configmust be empty. This is the only cheap detector. - On breakage, do NOT surgically patch:
git checkout -- src/ config/and replay from the on-disk drafts. Replay is deterministic (7/9 and 22/39 reproduced exactly), so recovery costs build time only.
§157 — the cheap-tier size cliff, measured (S45 p6)
Two controlled Haiku waves, same prompt, same pool construction, same independent verification — the only variable was function size:
| band | close rate | tokens/match |
|---|---|---|
| 4–27 ins | 43/50 = 86% | ~44k |
| 30–39 ins | 9/17 = 53% | |
| 40–49 ins | 5/8 = 62% | |
| 50–59 ins | 2/9 = 22% | |
| 60–69 ins | 1/8 = 12% | |
| 70–85 ins | 2/8 = 25% | |
| ≥50 combined | 5/25 = 20% | ~177k (4× worse) |
The documented "Haiku ≤~50 ins" band is optimistic. The cliff starts around 30 and collapses past 50. Route ≤30 → Haiku (unbeatable cost), 30–50 → Haiku only when targets are plentiful, ≥50 → Sonnet.
Agent honesty at the cheap tier is excellent and should be relied on as a FILTER (never as the
gate): across 100 drafters, 63 MATCH claims, 63 confirmed by independent match_one re-runs,
0 false. The one apparent false claim was the verifier's own fault — an -O0-cluster function
(func_8013C360) checked without --o0. Always retry a failed verification with --o0 before
calling an agent wrong (§116).
§158 — The RANGE-EXTENDER: a fourth zero-emission asm lever completes the allocno toolkit (P30 S46 tier-3, func_8017CE58, 733 ins)
Fourth member of the zero-emission-asm family: §148-C moves the priority NUMERATOR (refs),
§47 slides the DENOMINATOR window (+1 slot, no refs), §153 launders an address out of cse,
and this one MOVES A DEATH — __asm__ __volatile__("" : : "r"(var)) placed AFTER var's
natural last use extends var's live range to the asm, growing its reg_live_length by the whole
gap (every insn crossed, +1 each) at the cost of +loop-depth-weighted refs. It reaches the one fork
§47 declares one-directional: a chained order↔registers fork on a USER-VARIABLE pair.
Symptom
"Source order buys the ORDER or the REGISTERS but never both" on a same-source pair (here
mny = wy; my = wy >> 16; — copy/shift of one loaded word straddling the next load), with the rest
of the function matching. SCHEDULE-REORDER/2 by the classifier, but the resolution lives in
global.c, not sched.c.
Why the fork is chained (gcc source, validated insn-by-insn against -dS/-dR dumps)
- sched.c schedules each block BACKWARD; rank = priority → independence-from-last-scheduled
(an insn ANTI-dependent on the just-placed one is class 2 and waits a tick — that is what slots
the next
lwBETWEEN the pair) → INSN_LUID = source order.adjust_priority's REG_DEAD cases are dead code (the???comment is accurate); only the birthing boost runs, and only forreg_n_sets==1pseudos — single-set load temps get boosted, multi-set user vars never do. Net: the pair's final order follows source statement order, full stop. - global.c
allocno_compare:pri = floor_log2(refs)·refs·10000·size / live_length; equal refs → shorter life allocated FIRST → LOWER free register; exact tie → lower allocno = declaration order. Live length is counted on SCHED1's output order (sched.csometimes_live, +1 per insn per segment) — so the earlier-born pair member is always longer → always loses the low register. Order and allocation are chained to the same source order; inside the block the fork is unwinnable.
The method (dump-arithmetic first, then place — no probing)
-dlthe current best: read both pseudos'Register N used R times across L insns. Compute both priorities. Work out the needed inequality (here: extendmysopri(my) < pri(mny)).- The extender adds ~
+depth+1refs and+gaplength; solve for the required L before placing (here refs 46→49 ⇒ need L ≥ 111 from 102 — the first placement gave 110 and flipped back; ONE slot deeper was exact). Knife-edge is normal, the dump re-check costs seconds. - Placement rules: after the pair's rival is DEAD (so only the extended one grows); adjacent to an existing volatile asm (§47's rule — no new barrier); NEVER before a call-crossing gap or a loop-entry (an upward-exposed use makes the var live-on-entry/call-crossing → it loses its caller-saved register entirely and the cascade is catastrophic).
- Expect §47-plateau collateral and compose the levers. The extender's +1 slot shifts EVERY
pseudo spanning it; here it pushed the two loop-invariant
&gszaddress pseudos (refs 7, 560/559) off theirint(140000/L)=250plateau → $s5/$s6 swapped. A bareasm volatile("")immediately after the extender (+1 slot, ZERO refs) put them back on a tie (562/561 → 249==249 → allocno order). One lever per priority relation: refs-carrying extender for the pair, ref-free slider for the plateau.
Bonus facts worth keeping
- USE insns are the mechanism, and gcc plants its own: loop.c leaves
(use (reg))insns in the stream that exist at sched1/alloc time and vanish before final — live length is measurably a function of insns final never emits. The keep-alive asm is just a plantable one. - The sched1-vs-sched2 uid-sequence diff (
-dSvs-dR, ~14 windows on this function) is a fast map of WHERE zero-byte freedom exists; both my failed "swap the min/max arm statements" probes sat OUTSIDE any window and predictably broke bytes (arm order + the y0/y1 $a0/$a1 tie flip together — measured, 6 mismatches). - Declaration order = allocno tie-break is a real lever (
s16 mny, my..won a 103/103 tie in a probe) but it only fires on EXACT length ties; the extender makes the inequality strict instead.
Symptom lines for the index: "order or registers, never both" · "copy/shift pair swapped around a load, registers correct" · "first-born always gets the higher register" · "keep-alive flipped an unrelated $sN pair" (→ compose with §47's bare slider).