36 KiB
§136 — The LOCAL-VARIABLE lever: how many C locals, at what scope (P30 wave 4a, 25 byte-verified banks)
Distilled from the wave-4a index-gap reports (33 drafted, 23 banked whole-binary first pass). The wave's dominant finding, and the reason this section exists as a class rather than a list:
In the 60–120-instruction band, most "regalloc/scheduling residuals" are decided by HOW MANY C LOCALS YOU DECLARE AND AT WHAT SCOPE — not by register pins. gcc-2.7.2 allocates one pseudo per C local;
local-alloc.c:472refuses a local allocno whoseREG_N_DEATHS > 1, promoting it to a global allocno that is ranked by density and loses the low register. So splitting one reused local into two, or merging two into one, moves whole register assignments — deterministically, at zero blast radius. Reach for the local-count lever BEFOREregister __asm__pins.
One report makes the anti-case explicit: for a redundant move $sN,$sM the pin is the wrong lever
— it acquires the register but lets gcc reuse it destructively as the sign-extension scratch. The
right lever was hoisting the assignment above the call (func_80183394).
The splitting/merging rules (each closed a residual, byte-gated)
- One local reused across N arms/repetitions ⇒ SPLIT it per arm. A function-scope pointer used
in 3 if/else arms dies in 3 places, fails the local-alloc gate, becomes a global allocno and loses
the low reg to a block-local constant. Per-arm block-scoped locals fix it in one edit. Symptom:
the same
$v0/$v1pair swapped in ONE arm only, siblings byte-correct. (func_8018C96C; same mechanismfunc_801909D8,func_80183A14.) - A compound initializer holding two values ⇒ SPLIT into two statements when you have one
callee-saved register too many.
x = *(u8*)p << k;makes a load pseudo and a shift pseudo, both live across the call ⇒ two$sregs;x = *(u8*)p; x = x << k;reuses one ⇒ one. Symptom: an extrasw $sNin the prologue, frame size otherwise identical. (func_80184FA4.) - Two variables where you wrote one ⇒ the target keeps a copy you cannot reproduce. An extra
addu $vX,$v0,$zeroright after ajalplus a later copy of the same value = the source hadv1 = f(); rnd = v1;withv1pinned. An unpinned pseudo always coalesces the pair away — verified: removing the pin merged them and cost 27 instructions of drift. (func_8017BF88.) - A local's address in a callee-saved base register ⇒ write a POINTER local, assigned before the
loop and used only inside it. Writing
local.fieldeverywhere addresses$sp-relative, allocates no register and shrinks the frame. Symptom:LENGTH-DRIFTshort by anaddiu $sN,$sp,Kplus one save/restore pair. (func_8017DAC4.) - A symbol read at a constant offset, but built into
$s1bylui/addiu⇒ cache it in a pointer local (u8 *p = D_80078E78;). A directD_xxx[0x1A]folds%loper use, loses the pin and shrinks the frame 0x20→0x18. (func_8017C61C; extends §17.) - Local stack slots are assigned in DECLARATION order, ascending from the outgoing-arg area
(0x10) — independent of use order. Frame-offset drift with correct code is a declaration-ORDER
problem. (
func_8017D2B8; complements §135-6, where the delta is a dead local.)
The type-form rules
- A real
mult $rX,$rYwith a small constant ⇒ the multiplier is a NON-CONST LOCAL, not a literal. A literalx * Kalways goes throughsynth_mult(sll/addu/subu chain). Assigns32 r = K;as its own statement before the first multiply:expand_multthen sees a REG, and CSE cannot fold it back becausemulsi3has no immediate form. Bonus — thelilands in whichever basic block the assignment is in, so its position in the.stells you where to put the statement. (func_80183B04. The inverse of the existing synth_mult entry.) - A negative addend on a narrow field coming out as
li $sN,0xfff0+addu(target:addiu $vN,$vN,-0x10) ⇒ gcc narrowed the whole expression to HImode, where the negative constant is its 16-bit unsigned image and no longer fitsaddiu. Fix: hoist the call to its own statement and put the load+subtract in a block-scopeds32temp. A signed*(s16*)load does NOT fix it (it reassociates); a function-scope temp does NOT fix it either. (func_801840FC.) lhu+sll 16+sra 16+Non a stack local an out-param call wrote ⇒ the local isu16, read as(s16)x >> N— the combiner folds the sign-extendingsra 16into the user shift. Through a PsyQSVECTOR(short vx) you getlh+sra Ninstead. **A scratch vector read back
⚠ RULE 9's CURE IS BYTE-REFUTED — see §197-A (P31 S54).
u16 v[4]andSVECTOR vcompile BYTE-IDENTICALLY in rule 9's own stated context (lhu ; sll 16 ; sra 23in both); the declared type is inert. The residual is real, the tell is right, the cure is a zero-byte asm re-tie. sign-extended-then-shifted must beu16 v[4], NOTSVECTOR.** (func_8017D2B8.)
- An unexplained
addu $vX,$aY,$zerobefore a conditional branch, with the two feedinglhloads in the wrong order ⇒ the value is ans16LOCAL, nots32.LOAD_EXTEND_OPfolds the sign-extend of an already-lh-loaded HImode pseudo into a plain move, giving a second pseudo for the arithmetic while the comparison keeps the original. (func_80183EF8.) LENGTH-DRIFT +1with a narrow load of the SAME stack slot (lh 0x12($sp)besidelw 0x10($sp)) ⇒ gcc-2.7.2 narrowed a memory-operandlocal >> 16into a sign-extending halfword load at +2. Bind the local to ans32temp used twice to force onelw+sra. (func_801834BC.)andi $vN,0xffffright after ajalthat the target lacks ⇒ the TU declares the calleeu16/s16-returning and gcc re-extends the return value. Do NOT change the declaration — call through a cast. This is §135-9 applied to the RETURN axis, not the arguments. (func_8017E26C.)
The scheduling rules (refining §135-2 and §135-4)
- §135-2's
MEM_IN_STRUCT_Plever does NOT apply when the blocking store has a VARYING address.true_dependence()only drops the edge for a non-varying (constant-address) store. If the load must hoist above stores through a different register base, the only lever is source order — assign the load to a temp ABOVE the stores. This is the load-side dual of §135-4. Tell: one extranopin a load-delay slot plus constants landing in$v0instead of$v1. (func_80182F00.) memrefs_conflict_ptreats$sp-based and register-based MEMs as CONFLICTING, so a register load cannot hoist past an$spstore — but two$spstores at different constant offsets ARE disambiguable and reorder freely. Reading store order as literal source order will send you down a wrong path; the load/store base-class asymmetry is the discriminator. (func_80183B04.)- An unfilled load-delay
nopwhere the target fills it with a trailing call's argument setup ⇒ hoist a LOAD, don't chase the arg setup. Split a read-modify-write (*p = *p + 1) intov = *p + 1; … *p = v;so its load rises above an intervening pointer chase; the chase'slwthen fills the slot and the freed arg-setup instructions cascade into the earlier nops. Store order is unchanged, so it is byte-safe. (func_8019064C.) - Prologue
sw $sNstores in REGNO order where the target has DEF order ⇒ the saves are anti-dependent on each register's first def, so emission order tracks def order. Aregister __asm__pin on the incoming parameter turns itsmoveinto a schedulable body instruction that loses priority to the%hiaddress chain and reshuffles the whole prologue. Pin loop variables; NEVER pin the incoming parameter. (func_8018613C.) - An extra induction register (3 IVs where the target has 2) ⇒ do NOT write the second pointer.
Write ONE pointer and address every field as
p + const;combine_givsmanufactures the representative itself. And when the preheaderaddiu rIV,rBASE,Kbuilds the WRONG K, that is the combined-giv ANCHOR choice:record_givprepends andcombine_givstakes the list head, so the last-emitted reference in the body anchors — move the statement whose final reference sits at the target's K to the END of the body. (func_801862A8,func_8018B128; §3-Giv/§70 keyed by symptom rather than by the word "induction".)
The declaration surface (integration, not codegen)
conflicting typesfor aD_symbol whose declaration you cannot find in the split.c⇒ it lives inside aDEFINE_func_*()macro body insrc/shared/engine_core.h. Grep the macro bodies for everyD_symbol your draft names and reuse the canonical type verbatim. In particular an 8-byte-stride table declared there ass32 D_x[][2]must be indexed[i][0]/[i][1]— do NOT declare splat's interior label (D_x+4) as its own extern; that both conflicts and duplicates. Extends §135-7 to the shared-header declaration surface. (func_801854C4.)- A
lui+oripair whose halves are both small (e.g.0x8000A8) and which resolves to no symbol is a PACKED COORDINATE LITERAL, not an address. Confirm against a sibling TU's call. (func_8017CDB0.)
Wave economics (measured, for the next batch's sizing)
33 targets · 46 agents · 4.44 M tokens · 29 min → 29 claimed MATCH, 23 banked whole-binary (70%).
Bank rate by the tier that produced the FINAL draft (derived per-function from the journal + the
gate, NOT read off the workflow's by_tier, which counts claimed matches and therefore sums to 29
rather than 23 — R37):
| tier | banked / attempted |
|---|---|
| Opus direct (≥90 ins) | 10 / 14 |
| Haiku direct (≤89 ins) | 3 / 8 |
| Opus escalation after a Haiku miss | 10 / 11 |
The operative number is the escalation rescue rate: 10 of 11. On a 60–120-instruction pool the
cheap tier closes outright only ~3/8, so Haiku here is a triage stage, not a substitute — it is ≡
Opus at ≤~50 ins (cheap-tier-ab-validated), and this band is above that line. The two-lane shape
still pays because the escalation almost never fails; route ≤50 ins to Haiku and expect to pay for
an Opus pass on most of the 60–120 band.
index_hit was 13 true / 18 false — the index is now the bottleneck the cookbook itself was in
wave 1, which is why the 31 gap reports above are worth more than the matches.
§136a — Blocker capture: classify on the OUTPUT, never on the exit status
The reconcile lane only runs 12/12 because each agent is handed the compiler's own error line (§135, the S29 law). Capturing those lines needs one care point, learned the hard way this session:
make build runs check, so a draft that COMPILES PERFECTLY and merely produces different bytes
also exits non-zero. A capture tool that branches on returncode == 0 to mean "compiled fine ⇒
byte DIFF" therefore has an unreachable branch, and silently files every genuine byte-DIFF under
"unknown". Classify on what the build PRINTED:
| what the output shows | class | route |
|---|---|---|
a non-warning line matching error / conflicting types / undefined |
PLUMBING | reconcile lane — hand the agent the line verbatim |
no compiler error, but [FAIL] / got <sha> / want <sha> |
DIFF | redraft lane — the C is wrong, not the declarations |
| neither | UNKNOWN | investigate; do not route |
Note the filter must exclude warning: lines: the same conflicting types for … text appears as a
warning for built-ins (memcpy) and for benign external-decl mismatches, and those do NOT block
the bank. Only the hard-error form does.
Measured on wave 4a's 10 gate failures: 7 PLUMBING / 3 DIFF. That ratio is why the capture step is worth its ~10 builds before any reconcile fan-out.
⚠️ CORRECTION (earned the hard way — do not repeat my error). I first wrote that this meant "70% of the refusals were paperwork, not codegen." That is wrong, and a reconcile agent refuted it against the bytes. A PLUMBING verdict means only that a declaration conflict EXISTS — the conflict aborts the compile, so the byte question is never reached and the capture says NOTHING about whether the body is correct. Two of the three second-round PLUMBING drafts had a real codegen residual hiding behind the declaration conflict:
func_80188694wasDIFF/4 SCHEDULE-REORDERon the untouched draft (closed with a §21 zero-byte re-tie barrier after six other variants failed), andfunc_8018C638wasDIFF/6 ADDRESSING/cse(closed by hoisting a store above a call). Both agents ranmatch_oneon the unmodified draft FIRST, found the body defect, and said so instead of accepting my premise.So: PLUMBING ⇏ byte-correct. Route it to the reconcile lane, but tell the agent to re-verify the BODY before assuming only declarations are wrong — the wave-4b reconcile prompt's "do not rewrite the body unless you prove it is actually wrong" is the right instruction precisely because it leaves that door open. An agent that rejects your premise is working correctly (§135).
The classification was EXACTLY predictive, which is the point: all 7 PLUMBING banked through the reconcile lane; all 3 DIFF stayed stubs. Wave 4a therefore closed at 30/33 = 91% (23 first-pass
- 7 reconciled), and the reconcile lane is now 19/19 across three waves at ~13× lower token cost
than drafting (329 K for 7 fixes vs 4.44 M for the wave). Every agent found the reported conflict
PLUS a hidden second one cc1 never reached — which is the mechanical reason the "grep the whole TU
in one pass" instruction (§135-8) has to be in the reconcile prompt, not just the drafting prompt. Tool:
.run/s7_capture.py(any overlay, any draft dir; the ov_SC01_077-only.run/uc_capture.pyis its ancestor). It reverts the TU in afinally:— a killed process performs no undo (the S27 law).
Corollary — mandate a PRIVATE scratch path in the agent prompt. A wave-4a reconcile agent chose
.run/s7/scratch/spliced.c (not kept) on its own initiative; a concurrent agent in the same wave overwrote it
mid-run, so its first verification compiled another agent's TU and returned a meaningless rc=0.
It caught the swap only because the emitted .s did not contain its own function. This is the
Phase-28 match_one fake-isolation defect recurring one level up — at the AGENT layer, where no
tool fix reaches it. Any parallel wave whose agents may compile must tell them to use a
process-unique scratch path ($$/pid-suffixed) and must never suggest a fixed shared one. A
shared scratch path does not produce an error; it produces a CONFIDENT WRONG VERDICT.
§136b — A prior wave's "genuine byte-DIFF" verdict is NOT reliable evidence (4 of 4 refuted)
FINAL TALLY: 8 of 8 DIFF-ledgered functions banked on redraft — wave 3's four, plus the three I classified DIFF from wave 4a's capture, plus one from wave 4b. The classifier is not wrong about what it measures ("this draft compiles clean and produces different bytes" is true and useful); it is wrong to read that as "this function resists matching." A DIFF verdict is a fact about one draft.
Phase-30 wave 3 ledgered four functions as genuine byte-DIFF — the class we treat as "real codegen residual, redraft is unlikely to help." Wave 4b re-drafted all four with fresh agents. All four banked. The recorded causes were not codegen at all:
func_801845B0— the prior draft readbeqz $v0, .L8018467Cas an inner early-exit when.L8018467Cis the epilogue, so it hoisted the whole tail out of the enclosingif. A control-flow misread. Resolve every branch TARGET LABEL to its actual instruction before trusting a prior draft's nesting: a branch to the label that beginslw $ra,K($sp)is a RETURN, not a join.func_80184A94— a D2 declaration conflict (D_801BBB78declared scalar in the draft, array in the TU at a line below the splice point). Fixed by copying an already-banked family sibling's declaration forms verbatim (§71 sibling-first).func_8017BEBC— the cached Ghidra seed was an entirely different body; the prior draft had followed it. The.swas the only usable source.func_8018480C— likewise re-derived clean.
The rule this establishes: a DIFF verdict describes the draft that was attempted, never the function's matchability. It is a statement about one attempt by one agent at one moment. So:
- Never retire a target on a DIFF verdict. Route it to the REDRAFT lane, not to a wall ledger.
- A RETRY note must be handed to the agent as a data point, explicitly labelled as one — the wave-4b prompt said "treat that as a data point, not a verdict; re-derive from the .s", and every retry agent did exactly that and refuted it.
- Re-GATING an unchanged draft is not a retry. Wave 4a's three DIFFs stayed stubs through a second gate purely because the same bytes were resubmitted; they still owe a redraft.
- Corollary for the backlog generally:
docs/backlog.mdentries carrying an old closeness/class are stale by construction (Phase-29 measured 77% of stored drafts had decayed). Re-verify before valuing one.
§136c — SIBLING-FIRST is a DERIVATION shortcut, not just a conflict fix (the fastest route in a family wave)
§71/§D1 are written as remedies for conflicting types. Wave-4b agents found their far higher-value
use: before deriving anything from the .s, grep src/shared/engine_core.h's DEFINE_func_*
macro bodies — and the target's own TU — for a byte-verified NEAR-TWIN. In a family wave the
twin usually exists, because that is what a family IS.
Measured instances this wave:
func_801859D8—DEFINE_func_80185978()inengine_core.his a near-twin: identicala0layout, identicalD_801B8748[D_801B8788[*(s16*)(a0+0x70)]][0]chain, differing only in three store values and a trailing call. Reusing its expression forms verbatim reproduced the schedule with no intervention — first-draft MATCH, and the same twin generalizes to the whole 10-member family.func_80184A94— copying an already-banked family sibling's declaration forms verbatim (extern u8 D_x[]used as(s32)D_x; the__asm__data alias) is what made it bank after a prior wave had ledgered it a genuine byte-DIFF.func_8018CB18—func_80180CC0/func_80185C6Cin the same TU fixed the whole tail shape and the u16-compare form before a line was written → first-attempt MATCH.
The rule: a byte-verified sibling is stronger evidence than the decompiler seed AND cheaper than
deriving from the .s, because its expression forms are already proven to produce the gcc-2.7.2
schedule and register assignment you need. Search order for a family target:
engine_core.h DEFINE_* near-twin → same-TU banked sibling → the .s → the Ghidra seed (last:
this session it was byte-proven to be an entirely different body twice).
§136d — Four gcc-2.7.2 levers the redraft lane found (each closed a residual no other lever moved)
These came from re-deriving four functions that had been ledgered "genuine byte-DIFF" (§136b). Each is byte-gated and none was reachable from the symptom index at the time.
-
RC-12, the
$0-add OPAQUE COPY — for a copy-pair whose COMPARE reads the wrong register. Symptom:REGALLOC-PERM, one register, on a copy pair — the target'sbeqzreads the SOURCE's register while every plain-Cb = a;spelling makes the compare read the COPY's. Cause:cse.c make_regs_eqvpromotes the longer-lived copy to canonical andcanon_regrewrites every use. Lever:register s32 zr __asm__("$0"); b = a + zr;— an opaque copy CSE cannot see through. Do NOT pin the copy's source or dest to a real hard register: pinning the source perturbs the prologue's sign-extend temp, pinning the dest lets gcc propagate the hard reg forward and delete the copy — both cost 2 instructions elsewhere. (func_8017E978, the last 1-instruction residual.) -
gcc-2.7.2
jump.cCOLLAPSES an if-then-else into a conditional overwrite.if (c) t = A; else t = B;— both arms single SETs of the SAME pseudo — becomest = B; if (c) t = A;, hoisting the else-arm's%hi/%lopair ABOVE thebeqzand shifting the whole tail (LENGTH-DRIFT -3, 20 mismatches). Symptom to look for: one arm's%hi/%lopair appears BEFORE the branch, and the target'sj-over-arm shape is missing. Lever: write the selector as TWO SEPARATE CALLS, not a ternary/select feeding one call — a call is not a simple SET so the transform cannot fire, and post-reloadcross_jumpthen merges only the common[move $a0,$s0; jal]tail, which IS the target shape. Corollary: a branch-delay slot holdingmove $a0,$s0that sits BEFORE the two arms is the fingerprint of cross_jump tail-merging, not of a hoisted argument. (func_80184494.) -
When a load hoists above a CONSTANT-ADDRESS store of a
D_global, fix the STORE, not the load.true_dependencedrops the edge between a/svarying-address load and a non-/sfixed-address store. Force/sonto the STORE via a COMPONENT_REF:((struct { s32 w; } *)&D_801E7998)->w = 1;— unconditionalMEM_IN_STRUCT_P, address stays constant, identicallui $at/sw %lo($at)codegen. REFUTED axis, recorded so nobody repeats it: §135-2's "reshape the LOAD to an INDIRECT_REF" is the wrong half of the lattice here —*(a[i] + j)still earns/s(a top-levelPLUS_EXPRgrants it) and merely re-folds the symbol into the load's%lo, going 2 → 32 mismatched. (func_80184960.) -
A branchless flag is
-(a != b) & 0xFF, never a ternary.xor / sltu $zero,x / negu / andi 0xFFisstore_flagnormalised to −1 followed by a u8 truncation. Acond ? 0xFF : 0ternary emits a BRANCH and can never reach that shape. (func_8017E978.)
Also confirmed here (§76's inverse, previously unindexed): when a narrow load lands in $v0 but
the target uses a mid scratch register, reuse an EXISTING global allocno as the destination —
but reuse one whose live range ALREADY spans the arm. Extending a SHORT allocno into the arm
lengthens its live range, drops its global.c:594 allocno_compare rank, and swaps two grants
instead of fixing one.
§136e — §136c's PRECONDITION, and two more symptom keys (wave 4b batch 3)
Sibling-first has a precondition, and an agent hit it honestly: func_801899AC's family has
all 13 members still unmatched and no DEFINE_func_801899AC in engine_core.h — so there IS no
byte-verified twin, and the search is a pure cost. Check that a banked sibling exists before
spending the greps; in an all-nonmatchings family, go straight to the .s. §136c is the fastest
route when the family has already been opened, which in a family wave is usually but not always.
Two symptom keys that had no index entry:
-
LENGTH-DRIFT -2, where MINE returns via a bare branch to the epilogue but the TARGET emitsj+addu $vX,$vY,$zero⇒ this is §136-L1 on the RETURN axis. An over-scoped function-level temp became a global allocno and swapped$v0/$v1with the returned local, so the return no longer needed a move. Scope the temp inside the loop. (func_801899AC.) -
A loop increment sitting in the loop-back DELAY SLOT plus a compensating negative
addiuon the fall-through is a SOURCE SHAPE, not areorgartefact — MIPS1 has no annulling, soreorgcannot invent the compensation. Write it asp += 2; if (t == cur) break; … p -= 2;.combine'sreg_n_sets == 1guard is what stops theaddiu -8folding into the followinglw 4($a1). The index's delay-slot entries point atreorg, which is a dead end for this one.
Also demonstrated (composition, func_8017D5F4, 46 ins): flat early-returns instead of nested
ifs to get the cross-jump layout → s32 pad[2] dead locals to sweep the frame size → s16 locals
so each load emits its lh+addu copy pair → mask-first or operand order → three register __asm__ pins on the mask temps (local-alloc otherwise takes $v0/$v1/$a3 and shifts the whole
global assignment) → two zero-byte __asm__ re-ties from the §30 toolkit → and finally
tools/permuter/run_masked.py on a pin-carrying base for the last 2. The toolkit composes; the
permuter is the LAST step on an already-pinned base, not the first.
Second correction to the capture classifier — DERIVE the class, do not pattern-match error prose.
The output-based rule above was still wrong in a third way: it decided PLUMBING by matching a regex
against cc1's diagnostic text, and cc1's vocabulary is open-ended. too many arguments to function 'func_80146C3C' — a plain arity conflict — matched none of error|conflicting|undefined|previous declaration|redeclar, so a trivially reconcilable function sat classified UNKNOWN through two
gate rounds. The closed, true fact is whether the compile produced an object:
compile_failed = bool(re.search(r'^make: \*\*\* \[.*\] Error \d+', txt, re.M)
and 'Deleting file' in txt) # .DELETE_ON_ERROR, added Phase 30 S29
Decide the class from that; keep the diagnostic lines only to hand the agent verbatim. Re-running the
fixed tool over six stubs moved the population from 5 PLUMBING / 1 UNKNOWN to 5 PLUMBING / 1 DIFF with no other change. That is three defects in one small tool in one session — an exit-status
branch that was unreachable, a regex that missed a common phrasing, and the prose-matching design
that made both possible — and each one silently mis-routed real work. R33 in one line: if an
invariant answers the question, never re-parse the output.
§136f — Two declaration sub-cases the reconcile lane surfaced (lane now 15/15 lifetime)
-
A symbol you are calling may be DEFINED — not merely declared — below your splice point.
func_8017CD9C's draft forward-declaredextern void func_8017D540(s32);, guessed from the asm (a barejalwith$a0 = 0and an unused$v0is consistent with several signatures). Butfunc_8017D540is defined in the same TU ~275 lines BELOW the splice, asint func_8017D540(int). cc1 took the draft's prototype first and rejected the later definition. D2's "grep the whole TU below the splice point" must look for DEFINITIONS, not justexternlines. The fix was to copy the definition's own signature; it was byte-neutral because the argument is a literal0and the return is discarded. -
An ARITY clash on the symbol you are DEFINING cannot be fixed by a cast — use lever (B).
func_801848DCis forward-declaredextern s32 func_801848DC(void);at three places above the splice, each used by a banked caller invoking it with no arguments, while the byte-true signature takes a pointer in$a0. Cast-at-use fixes a callee's type; it cannot change how your own definition is declared. The §37/§124 asm-label alias is the lever:s32 aF801848DC(void *a0) __asm__("func_801848DC");— define under the alias identifier, emit under the real symbol, zero blast radius, no shared header touched. In-TU precedent for this exact TU atov_SC04_018_jr_8017AE2C.c:8872(aF8018CB18). Verify the alias is byte-neutral by gating with and without it.
Lane record: 15/15 across five waves. The reconcile lane remains the most reliable stage in the pipeline and the cheapest per bank — but see §136a: it is not purely paperwork, and several of those 15 also carried a real codegen residual behind the declaration conflict.
§136g — When the index points at the WRONG lever: two byte-refuted routings (func_801863B4)
The last redraft of the session is the clearest case yet of why an agent must byte-test the index's
suggestion rather than trust it. It read the real tools/reference/gcc-2.7.2 source and refuted two
entries that the index confidently routes:
1. BRANCH-POLARITY where the return K block RELOCATES to the function tail.
The index routes BRANCH-POLARITY to §3-T4 ("invert the source condition") and §34 (the zero-byte
__asm__("") fence). Both were tried and byte-refuted here. The actual transform is
jump.c:1806 — /* Look for if (foo) bar; else break; */ — which SWAPS range1/range2 and
inverts. It runs long before reorg, so a fence instruction cannot block it. Its real
precondition is label2 = next_label(label1) being the RETURN label with
JUMP_LABEL(range1end) == label2. C lever: put ANY label between the if-join and the return
label — wrap the loop inside the guard, if (…) { loop } with ONE trailing return 0, instead of
an early-return 0 guard. The if-join label disarms the swap. (Closed the last 8 instructions.)
2. lh AND lhu of the SAME address, feeding an sll 16/sra 16 pair — not a weird cast.
MIPS LOAD_EXTEND_OP == ZERO_EXTEND (mips.h:1163), so a plain HImode local load emits lhu and
its later int use costs sll/sra, while an SImode use of the same lvalue (*p == -1) emits
lh. Source form: s16 v = *p; plus a separate *p == -1 test. combine collapses the pair
back into ONE lh unless the HImode pseudo has TWO reaching defs — so the shape only survives with
a hand-rotated guard (v = *p; if (*p != -1) { … do { …; v = *p; } while (*p != -1); }), which is
what makes both loads appear in the guard AND the loop-bottom block.
The generalizable point: the index is a starting hypothesis, not an answer. This agent tried both indexed levers, measured them at zero, went to the compiler source, and found the transform in a pass earlier than the one the index named. Record the refuted routing next to the correct one — otherwise the next agent re-runs the same two dead ends. (§136b tally: 9 for 9.)
§136h — CORRECTION: the zero-crack pool does NOT "refill with cheap work" (my error, byte-measured)
At the S7 close I recorded that the zero-crack (propagation-only) pool grew 120 → 147 families /
62,232 → 70,924 templatable ins even after ~1,400 members were propagated through it, and framed
that as a compounding cheap lever — "run the zero-crack sweep FIRST next session." I priced it off
the family map's byte_weight_templatable and did not probe a single member first.
Measured: the sweep banked 1 of 1,781. Diagnosing the top families (.run/s6_diag.py): two
compile clean and produce a byte DIFF, one fails at link (undefined reference to tail_8012F274). These are genuine per-member residuals — the remapped exemplar body does not
reproduce in the sibling.
Why the pool grows, correctly stated:
- It accumulates members that already failed earlier sweeps (S6a/S6b banked 1,582 out of this same population and left the rest).
- A fresh crack adds its family's members to the pool — but if you propagate behind every crack (which you should), those members are harvested at crack time. What accrues afterwards is the fraction that refused to propagate.
⇒ A growing zero-crack count is a residue signal, not an opportunity signal. Price this pool by
probing one member per family, never by summing byte_weight_templatable — that column counts
what could template if the bodies reproduced, which is exactly the thing in question. This is the
S28 worklist.md mis-pricing (§133) recurring on a different column: a weight column is a
prediction; the gate is the fact.
(Recorded against myself: this is R37 — probe before costing — violated in the scoping step of the very session that was correcting R37 violations elsewhere. The byte-gate cost was ~1,780 build cycles and zero tokens, and nothing wrong entered the tree; the loss was wall-clock and a wrong line in a checkpoint that a fresh session would have acted on.)
§136i — The drafter model LADDER: Haiku → Sonnet → Opus → Fable5 (Drew, 2026-08-03)
The two-tier rule from the 2026-06-29 A/B (cheap drafter ≤~50 ins, Opus for the 90+ tail) left the ~50–120-ins band unassigned, and every wave since defaulted it to Haiku-with-Opus-escalation. P30 S7 measured what that costs:
| tier that produced the FINAL draft | banked / attempted |
|---|---|
| Haiku direct (≤89 ins as routed) | 3 / 8 |
| Opus escalation after a Haiku miss | 10 / 11 |
Haiku on that band was expensive triage — a wasted draft plus a full Opus redraft — not a cheap drafter. The original A/B only proved parity ≤52 ins; everything above was extrapolation.
Route drafters by size:
| band | model: |
|---|---|
| ≤ ~50 ins | haiku (measured ≡ Opus, ~4.8× cheaper) |
| ~50–120 ins | sonnet ← the rung this section adds |
| ≥ ~120 ins, or escalation after any lower rung returns non-MATCH | opus |
| a genuinely NEW wall class nothing else cracks | fable (discovery only — never for applying known idioms) |
Never route Haiku → Opus directly, and never default a whole wave to Opus because the band "looks hard" — that is the same extrapolation in the other direction. Escalation is unchanged: any rung returning non-MATCH escalates one step up. The whole-binary byte-gate remains the sole arbiter, so a weaker drafter is a throughput risk, never a correctness risk (G3/P9).
Treat ~50 and ~120 as current best estimates, not constants — re-measure the boundaries whenever
a wave gives a clean per-tier signal (derive the split per-function from the journal + the gate, not
from the workflow's by_tier, which counts claims — §136).
§136j — The failure MIX flips with function size (measured across four bands, one session)
Blocker-capture classifications from P30 S7/S8, same tooling, same gate, four size bands:
| band | drafted | PLUMBING (declaration) | DIFF (genuine codegen) |
|---|---|---|---|
| ≤60 ins (volume lane) | 111 | majority | few |
| 60–120 ins | 33 | 7 of 10 failures | 3 |
| ≤120 aggregate, second round | — | 5 of 6 | 1 |
| 121–328 ins | 23 | 1 of 7 | 6 of 7 (86%) |
Small functions fail on paperwork; big functions fail on the compiler. The mix inverts almost completely across the range. Two operational consequences:
- Budget the lanes by band. Below ~120 ins, expect the reconcile lane to be the workhorse (it ran 15/15 lifetime and costs ~13× less than drafting). Above ~120 ins, expect the redraft lane and real gcc-source work — reconcile will have little to bite on.
- Do not read a low bank-rate on a big-function wave as a tooling problem. 16/23 (70%) on the 121–328 band with 86% of the failures being genuine byte-DIFFs is the expected shape, not a sign the pipeline is broken. The equivalent 70% on a ≤120 wave WOULD have been a tooling signal, because there the failures should be declarations.
This also re-frames §136a's correction: "a PLUMBING verdict says nothing about the body" is true everywhere, but the prior probability that a failure is paperwork at all is strongly size-dependent.
§136f addendum — the collider is often an ALREADY-BANKED SIBLING below the splice, and you can
locate it by arithmetic. func_801832A8's draft audit scanned only above its INCLUDE_ASM at
TU:4680 and concluded two callees "appear nowhere in the TU". They were declared at TU:4846-4847 —
inside the already-banked sibling func_8018389C, below the splice. The proof is arithmetic, and
it is worth doing before hunting: the draft grows the file by N lines, so a pre-splice TU line L
appears at L+N in the error output. Here N=184, and 4846+184 = 5030, 4847+184 = 5031 — exactly the
two conflicting types lines cc1 reported. If the reported line number exceeds the splice point,
subtract the growth and look there.
And a self-verification an agent can run WITHOUT the gate (stronger than match_one): splice into
a scratch copy of the real TU, run the Makefile chain cpp → cc1 → maspsx → as, then objdump the
function out of the resulting object and compare word-by-word against the target .s. The words that
differ should be exactly the unlinked relocation slots — set-compare the differing indices against
objdump -r, and require the differing-but-not-relocated set to be empty. That proves both the
declaration surface AND the codegen in real TU context (func_801832A8: 237 words, 25 differing,
all 25 relocations). It is the closest an agent can get to the whole-binary gate on its own.