Commit Graph

114 Commits

Author SHA1 Message Date
Christopher Williams eeaaf5b059 phase11: merge 58 — worker E's 0x80027D88 -> 598 bodies / 607 regions, TWO to the milestone 2026-09-24 11:30:55 -04:00
Christopher Williams 7a93946795 phase11: merge 57 + cookbook 170-171 — 597 bodies / 606 regions, THREE from the milestone
Worker E's 0x8005E17C and 0x8002FAB8; worker F's first two claims 0x800FBE84 (216 B, FIRST
SPELLING with worker A's derivation) and 0x80026274 (108 B).

170 generalises the argument-evidence levers (157/164) into a mechanism: a redundant ENTRY-BLOCK
copy of an argument means that value is still live at a call whose argument setup CLOBBERS that
same register. The copy is materialised in the entry block because the tie to a0's home is
illegal. The test that nailed it: the same body with a 2-arg call is 104 B LENGTH-MISMATCH; with
the 3-arg call it is 108/0. Two prior corpus instances had the copy AT the call; this is the
hoisted-to-entry variant.

171: worker F confirmed EXHAUSTIVELY that the constant-division divisor is unique per magic --
(n*M)>>(32+s) == n/D has exactly one D. So finding 67's identity is not an approximation.
2026-09-24 11:27:41 -04:00
Christopher Williams c091483083 phase11: merge 55 + cookbook 123 SOLVED + 167-168 — 592 bodies / 601 regions
Worker E's 0x8002311C (160 B) CLOSES COOKBOOK 123'S OPEN QUESTION. Finding 123 recorded the
branchless MAX0 (x & -(x > 0)) as unreached -- 'no ternary and no bitwise spelling reached it'.
Worker E solved it: the lever is NAMING THE BOOLEAN.

  return s & -(s > 0);        -> BRANCHES
  return s > 0 ? s : 0;       -> branches
  flag = s > 0; return s & -flag;  -> EXACT (slt / negu / and)

Mechanism: naming the comparison forces cc1 to materialise it as a VALUE (slt) rather than a
test feeding a branch. That is finding 44's 'name the boolean' lever applied to the MAX half --
finding 44 previously had only the cond-into-&& direction for this family.

167: a 4-byte store cc1 DELETES means the object's address is never taken -- fold the word into
the array whose address IS taken by a call.

168: s = f(); s += f(); s += f(); loses one instruction vs three named results summed.
2026-09-24 11:22:01 -04:00
Christopher Williams db6022c9f7 phase11: merge 54 + cookbook 166 — 590 bodies / 599 regions
0x800F3DC0 (88 B) — a ONE-WORD sibling of the matched 0x800F3E18, found by worker E via
sf3_family at ratio 1.000 and confirmed by raw-word diff: identical in all 22 words except the
COP2 command field (0x4B70000C vs 0x4B78000C). The route was one copy, two renames and one field
change; every __asm__ and register binding carried over untouched.

166 records it, and notes it is the MIRROR of finding 161: on 0x800F3E18 the field 0x178000c was
the WRONG answer (one byte off, 0x170000c correct); on 0x800F3DC0 0x178000c IS correct. A count
tells you a field is COMMON, not that it is right -- and a ratio-1.000 sibling is the cheapest
place to learn which one a row wants. When the family tool reports one, diff the raw words FIRST.
2026-09-24 11:18:58 -04:00
Christopher Williams f4e14569bd phase11: merge 53 — worker E's 0x800B0E64 -> 589 bodies / 598 regions 2026-09-24 11:13:40 -04:00
Christopher Williams 22cfdc874d phase11: merge 50 — worker E's 0x80107DE8 -> 587 bodies / 596 regions 2026-09-24 11:06:56 -04:00
Christopher Williams a74bc32306 phase11: merge 49 + cookbook 162 — 586 bodies / 595 regions
Worker A's final row 0x8010A6C4 (132 B, first attempt, maspsx=epilogue) -- its ninth epilogue
row and its 46th claim.

162: a callee called with DIFFERENT argument counts needs a NON-PROTOTYPE declaration --
func_8010A444(1) / (2, x) / (3, s1, s0) is only expressible as 'void func_8010A444();', the C89
empty-parameter form, not '(void)'. Same constraint that cost worker A a compile on 0x8002DD14.

Worker A's final totals: 46 claims (33 first-attempt), 95 evidence rows, 39 levers, 3 deferred
rows with derivations, 1 blocked row, 9 rows carrying maspsx=epilogue.
2026-09-24 11:05:36 -04:00
Christopher Williams f27691f6c3 phase11: merge 47 + cookbook 158 — 584 bodies / 593 regions
Worker A's 0x800F4B88 (128 B, first attempt) -- its eighth epilogue-class row and its last.

158: two type views over the same halfwords are DELIBERATE. The first helper call loads with lh
(signed) and the second with lhu (unsigned) over the SAME pointer, so the source declared a
short* view for one expression and an unsigned short* view for the other. Writing the whole row
as short* gives lh for the second call too and changes the bytes. When one function reads the
same field both ways, the mixed lh/lhu pair over one pointer is the evidence.
2026-09-24 11:01:49 -04:00
Christopher Williams eb4b23b219 phase11: merge 46 + cookbook 155-157 — 584 bodies / 593 regions
Worker A's three epilogue rows (0x800F452C, 0x800F6DD0, 0x800F6E50).

155 is a DISPATCH finding: the epilogue list is ALSO a family list. 0x800F6DD0 and 0x800F6E50 are
siblings differing in exactly two ways, and worker A read one and got the second for free, both
first try. Adjacent pairs already identified: 0x800F6DD0/0x800F6E50, 0x800F42AC/0x800F452C,
0x800FFFEC/0x80100038. A worker taking an epilogue row should read its NEIGHBOURS first -- the
class was selected on a TAIL SHAPE, and tail shape correlates with the translation-unit layout
that makes neighbours siblings. Generalised: any class selected by a structural feature clusters
its results by address.

156: the three writes are ASSIGNMENTS not accumulations -- the original never loads the old
destination value, so writing += adds three loads.

157: fewer argument registers set than parameters means the source passes its OWN LIVE parameters
directly. Now confirmed on three rows.
2026-09-24 11:00:24 -04:00
Christopher Williams 57c1cd22f2 phase11: merge 44 + cookbook 151-153 — 580 bodies / 589 regions
Worker A's 0x800F4098 and worker D's 0x800FB54C (104 B, first attempt, maspsx=epilogue).

151: the 2^k-1 add-back rule is CONFIRMED on two independent divisors -- worker C derived it
from 63 (0x800FEE3C) and worker D found it again on 127 (0x800FB54C, magic 0x81024409). Same
structure, two divisors, so finding 67's decision table is complete and not hypothesised.

152: FIVE finders each produced bodies over worker D's 20 -- redundancy rank 6, size rank 5,
adjacency 4, epilogue class 2, constant-division census 1, family 1. No single finder dominates.
This broadens finding 109: 'five different finders each produced bodies, and the price was set
by the LEVER, not the finder.' The tools cover different parts of the population, so keep every
finder running rather than consolidating onto the current best.

153: a saved register can force a local to be SMALLER than the data written through it, and
enlarging it to fix that breaks the frame.
2026-09-24 10:55:29 -04:00
Christopher Williams 150e672b5a phase11: merge 42 + cookbook 144-146 — 574 bodies / 583 regions
Worker D's 0x800FAF84 (104 B), its first maspsx=epilogue match.

144: THE EPILOGUE CLASS NEEDS ONLY ONE TOKEN. Worker A asked for a second one; it does not need
it. The 120 rows split into two shapes -- A) lw $31 immediately before the release, which needs
the release moved AND a nop inserted after lw $31; B) other loads in between, where the release
moves and the trailing nop is DROPPED. My first implementation did A only and left every B row
4 bytes long. Verified on both: 0x800FFBEC (80/0/MATCH) and 0x800F44D0 (92/0/MATCH, a row worker
A had released as unfixable).

145: read the frame arithmetic and the saved-register offsets TOGETHER -- worker D's local had to
be 8 bytes not 12 because the saved s0 sits at sp+24 and the callee writes through sp+16. Third
instance of the size family, first where the constraint came from a saved register.

146: the SAME expression at two divisors produces two unrelated code shapes (/64 branchy bias vs
/63 add-back magic), which is why worker D's divisor sweep missed it.
2026-09-24 10:52:37 -04:00
Christopher Williams dce9c40896 phase11: merge 41 — 0x800F44D0 closes on the CORRECTED transform (shape B) — 573 bodies / 582 regions
Worker A released this row as 'the three-load-with-nop shape that the swap cannot fix' and
requested a --no-load-delay-nop token. It does not need one: the corrected transform already
handles it, by DROPPING the trailing nop for shape B rather than moving it. Verified: 92 bytes,
differing_bytes=0, MATCH, with maspsx=epilogue. It was 96 bytes before the fix.

So the epilogue class does NOT need a second token -- it needed the transform to distinguish the
two shapes, which worker A's report 24 is what revealed. The 120 rows should now be attemptable
with maspsx=epilogue alone.
2026-09-24 10:51:28 -04:00
Christopher Williams ea51ac9629 phase11: merge 40 + the epilogue transform now handles BOTH shapes — 571 bodies / 580 regions
Worker A's two epilogue-class rows (0x800F42AC 96 B, 0x80100038 104 B), both carrying the
maspsx=epilogue token -- the first rows closed through the new mode.

AND THE TRANSFORM IS NOW CORRECT FOR BOTH SHAPES, which worker A's report 24 showed was
necessary. The 120 rows split:
  A) lw $31 IMMEDIATELY before the release -> the release moves into the slot AND a nop must be
     inserted after lw $31, or j $31 lands in its load-delay slot.  0x800FFBEC.
  B) other loads between lw $31 and the release -> the release moves into the slot and the
     trailing nop is DROPPED; no load-delay nop is needed.  Worker A's 0x800F44D0.
My first implementation did A only and left every B row 4 bytes long. Both are handled now, and
the discriminator is whether the jump's own register was loaded immediately before the release.

A BUG WORTH RECORDING: reading out[-1] to find that preceding instruction saw maspsx's own
'#nop # DEBUG: ...' comment instead of the lw, silently producing the shape-B answer for a
shape-A row and turning a MATCH back into a LENGTH-MISMATCH. The scan now skips comments.
2026-09-24 10:50:09 -04:00
Christopher Williams a72a8d4127 phase11: merge 39 — 4 rows (A's 0x80038D48, 0x800F6D60; D's 0x8009B56C, 0x80018458)
Worker A's two: the addition operand-order row (a1[i]+a0[i] vs a0[i]+a1[i] -- same length,
16 bytes apart, because cc1 evaluates the right-hand operand first) and the unconditional
p[0]=0 that lands in a branch delay slot.

Worker D's two: 0x8009B56C closed on cookbook 43 trigger 1 after D had nearly written the row
off, and 0x80018458.
2026-09-24 10:44:06 -04:00
Christopher Williams 05be974ce2 phase11: THE EPILOGUE POST-PASS SHIPS (maspsx=epilogue) — 565 bodies / 574 regions
Finding 84 named the transform; it is now implemented and 0x800FFBEC matches (80 B, 0 differing)
where it was 6 differing bytes without it.

IT IS A SWAP, NOT A MOVE, and getting that wrong cost one implementation: the candidate is
lw $31,16(sp) / addiu sp,sp,24 / jr $31 / nop and the original is lw $31 / nop / jr $31 /
addiu sp,sp,24 -- SAME instruction count, two words swapped. My first version moved the release
after the jump and dropped the nop, producing 3 instructions instead of 4 and turning an 80-byte
row into a 76-byte LENGTH-MISMATCH. A 'small mechanical transform' still has to be checked
against the bytes.

SCALE: 120 unclaimed rows have the filled epilogue in the ORIGINAL (scanned every worklist row's
tail for jr $31 followed by a positive addiu sp,sp,N). They are mostly SMALL -- 76, 76, 80, 92,
92, 96, 104 B -- so this is a large class of cheap rows that were blocked on a HARNESS GAP rather
than on source shape. 770 other rows have the unfilled shape and need nothing.

The tracked patch is regenerated and verified to reproduce both modified maspsx files from the
pristine checkout.
2026-09-24 10:42:29 -04:00
Christopher Williams 686e906b97 phase11: merge 37 + calibrate sf3_family + cookbook 136-137 — 563 bodies / 572 regions
Worker A's 0x80036F70 (460 B, first attempt, family score 1.000 AND adjacent to its own
0x80036DA4). Its family run finished 7 for 7 with five first-spelling matches.

136: worker A CALIBRATED the family tool. It checked the two 0.97-scoring entries and NEITHER
shares its sibling's body at all -- one is a table-allocation routine, the other a summing
loop. '1.000 is the useful band; below ~0.99 the histogram is matching common idioms, not
bodies.' That is the same false-positive mode as the redundancy ranker (finding 110). The
default threshold is now 0.99.

137: a family's signature can be a CONSTANT TRIPLE -- worker A's 0x80036F70 differs from its
sibling only in six constants, whose signature is (A, A+12, A-58). Searchable in a way no
similarity metric can be, because the shapes are identical and only the immediates differ.
2026-09-24 10:37:35 -04:00
Christopher Williams bd3619d41e phase11: merge 36 — FIVE family-list rows in one pass -> 561 bodies / 570 regions
Worker A closed 0x800259A0, 0x80012918, 0x8006B2D4, 0x8003022C and 0x800506E4 -- every one a
sibling found by tools/sf3_family, which was built an hour ago from worker D's insight that
'the finder varies, the price does not' and therefore families should be SEARCHED for rather
than waited for.

That is the tool's first harvest and it is 5 bodies from one list. The family scores were
1.000/1.000/1.000/1.000/0.998 -- exact opcode-histogram and size matches against rows worker A
had already matched, so the levers transferred unchanged.
2026-09-24 10:35:06 -04:00
Christopher Williams 887155a733 phase11: merge 35 + 5-way re-partition — 556 bodies / 565 regions
Worker D's 0x80025ADC (136 B). Partitions re-interleaved 5 ways because workers B and C
are both at ~94% context and effectively exhausted, leaving 2 active workers against 45
remaining bodies. A fifth worker restores capacity.
2026-09-24 10:32:46 -04:00
Christopher Williams bc4c046625 phase11: merge 34 + cookbook 127-129 — 556 bodies / 565 regions
Worker C's 0x8009F4B4 (248 B) and 0x80068874 (156 B), both first attempt.

127 CLOSES WORKER D'S OPEN QUESTION. D left 0x800FEE3C's magic 0x82082083 unexplained; worker C
solved it and the answer is a general rule: the divisor 63 is of the form 2^k-1, which is why
cc1 uses that magic with an ADD-BACK (mfhi; addu; sra 5) instead of a plain shift. An add-back
magic is the tell for a 2^k-1 divisor, NOT for a large one. Finding 67's decision procedure is
now complete: no mflo -> constant division D = 2^(32+s)/M; mfhi+addu+sra -> a 2^k-1 divisor;
mfhi AND mflo -> a genuine 64-bit multiply.

128: worker C classified a division-by-constant row on decode WITHOUT attempting it, because
'every division expression has several equally-plausible spellings, so it is idiom-redundant by
construction'. That characterises the ranker's false-positive class from the SOURCE side for the
first time -- exactly the class finding 110 showed cannot be separated by operand comparison.

129: adjacency is now 9-for-9 across three workers (A 3/3, C 5/5, D 1/1).
2026-09-24 10:30:10 -04:00
Christopher Williams a3b5db4f61 phase11: merge 32 — worker A's three roving-list rows -> 555 bodies / 564 regions 2026-09-24 10:27:40 -04:00
Christopher Williams bbe342d5dd phase11: merge 31 — worker D's 0x80018210 -> 552 bodies / 561 regions 2026-09-24 10:26:31 -04:00
Christopher Williams bac9ab0d89 phase11: merge 30 + cookbook 120-122 — 550 bodies / 559 regions
Worker A's 0x800689DC and worker C's 0x8009F4B4.

120: worker B's justification for why the ranker works -- 'the allocator makes copies
non-identical, so OPCODE repetition survives while WORD repetition does not'. That is exactly
why finding 110's full-word metric failed and why the opcode metric works. A repeated source
block produces the same opcodes with different registers; requiring operands to match destroys
the signal rather than sharpening it.

121: the filled-delay-slot class has TWO sub-cases with DIFFERENT fixes -- reorg fills the slot
(source-shape hunt) versus maspsx mode changing WHICH instruction lands in the slot (a harness
token choice). Same diagnostic, different remedy. Check whether toggling maspsx changes the
fill before hunting a source shape.

122: NEW BLOCKED CLASS -- a GTE coprocessor body needs a harness token, not more spellings.
Worker B's 0x8001FAFC reads mfc2 $12/$13/$14 and branches on t7/s6 which are NOT the o32
argument registers, so the inputs arrive through a non-standard convention. Team rule: if a
body contains mfc2/mtc2, do not spend spellings on it -- these are tooling-blocked rows to be
worked as a batch once a token exists.
2026-09-24 10:24:53 -04:00
Christopher Williams e2bdada67b phase11: merge 29 + cookbook 105/119 — 548 bodies / 557 regions
Worker D's 0x80106AA8 (136 B, first attempt) -- found by the REDUNDANCY filter, not
adjacency, which is the first row where the ranker did the finding alone. Eight stores
through four global pointers, each re-materialised per store.

Cookbook 105's dial now has THREE measured settings: per statement (0x8006BC74 46x and
0x80106AA8 8x, both matched), once per block (matched), once per function (does not match).
So per-statement re-reads are the NORMAL shape, not an extreme.

119: worker D ran the fragment check, called 0x80058BA0 a confirmed fragment, then
SELF-CORRECTED -- it is legal, because in o32 a frameless leaf may both read and write the
caller's outgoing argument area (sp+0..sp+31). All three of D's suspects are legal. Worker B
found the read side, worker D the write side, and both had to read the row to do it: a
heuristic keyed on shape must state its exclusions, and only the worker reading the row can
find them.
2026-09-24 10:23:08 -04:00
Christopher Williams fc7ebc4b2d phase11: merge 28 + cookbook 115 — 547 bodies / 556 regions
Worker B's 0x80050CA8 (120 B, first attempt).

Lever: the status word is masked by TWO separate statements (&= -3; &= -5;), and the original
emits one load, two ands against two different constants, one store. Combining the masks
folds to a single and and LOSES an instruction -- the same principle as finding 81 (a slot
stored twice is two statements) applied to read-modify-write. Companion: the status load is
hoisted above nine halfword clears, so the clears' source order is only observable through
the store order.
2026-09-24 10:20:18 -04:00
Christopher Williams b7ccad87a5 phase11: merge 26 — 545 bodies / 554 regions
Worker A's 0x8006B7C0 (420 B) and worker D's 0x800307FC (92 B).
2026-09-24 10:14:45 -04:00
Christopher Williams 14e927fca6 phase11: merge 25 — 543 bodies / 552 regions
Worker A's 0x800914E4 (400 B), closed on the row assigned under the revised picking order
(adjacency first, then redundancy, preferring the smaller of similar-scored rows).
2026-09-24 10:13:18 -04:00
Christopher Williams 6b7239d83d phase11: merge 24 + cookbook 104 + workflow protocol — 542 bodies / 551 regions
Worker A's 0x80033DC8 (360 B) and 0x8006A98C (132 B), plus two gp symbol rows
(D_80122724, D_80122728) that unblock 0x800A4CA8.

PROCESS DEFECT FOUND AND FIXED. Worker C and worker D both matched 0x800320D8
independently and D overwrote C's source file. Nothing corrupted -- both spellings match
and the region still reports 276/0/MATCH -- but one worker's effort was duplicated. The
partitions are genuinely disjoint (273/279/278/271, union 1101 = sum), so there was NO
assignment error: the gap was that no worker could know another had started a row, since
the registry only knows about MERGED claims and both started before either merged. The
root cause is the adjacency rule (cookbook 99, 4-for-4) crossing partition boundaries --
the best dispatch heuristic found so far invalidated the assumption the assignment rested on.

Fix: .run/p11/inflight.tsv (write-ahead log alongside the merge registry's commit log),
with the protocol written up in docs/ORCHESTRATOR_WORKFLOW.md so the next orchestrator
inherits it.
2026-09-24 10:07:56 -04:00
Christopher Williams 73b706e667 phase11: merge 22 + cookbook 99-102 — 538 bodies / 547 regions
Worker A's 0x800556E8 and 0x80055654 (both first/second attempt).

Cookbook 99 is a DISPATCH rule, not a codegen one: take the row ADJACENT to one you just
matched. The binary is laid out by translation unit, so neighbours share the author's habits.
3 for 3, all first or second attempt, and it beat both the size ranker and the LRS ranker.

Cookbook 100 puts the three branch-shaped diagnostics side by side -- each maps a residual
shape to exactly one cause and each is a glance rather than a spelling:
  branch displacement words only -> block NESTING (95)
  first few instructions, right length -> then/else ORDER of a single-statement arm
  whole prologue, same multiset -> declaration vs assignment order
Vector copies are now confirmed on FIVE independent rows.
2026-09-24 09:51:10 -04:00
Christopher Williams 2472e2e2c1 phase11: merge 21 + cookbook 95-98 — 537 bodies / 546 regions
Worker A's two adjacent claims (0x80036B14, 0x80036DA4).

Cookbook 95 is the cleanest diagnostic of the phase: a correct-length candidate whose residual
is a handful of BRANCH WORDS means the block NESTING is wrong, not the code inside the blocks.
Worker A got exactly 656 bytes (correct length) with exactly 2 differing bytes, both branch
displacements, by writing two guards as siblings instead of nested. Residual = 2 bytes at a
branch displacement => go look at your braces.

Also: the project's 4-int vector type is identifiable from the frame (multiple of 16 with
offsets stepping by 16); vector copies are struct assignments (third independent confirmation);
and an OPEN question is recorded -- 'the original spills everything, cc1 promotes' -- with a
request for a recipe from any worker who has solved it.
2026-09-24 09:49:26 -04:00
Christopher Williams a1c41771e7 phase11: merge 19 + cookbook 92-94 — 535 bodies / 544 regions
Worker B's four first-attempt claims (0x800B255C, 0x8004857C, 0x80030858, 0x800909D8).

Cookbook 94 is the strategic one: worker B's failures cluster into exactly TWO mechanical
classes -- reorg slot-fill choice and rare-epilogue fill -- and neither is a shape problem.
Both are the post-pass family, which two workers have now independently arrived at and
stopped on. That is the strongest argument yet for writing the post-pass rather than
grinding these rows with source spellings.
2026-09-24 09:46:48 -04:00
Christopher Williams 6862af1d0b phase11: merge 17 + cookbook 84-86 — 526 bodies / 535 regions
Worker B's 0x80069580 (88 B) and 0x8007E7FC (96 B), plus two gp symbol rows
(D_80122168, D_801221D0).

Cookbook 84 is the harness row for the post-pass: worker B isolated the rare-epilogue
transform exactly (move the frame release into the jump slot AND insert the load-delay nop
after lw ra), and established the load-bearing detail that as will NOT perform this fill
because doing so would put jr ra in the lw ra load-delay slot. So a post-pass that merely
moves the release into the slot produces wrong code. Also measured: maspsx=off is WORSE on
this row (72 bytes) because it strips nops from the beqz/jalr slots the original keeps, so
the two mechanisms are not substitutes.

85: cc1 folds SYM+N into a single la and SIX spellings do not defeat it.
86: cc1 cross-jumps identical guards; goto to a shared return label is the named lever.
2026-09-24 09:41:56 -04:00
Christopher Williams 6fdcaf3740 phase11: merge 16 + cookbook 83 — 525 bodies / 534 regions
Worker C's 0x800320D8 (276 B), matched on the FIRST spelling where its sibling 0x80031FC4
took 5 -- the family lever measured, on one family, both ways. Finding 55's limit confirmed
on the same family: a third row calling the same callee is NOT the same body and sits at
+16 instructions. The family transfers the derivation method and the stable positions,
never the body.

Also recorded: an OR nested inside an && chain is observable from the branch DIRECTIONS --
bne to the call block on one test and bnez to the manual-copy block on the other is
if (x == 0 && (a != 6 || b == 0)) call; else manual;
2026-09-24 09:40:40 -04:00
Christopher Williams c22ef879e8 phase11: merge 15 + amend cookbook 67, add 82 — 524 bodies / 533 regions
Worker D's 0x800910BC (280 B), its 8th match.

TWO CORRECTIONS TO THE COORDINATOR'S OWN COOKBOOK ENTRY, both from measurement:
 - 67 was INCOMPLETE and cost worker D a spelling. The magic alone is AMBIGUOUS: D = 2^(32+s)/M
   where s is the shift of the sra after the mfhi. 0x2AAAAAAB is /6 at s=0, /12 at s=1, /24 at
   s=2, and worker D read it as /6 when the shift was 1. The corollary is worth having too: the
   same magic twice in one function is not a contradiction (0x66666667 serves both /10 at s=2
   and /5 at s=1, materialised once into a callee-saved register).
 - The named-local rule is PER-SITE within one function. Naming a result the original consumes
   immediately costs 2 words; naming one the original reuses is free. Apply the decision once
   per VALUE, not once per function.

That is now the fourth correction to coordinator work this phase, and every one came from a
worker measuring something the coordinator had asserted.
2026-09-24 09:39:25 -04:00
Christopher Williams 1a6a00970f phase11: merge 14 + cookbook 79-81 — worker A's redundancy ranker
Worker A built a repetitiveness score (repeated 2/3/4-instruction opcode subsequences,
normalised by body length) and produced the cleanest controlled comparison in the phase:
3 spellings on a 248 B repetitive row vs 9 failures on a 176 B tie-break-dense one. That
converts 'prefer a repetitive body' from a hunch into a sortable number, so size is
deprioritised as the ranking signal.

Also recorded: a transposed temp array is byte-required (int m[3][4] used as m[c][r]) with an
exact diagnostic -- right length + right instruction multiset + residual only on sp-relative
offsets means the frame LAYOUT is wrong, not the code; and when the original stores the same
slot twice, suspect two source statements rather than a scheduler quirk (GCC 2.7.2 has no DSE).
2026-09-24 09:38:05 -04:00
Christopher Williams 61be53994b phase11: merge 13 + cookbook 73-78 — 522 bodies / 531 regions
Worker C's 0x80031FC4 (276 B). Cookbook gains six entries, the most important of which is
worker C's correction of the COORDINATOR: a DEPENDENT row is one you cannot VERIFY, not one
you have MATCHED. C's 0x800A613C and 0x800FD120 had symbol rows outstanding, but
re-verifying against the tracked registry gave byte-identical results to the overlay runs --
both are still near-matches blocked on an ALLOCATION lever. Adding a symbol row unblocks the
verification, not the match; conflating the two would have had a worker stop working a row it
had not solved.

Also recorded: the address-taken value may be a PARAMETER not a local (frame 8 too big with
all offsets shifted by 8 is the tell); address-taken form forces a register; the struct
assignment is what BATCHES the loads where element stores serialise behind maspsx nops; and
the cop2 operand is the 25-bit field (0x486012 -> 0x4A486012).
2026-09-24 09:36:47 -04:00
Christopher Williams 96f41be66a phase11: merge 12 — worker B's P1 band (8 claims) + worker A's 0x800319F0
+9 bodies: worker B's 0x8005E340, 0x8002C7EC, 0x800A8224, 0x8006B6BC, 0x80045F1C,
0x800F8A0C, 0x8002E9AC, 0x800196B4 and worker A's 0x800319F0 (whose dependent gp row
D_80122320 landed in the previous merge).

One new gp symbol row: D_801226E0 (append-only; the registry now has 395 rows).

Worker B's P1 band is finished: 10 rows, 5 matched, 4 near-matches with exact residuals,
1 blocked. THREE of the four near-misses failed on scheduling/allocation with the control
flow already EXACT, and one is a reorg slot-fill choice -- so that band's remaining yield
is in that class, not in shape work.

Two levers recorded from it:
 - The address-taken value may be a PARAMETER, not a local. 0x80045F1C's frame is only 40
   bytes yet it touches sp+56 and passes &a4 -- the FIFTH parameter, whose home is the
   caller's outgoing-argument area at frame+16. Modelling it as a local reproduces the same
   instruction SHAPE with a 48-byte frame and every offset +8 (22 differing bytes). So
   'right shape, frame 8 too big, all offsets shifted by 8' => check for a parameter first.
 - Address-taken form forces a register: 'int *p = &SYM;' gives la into a saved register
   plus indirection, where reading the symbol directly gives the macro pair and no save.
2026-09-24 09:35:26 -04:00
Christopher Williams 79765257ad phase11: merges 9-10 — 511 bodies / 520 regions, MAX 1232 B (worker D)
+10 bodies in one gated pass from all three active workers: worker D's 0x8010AF50
(1232 B — the >800 B band broken on the FIRST attempt, in TWO spellings, 5.0x the old
244 B ceiling) and 0x80031BBC (260 B); worker C's 6 claims including the re-tested
0x800AFDBC which my nopmarker correction closed with NO source change; worker A's
0x8006B214 and 0x80027CA0.

Four gp symbol rows added (D_80122320, D_80121F2C, D_80122128, D_801226DC), APPEND-ONLY.

PROCESS BUG FOUND AND FIXED IN MY OWN FLOW: the merge chain piped the gate into grep and
chained with &&, which PROMOTED A DIFFED REGISTRY -- grep succeeds whenever it finds the
word 'result=' regardless of the verdict. Caught by the following make check (806734
differing bytes), reverted, and the registry restored from the last green commit. The
merge flow now lives in .run/p11/merge.sh, which gates on the gate's EXIT CODE and
refuses to promote on failure.

Worker D's recognition, which is worth more than the row: mult + mfhi + sra with NO mflo
is a CONSTANT DIVISION, not a 64-bit multiply. 0x4BDA12F7 is ceil(2^45/27648); a genuine
64-bit multiply emits mfhi AND mflo in every available cc1. Recover the divisor from the
magic as D = ceil(2^(32+s)/M), never from the constant's face value.
2026-09-24 09:31:08 -04:00
Christopher Williams 55e59b3a4a phase11: merges 6-7 + the maspsx=moves mode — 501 bodies / 510 regions, MAX 700 B
+5 bodies: worker A claims 5-8 (0x800507A0, 0x80017B50, 0x800BBAC8, 0x800AFACC) and
worker D's 0x8006BC74 (700 B). Every candidate gate MATCH before promotion; all md5s
verified on disk.

*** 700 B IS THE LARGEST BODY EVER MATCHED IN THIS PROJECT *** — 175 instructions,
2.9x the old 244 B ceiling, and the FIRST match in the 401-800 B band. It cost THREE
spellings, fewer than worker D's own 248 B row (four). Both residuals were mechanical:
a missing `li 4096 / sw` pair hidden inside a run of 46 zero stores ("a run of repeated
stores is not a run of identical stores -- read every immediate"), and four extra
pointer reloads fixed by naming the sub-object pointer ONCE for the three byte stores
of 255 while leaving the fourth store its own re-read (cookbook 45's named-locals
family at its cheapest). Nothing about 700 bytes was hard: the body is large but highly
REDUNDANT, and redundancy is what a matcher keys off.

NEW HARNESS MODE `maspsx=moves` (worker B's oracle result, developer-authorized).
Worker B ran all five SDK assemblers (ASPSX 2.56/2.67/2.79/2.81/2.86) and every
supported option as a read-only oracle and found that **ASPSX does NOT fill delay slots
at all** -- it produces maspsx's exact shape. So maspsx is FAITHFUL to ASPSX, and the
fills in the original did not come from ASPSX. That overturns the "model ASPSX's fill"
framing: what fills the slots is GNU `as` in REORDER mode, i.e. maspsx OFF, and the only
real gap is ONE MNEMONIC -- `as` expands cc1's `move` to `or` where ASPSX emits `addu`.

So the mode is `maspsx=off` plus a single `move`->`addu` rewrite, letting `as` fill
exactly the slots cc1 left empty while cc1's own `.set noreorder` windows are preserved.
DEMONSTRATED: 0x800FA5D8 now reports 132 bytes / differing_bytes=0 MATCH where default
maspsx gives 148 LENGTH-MISMATCH. 7 new tests; suite 246 -> 253.

REGRESSION-VERIFIED: make check green at 510 regions / 253 tests with the mode OFF, so
every one of the 510 regions is byte-identical. The mode stays opt-in per region --
worker B measured the counterexample 0x8002D2BC, which has the SAME cc1 shape but whose
original keeps the store before the jr with a nop, so the original's assembler behaves
differently in different files.
2026-09-24 09:13:19 -04:00
Christopher Williams 7d7f41fbaa phase11: merge 3 — 493 bodies / 502 regions, TWO bodies above the old 244 B ceiling
+8 bodies: worker A claims 1-4 (0x8005E820, 0x80012DE8, 0x8001644C, 0x800A6880)
and worker C claims 1-4 (0x80048180, 0x800B107C, 0x80082868, 0x80036134).
Candidate gate MATCH before promotion; all 8 md5s matched on disk.

FOUR gp symbol rows added for 0x80017C6C (D_80121A2C/34/3C/44), coordinator-verified
against the payload: the original materialises them with addiu $2,gp,244/252/260/268.

C's claim 4 (0x80036134, 248 B) is the SECOND body above the old ceiling, matched on
lever-c-large row 1 — so two independent workers have now matched above 244 B, and the
dispatch-artefact verdict is confirmed by result rather than by inference.

sf3_merge format fixes, both triggered by real worker files:
  - a bare header row is now rejected with "looks like a column HEADER" instead of a
    confusing "not a hex address: 'start'"
  - a lone `-` in the override column means "no overrides", matching the absent-value
    convention the other tracked tables use
Suite 246 tests OK; make check green: regions=502 AGREE, differing_bytes=0 MATCH.
2026-09-24 08:59:21 -04:00
Christopher Williams 5d41f97421 phase11: *** CEILING BROKEN *** — 485 bodies / 494 regions, new max 248 B
Worker D's claim 0x8009F6A0..0x8009F798 (248 B) MATCHES. Verified independently by the
coordinator on a fresh work dir (candidate_bytes=248 differing_bytes=0 MATCH) and
gated on the whole binary before promotion. This is the FIRST body above 244 B ever
matched, and it sets a new corpus maximum (previous max 244 B at 0x80099078).

THE LEVER (worker D, 4 spellings): the local working buffer must be a 3x4 word array
(`int t[3][4]`, only columns 0..2 used), NOT `int t[9]`. The 4-WORD ROW STRIDE IS
BYTE-LOAD-BEARING: it moves the 2nd and 3rd triples to 0x10 and 0x20, makes the frame
48 B instead of 40 B, and leaves the unused 0x0C/0x1C slots the original shows. New
instance of cookbook 54 (a 2-D array's row stride is byte-load-bearing). The element
type is the other half: `short` locals let cc1 drop the sign extension (lhu/subu, no
frame); `int` locals keep it (lh/negu).

Diagnostic broadcast: correct length + right instruction multiset and order + residual
concentrated on the FRAME ADJUSTMENT and every sp-relative offset => suspect a local
aggregate's row stride / element size, not the control flow. D's variant (c) was a
textbook case: 19 differing bytes, all of them the frame size and the address shift
that follows from it, closed by one array-shape change.

Also in this commit — a fail-fast fix to sf3_merge. Worker D placed the source md5 in
the claim row's 4th column, which sf3_merge passed through as a region override, so the
row MERGED and only `sf3_match gate` failed later with "unknown override key 'md5'".
sf3_merge now validates override keys at merge time and rejects the row with a message
naming the valid keys and pointing at report.tsv for per-claim metadata. 4 new tests,
suite 242 -> 246, OK. The candidate gate caught it; the tracked registry was untouched.
2026-09-24 08:56:26 -04:00
Christopher Williams 9aae442e64 phase10: merge 27 + negatives extraction — 484 distinct bodies / 493 regions
+2 bodies (worker C2 claims 3-4). Candidate gate MATCH before promotion.

NEGATIVES EXTRACTION (the closing-checklist step that protects worker findings
from ignored staging being lost): imported 28 new negatives from all five workers'
staging into the tracked index, and dropped 28 rows that had since been REGISTERED
(the reconcile step working as designed). Index 166 -> 194 rows, address-ordered,
0 duplicates, 0 registered. Worklist 1193 rows; excluded_recorded_negative=170;
0 unregistered negatives survive into the worklist.

make check green: regions=493 AGREE, differing_bytes=0 MATCH, 237 tests OK.

Worker C2's cross-cutting finding recorded: three of its four negatives are the -O2
SCHEDULER, not source shape. For 0x800A6C34 (16 differing) and 0x80016F80 (45
differing), both at correct length, the identical source with
`cc1 -quiet -O2 -G0 -fno-schedule-insns` produces the original's instruction ORDER
byte-for-byte -- and both then show the same two-sided signature: sched OFF gives the
original's order but the wrong allocation (or the wrong delay slot), sched ON gives
the right allocation and the wrong order. Neither alone matches. C2 correctly did NOT
take the scheduler override (out of scope); both stay recorded as negatives with the
lever named. Diagnostic broadcast: correct length + a residual that is a permutation
of a few instructions + the unscheduled build matching the original order => the
residual is sched.

Also recorded: the commutative-operand lever has a COUPLED ALLOCATION side-effect
(0x80027744 at 4 bytes, 0x80050674 at 1 byte) -- every spelling giving the original's
addu operand order also flips which value takes v0 vs v1.
2026-09-24 08:22:55 -04:00
Christopher Williams be2ee1874c phase10: merge 26 — 482 distinct bodies / 491 regions
+2 bodies (worker C2 claims 1-2: 0x80051864 88 B, 0x8006AA10 120 B). Candidate
gate MATCH before promotion. make check green: regions=491 AGREE,
differing_bytes=0 MATCH, 237 tests OK.

C2 also confirmed the paired named-boolean rule exactly as broadcast (one shared
li a3,1 across four acceptance paths; nested ifs give two blocks and 124 B
LENGTH-MISMATCH; the single combined condition is byte-identical), and caught the
mask 0x00400000 vs 0x40000000 by computing the lui high half.

CROSS-CUTTING FINDING (worker C2, recorded for the cookbook): THREE of C2's four
negatives are the -O2 SCHEDULER, not source shape. Diagnostic: when the length is
exactly right and the residual is a permutation of a few instructions whose
-fno-schedule-insns build matches the original, stop hunting for a source shape.
Trap: on 0x80016F80 sched OFF gives the original's ORDER but the wrong allocation
while sched ON gives the right allocation and the wrong order, so neither alone
matches. C2 correctly did NOT take the scheduler override; all three stay recorded
as negatives with the lever named.

Also: the commutative-operand lever has a COUPLED ALLOCATION side-effect
(0x80027744, 0x80050674) -- every spelling giving the original's addu operand order
also flips which value takes v0 vs v1, because operand order changes pseudo creation
order. Finding 22's re-spell advice is necessary but not sufficient.
2026-09-24 08:20:05 -04:00
Christopher Williams c1dbf9d64e phase10: *** MILESTONE MET — 480 distinct bodies / 489 regions (target 475) ***
+6 bodies (worker A claims 23-26, worker B2 claims 8-9). Candidate gate MATCH
before every promotion.

FULL CLEAN AUDIT GREEN at the milestone:
  make clean && make all  exit 0
  cmp                     exit 0
  SHA-1 both files        e173426c157384ebf1b6caf8c6fea18a85a14af9
  registry                489 rows, 0 overlap, 0 unsorted, 0 bad extents,
                          0 missing sources, 480 distinct sources
  firewall                0 tracked paths under any prohibited root (583 files)
  suite                   237 tests OK

The phase goal (475 from the 400 baseline) is met with 5 bodies to spare.
Per the plan, phase close requires the developer's explicit confirmation of the
milestone; this commit records the state, not the close.

SIZE-BAND FINDING now confirmed a third time, within-worker: worker B2's nine
matches cost 1, 4, 6, 4, 1, 2, 2, 1, 1 attempts -- the three <=120 B frameless leaves
all cost exactly ONE attempt, while the two >200 B P1 rows it opened with cost 6 and
4 attempts and produced ZERO matches.

Two more levers recorded from the closing rows:
  - RECORD IDENTITY FROM WIDTHS (0x80041610): the two arms read three shorts at
    +264/+266/+268 versus three ints at +20/+24/+28 through one extra indirection,
    so they are TWO record types; declaring one shared type would have been wrong.
  - a1[1] = -a1[1] RELOADS FROM MEMORY and its source order matters: the negation
    reloads 4(a1) because the intervening store to 8(a1) may alias it.
  - A 2-D ARRAY'S ROW STRIDE IS BYTE-LOAD-BEARING (0x800AC7A0): D[a3][a4] emits
    sll a3,4 + sll a4,2 + add (correct length); D[a3*4 + a4] folds the outer *4 into
    a second sll and comes out 4 bytes SHORT. When a scaled index is one sll short,
    the source is a 2-D array, not a flattened index -- a LENGTH-class lever.
2026-09-24 08:15:39 -04:00
Christopher Williams ef29df9b58 phase10: merge 23 — 474 distinct bodies / 483 regions (ONE from the milestone)
+3 bodies (worker B2 claims 5-7). Candidate gate MATCH before promotion; md5 drift
check clean on all three (second use of the guard). make check green: regions=483
AGREE, differing_bytes=0 MATCH.

Region option granted: gp=-D_80121BFC on 0x80048128. This one carries PER-ACCESS
evidence in a single merge: D_80121BFC is read gp-RELATIVELY in worker B2's claim 7
row (lw v1,708(gp)) and ABSOLUTELY in claim 5's row (lui v1,0x8012 + lw v1,7164(v1)).
So the same symbol needs the override in one region and not in the other -- direct
confirmation of worker B's lever 12 that the access form is per-SITE, and an argument
that the override is a region property rather than a symbol property.

THREE NEW LEVERS (worker B2), each with a control:
1. A LOCAL SHARED BY TWO GUARD BLOCKS GETS COALESCED; TWO BRACE-SCOPED LOCALS DO NOT.
   0x8008BA80: the original's first guard loads the state byte into a0 (the
   parameter's own dead register) while the second guard loads *a2 into a FRESH v1.
   One function-scope local makes cc1 coalesce the live ranges into a0 and DIFFs;
   brace-scoping each reproduces it. A SCOPING lever, distinct from the
   named-locals family.
2. THE EVALUATION ORDER OF TWO SCALED TERMS IS BYTE-REQUIRED. 0x800504E4:
   base + a0*384 + a1*3072 scales a0 first; base + a1*3072 + a0*384 scales a1 first,
   which is the original. Worker B's claim-1 lever extended from the operands of one
   '+' to the ORDER OF TWO INDEX COMPUTATIONS.
3. AN UNSIGNED LOOP COUNTER SHOWS AS sltiu vs slti -- ONE BYTE (opcode 0x0b vs 0x0a).
   0x800504E4's i < 16 is sltiu, so the counter is unsigned int. The loop-test form of
   the signedness trap.

Also independently reproduced: worker C's lever 6, (unsigned)(c - 58) < 2 giving ONE
addiu+sltiu pair where c >= 58 && c <= 59 gives two tests (on 0x80048128).
2026-09-24 08:11:03 -04:00
Christopher Williams 62e549b30f phase10: merge 22 — 471 distinct bodies / 480 regions (4 from the milestone)
+3 bodies (worker B2 claims 2-4; claim 1 was already merged). Candidate gate MATCH
before promotion. make check green: regions=480 AGREE, differing_bytes=0 MATCH.

FIRST MERGE USING THE MD5 DRIFT GUARD (adopted after the 0x800A9C24 collision).
Worker B2's claim rows carry the source md5 and all three matched on disk at merge
time, so no re-verification was needed. B2 folded its gp=-D_80121B88 region option
into the claim row itself, which is the standing rule.

Region option granted: gp=-D_80121B88 on 0x800A6998. Byte-required and verified both
ways (with the override 128 B / 0 differing; without it 120 B LENGTH-MISMATCH).
Third instance of the per-site gp form. Worth recording that on THIS row BOTH the
symbol spelling and the literal spelling are wrong for two DIFFERENT reasons: the
symbol gives the gp-relative encoding the original does not use, and the literal
makes cc1 cache the address in s0 so the frame grows 24 -> 32 bytes with an extra
saved register.

TWO NEW LEVERS (worker B2):
1. A SHARED CALL PAIR HAS TO BE A NAMED goto LABEL. 0x800A6998 is an infinite loop
   whose two func_800A6FAC(0) calls appear EXACTLY ONCE in the original, reached
   both by fall-through and by the w != 0 branch. Only a source naming that block
   reproduces it; while/for spellings either invert the loop (156 B) or make cc1
   CROSS-JUMP the two ==0 tests so the x-load goes dead (124 B). Same goto
   statement-form family as the 0x800A9C24 guard, but the lever here is SHARING ONE
   BLOCK rather than branch polarity.
2. && KEEPS A BRANCH THAT EARLY RETURNS FOLD AWAY -- the exact INVERSE of the
   named-boolean lever broadcast earlier. Three separate if (...) return 0;
   statements make cc1 fold the third guard into a branchless boolean
   (xor/sltiu, 124 B); ONE && chain with a single trailing return 0 keeps the branch
   (byte-identical). A goto fail; ... fail: return 0; spelling is also identical, so
   the distinguishing fact is "one combined condition with one trailing return".
   PAIRED RULE: when the original has a BRANCH on a boolean expression, combine the
   conditions into one && chain; when it has the BRANCHLESS form, name the boolean.

Also recorded: 0x800B67B8 is the CALLER of the registered func_80082750, and the
callee's source gave the D_80121C00 spelling directly (it holds a base POINTER, not
the table address), so the argument map transferred with zero guessing -- the
family lever working as advertised.
2026-09-24 08:05:01 -04:00
Christopher Williams 80b8b33eb4 phase10: merge 21 — 468 distinct bodies / 477 regions (7 from the milestone)
+4 bodies (worker A claims 19-22, all first-attempt). Candidate gate MATCH before
promotion. make check green: regions=477 AGREE, differing_bytes=0 MATCH.

THE STACK-SWITCH IDIOM IS NOW PROVEN TWICE, BYTE-EXACT, WITH NO OVERRIDE.
0x8006B9E0 matches on the first attempt with no maspsx override, no clobbers and no
operand declarations -- the three-macro form reproduces it exactly, including the
bare filler nop between the second restore and the epilogue's lw ra. Contrasting it
with A's 0x800BC658 near-match (same macros, off by 2 nops) isolates the variable:
0x8006B9E0's two calls have EMPTY delay slots while 0x800BC658's non-fast-path calls
set up an argument cc1 schedules into the slot. So the asm is not the variable -- an
ARGUMENT MOVE AT THE CALL SITE is. That narrows the 0x800BC658 remainder.

TWO NAMED-LOCAL FINDINGS, BOTH DIRECTIONS NOW OBSERVED (the family has seven
instances across the team):
  - 0x8005E538: the HANDLE must be a named local loaded before the first call.
    Inline, cc1 keeps only the object pointer in a callee-saved register and
    RELOADS the handle after the first call (76 vs 88 bytes). Generalisation: a value
    that must survive an intervening call has to be a named local.
  - 0x800F66B8 is the MIRROR CASE and a useful NEGATIVE: d[0] and d[1] are each
    loaded twice with neither held in a register, and the two loads of d[0] go to
    DIFFERENT registers. Naming a local would have been WRONG. So naming a value can
    be byte-required and NOT naming it can be byte-required; the diagnostic is the
    load count, not a rule about locals.

Inline asm in src/func_8006B9E0.c is documented per the Phase 10 convention
(header states the observed sequence, the reason and the limits).
2026-09-24 08:02:57 -04:00
Christopher Williams 82c65ed468 phase10: merge 19 — 464 distinct bodies / 473 regions
+8 bodies (worker A claims 11-18, all first-attempt; 16 of A's 18 claims needed
no override at all). Candidate gate MATCH before promotion.

TWO NEW LEVERS (worker A):
1. A two-arm selection's POLARITY is byte-required. For a1 = (x<2) ? a3 : saved,
   the natural order and the ternary both emit beq v0,zero with the arms swapped
   (right length, 5 differing bytes). Writing the larger-than arm first
   (if (x >= 2) a1 = saved; else a1 = a3;) makes cc1 emit bne v0,zero with
   a1 = a3 in the BRANCH DELAY SLOT, so that assignment runs on both paths and is
   overwritten on the >=2 path -- which is the original exactly. Same family as the
   mirrored comparison load order; second polarity case in A's partition.
2. A NAMED BOOLEAN LOCAL forces the branchless compare. rec[6] = ((a3 & 0xff) != 0) << 1
   is branchy (92 vs 88); only naming the boolean first,
   int flag = (a3 & 0xff) != 0; then rec[6] = flag << 1;
   reproduces the original's andi / sltu / sll. Six spellings measured. cc1 will
   un-do a boolean you inline when the VALUE (not the branch) is what you need.

make check green: regions=473 AGREE, differing_bytes=0 MATCH, 237 tests OK.
Worklist 1241 rows; excluded_already_registered=473.
2026-09-24 07:57:42 -04:00
Christopher Williams a087b7d936 phase10: merge 18 + negatives — 456 distinct bodies / 465 regions
+1 body (worker B2 claim 1, 0x8008A758, first attempt via the registered callee
func_80073F78). Candidate gate MATCH before promotion.

Worker C's cleanup-audit catch recorded: 0x800A6C34 (88B) is an OPEN near-match at
the correct length (16 differing bytes at 0x800A6C54) that fell out of C's staging
after a DIFF in its first batch. Added to the tracked negatives index. Coordinator
verified the residual exactly as C reported AND tested C's named lever (a) — the
gp=-D_80121B88 override does NOT close it (identical 16 bytes) — so the index
records the untried lever (a source order that stops the hoist) rather than the
disproven one. C's process note is recorded: run the cleanup audit BEFORE the final
report, not after.

make check green: regions=465 AGREE, differing_bytes=0 MATCH, 237 tests OK.
Worklist 1249 rows; index 166 rows, ordered, 0 registered.
2026-09-24 07:53:00 -04:00
Christopher Williams 7f4553c948 phase10: merge 17 — 455 distinct bodies / 464 regions
+2 bodies (worker C claims 21-22; 19-20 were in the previous batch).

SIZE-BAND MEASUREMENT CONFIRMED (worker C, strongest evidence yet for the
size-first re-rank): after switching to the <=200B band, four consecutive rows cost
1, 1, 1 and 2 attempts, against 1-in-12 on the 200B-800B band. Same worker, same
levers, same day -- the only variable is the size band. The <=200B P2 band is where
the remaining bodies are.

make check green: regions=464 AGREE, differing_bytes=0 MATCH, 237 tests OK.
2026-09-24 07:50:45 -04:00
Christopher Williams 091a9020f7 phase10: merge 16 + new worklist exclusion class — 453 bodies / 462 regions
+2 bodies (worker C claims 19-20) and 3 gp symbol rows (D_80121F90,
D_8012196E, D_80121964). Candidate gate MATCH before promotion.

NEW EXCLUSION CLASS `restores_unsaved` (worker A, coordinator-verified). A body
that restores a callee-saved register it never saves cannot be a whole function:
the register it restores was established by an enclosing prologue that the derived
extent cut off. These are jal targets INSIDE a real function, so the walk began
mid-body — distinct from bad_extent_start, which flags starts that are not
function entries at all.

Verified disjoint from the matched corpus before acting, as the project requires:
0 of 462 registered regions trip the rule; 11 worklist rows do.
excluded_restores_unsaved=11; worklist 1265 -> 1253 rows. 5 new tests; suite
229 -> 237, OK. make check green: regions=462 AGREE, differing_bytes=0 MATCH.

Independent corroboration worth recording: the new rule re-derives 0x8010080C,
the false extent start worker C reported earlier via a completely different signal
(the first instruction reads a register the range never defines). Two independent
detections of one defect class. bad_extent_start is False for that row, confirming
C's observation that the older rule missed it. Its Makefile --exclude entry is
retained only as the provenance record for that defect.
2026-09-24 07:49:34 -04:00