11 KiB
§60 — Classify the residual, don't rank it: the deterministic residual→class classifier and what it measured about the backlog (Phase 29 Task-13A, 2026-07-21)
The instrument. tools/residual_class.py decides a near-miss's class FROM THE BYTES. It decodes each
mismatching MIPS word into (operation-skeleton, register-fields, immediate) and runs a decision tree:
drift first (a single inserted instruction desynchronises the tail and inflates closeness by the tail's
length — a 1-instruction structural delta wearing a 200-mismatch costume), then a consistent-injective
register map (⇒ REGALLOC-PERM, the §31 S11/RC-3 class), same-multiset-different-order (⇒
SCHEDULE-REORDER), nop-vs-instruction (DELAY-SLOT), then operation-family splits (WIDTH lw↔lh =
the §18/§43 idioms, BRANCH-POLARITY, STRENGTH, ADDRESSING), then immediate-only (IMM-OFFSET
constant delta = a frame/struct-layout shift, IMM-VALUE). Every path ends in a NAMED class; an opcode the
decoder does not cover is UNKNOWN and COUNTED (R32). Each class routes to a bucket — permuter /
structural / integration / redraft — which says WHICH TOOL the failure wants.
tools/autopsy.py collect materialises the corpus by recompiling every open backlog draft through the
EXISTING match_one path (R33 — never a second copy of the pipeline), deriving the two silent-artefact
inputs rather than guessing them: the asm subdir (from the stub's self-describing INCLUDE_ASM line) and
the -O0 flag (corpus.is_o0, parsed from the Makefile's own -O0 rules with a coverage assertion). 1,752
drafts in ~21 s at -j12.
Cross-check (R34). The classifier's closeness is computed by a different route than
masked_diff.structured_diff's; collect asserts equality on every row and refuses the corpus on any
disagreement. 1,673/1,673 agreed, 0 classifier errors — so the MIPS decoder covered every opcode in the
real corpus.
What it measured — the whole open backlog, byte-grounded
| bucket | fns | reach-wtd | meaning |
|---|---|---|---|
redraft |
699 | 2,162 | the stored draft is not this function (a 15-ins body vs a 132-ins target) |
structural |
578 | 7,761 | local mutation cannot introduce it — wants a C idiom, not CPU |
integration |
306 | 2,303 | byte-correct standalone; blocked on plumbing (§58/§59) |
permuter |
75 | 740 | a search-closer can actually reach it |
unknown |
2 | 2 | the LLM tier's residue |
THE FINDING: of the 972 records the grinder's own filter admits, 75 (7.7%) are permuter-shaped. 547 are
structural and 348 are junk drafts. The daemon has been spending ~92% of its CPU where the permuter provably
cannot win — which is the byte-grounded explanation of "7 banks all-time, all in Phase 21, 0 since"
(Phase-22 audit). It was never a missing transform; it was targeting. grinder.candidates() now filters
on the measured bucket (1,303 → 78 candidates) and takes its directed permuter_weights profile from the
measured class instead of the logged label — 91% of records carry NO label, so classify() returned None and
the search ran on gcc defaults. Degrades to the old undirected behaviour if the corpus is absent, and says
which mode it is in (--no-targeting A/Bs it).
Two corollaries worth remembering
- A large
closenessand a hard function are different things. 699 records rank as "near-misses" with closeness up to 278 purely because a stub-sized draft was scored against a large target. Ranked by closeness they look like a wall of nearly-done work; they are un-attempted work misfiled as near-misses — fresh crack fuel, not a backlog of hard functions. Hence the separateredraftbucket: the routing is opposite (re-draft vs seed-tweak). match_oneMATCH still ≠ bank (§58), measured. A 12-draft gate probe of theintegrationbucket (reach-134, ov_SC01_077) banked 1 of 12; the other 11 failed PLUMBING. So the 306 is a pool of integration candidates whose conversion depends on the reconcile ladder — it prices Task 14, it is not 306 free banks. (Note the harvest_verify failure LABEL is the §58 red-herring: 10 of the 11 reported the sameconflicting types for built-in functionline from an unrelated TU position.)
The parallel-probe race this surfaced
masked_diff._common_typedefs() wrote, read and deleted ONE shared path src/.masked_diff_probe.c. Under N
concurrent match_one/permuter processes, whoever unlinked first made another's open/parse fail, and that
process died with a traceback instead of a verdict: 14 of 1,752 drafts lost in a single 12-way run (0.8%)
— and every parallel wave has paid it invisibly, because a drafter that crashes on its self-check merely
looks like a drafter that failed. Now per-PID. Same defect class, one level down, as the Phase-28
match_one --work shared scratch whose docstring promised the isolation its default contradicted.
§60a — What the first DIRECTED grinder run exposed (Phase 29 Task-13B, 2026-07-21)
Turning the targeting on and running the grinder bounded (--once --batch 8 --permute-secs 90) produced a win
on the FIRST candidate — func_80181F78, close=1, classified DELAY-SLOT/schedule, banked in ~6 min — and
then immediately surfaced three latent defects that had been unreachable because the daemon had not banked
anything since Phase 21. All three are the same shape: a step whose REPORT and whose WORK had quietly
diverged.
gate_stage's commit path crashed onsrc=None.srcis deliberately never defaulted (the Phase 26-A audit: a default would silently PIN the gate to the main.c), but the commit didgit add src …unconditionally. So every caller that omitssrc— grinder, orchestrator, idiom_hunt — crashes the moment it banks. Fix:git add -u src/(every modified tracked file under src/), which also retires thesrc/ov_*/*.cfilename glob that once omitted 4 R22-verified banks from a commit because a family's members do not all live in the same-named split. COMPLEMENTARY HOLE (2026-07-22):git add -umisses the NEW files an ISOLATION creates. A jtbl sweep cuts a fresh region file per sibling (src/<ov>/<ov>_jr_<ADDR>.c), which is UNTRACKED — so-ucommits the modified TU and drops the file holding the banked body, i.e. a tree that cannot clean-rebuild. For any carve/isolation bank usegit add -A src/ config/.jtbl_family_bankalready refuses to sweep on an uncommittedconfig/+src/(its per-sibling revert restores from HEAD), and that guard is what caught this — a fail-loud precondition doing exactly its job.- The
_xformladder dirs accumulate.<drafts>-cn/-cast/-rc/-uniare reused across runs and the transform tools only write the drafts they are handed, so every stale draft from every previous run survives and is re-submitted to the byte-gate. Measured: the grinder submitted 1 draft, the gate processed 34 and banked 2. Nothing wrong entered the tree (G3/P9: the gate banks only byte-identical output) — but a run banked a function it was never asked to try, and would have committed it under a message naming a different one. Fix: clear the out dir per run. Note the symmetry with R32: a scanner that silently NARROWS its input hides work; a stage that silently WIDENS it fabricates provenance. - The grinder fired the fleet-wide propagate from inside the gate.
gate_stage(propagate=True)runsdedup_propagate --auto-from— the §55b path that timed out at 3600 s and left 90/140 overlays broken, and which, being inside the gate, takes the banks down with it when it fails. The unattended caller must never fire it: bank withpropagate=False, commit the cheap verified banks, then run ONE targeteddedup_propagate --addras its own batch.
The transferable point: a tool that has been failing for a long time accretes latent bugs on its success
path, because nothing exercises it. Before trusting an unattended fix-and-run, budget for the first success
to fail — and check the tree state, not the exit code (here the durable Task-12 winner save in
.run/permuter-winners/ is what made the crash a non-event).
§60b — The plateau autopsy's verdict: a partial drift is a WRONG DRAFT, not a missing transform (Phase 29 Task-13B close, 2026-07-21)
hindsight-study §7 predicts that a search-closer's plateaus decompose into missing-transform (extend the mutation set — "the highest-value bucket and the whole point"), seed-structural, and genuine-wall. Run against real plateaus, this class produced ZERO missing-transforms. The autopsy is worth recording because the answer was legible in the bytes and needed no LLM at all.
The measurement. A 20-target probe of the length profile: tail drifts (a single shift point explains
the whole tail) converted 1/6; partial drifts (length differs AND other positions differ) converted
0/12. Reading three partial plateaus directly:
func_8017F0C0,func_801806C8— target containssltiu $v0,$v0,0x1. That is gcc's codegen for!x/x == 0. The drafts wrote(u32)(D_x ^ 1), which emitsxori. No local mutation rewritesxoriintosltiu: it is a different operation, chosen by the front end from a different C expression.func_8017FF90— the draft stores toarg0 + 8; the target stores to a global (lui $at,%hi(D_…)/sw $zero,%lo(D_…)($at)). Not the same function at all.
Two permanent fixes, both NARROWING what the offline tool is allowed to attempt:
_drift_routenow admits a drift to the permuter only when|Δ|<=2ANDexplains == "tail". Length-profile pool 339 → 34; permuter bucket 389 → 84. Apartialdrift is seed-structural by construction — the count is wrong and other positions are wrong, which is not one local edit.SIZE-MISMATCHgained a PROPORTIONAL test (|Δ| >= 0.5*nt) alongside the absolute one.max(2, 0.15*nt)is far too permissive on a tiny target: a 2-instruction draft against a 4-instruction target is|Δ|=2and read as a near-miss when it is a wholesale mismatch.
The transferable lesson. The §7 taxonomy tacitly assumes the plateaus are near. Ours mostly were not —
they were bad drafts wearing a small closeness. So the highest-value autopsy outcome was not a new
transform but a tighter admission rule: the way to raise a search-closer's yield is at least as often to
stop feeding it unreachable work as to widen its mutation set. Same knife as Task-13A's targeting fix, one
cut finer.
Cookbook idiom for drafters (recurring): sltiu rd, rs, 1 ⇒ the C is !x / x == 0, NOT x ^ 1.
The XOR form emits xori and can never match.