SS127/SS127a/SS127b distil what the wave's agents kept re-deriving, because the index
fired on only 3 of 15 targets:
- the -O0 CONSTANT-OFFSET FOLD: `p->f` folds to `lbu 3(r)`, `p[i]` does NOT (addiu +
0-displacement load). At -O2 these converge, which is why nothing in SS1-SS126 covers it.
- the -O0 regime generally: spill/reload pairs are REAL named locals; load-delay nops and
redundant copies are normal; write plain C, the -O2 steering levers are inert here.
- SS127a: SS71 sibling-first is the STRONGEST -O0 lever — an -O0 TU is a near-uniform code
regime, so a banked sibling's shape transfers far better than at -O2.
- SS127b: two agents' decisive levers came from a SOURCE COMMENT in ov_SC01_077_o0.c, not
from docs/. Promote levers out of source comments or every future agent re-buys them.
Checkpoint records the wave AND its honest ROI: 1.33M tokens for 12 banks and +0.00pp
headline. The value is contingent on three reach-138 functions, and all three are
currently unpropagated (func_8013C08C 0/137, SS94 type-carry) or gate-failed
(func_8013BD74 CARVE-REFUSED, func_8013B83C CC1-FAIL). Fix propagation before wave 2 —
drafting more x2-reach targets is not where the leverage is.
First Ultracode wave against the population the -O0 routing made draftable. 30 agents
(15 drafters + 15 adversarial verifiers), 1.33M tokens, 8.8 min. Every drafter self-checked
with match_one --o0 AND rtu_match --o0; every MATCH claim was then re-run from scratch by an
independent skeptic instructed to default to REFUTED. Result: 15/15 confirmed, 0 disputed.
WHOLE-BINARY GATE (the sole arbiter, G3/P9): **12 banked / 3 failed** — a textbook SS52b
outcome (an rtu MATCH is a CANDIDATE, not a bank). All three failures are NAMED INTEGRATION
classes, none a compiler wall:
func_8013BD74 CARVE-REFUSED — it is a jr function; needs the SS81 carve chain (reach 138)
func_8013B83C CC1-FAIL — real-TU compile, error not yet read (reach 138)
func_80184058 PLUMBING — recovery ladder
BANKED: func_8013C08C + the 11 fourth-region fns (func_80183CF0/D50/F28, func_80184028/264/
2E0/354/474/538, func_801847EC, func_80184868).
Bank truth read from the SOURCE (INCLUDE_ASM absence), never the gate report (SS55b trap 4).
R22 CLEAN-FLEET: extract-all 139/139 (+main); check-all 140 passed, 0 failed of 140.
PROPAGATION OF func_8013C08C (reach 138) IS NOT DONE: the first sweep returned "0 families"
because the family map still listed it as a stub — regenerated it (the documented
crack-wave-sweep-map-regen path), after which the sweep found 137 candidates and banked
**0/137**. Per SS94 a family 0/N is a TYPE-CARRY failure until proven otherwise, and this
body carries a SS100 body-scoped typedef, so that is the first hypothesis to test. Recorded as
open, NOT as a wall.
FLYWHEEL FEEDBACK (R16), the honest read: the cookbook index fired on only 3/15 targets.
Agents independently re-derived the SAME undocumented idiom — the -O0 CONSTANT-OFFSET FOLD
(`p->f` folds to `lbu 3(r)`; `p[i]` does NOT, it emits `addiu; lw 0(r)`) — and two found their
decisive levers in a SOURCE HEADER COMMENT in ov_SC01_077_o0.c rather than in the cookbook.
The -O0 regime is under-documented relative to how much of the frontier now lives in -O0 TUs.
SS71 also generalises to -O0: several agents cracked their target off an already-banked sibling
in the same TU (func_80184868 came straight off the shape banked earlier today).
(1) --fix-def-sig POSTURE: AUDITED CLEAN. `action="store_true"` (defaults False), one
consumer via getattr(a,"fix_def_sig",False), and NO caller anywhere passes it — checked
tools/, .run/ scripts, docs recipes and the Makefile. The flag help already carries the
SS119 warning.
BUT the audit surfaced a live hazard the earlier pass missed: docs/decision-log.md
still recommended "--fix-def-sig should likely be default-on for the h_seq path".
That was byte-REFUTED by T84/SS119 — the flag is a REPAIR, not a default; on 0x80161c98
it imposed a signedness-wrong `s32 a1` over the true `u32`, turned a byte-correct draft
into a 1-instruction DIFF (slti vs sltiu), and held 137 members at 0 until DROPPED.
Struck through in place with a superseding note rather than deleted, so the original
reasoning stays legible (R31) — but a forward-looking "should be default-on" sitting in
a doc a fresh session reads FOR DIRECTION is a hazard, not a historical note.
(2) GRINDER WARM-START: tools/permuter_ils.py has sat beside grinder.py since Phase 24 and
was never wired in, so every grind was a COLD search that burned its whole time box
re-descending ground the previous run had already covered. grinder.py now runs `--cycles`
(default 4) timeboxed permutes, each warm-restarted from the previous cycle's best byte
waypoint, stopping early on no gain. `--cycles 1` reproduces the old cold behaviour exactly,
so it is opt-out. --permute-secs is now documented as the PER-CYCLE box.
JUSTIFIED BY MEASUREMENT, not by the task list: the lane looked dead (Phase-22 audit: 7
all-time banks, all Phase 21, 0 since), so I checked for live fuel before building. The
backlog holds 665 open near-misses in the permuter-tractable band (close 1-20), 157 of them
close 1-4, including func_8016BA68 at close=1 with reach=134.
HONEST LIMIT: this is a WIRING change whose yield is UNPROVEN. The Phase-24 evidence for ILS
is one function (func_80148094, 72 -> 36 over ~8 restarts); I have not run it on this
backlog. A winner remains a CANDIDATE — the whole-binary byte-gate is still the sole arbiter
(G3/P9), and an intermediate waypoint is only ever re-seeded, never banked.
The second and last has_mid_jr family of the -O0 cluster. Same route as func_8013C0F8:
jtbl_family_bank (the SS81 carve chain per sibling) -> 137 BANKED / 0 failed.
Both jr families together: 274 members, 483 ins x137 = ~66k instructions, 0 failures.
Neither was bankable before commit:1270 routed the destination TUs to -O0.
R22 CLEAN-FLEET: extract-all 139/139 (+main); check-all 140 passed, 0 failed of 140.
The first of the two has_mid_jr families SS53 correctly refused from the carve-less
family_sweep. Routed through the tool its tier needs (jtbl_family_bank, the SS81 carve
chain per sibling) it banks CLEAN:
jtbl_family_bank func_8013C0F8 ov_SC01_077 0x8013c0f8 -> 137 BANKED / 0 failed
~7s per sibling (carve + extract + remap + whole-binary gate), ~16 min total
Only possible now because commit:1270 routed each destination TU to -O0; before that the
member's home file compiled -O2 and no body could ever match there (SS116).
INDEPENDENT CONFIRMATION OF THE SS125 DIAGNOSIS: this is the SAME tool that returned
gate-fail on every group-B (func_8017BEBC) probe earlier today. 137/137 here versus 0/4
there, same jr machinery, is exactly what the carve-vs-body diagnostic concluded --
group B's failures are its BODY, not the carve. I had first blamed the carve; the
body-free probe corrected it, and this run corroborates the correction from the other side.
R22 CLEAN-FLEET: extract-all 139/139 (+main); check-all 140 passed, 0 failed of 140.
The payoff of routing the cluster to -O0 (commit:1270). These functions were ALREADY
CRACKED in ov_SC01_077 and could not be banked anywhere else purely because every
destination file compiled -O2. With the destinations now -O0, they template in
deterministically -- no drafting, no agents.
dedup_propagate --recover 0x8013C360 (h_exact x138) -> 137 overlays byte-identical
family_sweep --hseq 10 variant families -> 1,227 banked / 133 failed (90%)
1360 staged across 136 groups
------------------------------------------------------------------------------------
1,364 new banks
FLEET: fn-count 92.71 -> 93.09% · instr 88.3 -> 88.6% · distinct-code 78.7 -> 79.3%
(71,756 / 87,459 unique fns; +1,162 unique). dedup 1904 -> 1905 groups, 0 failed;
C1 coverage 240496/240496. 0 NON_MATCHING in any default build (G4).
R22 CLEAN-FLEET: extract-all 139/139 (+main); check-all 140 passed, 0 failed of 140.
--recover WAS LOAD-BEARING (SS75): without it dedup_propagate took its historical
all-or-nothing branch -- one failing overlay (the SOURCE, ov_SC01_077) dropped the whole
function and it printed "all candidates dropped", which reads exactly like a wall. Reading
the exclusion code instead of believing the message showed the remedy: --recover excludes
only that overlay (kept x1 with its own inline match) and propagates to the other 137.
The two has_mid_jr families in the cluster were REFUSED BY DESIGN, not attempted (SS53
interlock): 0x8013C0F8 (154 ins) and 0x8013C414 (329 ins), ~137 members each = ~466
members queued behind the jtbl carve path they actually need, rather than a fake 0% from
the wrong tool.
REMAINING in the cluster: the 133 sweep failures + the 2 jr families + the 3 addresses
never cracked anywhere (0x8013B83C, 0x8013BD74, 0x8013C08C) -- the last are genuine
drafting work, now finally possible since their TU is -O0.
I under-counted this cluster 8x (reported 275 stubs/18 overlays; truth 2,184/138). The
scan ran during a background rebuild AND wrapped corpus.stubs() in `except: continue`,
so every R32 coverage refusal became a silent skip and the total was taken over the few
overlays that happened to be re-extracted already.
Two of our own rules broken at once: a measurement taken during a rebuild is not a
measurement (caught EARLIER the same session, by the same assertion I then suppressed),
and R32 lives in the CALLER — an oracle only asserts coverage if the caller lets it raise.
It also cost credibility the other way: I used the bad number to declare the T0(f)
"2,192 open members" pin STALE. The pin was right. R35 applies to a re-measurement as
much as to the original measurement.
Checkpoint updated with the corrected population and the completed fleet-wide sweep.
Applies tools/o0_subsplit.py across every overlay whose 0x8013B568..0x8013C98C -O0
cluster was trapped inside an -O2 jr split. This is the population the phase opened
against, and it has never been buildable-at--O0 before.
135/135 sub-splits applied, 0 tool refusals
140 new -O0 region files; corpus.o0_sources() 137 -> 277
0 of them invisible to the -O0 oracle (verified explicitly -- a SILENT -O0 miss is
the exact failure SS126 warns about: the region would compile -O2 and every residual
it produced would be a pure artifact)
**2,200 open stubs now live in a genuinely -O0 translation unit**
R22 CLEAN-FLEET: extract-all 139/139 (+main); check-all 140 passed, 0 failed of 140.
make tools-health OK: corpus(+resident) 0 PHANTOM/0 TRUNCATED; cdecl ALL ORACLES GREEN;
audit-binaries 140 onboarded, every one a full citizen (R36); dedup-check 1904/0,
C1 coverage 240359/240359; cookbook-index 336 sections.
The transform is byte-neutral by construction (it only moves subseg boundaries and
repartitions source), so the whole batch was gated by one clean-fleet R22 rather than
135 individual builds -- after a single-overlay probe (ov_SC01_000) proved it (R37).
NOTE THE POPULATION CORRECTION (R14, mine): I earlier reported this cluster as "275 open
stubs across 18 overlays". That was WRONG -- the sizing scan ran while R22 was rebuilding
in the background, so corpus.stubs() raised for most overlays and my bare `except:
continue` SWALLOWED the very coverage assertion R32 exists to raise. The true figure is
2,184 open stubs across 138 overlays, which vindicates the T0(f) pin of "2,192 open
members" that I had called stale. The target list for this sweep was rebuilt with NO bare
except, so R32 can do its job.
Drafting fuel confirmed present: cached Ghidra-C seeds exist for the cluster's addresses.
Drafting the 2,200 is crack-wave work (T3), not T2.
- T0.5 was COMPLETE in SESSION-27 (124/124 programs, 7,716 files) but left unticked;
corrected, with the omission noted rather than silently fixed.
- T2 marked [~] part-done and its DEVIATION recorded against the phase-start plan:
the PRIMARY (two-file atomic o0b append) was byte-refuted, and the FALLBACK premise
was ALSO wrong — Arm-A does not bite. What actually blocked it is §126. Route proven,
tool shipped, 6 banked; the 18-overlay sweep of the 0x8013B568 cluster remains, sized
at 275 open stubs (re-derived; the T0(f) "2,192" pin is stale).
Records the T2 result as the phase's biggest unblock: the carve-within-a-carve is
byte-neutral (Arm-A does NOT bite), the real constraint is that an address range is
not an optimization region (SS126), and tools/o0_subsplit.py implements the correct
bound. Measured, not assumed, what it unblocks: 275 open stub instances across 18
overlays in the 0x8013B568..0x8013C98C cluster, homed in an -O2 jr split — plus a
note that the T0(f) "2,192 open members" pin is STALE and must be re-derived before
costing (R37).
Also flags my own under-count: the "15 contiguous -O0 fns" came from an asm scan that
cannot see matched functions.
Promotes the proven probe (commit:1266) into a real tool, and validates it FIRST-TRY on a
fresh overlay.
tools/o0_subsplit.py <ov> --lo <vram> --hi <vram>:
- derives the range's contents from the SOURCE ANCHORS (overlay_src_split.parse_overlay_c:
`asm` = unmatched stub, `define`/`def`/`nonmatch` = already matched), NEVER from an asm
scan -- a matched fn emits no .s, which is exactly the blindness that made the range look
like a clean contiguous run (SS126 / SS124's shape);
- computes the -O0 bound as (address range MINUS already-matched bodies) and emits ONE
sub-region per maximal run of unmatched anchors (K matched islands => K+1 regions);
- names each `<ov>_o0<letter>` picking free suffixes, so the widened Makefile glob selects
them; refuses loudly if it runs out or if the range spans >1 object or is already -O0;
- honours the one-carve-per-region law (forces a cut at every already-banked jr in the
object) and reuses jr_isolate_all's plan/build_new_config/ascending-unique validation
verbatim, so carve-repoint + source-repartition stay on the proven path;
- warns (does not refuse) when a stub in an -O0 run lacks the frame-pointer prologue --
the byte-gate is the arbiter, not the heuristic.
VALIDATION on ov_SC03_015 (untouched by the manual probe): the tool independently derived the
SAME structure found by hand on ov_SC03_014 -- 2 matched -O2 islands (func_80184440,
func_801848E4), 2 -O0 regions (8 + 7 fns), same 5 cuts. Sub-split -> BYTE-IDENTICAL. Then 3
drafts, each global DERIVED FROM THAT OVERLAY'S OWN ASM (%hi operand) rather than copied:
3/3 match_one --o0 MATCH (22 ins), 3/3 through the whole-binary gate.
BANKED this commit: func_801846E4 / func_8018473C / func_80184794 in ov_SC03_015 (6 across
the two overlays now). The other 24 stubs in the region are undrafted -- the route makes them
DRAFTABLE (they were un-bankable at any effort before); drafting them is crack-wave work.
R22 CLEAN-FLEET: extract-all 139/139 (+main); check-all 140 passed, 0 failed of 140.
cookbook SS126 (the address-range-is-not-an-optimization-region law + the probe ladder).
T2's central unknown is resolved, and it is NOT what the phase plan predicted. The plan
named the Arm-A splat `%lo +0x20` defect as T2's real substance. Four probes, each
isolating ONE variable, SHA vs config/check from a clean tree (SS125 rules):
1. sub-split the jr object at ARBITRARY addresses, pure -O2 -> BYTE-IDENTICAL
** the re-carve is NEUTRAL; Arm-A does not bite here **
2. same split, middle region routed to -O0 -> DIVERGED (2 vars at once)
3. same NAME as probe 1, only the -O0 flag added -> DIVERGED
** therefore the FLAG, not the subseg name **
4. -O0 regions cut to EXCLUDE the matched bodies -> BYTE-IDENTICAL
** route PROVEN end-to-end **
THE REAL OBSTACLE: an address range is not an optimization region. Interleaved among the
15 -O0 stubs at 0x80183CF0..0x80184920 are TWO already-MATCHED functions (func_80184440,
func_801848E4) that expand from engine_core.h and compile at -O2. Flipping the FILE to -O0
recompiles them too. Cut around them and it is byte-clean.
MY EARLIER "15 contiguous -O0 fns, clean cut" WAS AN UNDER-COUNT (R14 on myself): I derived
it by scanning asm/**/*.s for the frame-pointer prologue, and a MATCHED function emits no
.s -- so the scan was structurally blind to exactly the bodies that break the flip. Same
shape as SS124. Any T2 driver must derive -O0 bounds as (address range MINUS already-matched
bodies), never from an asm scan.
BANKED (3, whole-binary gate the sole arbiter): func_801846E4 + siblings func_8018473C /
func_80184794. Each global DERIVED FROM THE ASM (%hi/%lo operands), not taken from the
draft's comment; 3/3 match_one --o0 MATCH (22 ins) with a -O2 control showing the mismatch;
3/3 verified through harvest_verify; bank truth read from the SOURCE (SS55b trap 4).
MAKEFILE: the -O0 glob widened `ov_*_o0b.c` -> `ov_*_o0?.c` so ANY lettered -O0 sub-split is
covered by one rule. A MISSED rule is silent -- the region would compile -O2 and every
residual it produced would be a pure artifact (SS116). corpus.o0_sources() re-verified: 137
sources, resolves `?` via glob, both pre-existing rules intact.
CONFIG: ov_SC03_014_jr_8017EB7C sub-split into 5 regions (pre / _o0c / matched-O2 /
_o0d / post), reusing jr_isolate_all's plan+build_new_config+validation verbatim so the
carve-repoint and source-repartition semantics are the proven ones.
R22 CLEAN-FLEET: extract-all 139/139 (+main) ; check-all 140 passed, 0 failed of 140.
Max-effort re-measurement of the three jr refusals I ledgered earlier this session.
Two of the three verdicts were FALSE. Every number below is SHA vs config/check.<ov>.sha
from a clean tree, with the restore re-verified.
func_8018057C / ov_SC01_009 : jr_isolate_all is BYTE-NEUTRAL
-> "JR-ISOLATE-BREAKS-BYTES" RETRACTED; original failure not reproducible.
func_80191C50 / ov_SC06_018 : isolate NEUTRAL -> carve DIVERGED
(got 1b1667ea, want cbbc4f44) -> the ONE real instrument failure. CONFIRMED.
func_8017BEBC / ov_SC04_004 : carve is BYTE-NEUTRAL (body-free)
-> failure is the TEMPLATED BODY, the OPPOSITE of what SS125 first claimed.
Re-probed once more from a verified-clean tree: still gate-fail. Reclassified
BODY-TEMPLATE-GATE-FAIL.
So the tidy "two apparent walls are ONE tooling problem" conclusion was wrong: they
are two different problems, and the third target has no demonstrated problem at all.
ROOT CAUSE, and it is mine not the tools': a grep-of-the-build-log gate inside a driver
that did not revert on abort. config/overlays.mk is SHARED, so target 1's half-applied
isolate was still in the tree while target 3 was measured. Separately reproduced the
SS42b stale-object trap head-on: `git checkout -- config/` WITHOUT a re-extract turned a
byte-identical overlay into [FAIL] got 8f28aa77 / want 38a3d919 (Phase-20's R22
corollary, live).
SS125 rewritten. The METHOD (split the carve from the body, one build) is kept and is
what refuted this section's own first conclusion; what is added is the instrument rules
that make its answer trustworthy: compare the SHA against config/check, never grep the
log; re-extract after every config change AND every revert; a driver that aborts a
target must revert it before the next; verify the BASELINE against canonical too.
Meta-lesson recorded: SS53 says a 0% from the wrong TOOL manufactures a doctrine — this
is the same failure one level up, a verdict from the wrong MEASUREMENT, and my own
diagnostic script is an instrument subject to R35 like any other.
Ledger corrected in place (3 entries, superseding the earlier misattributions), so the
scheduled repair is the right one. No source/config change; no bank affected; the fleet
is untouched at 140/140 (last full R22 this session, HEAD commit:1263).
The session's most useful finding is an instrument ticket, not a match.
SS125 (new): before ledgering any jr residue, run jtbl_carve with NO body spliced
and rebuild. Byte-identical => the carve is neutral and the failure is the template;
NOT identical => the failure is the carve and the body was never fairly tested.
One build, and it collapses ambiguity that SS53 warns has twice steered strategy.
MEASURED: group B func_8017BEBC had gate-failed 3 probes in a row (default AND
--raw, cross-address ov_SC02_015 AND same-address ov_SC04_004). Carve-only on
ov_SC04_004 broke the bytes with nothing spliced — so all 3 probes were testing a
body that never got a fair run. The SAME stage had already refused behemoth
func_80191C50/ov_SC06_018. Two "unrelated walls" = ONE tooling problem.
func_8018057C/ov_SC01_009 fails at a DIFFERENT stage (jr_isolate_all, step 1) and
is deliberately NOT grouped with them.
Both tools reported SUCCESS on every failing target; only the whole-binary gate
refused (G3/P9). A tool's exit code is not the oracle.
Ledger: the three logged by STAGE (JTBL-CARVE-BREAKS-BYTES / JR-ISOLATE-BREAKS-
BYTES), not by function, so a carve fix auto-reopens every target it should.
None is diagnosed, so none is called a compiler wall — that guess has been wrong
four times running on this project (R35).
CURRENT_PHASE: SESSION-28 checkpoint refreshed; the carve diagnosis is now resume
item 1 (it gates 13 members + a 710-ins behemoth and every future jr family).
The 3 jr behemoths carried from SESSION-27 (staged .run/beh-gate/), each run through
the SS81 3-step chain SERIALLY under tools/treelock.sh (every step calls `make extract`).
Re-derived the stub set on a QUIESCENT tree first: 4 of the 7 staged fns were already
banked in S27 and the staging dir still held all 7.
BANKED (1):
ov_SC05_010 func_80181CDC (769 ins) — jr_isolate_all --only -> BYTE-IDENTICAL,
jtbl_carve --func -> BYTE-IDENTICAL, then harvest_verify --chunk 1. Bank truth read
from the SOURCE (INCLUDE_ASM absence), never the gate report (SS55b trap 4).
BYTE-REFUSED (2) — to the wall ledger WITH EVIDENCE, not forced, not labelled walls:
ov_SC01_009 func_8018057C — jr_isolate_all reported success (2 jr in 1 -O2 object,
2 region .c) but the post-isolate build was NOT byte-identical. Aborted at step 1.
ov_SC06_018 func_80191C50 — isolate clean; jtbl_carve reported success (48-piece
carve set, 22 tail pieces) but the post-carve build was NOT byte-identical.
Aborted at step 2.
Both TOOLS claimed success; the whole-binary gate refused (G3/P9 doing its job).
Neither is diagnosed, so neither is called a compiler wall — on this project that
guess has been wrong four times running (R35). Drafts + logs preserved under .run/.
R22 CLEAN-FLEET (config changed => T2 blast radius => mandatory):
extract-all: 139 extracted, 0 failed of 139 (+ main, serial)
check-all: 140 passed, 0 failed of 140 <- BYTE-IDENTICAL
DRIVER DEFECT FIXED (mine, SS105/SS61): v1 of .run/s28_beh.sh `continue`d past a failed
gate WITHOUT reverting, leaving two half-carved overlays in the tree; only a full
`git checkout -- src/ config/` recovered it. Nothing was ever committed, so the cost was
build cycles and zero work. v2 reverts that target's config+src on EVERY failure path.
- §124: a "not matched" verdict can mean the definition is there under a DIFFERENT
C NAME (the §37/§73 asm-label alias). The whole "no matched unit" skip class was
one exemplar x 137 members. Includes the two traps in the fix (re-derive the
pattern PER FILE; CARRY the alias declaration or the sibling emits the wrong
symbol and still links) and the law: when corpus.stubs and a source scanner
disagree, the SCANNER is wrong.
- §124a: `0 matched-exemplar families` from family_sweep may be the --band FILTER
(defaults to `substantial`), not a wall — same shape as §53 / §116.
- cookbook-index regenerated (tools/cookbook_index.py, R33).
- CURRENT_PHASE: SESSION-28 checkpoint. Fleet 92.70 fn-count / 88.3 instr / 78.7
distinct, R22 140/140. Records the 4th -O0 region VERIFIED + SIZED (30 instances
/ 1,504 ins, ov_SC03_014+015 only) and correctly BLOCKED on the T2 re-carve, and
two R14 corrections to my own SESSION-27 checkpoint (the 137 were skips not
failures; the cause was not "banked in the wrong binary").
The skip class that the SESSION-27 checkpoint carried as "~137 members, a cheap
sweep-routing gap" is closed: ONE exemplar, ALL 137 same-address members banked,
0 failed, in a single --only sweep after the extract_unit alias fix (commit:1260).
- family_sweep --hseq --only 0x8016191C --band all -j 8, under tools/treelock.sh
(the campaign-wide mutex, not a pgrep poll).
- BANKED 137 member-matches / 0 failed across 137 overlays; skipped {}.
- 24 ins x 137 = 3,288 instructions that were sitting behind a tool lookup miss.
- R37 probe before the sweep: rtu_match MATCH (24 ins) on 3 members in the REAL
TU (ov_SC01_000 / _001 / _004) before spending 137 gate cycles.
R22 CLEAN-FLEET: make clean && make extract-all && make check-all ->
extract-all: 139 extracted, 0 failed of 139 (+ main, serial)
check-all: 140 passed, 0 failed of 140 <- BYTE-IDENTICAL
0 NON_MATCHING in any default build (G4). The whole-binary byte-gate was the sole
arbiter for every one of the 137 (G3/P9).
Note --band: the sweep defaults to band=substantial and this family is band=mid,
so the first invocation returned "0 matched-exemplar families" — a silent-looking
zero that is a FILTER, not a wall. --band all is required for mid/tiny families.
The h_seq sweep's 137 "no matched unit for func" skips were ONE exemplar x 137
same-address members, not 137 distinct failures: func_8016191C @ ov_SC01_077,
band=mid, non-jr, 24 ins, ALL 137 members still INCLUDE_ASM stubs => 3,288 ins
left on the floor by a tool lookup miss (R35 again).
ROOT CAUSE (and it refutes the SESSION-27 checkpoint's own diagnosis, R14): the
exemplar is NOT "banked in a different binary". It is banked in ov_SC01_077 under
the SS37/SS73 asm-label alias — the byte-true body conflicts with the fleet-canonical
decl on BOTH SS73 axes (return void vs s32 AND params void*/s32 vs int/unsigned),
so it was banked zero-touch as
int aF8016191C(int, unsigned int) __asm__("func_8016191C");
extract_unit only ever matched a definition head literally NAMED func_<ADDR>, so
it returned None, and every caller reads None as "not matched".
FIX (tool-only, T0 blast radius):
- _alias_decl_for(): resolve `<ident>(...) __asm__("func_<ADDR>");` -> <ident>.
- extract_unit(): accept the alias identifier as the definition head, re-derived
PER FILE so one file's alias can never leak into the next.
- carry the alias DECLARATION into the unit (without it the sibling emits the
symbol aF8016191C and the body never lands at func_<ADDR>); the backscan now
walks PAST the alias line so the fn's own preceding externs are carried exactly
as for a plain definition, with a start<=alias_ln<=end guard against double-emit.
- R32: an alias decl with no findable definition now REFUSES LOUDLY instead of
falling through to _macro_unit and reporting "not matched" — the silent-skip
class this fix exists to delete.
Rejected the alternative the banking agent suggested (widen engine_core.h's decl
void->s32, drop the alias): it addresses only SS73's RETURN axis while decl and body
also disagree on PARAMS, and it is a T2 fleet-shared edit where the alias is T0.
VERIFIED: unit extracts with externs + exactly one asm label; remap_hseq produces a
sibling draft; rtu_match 3/3 MATCH (24 ins) in the real TU (ov_SC01_000/001/004).
make tools-health RC=0. The whole-binary byte-gate remains the sole arbiter (G3/P9).
The 12 agents killed by the usage-limit pause were resumed and ALL returned MATCH (2 had already
banked from their partial drafts, so 10 ran). h_exact propagation leg completed over all 112 banked
exemplars: 14 propagated, 42 benign skips (h_seq tier, correctly routed away per §123), 0 failures
— the 0x801466F0 'halt' was a third benign-refusal phrase, not a partial write.
fn-count 92.61 -> 92.67% | instr 88.2 -> 88.3% | distinct 70,581 -> 70,590 unique fns.
Every line is a symptom an agent HIT and had to re-derive from gcc internals because title-keyword
search structurally cannot surface it:
- SIZE-MISMATCH/short + frame-pointer prologue => the target is -O0, pass --o0 (the flag was
documented nowhere an agent would look; a 4th -O0 region also exists beyond the 3 known ones)
- rotated instruction window => sched1 order; brute-force all N! statement orders (24 runs, 2 min)
- if/else result in $v1 vs target's $v0, and load-hoisted-above-store => §76 variable reuse
(§76's title reads behemoth-only, so nobody finds it for a 48-ins function)
- ori 0xffd8 vs addiu -0x28 => negative const in an UNSIGNED narrow local; signed keeps the lhu
- LENGTH-DRIFT -1 as a missing jal-delay copy => narrow ANSI prototyped param (not just K&R §43)
- lwl/lwr+swl/swr is a delay-slot SPONGE (the inverse of the §5a fence case)
- a vanished param copy => cse.c make_regs_eqv live-range rule
- a ghidra_c seed may be a DIFFERENT function (overlays share VAs)
WAVE 3 (48 agents / 8 binaries, dealt across binaries so BANKING fans out): 48/48 match_one MATCH,
48/48 banked through 8 PARALLEL per-binary gates. WAVE 2: 15/19. BEHEMOTHS: 4 non-jr confirmed
(func_8017E120 884ins x14, func_8017FA5C 728, func_8017CAD4 755, func_8017E35C 719).
Tier-routed propagation (§123): family_sweep --hseq banked 911 members across 137 overlays.
fn-count 92.32 -> 92.59% | instr 87.9 -> 88.2% | distinct 78.3 -> 78.7% (70,506 unique fns)
R22 clean-fleet 140 passed / 0 failed, under one campaign lock (treelock.sh).
CORRECTION (R14): the 'per-binary bank-rate cliff' I reported from the pre-incident gate run
(SC03_014 1/6, SC04_018 1/6, SC06_018 2/6) was an ARTIFACT — those gates ran against a tree
propagation was concurrently rewriting. Re-gated clean: 6/6 everywhere. A measurement taken during
corruption is not a measurement; I should not have theorised a cause before re-running it.
I gated 8 binaries in parallel while wave-2's propagation loop was still running, then ran
'make clean' on top. check-all 77/140; the corpus denominator moved, so the apparent 91.4% instr
was a half-written tree, not a gain. Reverted to commit:1245 (last R22-verified) — 140/140 restored,
all 58 drafts survived because agents only ever write .run/.
ROOT CAUSE, and it was structural not unlucky: my guard was
while pgrep -f dedup_propagate; do sleep; done
A CAMPAIGN is a LOOP of short-lived processes (15 sequential invocations), so it has gaps where no
process matches. The poll sampled a gap and started. Presence-of-a-process cannot express 'a
campaign owns the tree'.
treelock.sh holds one flock for the WHOLE campaign, released by the kernel on exit OR kill, with
--status; both drivers refuse to run unlocked. LAW: guard the CAMPAIGN, not the process.
Corollary (twice today): a killed process performs no undo — a fleet-tier write needs a lock ABOVE
it, not cleanup inside it.
run_gate() always took per-worker result paths; the CLI never exposed them, so every CLI gate used
the shared .run/harvest_{verified,failed}.txt. Concurrent per-binary gates — safe on every other
axis, since the byte-gate IS per-binary — would have read each other's results and mis-attributed
banks (the §55b trap-4 shared-scratch defect that bit match_one in P28, one level up). Default is
now .run/harvest_verified.<binary>.txt: safe by construction, not by remembering a flag.
WAVE-2 MEASURED THE INDEX: 15/19 index hits and the bank rate went 57% (wave 1, no index) -> 79%
(wave 2, index-first) on the same gate. Agents also NAMED its gaps, which is the flywheel working.
Two real defects found and fixed:
1. COVERAGE. The parser required a '§' prefix, so 111 h2-h4 headers were invisible — including
'### T4 — Branch polarity', the fix match_one names by class (BRANCH-POLARITY) and which two
agents re-derived by hand, and the §1/§2 idiom-catalog entries (I1-I4, T1-T4).
2. THE ASSERTION ITSELF. My R32 check compared §-headers-parsed against §-header-CANDIDATES — a
tautology over a set I had already narrowed. R32 says the candidate set must OVER-approximate;
it now counts EVERY header and accounts for each as indexed-or-explicitly-skipped. The tool
written to stop silent skips had the silent-skip defect.
3. Keyword matching over titles cannot surface an idiom whose title omits the symptom, so the
index now opens with a hand-curated SYMPTOM -> section list, seeded from what agents actually
hit (branch polarity, (void)-canon conflicting types, asm-label alias, one-base-register reuse,
folded andi, slti/sltiu, sibling-first, delay-slot theft, void->s32).
Verified against code rather than the ledger: the split-blind lookup globs */ correctly, and the
churn was fixed by T5's input-signature gating. Both struck. asm_subdir_for was still a parallel
oracle (silent g[0] on multi-match) -> now corpus.asm_path. --fix-def-sig defaults off correctly
but advertised 'Byte-neutral; gate arbitrates' — the claim T84 refuted (signedness-wrong header
decl over a byte-correct draft; 137 members held at 0 until the flag was dropped).
A defect ledger nobody re-verifies decays into busywork — verify before scheduling.
Wave-1 measured the tax: three agents each reported a 'NEW idiom' that was ALREADY documented —
the asm-label alias (line ~2516, same 'address-of perturbs regalloc' mechanism) and the void->s32
non-neutrality (§41d, Phase 26; the agents cited the very entry §41d corrects). They consulted the
cookbook as instructed and could not FIND them. 716 KB / 226 sections with no index = a
discoverability failure, and every wave re-paying for prior waves' findings is the inverse of R16.
docs/cookbook-index.md maps SYMPTOM (what you see in the diff) -> sections, 14 buckets, a section
listed under every symptom it addresses. Derived by tools/cookbook_index.py (R33 — cannot drift),
--check wired into tools-health.
R32 on my own tool: the first regex required an em-dash separator and silently dropped 50 sections
— including §1 (idiom catalog), §2, §5a (cross-jump, cited by an agent today). An index missing its
most-cited entries turns 'I could not find it' into 'it is not there'. Now asserts extracted ==
candidate '§' headers and hard-exits on a gap.
The 4 cores dedup_propagate refused (h_exact tier) templated cleanly via family_sweep --hseq once
the family map was regenerated post-bank (a bank invalidates the map: sig-overlays + family_hseq
must run BEFORE the sweep — the standing wave-loop order). 548/686 banked, 138 failed (one
consistent per-overlay slice, diagnose next). Wave-1 total: 8 cores -> 1,121 instances.
instr 87.5 -> 87.9%, distinct 69,828 -> 70,094 unique fns.
Ultracode wave of 14 agents over fresh reach-138 cores: 14/14 match_one MATCH, 8 accepted by the
whole-binary gate (the §52b law reproduced exactly). Propagated per-function (the incident fix):
0x8012E014, 0x80151C54, 0x8012F49C, 0x80151B98 -> +573 instances. R22 clean-fleet 140/140;
instr 87.5 -> 87.7%, fn-count 92.00 -> 92.16%, distinct 69,828 -> 69,836.
The other 4 banked cores are h_seq (PURE/IMM) families: dedup_propagate is h_exact-only, so its
'reach<2' / 'not self-contained' refusals were statements about the TOOL's tier, not the functions
-> cookbook §123 (the §53 carve-law generalized to the propagation-tier axis) + a routing table.
They bank via family_sweep --hseq next.
A killed process performs no undo, so a fixed timeout on a fleet-tier write is a tree-corruption
mechanism, not just a delay. Now: timeout = min(6h, 1800+1800*banks); on expiry the driver reports
TREE DIRTY, REVERT REQUIRED and returns cleanly. Standing practice for multi-bank waves: re-gate
--no-propagate, then propagate PER FUNCTION (dedup_propagate --addr) — bounded and resumable.
Wave-1 result recorded: 14/14 match_one MATCH -> 8/14 banked (the §52b law); 3 new idioms owed.
39% prior did not generalize (S16 measured FRESH wave drafts; this is A10's stored-backlog class,
0/958 by plain re-gate) — the driver lifted ~16% over that 0%. Residue routed to T3 redraft lanes.
§61 orphan-carve residue reverted; two T3 pre-work gaps recorded (gate_stage commit add-scope for
new carve files; no tracked writes during tree-writing campaigns). Ledger pruned: 1,350 -> 1,332.