Found by the Fable blocked-pile audit. `jr_isolate_all.jr_inventory` resolves each
committed .rodata carve's owner through `family_remap.reloc_targets`, whose
`nins_of` reads the gitignored `.run/sig.<binary>.jsonl`. A fresh worktree has no
`.run/sig.*`, so inside a worker EVERY carve reads UNOWNED, jr_inventory
R32-aborts, harvest_verify prints `isolate FAILED`, and the draft is booked
CARVE-REFUSED.
That verdict was about the WORKTREE, not the function. Measured on
ov_SC02_000/func_8017F950 (a RELOC-ONLY twin whose body rtu-MATCHes 117/117):
dry-run isolation passes in the main tree and aborts in the worktree with 30
phantom UNOWNED carves. Linking one file is the whole difference. When the file
is absent it is now reported in missing_generated rather than silently skipped.
This invalidates the CARVE-REFUSED rows I quoted in the S69 census — they were
instrument verdicts, and the class is far smaller than recorded.
Also adds tools/asm_verbatim.py (new): .s -> §265 file-scope __asm__ block with
decimal immediates/offsets and comma-no-space operands (maspsx dies on
`sltu $v0, $s0, $v1`), derived .frame/.mask, R43 refusals for rodata/jtbl.
Ledger MATCH 12 / NEAR 1 / REFUSED 2 plus a non-wall control. Byte-equivalent to
the stub by construction — for genuine hand-asm only; §265 accounting applies.
Correcting my own guidance from earlier today. §378 gave the self-caller chain;
three more variants appeared within hours and two of them BREAK the chain.
Variant 3 (NEW, byte-proven ov_SC04_018/func_8017F35C, banked): conflicting
RETURN type on a decl that is ALREADY no-proto, where the symbol is
ADDRESS-TAKEN rather than called. --any-proto has nothing to relax and
cast_self_callers has no call site to cast; --sync-decls ALONE fixes it, and is
safe precisely because an address-taken site has no arguments to convert.
Variant 4 (REFUTATION of what I wrote in the playbook this morning): "run the
same chain on the callee the diagnostic names" is wrong at scale. Applied to
func_8012AD44 in ov_SC07_000 it no-protoed 60 caller decls and the binary went
RED (265b24bb vs 9dbe4241); reverted via journal. The self case is safe because
step 2 casts the call sites so the decl change cannot alter argument conversion;
for a callee, cast_self_callers correctly refuses and the decl change runs
unprotected. It banked main/func_80021D38 only because that callee had ONE decl,
not sixty.
Rule added: never --any-proto a symbol whose call sites you are not also casting;
count the sites first. The chain is a DECISION TABLE, not a sequence to run
blindly.
The hard gate caught me: m1/m2 (§379-§383) and the fable escalation (§385-§388)
were harvested, but o1-o4 and p1-p3 were not — 57 MATCH notes sat unbanked while
I was about to draw new waves.
§392 — seven byte-proven spelling levers, each of which closed a match on its own:
(a) a same-address dual-sign read is fixed by ORDER (emit the unsigned
store-source read first); cse merges lh/lhu for every cast spelling tried
(b) a narrow temp picks the narrow load — s16 vs s32 decides lh vs lhu, and a
signed decrement temp yields "sll 16" where unsigned yields "andi 0xFFFF"
(c) tbl[idx-2] folds -8 into the lw offset; hoisting the subtract forces addiu
(d) identical switch arms must be SEPARATE case blocks — the target duplicates
arg setup per case and cross-jump-merges only the shared tail
(e) distinct pseudos per repeated inline copy — one shared pair biases sched1's
tie-break for the first copy only
(f) split the widen into two statements to move the sll off a pinned register
(g) the RETURN TYPE alone closed a 7-ins schedule residual (s32 -> void)
§393 — the BIRTHING BOOST: a single-set local gets max scheduling priority and
sched2's backward pass pushes it LATE; a zero-byte re-tie gives it a second set
and kills the boost. The scheduler-side sibling of §380 — same trick, different
pass, opposite symptom.
§394 — two align-1 accesses in one function reserve a phantom 8-byte stack slot;
a frame 8 bytes too large with no spill to account for it is the tell.
tools/seed_ref.py gains --contained/--contained-control: an open stub that is a
banked body plus or minus WHOLE BLOCKS — the class edit distance ranks badly.
Branch-offset masking was required (unmasked offsets veto exactly the target
pairs) and a min-side-25 floor (89% of raw hits were prologue/epilogue vacuity).
Ranks by (substitutions+regions, cover), not by d. Controls: planted-deletion
positive 60/60, random-pair base rate 0/397, R32 population 346/346, and a
post-refactor --near regression reproducing the stored slice exactly.
Banked on first use: ov_SC01_077/func_80184D50 = banked ov_SC03_007/func_8018283C
minus its trailing `&= 0x7FFFFFFF;` — MATCH, closeness 0, 98/98.
* cookbook §390: minimum distance is not minimum work (rank by effort; a deletion
is free, a substitution is thought), the lookalike filter r = d/min(nins) ~ 0.3
(17 of 30 "cousins" were boilerplate coincidence), and the three fleet-wide
nulls that close the scanner question — 0 new / 9 / 2. Spend integration
effort, not scanner effort.
* cookbook §391: a byte-aligned struct copies in FOUR instructions (lwl/lwr/swl/
swr), a word-aligned one in TWO. Never invent an aggregate type to make a draft
compile — an invented word-aligned Blk8 lost exactly 8 ins across two copies and
read as a believable "near, closeness 70" codegen residual.
* accelerators #18: a claim derived from BYTES is not a claim verified by a
COMPILER. Every similarity/correctness claim must name the tier it reached
(stream containment / compiled standalone / whole-binary gate / clean fleet);
a report that says "verified" without one invites the strongest reading.
Non-reproduction is a finding — say so rather than assuming your own setup.
* playbook §2a-2: the twin ladder (exact -> RELOC-ONLY -> CONTAINED -> cousin ->
cold), take the cheapest tier available, widen only when the tier above is empty.
* SETUP inventory row; generic-decomp-package: rank by work, and stop building
scanners once the well is dry.
The exact-hash twin tier found 22 of 352 reachable open stubs (6%). The
edit-distance band added by `seed_ref --near` finds 75 of 352 (21%) — 3.4x — on a
corpus we believed fully mined. 31 of the new rows were PURE reloc-only twins of
already-banked bodies; 8 banked the same day at ~0 agent tokens, one 94-ins
exemplar serving five open copies.
* cookbook §389: the h_norm hole (norm_stream drops its pending lui-hi on an
intervening R-type, so indexed-global reloc twins hash differently and vanish
from seed_ref/twin_sweep/dedup/family-maps at once). Do NOT fix h_norm — every
stored calibration keys on it; the near tier reads through it.
* accelerators #17: the generalisable law. A similarity hash built for DEDUP
under-matches by design, which is correct for dedup and silently lossy as a
FRONTIER join — the two questions want opposite error directions, and the
frontier failure looks exactly like "this function is unique".
* generic-decomp-package §2b: build the near band at the same time as the exact
tier, with the three verifications. It pays from the first bank for a new
project, where we paid a session to recover the debt.
* SETUP inventory row + playbook §2a (run it before believing any "no twin"
verdict; never send a RELOC-ONLY row to a drafting agent).
Measured over 129 drafting agents in one session, per MATCHED instruction (the
only cost that matters, since a failed agent is billed in full):
sonnet 105 agents, 57 MATCH 4,289 tok/matched-ins (flat ~47% above 30 ins)
opus 24 agents, 11 MATCH 2,083 (m1 191-347: 1,291, 67%)
opus at 347-670: 1/9 7,158 <- the cliff, 2.92M tokens for ONE bank
fable escalation: 3/4 closed at ~1/3 the cost of the attempt it rescued
Sonnet's per-agent price was never the cost that mattered; cost per BANK is, and
it lost on that by 2.1x. The m2 wave should have been fable from the start.
Escalating SOONER is the standing finding — higher models crack harder functions
in fewer tokens. Tested twice now (S68 A/B, S69 measurement); do not re-derive a
cheap-tier argument from per-agent price a third time.
The 'false bank' in the S69 checkpoint was not one. Both instances verify
byte-identical after 'make extract BINARY=<b>'. §384 states the law (verification
must regenerate whatever the gate changed the inputs to), the trap inside it (a
src-only revert of a carve commit produces 'table-count drift vs the carve', which
reads like progress), and the give-away I ignored — the commit diffstat showed
config/overlays.mk and a splat yaml sitting next to the .c.
Reverts commit:3475. The bank is byte-identical; MY VERIFICATION WAS BROKEN.
A jtbl bank changes CARVE CONFIG (JTBL_PADS in config/overlays.mk + the splat
yaml). Those are splat INPUTS: asm/ and the linker script are regenerated FROM
them. I checked the binary with `make build` alone, so the build linked
newly-carved C against STALE extracted state and produced a mismatched SHA. That
is the R22 corollary ("a reverted config needs a make extract, not just a make
check") pointed the other way — a LANDED config change needs one too.
Proof, run on both binaries:
make extract BINARY=ov_SC06_025 && make build -> BYTE-IDENTICAL
make extract BINARY=ov_SC04_011 && make build -> BYTE-IDENTICAL
So: R40 against myself. I attributed the failure to the subject (the bank) when
the instrument (a build over stale extract state) was at fault — after writing
"it may not even be false" into the checkpoint and reverting without testing it.
The first revert also cost real work: it discarded a legitimate 96-line match.
STANDING FIX: a per-binary verify after any gate that touched config/ MUST be
`make extract BINARY=<b> && make build BINARY=<b>`. Build-only is a valid check
ONLY when the gate changed nothing under config/.
The lever existed but nothing downstream applied it. Proof it mattered: a wave
agent this session diagnosed its own blocker as "§378 THE SELF-CALLER CAST, a
TU-level fix (cast_self_callers.py) that requires editing src/, which I'm not
permitted to touch" — the knowledge propagated, the automation did not.
* recover_integration.py: NEW "self-cast" stage (tier=binary), so the driver can
run the whole chain as --stages arity,self-cast. The docstring states WHY the
order is not arbitrary: self-cast answers the error that "arity" CREATES.
* residual_rules_b.py: both decl-conflict tiers now prescribe the full chain
instead of "route to integration / budget for banking", and
NOCOMPILE-UNDECLARED-FIXED now says outright NOT to gate the autodecl arm (it
is a second conflicting declaration in the real TU).
* wave-playbook §4b: replaced the stale two-step recipe with the three-step
chain, the one-driver form, the callee variant, and the MANDATORY
--undo-journal.
* SETUP.md: full inventory row (R21) — it had zero mentions.
Not wired, deliberately: gate_stage's ladder rewrites DRAFTS via _xform, while
this edits the TU; a src-side edit inside the automatic gate needs
revert-on-failure, which recover_integration already owns.
Still open: a draft_prechecks rule to catch the self-decl conflict statically,
before a build is spent. The new stage's plumbing is verified (CLI + candidate
selection); its functional end-to-end run is NOT — gate12 held the tree.
cast_self_callers casts a function's call sites in PREPARATION for banking it.
When the draft then fails, the cast must come back out — the tool journals every
edit for exactly that, and I did not run the undo.
The cost was concrete: the leftover cast on func_8017F8B8 made ov_SC07_000 fail
to COMPILE at HEAD, so every subsequent gate verdict on that binary was measuring
a broken baseline rather than the draft. Two drafting agents reported it as
BASELINE-RED before I noticed.
24 casts reverted across 11 files in 7 binaries; all 7 rebuild green. This is the
discipline recover_integration already documents ('REVERTS the caller edits for
anything that doesn't bank') applied to the new tool.
Reverts commit:3472. The binary was RED at HEAD: sha1 9c94d36a vs expected
8bc09c42. The gate that banked it ran with --r22 disabled because 24 drafting
agents were live (R22 does make clean, which deletes asm/ under them), so the
one check that would have caught it was the one I had turned off.
The revert must carry the CARVE STATE, not just the C: the bank moved
JTBL_PADS 0,0,4,4 -> 0,0,4,4,4 plus the splat yaml, and a src-only revert left
4 tables against 5 pad specs ('table-count drift vs the carve'). Reverting the
whole commit restores BYTE-IDENTICAL.
Found only because two drafting agents independently reported their target's
binary as BASELINE-RED and I checked their claim against the bytes.
main func_80020A28
main func_80021284
main func_800221A8
main func_8002374C
main func_80026514
main func_8002D904
main func_800377D8
main func_8003DC90
next-session-triage-ladder.md was still written as a to-build spec. It now leads
with the shipped status, the acceptance numbers, and the three things the spec got
wrong (the '32 free banks' were 0/28; the autodecl arm is worse in-tree than the
raw draft; PRE and POST cannot be the same pass because residual_rules_b needs a
draft), plus the one found by building it — never classify on a moving tree.
accelerators #16: a 'verified, just bank it' claim must name the compilation it
survived. Day-one kit material for a new decomp: any per-function oracle compiles
in isolation, every real bank compiles in context.
The first version of this parsed 'failed by class:' from the worker's stdout and
was INERT: the worker is gate_stage, which never prints that line (harvest_verify
does, one level down). classes came back empty for all 17 binaries of a batch and
the retry gate that consumed it fired ZERO times — a field that is always empty
makes its consumer a silent no-op (R54). Verified the claim only after re-reading
the log; correcting it here.
Now parallel_gate copies harvest_verify's <stem>.classified.txt out of the
worktree before teardown (it lives in the worktree's own .run/, which is not
symlinked and dies with it) and derives the class summary from those rows. That
also PRESERVES the verdict layer, which until now survived only as a side effect
of gater_lane re-running the whole binary in-tree afterwards (R47).
gater_lane judges the retry on the rows: a class with no per-function diagnostic
is the blind-worktree signature; anything cc1 named is a real compile error and
the serial rebuild would only reproduce it.
Verified live on ov_SC07_000: 'NOT retrying in-tree' fired, and the verdict row
landed at .run/gate_lane/ov_SC07_000.pgate.classified.txt.
tools/triage_ladder.py — the zero-token pre-agent pass, split PRE (target-side:
BANKED/WALL-332/PARKED, no build) from POST (residual_rules_b, needs a draft).
--escalate refuses a walled or banked target; --acceptance is the R39/R32 harness.
Refuses on a non-quiescent tree: a merging gate makes the stub oracle wrong in
both directions (measured, ov_SC01_004:func_8017EB30).
Acceptance, on the whole corpus: false-skip 0/1367 open stubs, recall 426/426
matched, wall tier fires on exactly the 10 enumerated walls (0 extra, 0 missing).
The first wall control asked for evidence that CANNOT exist — it scanned banked
functions' .s, which splat never writes — and printed '0 scanned / 0 tripped',
indistinguishable from a pass. The R32 empty-denominator assertion caught it on
its first run; replaced with a two-sided sweep over all open stubs.
tools/cast_self_callers.py — the §378 lever + --sync-decls for the narrow-param
case C89 forbids no-proto from reaching (§378a).
Wiring: wave_args drops walled/parked targets at draw time via pre_classify (one
implementation, R33); escalate_fable.js refuses any target without triage:'DRAFT'.
Tool fixes found by measurement:
* fix_arity_callers was blind to main entirely (globbed src/main/main*.c; main is
src/*.c) — reported success over an empty file set through three gates. Now
refuses when --binary selects no files.
* parallel_gate records each worker's 'failed by class' line (was truncated out of
the 200-char tail); gater_lane retries in-tree ONLY on the diagnostic-free
blind-worktree signature — S69 ran 22 serial retries against real cc1 errors.
docs: cookbook §376/§377/§378 (index 1033), SETUP.md, wave-playbook §4b.
A no-proto decl is ILLEGAL against a definition whose parameter is affected by
the default argument promotions (s16 here): C89 requires the parameter types be
promotion-stable when one declaration has no prototype. So fix_arity_callers'
--any-proto cannot reach this case (it skips it as 'narrow-param').
With the call sites already cast (§378) the decls emit no code, so syncing them
to the draft's exact signature is byte-neutral: 3 decls in src/800.c rewritten to
extern void func_80036D58(s16). Byte-identical, main.
The §376 class's real blocker: after fix_arity_callers no-protos the conflicting
forward decl, the draft's definition becomes the prototype in scope and the TU's
own call site fails with 'too few arguments'. Casting THAT call site to a 0-arg
function pointer is byte-neutral (gcc-2.7.2 folds a cast of a known symbol back
to a direct jal, §20) and banks the function.
§374 a register __asm__($30) reservation is NOT honoured by move_movables under
pressure, and the corruption is SILENT (the build succeeds) -- audit the raw
objdump register uses before trusting a pinned build that compiles.
§375 an $a0-$a3 pin used LATE relocates an EARLIER outgoing-call use of that same
register ~26 slots early, identically across three structural variants; argument
pins are not local the way $s pins are.
With §368 and §373 these now form a usable four-way rule for when a pin helps,
when it fights the allocator, when it is ignored entirely, and when it acts at a
distance.
A tool nobody knows about is invisible work. Audit found neighbor_ref (built an
hour ago), residual_rules, lane_inflight and r22_verify in NEITHER doc, and
wall_sweep in the playbook but not the inventory.
SETUP.md gains a tooling-inventory row for all five with what each is FOR.
wave-playbook gains §2b: run neighbor_ref for EVERY card, placed right after the
seed_ref step because it answers the weaker and far more common question ('which
matched function should this agent READ?') that seed_ref structurally cannot. It
carries the measurement that justifies it -- a ~20x token swing on that single
variable -- and the failure it prevents: func_8017BEBC's card said 'no banked twin'
while a matched 755-instruction near-twin sat 3,700 lines up IN ITS OWN FILE.
Also states the two honest limits: an opt-level mismatch is PENALISED not merely
ranked low (§116 -- an -O2 example misleads an -O0 target), and a neighbour is a
worked example to READ, never a body to copy (§168 law 1, cousin-remap 0/26).
residual_rules.py (mine) and residual_rules_b.py (an independent Fable build,
forbidden from reading mine). Committed because the EXPERIMENT is the artifact:
mine b
classified 85/113 113/113
errored 28 0
any rule fired 18% 88%
certain/high 1% 63%
pure residual-SHAPE ~1% 1.8%
The last row is the finding. Two independent implementations CONVERGED at ~1-2% on
pure cookbook-shape rules, so that tier's ceiling is the POPULATION, not the code:
surgical single-mechanism residuals live at the END of escalations, not in
first-pass wave output. The shape tier belongs in escalation loops; the ladder's
value is everything above it (banked / wall / compile / autodecl / integration).
b also diagnosed my 28 errors exactly: they are functions banked DURING S68 after
the eval set was drawn, so corpus.stubs() no longer contains them and my resolver
raised IndexError on every one. It detects the same condition via corpus.matched()
and calls it ALREADY-BANKED — stale card, spend zero tokens.
Two things b did better that are worth copying: it never parsed disassembly TEXT
(every decision decodes the raw 32-bit word, so the two-disassembler formatting
disagreement that cost me two bugs never touched it), and it REMOVED three of its
own false-positive mechanisms found on held-out cases, all score-reducing, and
disclosed them.
Spec for finishing the ladder: docs/next-session-triage-ladder.md