- ran residual_rules_b over the WHOLE open frontier (1,312 cases, 0 errors, ~2min, $0)
instead of a 50-row sample; artifacts in .run/S70_*
- DENOMINATOR (Drew's correction, R41): main's 960 PsyQ LINKED stubs are not
matching targets; true frontier = 355 (67 main REAL + 288 non-main), partitioned
with progress.linked_subsegs() rather than a hand-rolled filter (R33)
- discriminating test settles population-vs-coverage: fire rate DOES rise as
residuals get clean (35.3% at <=8 vs 6.1% at >64) but 57% of the cleanest band
is still UNKNOWN -> coverage binds where rules are worth writing
- hand-label 4/4 labelable to existing cookbook buckets; WIDTH/lhu!=lh has its
discriminating sig already computed and still returns top=None
- 86 REAL standalone MATCHES (closeness 0) = 24% of the frontier, blocked on TU
plumbing only -- outranks the rule work (standalone-match-is-not-bankable)
- logs 4 instrument defects in my own probe, incl. one wrong answer reported to
Drew before checking: 4 of S68's 10 autodecl MATCH drafts are STILL OPEN
- tools/r22_verify.sh from a clean tree: clean rc=0, extract-all 212+main rc=0,
check-all 213 passed / 0 failed of 213 (2m49s). Clears the S69 --no-r22 debt.
- R38 read of the recorded measurement behind the "1-2% ceiling" (S68 eval set +
.run/rules_b/eval_results.jsonl) before designing the queued probe:
* citation fix: the design is Fable-1 (.run/S69_fable/report.md:93), not Fable-2 §7.7
* denominator fix (R41): shape rules can only fire on the 39 near rows, not 113;
real fire rate 3/39 = 7.7% (5/39 with REDRAFT), and 14 are UNKNOWN
* the probe as written is unrunnable: backlog has 125 rows / 20 with residual text
and the UNKNOWN pile is 14 -- sampling 50 would report a narrower world (R41/R32)
§400 — a baseline check that conflates "absent everywhere" with "changed under
us" silently drops new files. The general law: when a comparison uses two
different sentinels for "nothing" ("" from a failed command, None from a missing
file), it reports a difference that does not exist — and in a GUARD, a phantom
difference becomes a refusal, which looks exactly like the guard working.
Corollary recorded in both §400 and the carve-state memory: "never blanket-add"
covers SHARED carve state (overlays.mk, splat yamls). It does NOT cover a carve's
own new per-binary source file, which is named by a committed yaml and whose 31
siblings are tracked — that one must be adopted with the bank that created it.
Docstring correction: parallel_gate does NOT use `git add -u src/` (that is
gate_stage's form); it adds exactly the adopted paths. My first diagnosis of this
bug blamed `-u` on the strength of that stale line and was WRONG — the cause was
the baseline comparison. Noted in the docstring so the next reader is not
misdirected the same way.
Root cause of the 8 untracked src/ files. The merge-safety check compared:
base = sh(["git","show", pin:path]).stdout -> "" when the path is NOT at the pin
cur = open(path).read() if exists else None -> None when absent from the main tree
if cur != base: REFUSE
For a file that exists in NEITHER — exactly what a jtbl carve creates when it
splits a TU into src/<bin>/<bin>_jr_<addr>.c — that is `None != ""`, so every
carve-created file was refused as "main tree moved under them" and never added.
Nothing failed locally: the file is on disk and R22 passes. But config/splat.<bin>.yaml
names the subseg and IS committed, and 31 sibling _jr_ files in the same binary are
tracked — so a fresh clone (or a push) got the config without the source. Eight
accumulated in one session and only surfaced because the dirty-tree guard refused a
later run.
Fix: distinguish "not at the pin" from "empty at the pin" via git show's RETURN
CODE, so absent-in-both compares equal and the file is adopted. New adoptions are
reported explicitly ("N NEW file(s) created by a carve, now tracked") rather than
merged silently — adopting a brand-new source file should never be invisible (R32).
The `git add -- <adopted>` step was always correct; it simply never received these
paths.
`parallel_gate` commits with `git add -u src/`, which updates TRACKED files and
cannot add NEW ones. A jtbl carve SPLITS a TU, creating `src/<bin>/<bin>_jr_<addr>.c`
— so every carve landed its yaml change (tracked) while leaving the new source
file UNTRACKED.
Why this mattered: `config/splat.<bin>.yaml` is committed and names the subseg
(`- [0x577f8, c, ov_SC02_000_jr_8017F950]`), and 31 sibling `_jr_` files in that
same binary are tracked — so these are source by convention, not build artifacts.
R22 passed locally only because they exist on disk. A fresh clone, or Drew's
push, would have the yaml without the file.
Found because parallel_gate REFUSED to run with an unclean tree (rc=1) and listed
them — the guard did its job; the earlier `REFUSED 8 (main tree moved under them)`
line in the carve gate was the same eight files.
TODO for the tool: parallel_gate's commit step must add NEW files under
src/<binary>/ that its own carve produced (narrowly, per-binary — never a blanket
`git add src/`, per the carve-state discipline).
Caught by Drew asking whether the last waves were harvested. They were not: I
banked 1 of 5 (§398b) and left four lever sets in the notifications. Also found
two paid-for MATCHes that were never staged or gated.
(a) a fence BETWEEN two prologue loads, where source reorder does nothing —
the order is fixed before statement order matters (md_MAIN_013/func_800CB56C)
(b) SINK a call into BOTH arms and let cross_jump keep only the jal suffix;
88ins/close86 -> 92/13, then §3-T2 field order let each sh $zero fill an lhu
load-delay. Duplicate in source so the compiler merges, rather than writing
the merged form yourself (ov_SC07_001/func_8017EDC0)
(c) a $v0->$a0->$s3 DOUBLE COPY is a two-pseudo tell: SImode temp for the compare
+ separate HImode var for the tail (70->37); plus §195-N precondition 5 —
nesting `return 1` with ONE trailing `return 0` blocks jump.c's store-flag
transform so reorg fills both delay slots (18->0) (ov_SC02_017/func_8018347C)
(d) the re-tie as a BIV KILLER: a second set makes n_times_set>1 so loop.c
refuses the pseudo as a biv, killing the combined address giv. volatile was
worse, a dead read did nothing (ov_SC07_001/func_8017E4DC)
(d) makes THREE distinct uses of the zero-byte re-tie in one session — §380
un-hoists a move_movables invariant, §393 kills the scheduler's birthing boost,
§399d denies a biv. One line, three passes: when a single-set pseudo is being
treated specially, give it a second set.
Measured: 22 twin remaps gated as a batch -> 3 banked, 11 CC1-FAIL/PLUMBING, 8
DIFF. The eleven integration failures read 'syntax error before', 'undeclared',
'conflicting types', 'parse error' — the signature of a decl block written for
another TU.
family_remap rewrites the BODY correctly (per-overlay symbols, reloc targets) but
carries the source TU's typedefs/externs/callee prototypes verbatim into a
destination that already owns those names — §378c's fifth variant, at scale and by
construction. The 8 DIFFs are the h_norm class being 80%, not 100%.
Planning consequence (R41): the remap lane's realistic yield is ~15% straight
through and ~50% after the integration pass, NOT the 76-88% PURE-class rate.
Quote the straight-through number.
Tooling gap named: family_remap should emit the body with the DESTINATION TU's
decl environment; decl_prior already computes it for cards, and
cast_self_callers/fix_arity_callers already edit it.
Both rules were already written down (§384, §397) and both were violated anyway,
which is the argument for a tool: a habit you must remember at the moment you are
impatient is not a control.
tools/verify_binary.py — ALWAYS re-extracts before building, because a carve
rewrites splat inputs and a build over stale extract state produces a meaningless
SHA. S69 read three binaries as red on build-only checks; all three were
BYTE-IDENTICAL after extract+build, and two false reds cost legitimate work that
had to be restored (a 96-line match, and 23 declaration edits). --all-touched
sweeps everything with uncommitted src/ or config/ changes.
tools/twin_rescan.py — the twin oracle answers "is there a BANKED body like
this?", so an OPEN-OPEN cluster correctly reports "no banked twin" for every
member and that verdict is stale the instant one banks. Diffs the scan against
the previous snapshot so it reports what JUST became free, not the whole board,
with the ready-to-run family_remap command per row. Baseline: 318 open stubs, 37
already carry a banked twin at d<=5.
Memories added: rescan-twins-after-every-bank, check-against-a-known-true-case.
Measured the expensive way. A reach-6 cluster showed open-open, so seed_ref
correctly reported 'no banked twin' for all six. I cracked the exemplar (203k
tokens, five new levers) and then drafted four siblings at ~60k each — including
one that had already burned 257k plateauing at permuter-class NEAR.
They were EXACT clones. The agents' own diffs said so: 'label-stripped .s diff vs
the twin is EMPTY', 'an EXACT clone (asm diff = labels only)'. The moment the
exemplar banked, seed_ref returned it as a banked twin for every sibling, and
family_remap + the §378 chain banks them for ~0 tokens.
The law: a bank CHANGES THE TWIN GRAPH. The twin oracle answers 'is there a
BANKED body like this?', so its verdict for every sibling is stale the instant the
exemplar lands. crack-wave-sweep-map-regen applied one level down — the family map
is not the only stale artifact, and the twin oracle is the one the cards read.
Also: never draft two members of one cluster in parallel; if either cracks the
other is free.
§322b — the carve class is COMPLETABLE, and every worktree CARVE-REFUSED was an
instrument verdict (.run/sig.<b>.jsonl is gitignored, absent from worktrees, so
jr_inventory read every carve as UNOWNED). build_carve's refusal is EXACT, not
conservative — one object emits one contiguous .rodata — and the real fix
(isolate into its own subseg) already exists and harvest_verify already runs it.
Live census: 71 non-contiguous of 123 jtbl stubs; 21 of those are twins of
already-banked bodies (3,852 ins) free at ~25s each. End-to-end byte-proven in
23 seconds. Remaining blockers are 18 overlay_src_split plumbing defects (<=30
lines each) plus a jr_isolate_all port for main.
§332b — the §332 "walls" are a per-OBJECT assembler mode, not a C limit. A 3-line
maspsx reorder-passthrough + as -O2 is byte-INERT across the whole 800c3/800c2
objects and yields 0 diffs for SIX walls whose drafts already exist. That turns
"permanently unbankable" into a per-object Makefile switch and retires
oracle_reorder.py. Only 13 of the 15 listed walls are even reachable.
§378c — a FIFTH decl-blocker variant: the DRAFT redeclares a type/data/callee the
TU or a header already owns. Fix the draft to the TU's spelling (§367), never the
reverse. Two "integration-blocked" rows were phantoms, one of them my own
--any-proto pre-pass breaking a sibling TU (variant 4, second bite).
accelerators #19 — a verdict recorded inside an isolated environment describes the
ENVIRONMENT. Isolation exists so the worker sees less; every gitignored input is a
difference it cannot distinguish from a genuine rejection, and it writes that
difference down once per function. Negative-control the environment with a
known-good item; assert the worker's inputs; report a missing input as MISSING,
never as a verdict.
Found by the Fable blocked-pile audit. `jr_isolate_all.jr_inventory` resolves each
committed .rodata carve's owner through `family_remap.reloc_targets`, whose
`nins_of` reads the gitignored `.run/sig.<binary>.jsonl`. A fresh worktree has no
`.run/sig.*`, so inside a worker EVERY carve reads UNOWNED, jr_inventory
R32-aborts, harvest_verify prints `isolate FAILED`, and the draft is booked
CARVE-REFUSED.
That verdict was about the WORKTREE, not the function. Measured on
ov_SC02_000/func_8017F950 (a RELOC-ONLY twin whose body rtu-MATCHes 117/117):
dry-run isolation passes in the main tree and aborts in the worktree with 30
phantom UNOWNED carves. Linking one file is the whole difference. When the file
is absent it is now reported in missing_generated rather than silently skipped.
This invalidates the CARVE-REFUSED rows I quoted in the S69 census — they were
instrument verdicts, and the class is far smaller than recorded.
Also adds tools/asm_verbatim.py (new): .s -> §265 file-scope __asm__ block with
decimal immediates/offsets and comma-no-space operands (maspsx dies on
`sltu $v0, $s0, $v1`), derived .frame/.mask, R43 refusals for rodata/jtbl.
Ledger MATCH 12 / NEAR 1 / REFUSED 2 plus a non-wall control. Byte-equivalent to
the stub by construction — for genuine hand-asm only; §265 accounting applies.
Correcting my own guidance from earlier today. §378 gave the self-caller chain;
three more variants appeared within hours and two of them BREAK the chain.
Variant 3 (NEW, byte-proven ov_SC04_018/func_8017F35C, banked): conflicting
RETURN type on a decl that is ALREADY no-proto, where the symbol is
ADDRESS-TAKEN rather than called. --any-proto has nothing to relax and
cast_self_callers has no call site to cast; --sync-decls ALONE fixes it, and is
safe precisely because an address-taken site has no arguments to convert.
Variant 4 (REFUTATION of what I wrote in the playbook this morning): "run the
same chain on the callee the diagnostic names" is wrong at scale. Applied to
func_8012AD44 in ov_SC07_000 it no-protoed 60 caller decls and the binary went
RED (265b24bb vs 9dbe4241); reverted via journal. The self case is safe because
step 2 casts the call sites so the decl change cannot alter argument conversion;
for a callee, cast_self_callers correctly refuses and the decl change runs
unprotected. It banked main/func_80021D38 only because that callee had ONE decl,
not sixty.
Rule added: never --any-proto a symbol whose call sites you are not also casting;
count the sites first. The chain is a DECISION TABLE, not a sequence to run
blindly.
The hard gate caught me: m1/m2 (§379-§383) and the fable escalation (§385-§388)
were harvested, but o1-o4 and p1-p3 were not — 57 MATCH notes sat unbanked while
I was about to draw new waves.
§392 — seven byte-proven spelling levers, each of which closed a match on its own:
(a) a same-address dual-sign read is fixed by ORDER (emit the unsigned
store-source read first); cse merges lh/lhu for every cast spelling tried
(b) a narrow temp picks the narrow load — s16 vs s32 decides lh vs lhu, and a
signed decrement temp yields "sll 16" where unsigned yields "andi 0xFFFF"
(c) tbl[idx-2] folds -8 into the lw offset; hoisting the subtract forces addiu
(d) identical switch arms must be SEPARATE case blocks — the target duplicates
arg setup per case and cross-jump-merges only the shared tail
(e) distinct pseudos per repeated inline copy — one shared pair biases sched1's
tie-break for the first copy only
(f) split the widen into two statements to move the sll off a pinned register
(g) the RETURN TYPE alone closed a 7-ins schedule residual (s32 -> void)
§393 — the BIRTHING BOOST: a single-set local gets max scheduling priority and
sched2's backward pass pushes it LATE; a zero-byte re-tie gives it a second set
and kills the boost. The scheduler-side sibling of §380 — same trick, different
pass, opposite symptom.
§394 — two align-1 accesses in one function reserve a phantom 8-byte stack slot;
a frame 8 bytes too large with no spill to account for it is the tell.
tools/seed_ref.py gains --contained/--contained-control: an open stub that is a
banked body plus or minus WHOLE BLOCKS — the class edit distance ranks badly.
Branch-offset masking was required (unmasked offsets veto exactly the target
pairs) and a min-side-25 floor (89% of raw hits were prologue/epilogue vacuity).
Ranks by (substitutions+regions, cover), not by d. Controls: planted-deletion
positive 60/60, random-pair base rate 0/397, R32 population 346/346, and a
post-refactor --near regression reproducing the stored slice exactly.
Banked on first use: ov_SC01_077/func_80184D50 = banked ov_SC03_007/func_8018283C
minus its trailing `&= 0x7FFFFFFF;` — MATCH, closeness 0, 98/98.
* cookbook §390: minimum distance is not minimum work (rank by effort; a deletion
is free, a substitution is thought), the lookalike filter r = d/min(nins) ~ 0.3
(17 of 30 "cousins" were boilerplate coincidence), and the three fleet-wide
nulls that close the scanner question — 0 new / 9 / 2. Spend integration
effort, not scanner effort.
* cookbook §391: a byte-aligned struct copies in FOUR instructions (lwl/lwr/swl/
swr), a word-aligned one in TWO. Never invent an aggregate type to make a draft
compile — an invented word-aligned Blk8 lost exactly 8 ins across two copies and
read as a believable "near, closeness 70" codegen residual.
* accelerators #18: a claim derived from BYTES is not a claim verified by a
COMPILER. Every similarity/correctness claim must name the tier it reached
(stream containment / compiled standalone / whole-binary gate / clean fleet);
a report that says "verified" without one invites the strongest reading.
Non-reproduction is a finding — say so rather than assuming your own setup.
* playbook §2a-2: the twin ladder (exact -> RELOC-ONLY -> CONTAINED -> cousin ->
cold), take the cheapest tier available, widen only when the tier above is empty.
* SETUP inventory row; generic-decomp-package: rank by work, and stop building
scanners once the well is dry.
The exact-hash twin tier found 22 of 352 reachable open stubs (6%). The
edit-distance band added by `seed_ref --near` finds 75 of 352 (21%) — 3.4x — on a
corpus we believed fully mined. 31 of the new rows were PURE reloc-only twins of
already-banked bodies; 8 banked the same day at ~0 agent tokens, one 94-ins
exemplar serving five open copies.
* cookbook §389: the h_norm hole (norm_stream drops its pending lui-hi on an
intervening R-type, so indexed-global reloc twins hash differently and vanish
from seed_ref/twin_sweep/dedup/family-maps at once). Do NOT fix h_norm — every
stored calibration keys on it; the near tier reads through it.
* accelerators #17: the generalisable law. A similarity hash built for DEDUP
under-matches by design, which is correct for dedup and silently lossy as a
FRONTIER join — the two questions want opposite error directions, and the
frontier failure looks exactly like "this function is unique".
* generic-decomp-package §2b: build the near band at the same time as the exact
tier, with the three verifications. It pays from the first bank for a new
project, where we paid a session to recover the debt.
* SETUP inventory row + playbook §2a (run it before believing any "no twin"
verdict; never send a RELOC-ONLY row to a drafting agent).
Measured over 129 drafting agents in one session, per MATCHED instruction (the
only cost that matters, since a failed agent is billed in full):
sonnet 105 agents, 57 MATCH 4,289 tok/matched-ins (flat ~47% above 30 ins)
opus 24 agents, 11 MATCH 2,083 (m1 191-347: 1,291, 67%)
opus at 347-670: 1/9 7,158 <- the cliff, 2.92M tokens for ONE bank
fable escalation: 3/4 closed at ~1/3 the cost of the attempt it rescued
Sonnet's per-agent price was never the cost that mattered; cost per BANK is, and
it lost on that by 2.1x. The m2 wave should have been fable from the start.
Escalating SOONER is the standing finding — higher models crack harder functions
in fewer tokens. Tested twice now (S68 A/B, S69 measurement); do not re-derive a
cheap-tier argument from per-agent price a third time.