A pack carrying a previous attempt said only 'the gate did not accept it, so it is wrong somewhere'.
That discards the one datum that decides how the agent spends its budget. Now the builder runs
match_one on that draft and embeds the verdict:
* near -> the closeness, the verdict sig, and the residual rows (idx / mine / tgt, capped at 16
with an honest '... N more'), plus how to READ them: two adjacent rows with the same
instructions in the opposite order = a SCHEDULE swap; a register-only difference = the
value came from the wrong place (often the copy, not the pre-copy value); a beqz/bnez
row = invert the test and swap the arms, constants included.
* match -> NOT a drafting job. The body is byte-correct in isolation and the gate refused it for an
INTEGRATION reason, so the pack names the $0 recover_integration --probe-only instead of
letting an agent burn a wave slot redrafting a correct body. If the probe also says
MATCH the residual is outside the function (the section-8e JTBL_PADS class).
Measured on the two t5u seeds, which I injected BY HAND this session before automating it:
func_8017F234 = 3 mismatched of 202 (a schedule swap + an "andi" reading the copy instead of the
pre-copy value); func_8017E7E8 = 11 of 66 (inverted branch + a cast written back into the variable
instead of a temp). Told that, an agent edits one use site; told "wrong somewhere", it re-derives 202
instructions.
Cost: one compile per target that HAS a prior draft; --no-residual opts out. Failures are swallowed
into a "(residual not measured: ...)" line — measuring must never break pack generation.
Control: rebuilt t5u's 15 packs into a scratch dir — both seeds gained the block automatically with
the same numbers I measured by hand, and a target with no prior draft is byte-identical to before.
Three properties compose into tree corruption under concurrency:
(a) assert_write_set measures a GLOBAL git status, so a concurrent run's writes read as THIS
run's blast-radius violation and abort it;
(b) an abort does NOT restore the stage edits already on disk;
(c) gate_stage's commit is a deliberately broad 'git add -u src/' — and it must be, since
propagation touches many overlays and a narrower filename glob once DROPPED four R22-verified
banks — so a concurrent --commit sweeps the aborted run's half-applied edits into its commit.
Measured today: xargs -P 4 over 33 binaries put 696 broken lines of ov_MAIN_012 into md_MAIN_026's
+1 bank commit; check-all went 212/213 and the wave bank was blocked behind it (R59).
Narrowing the gate's git add was the WRONG fix (it would restore defect (c)'s predecessor). Instead
the driver enforces its own contract: flock on .run/recover/.driver.lock, refuse loudly (R43).
Control: with the lock held -> rc 1 REFUSED; lock free -> rc 0 and the probe runs normally.
Two defects in one filter, both measured on the t5s wave (24 banked / 29 transcripts):
1. FALSE POSITIVE. The keyword 'no cookbook lever' matched "MATCH on first compile, no cookbook
lever needed" — a note reporting a TRIVIAL function — and that was the ONLY selection out of 24,
while three genuine multi-lever notes went unpicked. A selector whose single hit is the one note
saying 'nothing to learn here' is inverted, not merely noisy. Keyword removed, NOT_NOVEL guard
added, and the phrasings agents actually use ('cookbook lacks', 'new lever', 'worth banking',
'levers not in') added. Same 24-transcript scope now selects func_8017E044 instead.
2. STRUCTURAL BLINDNESS. The distiller only ever considered BANKED functions — but the richest
idiom notes come from the HARDEST functions, which are the least likely to bank. func_8017EB30
(279 ins, four levers written up) and func_8017C014 (246 ins, two) both say 'NOT in the cookbook
and worth banking' and were never candidates. That defeats cookbook §52 — a model that FAILS to
crack a wall still distills the idiom that cracks its siblings — using the flywheel's own tool.
New --with-unbanked includes them, each carrying banked=False so the distilling agent knows the
lever is UNPROVEN by the byte gate (R14/G3).
R39 control: the previously-selected note is no longer selected (it was the false positive) and
nothing legitimately selected was dropped.
os.execv'd tools/blocker_probe.py with --drafts <run_dir>/drafts while that directory was still
created further down, so every non---draft-dir probe died with FileNotFoundError. Only --draft-dir
worked, because stage_drafts() had already populated the dir. Staging now happens first.
Control: the --draft-dir path returns the same verdict as before the move (ov_SC06_029
func_80185214 -> DIFF 52/52 ins, identical to the pre-edit run). --funcs now works: 10 backlog
candidates classified in one pass (2 real-TU MATCH, 3 conflicting-types, 2 too-few-arguments,
1 parse error).
Worth recording (R40): my own probe loop grepped for result rows and swallowed the traceback, so the
crash read as 'no blockers found' — a silently narrowed scope in the harness, not the tool.
mask_for(reloc_kind='26') returned 0, i.e. 'compare NOTHING at this position'. Both comparers pick
the mask from ONE side (diff_object_s from mine, diff_object_object from the target's), so a j/jal
there masked the OTHER side's instruction entirely. Reproduced on synthetic pairs of real encodings:
my 'j 8017e248' (0805f892) vs target 'bne v0,v1' (14430002) -> 0; vs 'nop' (00000000) -> 0; my 'jal'
vs target 'bne' -> 0; while the mirror (my 'bne' vs target 'j') -> 1. That asymmetry is the bug.
_j_mismatch cannot cover it: it fires only when BOTH sides carry an internal-j target, which a
j-vs-bne pair by definition does not.
Fix: return 0xFC000000 — the 26-bit target field stays masked (it IS link-time), the opcode never
is. diff_object_object's masked-slot test updated to match so the reloc symbol+addend check still
fires there.
R39 negative control (tools/stub_invariant_audit.py, the INCLUDE_ASM invariant): 2554 stubs, nonzero
3 before and 3 after — the same three known main length-delta survivors, same values. Zero new false
positives, over a population that exercises the changed path (812 stubs carry internal-j .text
relocs, 3164 such instructions).
Found by a t5s drafting agent on func_8017EB30 (reported as a one-sided internal-j check); verified
here to be broader than reported. No bank was ever at risk — the whole-binary gate is independent
(G3) — but every crack agent and the permuter scorer read this number. R35/R14.
recover_integration's macro-externs stage rewrote a draft's callee extern to the FLEET macro's
signature and then gated only the rewrite. func_ADDR names are per-address, not per-function, so
another overlay's 'extern void func_8017C338(void)' replaced this overlay's correct 4-arg decl and
manufactured the CC1-FAIL it reported as the draft's failure. The untouched draft banks
byte-identical (ov_SC03_012:func_8017BEBC, 246 ins, banked in the previous commit).
reconcile_and_gate(draft_rewrite=) now gates raw (pass 1a) then rewrites only what raw refused
(pass 1b), records the winning variant per fn, and re-gates that variant in pass 2.
harvest_verify.classify_fail kept the 'note:' half of a benign warning pair and labelled a built
draft CC1-FAIL with it; notes now drop with their warnings. Negative-controlled over 5 diagnostic
shapes — only warning+note-only changed (to the honest no-diagnostic label). R39/R57/R32.
Cookbook 920 -> 921 sections, index green.
aprop_autodraft: where the seed carries no decl, infer a minimal extern from the MEMBER'S OWN
target .s (decl_from_use, negative-controlled 97.4%/4,702) and place it at BLOCK scope via
insert_decls — file scope collides with the fleet's per-function loose-typing (the ov_SC04_018
lesson). Strictly additive: runs only where the old path refused; genuine refusals keep the old
behaviour with the class named.
integration_resolver: (1) a CC1/CPP verdict is billed to the DRAFT only after the split TU passes
a TU-alone compile probe WITHOUT the draft — else TU-BROKEN, no demotion, auto-reopened when the
TU's hash changes (S61: one broken TU was billed to 39 drafts; one-off sweep over the CC1 stock:
1 broken TU of 55, 39 verdicts reclassified). (2) dup-def→extern demotion: a draft that DEFINES
data a still-stubbed sibling .s in the same TU also emits dies at the assembler with 'symbol
already defined' — invisible to rtu (INCLUDE_ASM neutralized). Single-line file-scope defs whose
symbol a sibling .s emits are demoted to extern as an ADDITIONAL candidate (new sha, so the
ledger's unchanged-skip does not hide it). Smoke test: md_MAIN_003/func_800D3204 demoted
D_800D3200 → rtu MATCH (12 ins); the whole-binary SHA remains the sole arbiter.
42-case first run: 39 were one uncompilable TU (ov_SC04_018_jr_8017AE2C.c), not draft defects —
the resolver's 'undeclared' classification needs a TU-alone compile probe first (open follow-up).
gate_stage: a binary on .run/baseline_red.txt (+ fleet_red.txt) refuses its drafts with class
BASELINE-RED before any build — a RED binary rejects every draft gated against it, and 174 of the
resolver's 245 doubly-verified drafts were refused exactly that way (negative-controlled both
directions: RED refused without a build, GREEN still reaches the ladder).
jtbl_pads_fix: (1) OBJ_ERR required the make bracket to start with build/src but make prints
[Makefile:687: build/src/...] — find_drift returned None over failing builds ('no pad-count drift',
six binaries); (2) only the 'consumed N but M' phrasing was handled — the 'more rodata .align
directives than pad specs' direction (a NEW table from a banked switch) now searches declared+1/+2;
(3) write_pads split its line on ':' but ':=' contains a colon, appending a second ': JTBL_PADS'
per write, and wrote via a raw truncating open() — the fifth un-converted mk write site; now
path-split-once + mk_write.
mk_write: refuses a parse-poisoned registry line (R43) — one such line kills EVERY build of EVERY
binary at make parse time, strictly worse than the wipes the line-count floor guards.
rtu_shadow: a 0-bank wave has no commit — gatedness now comes from the ox ledger; baseline-RED
binaries are excluded from the prediction metrics with their count reported (R41).
integration_resolver: stub_removals() nets out carve moves (72 gross -> 63 real in the first pass).
tools/rtu_shadow.py: --wave X records rtu_match's verdict for every draft of a not-yet-gated wave
(mapped to its binary through the SHARD's targets file, never by bare name — 54 of wave fa's 470
cards share a name across binaries); --join X after the gate commit prints rtu-verdict x outcome,
P(bank | rtu MATCH), the false-negative rate and current-vs-inverted build counts (R41 denominators).
build_wave_atlas: a (binary, fn) whose latest resolver verdict is STAGED/BANKED/GATE-REJECTED is not
drawn — its body already matches at the real TU and on symbols; DIFF/CC1 verdicts stay drawable
with the CURRENT closeness the resolver demoted them to.
frontier-analysis-s60 §4 measured that ~571 open functions had FINISHED drafting (closeness-0 backlog
rows / reloc shape-MATCH rejects) and were being re-drafted wave after wave. tools/integration_resolver.py
treats those ledgers as an index: still-open? -> rtu_match at the real split TU (CC1: the gate ladder's
draft-side transforms, one retry) -> reloc_identity as the disagreeing oracle (rtu masks reloc fields)
-> aprop_symfix on MISMATCH/shape-MATCH -> stage -> sweep_parallel (whole-binary SHA, sole arbiter)
-> commit at once (R42). Refuses main by name (gate_main owns it), //@EDIT drafts, dirty trees, collapsed
registries; every drop is counted (R32); a negative control over recently-banked functions must pass
N/N before any verdict is trusted (R35/R39 — its first form picked carve moves as banks, 9/12 FAIL,
and was fixed before a single stock verdict was read). Ledger .run/resolver/verdicts.jsonl keyed by
(binary, fn, draft-sha, split-TU-sha) so unchanged rejects are never re-judged.
First pass (commit:2991): 1,352 nominated -> 901 already banked, 27 main -> 424 judged in 41 s ->
245 staged (57.8%; 242 raw, 3 via transforms) -> 63 banked (net INCLUDE_ASM delta; that commit's
subject says 72 = gross incl. 9 carve moves), 182 gate-refused, zero model tokens, ~10 min total.
Lane wrapper tools/lanes/resolver_lane.sh (holds .run/auto/draw.lock for judge+gate: rtu reads the
TUs a gate splices into).
campaign_status: 'today: N banked' summed '— N banked' commit subjects and missed every bank that
rode in a chore/maint commit (S60: 2,185 reported vs 2,644 net stubs removed); it now derives the
number from INCLUDE_ASM stub counts at last-commit-before-midnight / HEAD / working tree (R33), and
alive() is anchored so pgrep no longer matches its own wrapper (every lane read ok with 0 processes).
ox_campaign gater: the ledger's wall_min counts drafting + queue wait since the ready marker's t0;
gate_min is the gate alone (the '30-67 min gates' picture was this conflation).
maintenance lane: the fleet R22 sweep skipped whenever any gate was in flight, i.e. always (last
real sweep 12:54 08-25); it now takes .run/auto/draw.lock and waits its turn, skipping only for gate_main.
Every campaign process stopped deliberately at session end (0 alive, verified after settling).
.run/ox_campaign.stop and .run/auto/STOP are SET — delete both before relaunching, or every lane
exits immediately.
One dirty overlay TU left by a killed gate was BUILD-VERIFIED as an abandoned substitution (the
binary failed to build with it) and reverted rather than committed — R42's distinction between a
proven bank and mid-gate residue, decided by the bytes.
Two shutdown hazards recorded: pkill on a lane's shell leaves its python running (hit the
drafter, gater and main lane tonight — kill by PID, verify with ps -o lstart), and a bash case
pattern 'src/[a-z0-9_]*.c' matches ACROSS SLASHES, which classified an overlay TU as a main TU
and nearly reverted the wrong file.
Also committing the two lanes built today: tools/lanes/elastic.sh (starts serial idiom lanes when
the API window is idle and the gate queue is deep — it scales the work that is NOT gate-bound,
because adding drafters to a full gate queue makes the backlog worse) and
tools/lanes/grinder_lane.sh (runs tools/grinder.py, the Phase-21 LLM-free permuter, which had
never been run this campaign against 5,388 near-miss rows).
THE GATE WAS A BLACK BOX. sweep_parallel's stdout was captured and dropped, so a gate logged
"reloc_identity -> gating 216" and then THIRTY MINUTES OF SILENCE before its bank line — no
worker count, no per-binary progress, no phase-A/phase-B split. Gate times went 31 -> 37 ->
50 -> 67 min across ej/ek/en/eo with nothing to diagnose from, and I twice asserted things
about phase B that the log could not support (its absence measured LOG CAPTURE, not
behaviour). A lane that must run unattended has to leave evidence.
MEASURED WHILE DIAGNOSING, and it rules out the obvious suspects: load average 2.6 on 32
cores with 1-3 concurrent builds during a gate — the gate is NOT CPU-bound and is not
saturating its own -j 24. Raising to 32 is cheap given ~8% utilisation, but the real answer
will come from the log this change adds.
TAIL_DONE_FRAC 0.85 -> 0.80. 0.85 overcorrected: the fleet fell to 15 agents / 11 req/min
because the drafter parks between waves while the gater drains a deep queue. 0.75 was too
deep (26% 429s, draft completion sliding 94->91->73->47% across eq/er/es/et). Neither number
is really the lever: the drafter cannot start a wave the gater has no room for, so the gate
throughput is what bounds the campaign now.
Generational tiering confirmed already correct: the top-off orders by generation at both
assembly levels (group ranking and within-group) without FILTERING any tier out, so every
generation stays eligible and the scarce never-drafted work simply goes first.
A distill reviewer reported "match_one resolves targets by bare symbol name, not
(binary, address)" from 6+ observed target-confusion instances. VERIFIED AND THE CLAIM DOES
NOT HOLD as stated: match_one resolves '%s/%s.s' % (asm_subdir, fn) — an explicit path — and
the drafting path is safe because api_draft.match_one() always passes
dirname(card['asm']). The 675 "match_one MATCH but the whole-binary gate rejected" rows today
keep their real explanation: they landed in the window when config/overlays.mk was empty and
NOTHING could build.
The narrower hazard behind the report is real. --asm-subdir defaults to
asm/resident/nonmatchings/resident, function names are ADDRESS-DERIVED, and overlays share
the address space — so the same name is routinely a DIFFERENT function in another binary
(§238 homonym trap). Any caller that omits the flag gets a confident verdict about the wrong
target, and the failure is silent because the file exists.
It now warns loudly on stderr when the default is used, naming the fn and the directory, and
stays silent when the flag is passed (controlled both ways). A warning rather than a refusal:
resident-era callers legitimately rely on the default, and R43's "refuse what you cannot
handle" does not apply to a tool that CAN handle the input — it applies to one that cannot
tell whether the input is what the caller meant.
ROOT CAUSE of both wipes today. config/overlays.mk was rewritten in four places with
open(mk, "w").write(txt)
(jr_isolate_all.py:593, jtbl_carve.py:1077/1118/1165) — which TRUNCATES to zero first and only
then writes. Three ways that loses the registry: the process dies between truncate and write
(empty file); another process reads inside that window (sees an empty registry); two writers
interleave (a partial line lands after the last good one — this morning's file ended in a stray
`uto.txt` fragment, exactly that fingerprint). The jtbl carve automation runs AT THE GATE, which
is when all three wipes happened, and ONE_PER_GID=0 made it far likelier by putting many more
carve members in every wave.
BLAST RADIUS, measured twice: with no binaries registered, main's object glob sweeps every
overlay's nonmatchings/*.s into MAIN's OBJS and assembles them standalone, so main cannot build,
the main lane correctly refuses against a RED baseline, and every overlay gate rejects every
draft. Waves dn/do banked 0/224 and 0/236; waves ei..em banked 2 of ~1,100 with 675 backlog rows
reading "match_one MATCH but the whole-binary gate rejected" — the local oracle proving the
drafts were byte-correct while the tree could not build them.
tools/mk_write.py is now the only writer: atomic (tmp + fsync + os.replace, so no reader ever
sees a partial file and a crash leaves the original intact), collapse-refusing (a rewrite below
80% of the current line count raises), and flock-serialized.
TWO HONEST LIMITS, recorded rather than papered over:
* Callers still READ outside the lock, so two concurrent carves can each read-edit-write and
the second drops the first's line. That is a LOST UPDATE — a missing line, not a wiped file —
caught downstream by the fleet check and jtbl_pads_fix. Closing it means holding the lock
across read-modify-write in every caller.
* The guard now also refuses when the CURRENT file is under 100 lines. That case cost me
directly: my own verification control overwrote a registry a carve had truncated seconds
earlier, because the collapse check was skipped when the old file was empty. A control must
assert its precondition; mine did not, and now the tool enforces it instead.
Two ways to spend a free drafting window on a tail that is 92% walls.
ONE_PER_GID=0 — the sibling collapse exists because a same-gid sibling banks by mechanical
remap once its exemplar cracks, so drafting it pays for what the remap does free. That prices
AGENT TOKENS as the scarce resource. On the free ox window they are not, and the collapse is
what makes 3,271 open crackable functions look like 334 drawable skeletons — of which 308 are
gen6+ walls whose exemplars have already refused six waves each. A sibling drafted directly
can crack on its OWN terms instead of waiting on an exemplar that never will.
Measured on a live draw rather than argued:
uncollapsed 627 cards / 44,403 ins / 160 binaries / 215 gate groups = 2.9 drafts per rebuild
collapsed 334 cards / 30,926 ins / 104 binaries / 127 gate groups = 2.6 drafts per rebuild
The gate cost is per (binary, TU) group and chunked, so siblings landing in binaries the wave
already touches are close to free at the gate — the card count nearly doubles and the gate gets
MORE efficient per build, not less.
ATTEMPTS=K — K independent shots at each card. The gate cost does not multiply: reloc_filter
keys by fn and staging writes <binary>/<fn>.c, so a function still gets exactly one whole-binary
build per wave; the attempts compete to BE that build, ranked by match_one, which is local and
needs no build. Alternates stay on disk for a later recovery pass. Default 1 (no-op).
A BUG I CAUGHT IN MY OWN SELECTOR before it shipped: it passed binof[fn] (a BINARY NAME) where
match_one wants --asm-subdir (an asm DIRECTORY). Every attempt would have scored identically at
infinity and the picker would have silently degraded to first-seen while appearing to rank —
the same "true number about the wrong thing" class as the day's other defects. Fixed with a
subof map, and a missing subdir now returns neutral instead of a fake score.
idiom_serial passed `--cards .run/aprop_cards.json` to api_agent. That file is a FAMILY-card
file — rows are {family, members, seed, cls, reach}, with no per-function key — while
api_agent's --cards wants per-function wave cards and built its map with `c['fn']`. Result:
KeyError at line 647, the agent died before its first turn, and every target logged
"no-draft". The lane read as a model failure; its ledger holds 8 rows total and the tells
lane had, literally, "never yet run".
TWO FIXES, one on each side of the contract:
* api_agent REFUSES, never crashes (R43). It accepts 'fn' or 'name', skips rows with
neither, and says how many it skipped. The R32 warning two lines below — "ZERO matched —
wrong card file?" — existed to catch exactly this and was unreachable behind the crash. A
guard downstream of the failure is not a guard.
* idiom_serial passes no --cards at all. Its fuel is the --brief: the target plus every
idiom distilled so far in the run, which IS the compounding channel the 2,000-way fan-out
lacks. It never needed a wave card.
Verified: the lane now reaches `api_agent: stealth/ox-alpha -> 1 target(s), max 60 turns` and
drafts, instead of exiting in 0.1 s.
The serial lane's first-ever run on tells died 0 seconds in, on both counts:
[1/6] func_8017D1C0 @ ov_SC06_027 · 118 ins -> drafting
[2/6] func_8001382C @ main · 103 ins -> could not commit — refusing on a dirty tree
MAIN IS NOT A SERIAL-LANE TARGET. main gates through gate_main on the main lane's own cadence
(a clean whole-EXE rebuild with bisection), so this lane cannot bank it however good the draft
is — a main target burns a slot to learn that. md_* was already excluded for jtbl; main never
was, because the lane predates main being in the atlas at all.
A CONTENDED INDEX IS NOT A DIRTY TREE. Six campaign lanes plus the re-gate runner commit
continuously, so `git commit` here loses the index.lock race routinely — twice today it was a
STALE lock blocking every lane for 8 and 21 minutes. The lane refused over a tree that was
fine. It now retries the commit six times at 5 s before concluding, and the refusal is
reserved for FOREIGN DIRT it must not adopt — which is the property that rule was protecting.
Verified: pick_targets('extend-tell', 8, 80) now returns 8 overlay targets, no main.
Also in this commit: MAX_BINS 160 -> 50 in the drafter shell. Wave size was the right lever at
50% conversion (dd: 217 banked of 422 gated in 39 min); at 5% it is dead weight — dq banked 8
of 167 gated and took 98 MINUTES, while four drafted waves queued and the free-ox fleet sat at
14 agents / 10 req/min. At this conversion a 150-card wave banks what a 430-card wave banks,
in a third of the gate.
A tell-lever card named a lever and then made the agent go find its sites: it said
"extend-tell" and nothing about where or how many. atlas_features already counts the
detectors per function into .run/feat.<bin>.jsonl at atlas time, so this is a JOIN, not a
computation — one dict load per binary in the wave.
build_wave_atlas attaches 'tells' {extpair, dupselect, magic_div, sign_lh, sign_lb} to every
card that has a nonzero one — not just tell-lever cards, because a sll/sra pair site or a
repeated select is worth knowing whatever lever drew the card. Measured on a live draw: 85 of
185 cards carry counts.
api_agent._fuel renders them as a CHECKLIST rather than a hint, which is the point: the
counts come from the TARGET's own bytes, so a draft emitting fewer has provably missed sites
and should go looking before spending a turn elsewhere.
Verified: a card with {extpair 3, dupselect 2, sign_lh 1} renders all three with the zero
fields omitted; NEGATIVE CONTROL — a card with no tells renders no TELLS line at all.
Takes effect on the next draw + the next shard (api_agent is spawned per draft).
Neither extreme was right. Ignoring generations lets a card that has failed five waves
compete with one nobody has ever drafted; filtering to a single generation starved the fleet
to 46 cards. So generation becomes the PRIMARY ORDERING and the existing mass/count criterion
breaks ties, at both levels of the assembly:
* gate groups holding never-drafted cards rank ahead of all-retry groups (before the
--max-bins truncation, so an untouched group is never cut for a fat retry group);
* within a group, untouched cards are taken before retries.
Wave SIZE is untouched — only the order changes — so the fleet stays full while the scarce
never-drafted work always goes out first.
Verified on a live draw: gen0 2 · gen1 2 · gen2 14 · gen3 2 · gen4 3 · gen5+ 340. It took
EVERY card below generation 5 (all 23 available) and filled the remaining 340 slots from the
5+ pile, which is the whole point.
The gen0 count is 2 because wave dw drew the last 51 untouched skeletons an hour ago. That is
the campaign's real state: essentially everything drawable has now been drafted at least
once, and ~635 distinct skeletons have refused. The remaining work is levers, not draws.
--generational (or BFM_GENERATIONAL=1) draws ONLY the lowest generation present: no function
gets a 2nd draft while any drawable function still lacks a 1st. draw_count is derived from
the prior wave card files already being read for the already-waved filter.
MEASURED BEFORE SHIPPING, and it changed the plan. With every cap opened — no band, no
--max-bins, all levers, whole fleet minus main — generation 0 is:
232 candidates -> 46 distinct skeletons (186 are same-gid siblings the remap banks free)
against ~681 drawable skeletons in total. So "3,926 never-drafted stubs" was three illusions
stacked: ~961 are main's LINKED PsyQ stubs and data blobs (not decomp targets at all), most of
the rest are same-gid siblings that one exemplar banks mechanically, and 2,119 were already
drawn in earlier waves. Making generational the DEFAULT starved the fleet from ~640 cards to
46 — so it is opt-in, for a priority pass over the untouched population, not standing policy.
The real shape of the endgame, stated plainly: of ~681 distinct drawable skeletons, only 46
have never been drafted. The other ~635 refused at least one draft each. That is a LEVER
problem — distillation, A-prop, the o0/jtbl carves, the permuter — not a resampling problem,
and no amount of drafting throughput addresses it.
Fired as wave dw (51 cards / 1,990 ins across 36 binaries) so the untouched population is
drafted today rather than left as a policy.
The gate's dirty-tree committer swept an empty config/overlays.mk into commit:2863 and took
the whole fleet down with it. R42 says commit a dirty tree rather than revert — true for
src/, where a per-binary gate leaves PROVEN banks uncommitted and reverting destroys them.
A config file is the opposite case: it holds no proven state that exists only in the
worktree, and a collapsed one is never intended.
config_sane() runs at all three commit sites: if config/overlays.mk or config/dedup.us.yaml
has fewer than 80% of HEAD's lines, it is restored from HEAD, NOT committed, and the refusal
is logged loudly. Controls both ways — positive (min_ratio=1.5 makes the healthy registry
trip the same branch: detected, restore path runs, file intact) and negative (normal
threshold: silent, returns True). P28's registry died this way too (H5); now it is enforced
rather than remembered.
THE CLASS. JTBL_PADS is a per-object spec written by jtbl_carve at CARVE time — one entry
per rodata `.align 3`, each 0 or 4 — describing how many jump tables the object emits. That
is a DERIVED property of the current source stored as static config, so any bank carrying a
`switch` (or any bank being reverted) invalidates it and nothing re-derives it. Three of the
five REDs on 08-25 were this one design choice: ov_SC02_005 and ov_SC07_006 from wave dd
banking switch-bearing functions, ov_SC04_018 from the identical symptom with the opposite
cause — a reverted bank taking its table with it.
WHY NOT DERIVE IT. The COUNT is derivable from the assembly stream; the VALUES are not — a
pad records where the ORIGINAL image has an inter-table pad, which lives in the retail
layout, not in our source. Guessing shifts every downstream data symbol: silent corruption,
the worst outcome available. So tools/jtbl_pads_fix.py does not derive. It ENUMERATES the
2^(N-1) candidate specs (first entry 0, rest in {0,4}) and accepts one ONLY if it is the
UNIQUE candidate that rebuilds the binary byte-identical to config/check.<bin>.sha; zero or
two matches restore the original and refuse. R39 negative control: on a healthy binary it
reports "no pad-count drift" and changes nothing.
TWO INSTRUMENT BUGS THIS TOOL FOUND IN ITSELF:
* JTBL_PADS is a target-specific MAKE VARIABLE, so changing it does NOT make the .o out of
date. The first run reported "no drift" against a spec I had deliberately broken. It now
deletes the armed objects before every build — R22's incremental trap in config costume.
* A failed object build leaves the PREVIOUS binary in build/<bin>/<bin>, so
`make build; sha1sum build/<bin>/<bin>` reports the OLD artifact as if it were this
build's — a FALSE GREEN over a build that never linked, which briefly convinced me two
binaries were fixed. build_sha now deletes the output too and requires make to exit 0.
Same family as R49: an error inside something shaped like success.
CADENCE: the fleet sweep runs EVERY maintenance pass, not every 4th. A RED fails at BUILD,
so every draft gated against it is rejected regardless of quality and the wave reads as a
drafting failure — detection latency is the whole cost. Gates now finish in ~35 min rather
than 60, so the sweep is affordable each pass. It still FIXES NOTHING by design, with this
single exception, admissible only because it proves itself against the byte gate first.
1. MAIN LANE — the largest single block of unfinished work was drawing 32 cards a wave.
main_lane.draw() never passed --max-bins, so it inherited build_wave_atlas's default of
12 gate groups — a cap that exists because each group costs a whole-binary rebuild, and
main's own --only-bins docstring says the opposite applies to it: "main is gated ONCE per
SLATE, so main has no per-TU gate cost and --max-bins can be large". Nobody passed it.
Measured cost: main banked ~19 stubs/hour against 1,291 remaining while the overlay lane
ran 650-card waves beside it. Now --max-bins 400 (MAIN_MAX_BINS overrides), and the lane
shell draws 600 cards with 600 workers instead of 200/150.
2. TWO LANES GATE, SO READ BOTH LOGS — a defect I introduced this session. The in-flight
exclusion derived "this wave has been gated" from .run/gater.log only, but the main lane
gates its own waves into .run/main_lane.log. Every m## wave therefore looked permanently
in flight and main's draw lost 425 cards to an exclusion meant for work in progress.
3. TAIL_DONE_FRAC 0.80 -> 0.65. At 0.80 the fleet runs 2-3 overlapping waves at ~250
req/min; the residual troughs are the gap between one wave draining and the next ramping.
65% keeps 3-4 waves overlapping. Stragglers keep their full 700s grace in the finisher
thread — this changes when the NEXT wave starts, never what lands.
4. ATOMIC ATLAS WRITE. The lanes read .run/atlas.json at every draw and atlas.py dumped
straight onto it, leaving a truncated file readable for the length of the write. Now
written to .tmp and os.replace'd.
Context for 1-3: the atlas both lanes draw from is dated 08-23 01:13 — two days stale,
predating ~4,600 banks — and its regen chain is running now (its own R32 assertion caught a
stale family map first and named the fix).
MEASURED on a live gate: 24 workers, 32 cores, and 0-2 concurrent builds at load 2.2.
gate_stage takes the fleet-shared lock EXCLUSIVE whenever it might write shared state, and
`_writes_shared = propagate or not GATE_NO_ARITY`. sweep_parallel passes propagate=False but
never set GATE_NO_ARITY, so the arity pre-pass (default on) made EVERY worker a writer and
all 24 queued on one lock. The gate has been effectively serial for the whole campaign,
while the CPU it was supposedly rationing sat at 7% — and that gate time is what recycles
cards back into the draw, so it throttled the drafting fleet too.
bulk_harvest has documented the contract since P30 — "SET GATE_NO_ARITY=1 FOR THIS PHASE ...
route arity-needing drafts to the serial phase" — and this driver, the one the campaign
gater actually calls, was the one that did not.
PHASE A: parallel, GATE_NO_ARITY=1, workers are READERS and actually run concurrently.
ASSERT: `git status --porcelain src/shared config` must be empty afterwards — the only
cheap detector for a shared-state write escaping a worker (bulk_harvest's rule).
PHASE B: serial with the pre-pass on, for binaries whose drafts failed to COMPILE — the
backlog separates "won't compile standalone (loose-typing / missing decl)" from
"residual: N mismatch", and only the former is what fix_arity_callers fixes.
Recent rows are ~13% failed, so phase B stays small instead of handing back the
parallelism. The lever is kept, not traded away.
Tested on ov_SC07_009 with two deliberately wrong drafts: one that compiles and mismatches
(stays in phase A), one that cannot compile (routes to phase B). Both phases ran, neither
banked, shared state clean, tree clean.
Takes effect on the gater's next wave — sweep_parallel is a subprocess, no restart needed.