R30/R31 capture while hot: the resolver pivot (63 zero-token banks of 245 staged of 424 judged of
1,352 nominated), the RED-fleet finding (15/214 baseline-RED refusing 174/182 doubly-verified
drafts), the three byte-proven repairs so far, and rule candidates R56–R58.
Drew asked whether waves cracked better before MAXTOK went 8000 -> 16000. Recording both halves
of the answer so next session does not relitigate it from memory:
CLEAN: raising to 16k did cause a real regression — draft completion 84-89% (8k) -> 41% on wave
bt, 69% on bu — but the cause was a harness interaction, not the model. A 16k generation runs
~530 s at ~30 tok/s while STRAGGLER_GRACE was 120 s, so agents were cut off mid-thought with no
draft. Grace at 700 s fixed it; completion has run 97-99% since.
CONFOUNDED: on banks per draft the 8k era looks better (S59: 1,335 of 2,996 = 44.6%; today's best
16k waves dd 34.7%, de 29.7%) — but the populations differ completely. 8k waves had never-drafted
work; today's draw from skeletons that refused six times. Budget and exhaustion moved together, so
neither figure isolates the other. Neither should be cited as evidence about the budget.
AGAINST the simple story: truncated-turn rate is INVERSELY correlated with bank rate (cx 8.7%
trunc/43.9% bank, dd 8.3%/51.4% vs dl 1.3%/0.5%, ej 0.6%/0%). Budget exhaustion driving the
decline would produce the opposite relationship.
THE A/B: split ONE wave's card pool — half the shards at 8k, half at 16k, same generation mix,
same binaries, same gate, same tree, grace 700 s in both arms. Compare banks per DRAFT and per
GATE MINUTE. Holding the population constant is the whole point; every historical comparison
fails exactly there. If 8k matches 16k, the cheaper budget also buys more agents per unit time.
Every campaign process stopped deliberately at session end (0 alive, verified after settling).
.run/ox_campaign.stop and .run/auto/STOP are SET — delete both before relaunching, or every lane
exits immediately.
One dirty overlay TU left by a killed gate was BUILD-VERIFIED as an abandoned substitution (the
binary failed to build with it) and reverted rather than committed — R42's distinction between a
proven bank and mid-gate residue, decided by the bytes.
Two shutdown hazards recorded: pkill on a lane's shell leaves its python running (hit the
drafter, gater and main lane tonight — kill by PID, verify with ps -o lstart), and a bash case
pattern 'src/[a-z0-9_]*.c' matches ACROSS SLASHES, which classified an overlay TU as a main TU
and nearly reverted the wrong file.
Also committing the two lanes built today: tools/lanes/elastic.sh (starts serial idiom lanes when
the API window is idle and the gate queue is deep — it scales the work that is NOT gate-bound,
because adding drafters to a full gate queue makes the backlog worse) and
tools/lanes/grinder_lane.sh (runs tools/grinder.py, the Phase-21 LLM-free permuter, which had
never been run this campaign against 5,388 near-miss rows).
Re-gate of the false-verdict waves finished 21:42: ei 34 · ej 0 · ek 4 · el 6 · em 3 · en 8 =
55 recovered from 2,814 pre-paid drafts for zero model tokens. Only ei paid well (18% of gated);
the rest returned 0-4% because the live lanes had already banked those functions in the interim,
so they come back NOT-A-STUB rather than as banks.
IMPORTANT FOR PLANNING: this does NOT confirm the uncollapsed-wave thesis. eh's 129/380 (34%)
stays an outlier with ei's 18% as its only corroboration — do not plan on sibling drafting
reproducing eh without more evidence.
Session close: 2,238 banked by the commit-message count (the stub invariant is higher — the
A-prop lane's banks ride in chore commits the regex cannot see), open crackable 2,981, fleet
98.2% instruction-weighted and 96.4% distinct-code, up from 97.4%/94.6% this morning.
Next session starts from docs/tool-designs/frontier-analysis-s60.md: the wall is an INTEGRATION
wall, and the first build is the zero-token integration-resolver lane.
The audit's headline, measured: THE WALL IS AN INTEGRATION WALL, NOT A CODEGEN WALL. Of the 292
functions the gate has refused 6+ times, 178 (61%) have ALREADY produced a closeness-0 draft —
match_one byte-equality, whole-binary gate rejection. The blocker is symbols/decls/TU plumbing,
and the fleet keeps re-drafting them: 10,049 reject rows over 574 distinct functions. Highest-EV
build is a zero-token integration-resolver lane, not more drafting.
CORRECTIONS TO MY OWN NUMBERS, verified against the tree before accepting:
* siblings are 1,334 behind 480 multi-member groups, NOT ~3,900. 1,292 groups are SINGLETONS
carrying 57% of open instruction mass. I conflated the never-drafted stub count with the sibling
count and overstated remap leverage ~3x, in this checkpoint and repeatedly in conversation.
* 'everything drawable is gen6+' holds only for the collapsed wave-eligible view; whole-pool
generation is 53% gen0/1, 25% gen6+, and only 292 fns are 6+ GATE-refused.
* '30-67 min gates at 8% CPU' conflated wall_min (includes drafting/queue) with gate wall (12-31
min healthy). Gate cost is proportional to FAILURES, not drafts: ~3 whole-binary builds per
failing draft, so banks/gate-min fell 17.5 -> 0.10 as conversion fell.
* the 5,388 closeness<=2 rows de-dupe to ~543 open functions; my own 19:40 re-measure found 290
still open, down from its 470 — the re-gate and grinder are draining that pool now.
* campaign_status's 'banked today' undercounts: the stub invariant says ~2,644 net, because the
A-prop lane's 357 rode in a chore commit its regex cannot see.
One documented counterexample to 'model quality is not a bottleneck': func_80181714, where
ox-alpha plateaued at closeness 4 while Opus/GLM/DeepSeek each reached reloc-verified MATCH —
argues for a small escalation tier AFTER the resolver drains the fake walls.
Taken on trust and flagged as such: the A-prop residual split (169 STRUCT / 121 no-seed-decl /
73 IMM / 12 void) — the refusal mechanisms exist in aprop_autodraft.py but no file carries those
counts; re-derive before building the decl-inference tool.
~2,200 banked today. Throughput went 65 -> 554 req/min peak by removing harness defects, not by
changing models. The registry was wiped THREE times by four non-atomic truncating writes, now
routed through tools/mk_write.py; each wipe made every gate reject every draft.
The strategic picture for next session: 3,062 open crackable collapse to ~334 drawable skeletons,
~308 of them generation 6+, with ~3,900 siblings behind them that bank by remap. Wide waves
convert at 1-5% and the GATE is the bottleneck (30-67 min at 8% CPU). Optimise banks per gate
minute. The reasoning budget is NOT the cause of the decline — truncation is inversely correlated
with bank rate.
A Fable analyst is writing docs/tool-designs/frontier-analysis-s60.md, briefed that we are not
married to the ox-wave model; that document is the first thing to read next session.
Five rule candidates (R51-R55), each earned by a defect that fired today.
1,947 banked today. Throughput went 65 -> 341 req/min peak and gates 63 -> 39 min, all by
removing harness defects rather than changing models. The registry incident (config/overlays.mk
committed EMPTY, taking main and every overlay gate down) is written up with its blast radius
and the config_sane guard that now prevents it.
The strategic finding is the part that matters for planning: 3,652 open crackable functions
collapse to 334 DRAWABLE skeletons, of which 308 are gen6+ walls — the ~2,900 untouched
functions sit behind those skeletons and bank by mechanical remap, not by drafting. Wide
drafting now converts at 5%. main is 327 crackable, not 1,288.
Four rule candidates for PhaseEnd (R51-R54), each earned by a defect that fired today.
1,342 banked, 140 commits, open stubs 6,575 (main 1,493 / overlay-md 5,082).
The session's one lesson, measured six times: every lane that looked like the models
underperforming was a harness defect — an -O0 oracle nothing ever passed, a lane
retired on a card-size verdict, carve machinery nothing fed, a poisoned main baseline
that made 737 drafts read as bad, a soft 429 killing 44-72% of shards at turn 1, and a
stager consuming one bit of one verdict.
Records what landed (jtbl island split + gate automation, the -O0 census and unlock,
the new main and distill lanes, two RED binaries fixed, the throughput settings with
their probe evidence, the portable-workflow doc), seven rule candidates for PhaseEnd
approval, the ranked open threads with the A-prop residual named and sized, and a
resume procedure that starts from campaign_status.py and verifies from the process
rather than the file.
Records the per-type answer to 'can the waves draw and bank this now', the four
commits that landed after the agents returned, three rule candidates for PhaseEnd
(a card may not name a lever the knowledge base lacks; draw-time bankability; a
budget is part of the harness), and the ranked open work from the agents' docs.
jtbl: the 154-A island split is byte-proven (one config line + jr_isolate_all --only),
with the object-level sh_size control a green SHA cannot give; four md_*/main tool
blindnesses fixed; one delay-slot instruction short of the first bank, logged.
o0: the census (167 real -O0 of 14,400; 116/14,148 ins stranded), the handoff's
md_MAIN_003/011 refutation shown to be itself wrong, the never-wired -O0 oracle fixed,
and 79 unbankable card-draws across 19 waves stopped at the source.
tells: restored and pinned to band 5-80 — the lane gap was a card-size gap, and 235
is refuted as the cause by the recorded pre-gate verdicts.
Three Fable agents in flight; their briefs and every fact they were given are recorded
here so a crash costs a re-spawn, not the knowledge.
jtbl (36,685 ins, but 150 of 177 groups are singletons): Fable's review says the fix is SMALLER
than proposed — ONE inserted .rodata carve line plus jr_isolate_all --only. No _pre piece (it
cannot build), no ld_interleave change (the native script is already rodata-first). Harden
parse_config FIRST: it corrupts md_*/main configs on disk before erroring, which is why
jtbl_carve now hard-refuses them. The ox study's negative control is misattributed — build a
fresh one.
o0/cc1 (6,564 + 6,511 ins): the study is half refuted, and the doc header says which half.
md_MAIN_003/011 do NOT carry the -O0 fingerprint; the 311 files that contain $fp are the real
population. The unanswered load-bearing question for both is whether the EXISTING gate can bank
them unchanged — a lane that drafts what the gate cannot accept has already cost two sessions.
tells (86,602 ins): removed from drafting on four waves of evidence. aprop_autodraft is NOT the
destination (4.2% overlap, checked after I asserted it three times). The live hypothesis is
cookbook §235, the phantom symbol — one wave with it in the brief tests it cheaply.
CURRENT_PHASE.md gains a CRASH-RECOVERY checkpoint (not a fresh-session handoff): what is
running, restart order, the measured fleet/scaling facts, the fixes that must not regress,
and the ordered work queue.
Lanes: drafter (never stop it), gater (restartable), maintenance (free A-prop sibling lane),
stallguard (60s auto-repair). Drafting holds no lock; one narrow draw-vs-gate lock exists
because build_wave_atlas reads corpus.stubs and misreads substituted drafts mid-gate.
main is off the wave critical path — 157 drafts parked to .run/main_queue/ rather than
stalling the gater for another hour on a bisecting whole-EXE rebuild.
api_agent: 5xx retried like 429 (a 502 was abandoning functions at near-19), HTTP_TIMEOUT
420s not 1800 (a hung request parked an agent 30 min), EXTRA_READABLE for tooling briefs,
and bare-directory paths no longer refused against their own granted root.
R42: gate_main reverted 61 byte-proven overlay banks it could not distinguish from its own
substitution (sweep_parallel gates commit=False by design). Fixed by committing overlay banks
before the main batch, chunking main at 8 to bound bisect cost, and replacing every blind
'git checkout -- src/ config/' with commit-or-refuse in ox_campaign and idiom_serial.
R43: sweep_parallel had an explicit branch admitting main, which cannot be gated incrementally
— wave ab banked 0/105 main cards while its non-main cards banked 94/115 (82%), and the wave
read as a drafting failure. sweep_parallel now refuses main and names gate_main.py.
Also: validate_targets now prefers the card's own addr field (named symbols like SYS_OBJ_F00
were MALFORMED and discarded whole 220-card waves); ox_campaign deals model lanes by
smallest-ratio scheduling (a 73-card wave had put 73 shards on ox and 0 on deepseek);
docs/accelerators.md gains the four vacuous-check defects.