Commit Graph

883 Commits

Author SHA1 Message Date
Drew T f84e8272ef docs(phase-31): S65 FINAL checkpoint — 42 banked (t5s 24 + t5t 13 + 5 recovery), 213/213 green, 5 tool defects fixed, the 1x15 cap, my own -P4 corruption + repair, the distill's 2-of-3 refutation, next-4-steps 2026-08-29 16:11:51 -06:00
Drew T d50cd66401 docs(phase-31): S65 checkpoint — 27 banked (t5s 24 + 3 recovery), 4 tool defects fixed with negative controls, the 1x15 wave cap, the JTBL_PADS class, next-4-steps 2026-08-29 13:43:27 -06:00
Drew T 18812211ca docs(phase-31): S64 FINAL CHECKPOINT for a fresh session — fleet 1,616 stubs / 213-213 GREEN / 98.8%-97.4%, the committed recipe, next-4-steps, 258 K+L remaining, T6 wall ledger with 21 diagnosed residues, standing hazards + rule candidates (P31 S64) 2026-08-29 11:52:19 -06:00
Drew T a51c0785cf docs(phase-31): S64 results — 11 waves, 2,380->1,616 stubs across S63+S64 (764 closed), 98.8%/97.4%; the R14 self-correction on func_8017BEBC; fix_decl_mirror's honest 0-of-17 verdict + its R39 saves; killed-gate adjudication 2026-08-29 11:51:45 -06:00
Drew T f25b890480 docs(phase-31): S63 FINAL checkpoint — 412 stubs closed (2,380->1,968), 98.6%/97.1%, 213/213 GREEN, cookbook 917; t5j pre-packed for resume; 5 proven findings + 2 rule candidates; machine quiesced for reboot 2026-08-27 11:32:45 -06:00
Drew T 963519799d docs(phase-31): S63 SESSION CHECKPOINT — fleet 2,068 stubs / 213-213 GREEN, the committed wave recipe, 690 K+L remaining, next-step order (P31 S63) 2026-08-27 09:44:58 -06:00
Drew T bd472c5fc9 docs(phase-31): S63 T5.6 results — 312 stubs closed (2,380->2,068), fleet 213/213 after every batch, 98.6%/97.1%; falsifier dead 6x; residue-on-Opus proven; A-prop lane DRAINING (23->3 on more exemplars); card (binary,fn) defect + fixes; rule candidates 2026-08-27 09:44:26 -06:00
Drew T 01d60955c7 docs(phase-31): S63 T5.5/T5.6 — t5a distill (§164-55 addendum, 913 sections); wave t5d 42/48; 5 of 6 misses were INTEGRATION not codegen; residue-on-Opus 4/4 2026-08-26 21:13:42 -06:00
Drew T 99ae0afe17 docs(phase-31): S63 T5.2/T5.3 — lane G 23 banked free; wave t5a 43/48 (89.6%) incl. 2 recovered; the (binary,fn) card defect measured at 48/48 and fixed; falsifier cleared 6x 2026-08-26 20:21:07 -06:00
Drew T f8cc4c77d3 feat(tools): T5 wave draw/bank halves — t5_targets.py (routed, ledgered, band-stratified draw) + t5_bank.sh (judge → R22 sweep → commit); claude_wave_draft.js generalized; packs/judge refuse duplicate fn names (R43/R48); sweep #8 213/213 (P31 S63 T5.1) 2026-08-26 19:35:54 -06:00
Drew T 20e7005855 docs(phase-31): S62 FINAL checkpoint — no-DeepSeek decision, T5 first-wave recipe on the promoted harness, residue ledger, rule candidates 2026-08-26 19:18:23 -06:00
Drew T 44f062a978 docs(phase-31): S62 checkpoint refresh at the T4/T5 boundary 2026-08-26 18:32:51 -06:00
Drew T 194cce7570 docs(phase-31): T4 DONE — routing rule (<=50: Sonnet+DeepSeek parallel, Opus residue; 51-120: Sonnet, Opus escalation; >120: Opus; haiku retired; M-extend-tell -> wall); frontier-s61 Addendum 5; judge artifacts (P31 S62) 2026-08-26 18:32:51 -06:00
Drew T 02c5d757c2 docs(phase-31): T4 Claude-arm table (haiku 8 / sonnet 14 / opus 17 of 20 at the whole-binary gate), the warm-start homonym confound, distill yield (P31 S62) 2026-08-26 18:14:48 -06:00
Drew T 1392bf6c93 docs(phase-31): checkpoint — fleet sweep #7 213/213 GREEN at T3 close 2026-08-26 16:02:20 -06:00
Drew T e8549eaf10 docs(phase-31): S62 session checkpoint — T0–T3+T5pre done, paused at the R27 gate for T4 2026-08-26 16:00:43 -06:00
Drew T 16d1d01b56 docs(phase-31): T3 log (75 banks; six carve-lane defect classes fixed; residue named) + frontier-s61 Addendum 4 (P31 S62) 2026-08-26 15:57:18 -06:00
Drew T dfd42d4586 docs: cookbook §304 (self-defining rodata) + T3 log incl. fleet sweep #3 = 213/213 (P31 S62) 2026-08-26 14:56:06 -06:00
Drew T df8f369eea fix(tools): jtbl_rodata_pads --derive — empty spec when the TU has no C tables (never refuses what it need not pad), sub-4-byte all-zero gaps before an anchor accepted as assembler alignment, .float/.double sized; 70/70 module binaries byte-identical through the derive stage; T3a log (P31 S62) 2026-08-26 14:50:12 -06:00
Drew T d718cecf9d docs(phase-31): S62 log — clean fleet sweep found 6 hidden failures (5 healed, md_SC07_003 re-measured); R60 candidate (carve-state files); SETUP: interleave_check --fix 2026-08-26 14:35:06 -06:00
Drew T b3e879a95e feat(decomp): ov_SC04_018 — 18 resolver-held drafts banked after the T2d surgery; re-verified byte-identical from a clean rebuild; T2d/T2e log — red list EMPTY (P31 S62) 2026-08-26 14:28:24 -06:00
Drew T 48a42b9ed3 feat(decomp): ov_SC06_022 — 9 resolver-held drafts banked after the T2c surgery; re-verified byte-identical from a clean rebuild; T2c log (P31 S62) 2026-08-26 14:21:24 -06:00
Drew T 0060dfd976 feat(decomp): ov_SC03_024 — 4 resolver-held drafts banked after the T2b surgery (func_8017D890, func_801812C4, func_801826E8, func_801848F4); re-verified byte-identical from a clean rebuild; T2b log (P31 S62) 2026-08-26 14:14:39 -06:00
Drew T c50e9578ba feat(decomp): ov_SC03_015 — 7 resolver-held drafts banked after the T2a surgery (func_8017C104, func_8017D890, func_8017DCC0, func_80180638, func_80180CF8, func_80188FB4, func_801893BC); gate-time §8a carve (yaml+mk) re-verified byte-identical from a clean rebuild; T2a log (P31 S62) 2026-08-26 14:11:20 -06:00
Drew T 3a5caa65b1 feat(tools)+docs(phase-31): T1 — resolver declfix wiring; masked_diff/rtu_match compare internal-j targets (jrel; §195 premise refuted, cookbook §301); diff_autopsy.sh + stub_invariant_audit.py promoted; frontier-s61 Addendum 2 (step-1 falsifier fired: 54 carve-refused, 2 DIFF autopsied+banked, 1 plumbing); S62 log 2026-08-26 14:04:54 -06:00
Drew T d79cee5f9c docs(phase-31): handoff item 2 done — cookbook distill committed, planner notified 2026-08-26 12:52:28 -06:00
Drew T c94fbe0d17 docs(phase-31): handoff item 1 done — frontier doc verified and committed 2026-08-26 12:45:16 -06:00
Drew T fa475da0f8 docs(phase-31): S61 FINAL-2 checkpoint — machine quiesced, 401+ banked today, ox era closed, budget-unconstrained finish directive recorded, two closing agents' handoff protocol for the fresh session 2026-08-26 12:31:35 -06:00
Drew T 5c4d090537 docs(phase-31): DeepSeek push paused by the key's $60 lifetime cap at $60.21 (account credit remains); resume script staged; theories corrected (R40) 2026-08-26 10:31:01 -06:00
Drew T 3fce1c76e8 docs(phase-31): ov_SC02_005 verdict — two stacked defects (older extract rot + the 05:38 shift); 5 byte-checked attempts, stable at HEAD, day-session surgery; attribution corrected (R40) 2026-08-26 09:44:04 -06:00
Drew T b46fc5488b docs(phase-31): morning red-set verification — 5 verified (4 known + ov_SC02_005 fresh drift); ov_SC05_010 self-healed by the maintenance lane 2026-08-26 09:34:52 -06:00
Drew T 04e685e8dc docs(phase-31): correction — true morning fleet numbers 98.4% iw / 96.6% dc / 2,540 stubs (the earlier line quoted a stale file; --fleet was required) 2026-08-26 09:06:46 -06:00
Drew T 96847afcf0 docs(phase-31): S61 overnight close — 378+ net stubs; gen0 goal beaten (392/421 drafted, 196 banked); ox window closed 07:55; fleet regenerated 2026-08-26 09:05:05 -06:00
Drew T 7fd61c8060 docs(phase-31): THE OX WINDOW CLOSED 07:55 — gen0 goal beaten before the door: overlays 196/196, main 196/225 drafted; remaining value is pre-paid drafts + zero-token lanes 2026-08-26 07:58:34 -06:00
Drew T 3b42747491 docs(phase-31): CORRECTION — ABCD judge had an R53 false-green hole; corrected table 8k 6 · 16k 7 · 24k 5 · 32k 3 (ordering stands, winner flips by one fn); fleet_red.txt reconciled 2026-08-26 03:09:14 -06:00
Drew T 80aad51098 docs(phase-31): S61-6 ABCD verdict — 8k/16k 10/10 beat 24k/32k 8/10; overlays sweep at MAXTOK=8000; gen0 sweep launched 2026-08-26 03:06:32 -06:00
Drew T daecc94ef0 docs(phase-31): S61 FINAL checkpoint — wave machine measured finished (17/~5,000 pre-paid), inversion closed NO (0/92 clean), residual classes named, 4 REDs documented, record corrections logged 2026-08-26 01:07:51 -06:00
Drew T eb132b719a docs(phase-31): S61 interim checkpoint + cookbook §293 (the sibling law decomposed: the baseline, not the siblings) + decision-log entries + accelerator #12
R30/R31 capture while hot: the resolver pivot (63 zero-token banks of 245 staged of 424 judged of
1,352 nominated), the RED-fleet finding (15/214 baseline-RED refusing 174/182 doubly-verified
drafts), the three byte-proven repairs so far, and rule candidates R56–R58.
2026-08-26 00:32:08 -06:00
Drew T 605a4b77c8 docs(phase-31): record the 8k-vs-16k question as UNRESOLVED, with the A/B that settles it
Drew asked whether waves cracked better before MAXTOK went 8000 -> 16000. Recording both halves
of the answer so next session does not relitigate it from memory:

CLEAN: raising to 16k did cause a real regression — draft completion 84-89% (8k) -> 41% on wave
bt, 69% on bu — but the cause was a harness interaction, not the model. A 16k generation runs
~530 s at ~30 tok/s while STRAGGLER_GRACE was 120 s, so agents were cut off mid-thought with no
draft. Grace at 700 s fixed it; completion has run 97-99% since.

CONFOUNDED: on banks per draft the 8k era looks better (S59: 1,335 of 2,996 = 44.6%; today's best
16k waves dd 34.7%, de 29.7%) — but the populations differ completely. 8k waves had never-drafted
work; today's draw from skeletons that refused six times. Budget and exhaustion moved together, so
neither figure isolates the other. Neither should be cited as evidence about the budget.

AGAINST the simple story: truncated-turn rate is INVERSELY correlated with bank rate (cx 8.7%
trunc/43.9% bank, dd 8.3%/51.4% vs dl 1.3%/0.5%, ej 0.6%/0%). Budget exhaustion driving the
decline would produce the opposite relationship.

THE A/B: split ONE wave's card pool — half the shards at 8k, half at 16k, same generation mix,
same binaries, same gate, same tree, grace 700 s in both arms. Compare banks per DRAFT and per
GATE MINUTE. Holding the population constant is the whole point; every historical comparison
fails exactly there. If 8k matches 16k, the cheaper budget also buys more agents per unit time.
2026-08-25 23:01:53 -06:00
Drew T 065e122c64 docs(phase-31): S60 shutdown — all lanes stopped, tree clean, sentinels set
Every campaign process stopped deliberately at session end (0 alive, verified after settling).
.run/ox_campaign.stop and .run/auto/STOP are SET — delete both before relaunching, or every lane
exits immediately.

One dirty overlay TU left by a killed gate was BUILD-VERIFIED as an abandoned substitution (the
binary failed to build with it) and reverted rather than committed — R42's distinction between a
proven bank and mid-gate residue, decided by the bytes.

Two shutdown hazards recorded: pkill on a lane's shell leaves its python running (hit the
drafter, gater and main lane tonight — kill by PID, verify with ps -o lstart), and a bash case
pattern 'src/[a-z0-9_]*.c' matches ACROSS SLASHES, which classified an overlay TU as a main TU
and nearly reverted the wrong file.

Also committing the two lanes built today: tools/lanes/elastic.sh (starts serial idiom lanes when
the API window is idle and the gate queue is deep — it scales the work that is NOT gate-bound,
because adding drafters to a full gate queue makes the backlog worse) and
tools/lanes/grinder_lane.sh (runs tools/grinder.py, the Phase-21 LLM-free permuter, which had
never been run this campaign against 5,388 near-miss rows).
2026-08-25 22:53:09 -06:00
Drew T 9a5a4f1b6a docs(phase-31): S60 close — re-gate verdict and final state for a fresh session
Re-gate of the false-verdict waves finished 21:42: ei 34 · ej 0 · ek 4 · el 6 · em 3 · en 8 =
55 recovered from 2,814 pre-paid drafts for zero model tokens. Only ei paid well (18% of gated);
the rest returned 0-4% because the live lanes had already banked those functions in the interim,
so they come back NOT-A-STUB rather than as banks.

IMPORTANT FOR PLANNING: this does NOT confirm the uncollapsed-wave thesis. eh's 129/380 (34%)
stays an outlier with ei's 18% as its only corroboration — do not plan on sibling drafting
reproducing eh without more evidence.

Session close: 2,238 banked by the commit-message count (the stub invariant is higher — the
A-prop lane's banks ride in chore commits the regex cannot see), open crackable 2,981, fleet
98.2% instruction-weighted and 96.4% distinct-code, up from 97.4%/94.6% this morning.

Next session starts from docs/tool-designs/frontier-analysis-s60.md: the wall is an INTEGRATION
wall, and the first build is the zero-token integration-resolver lane.
2026-08-25 22:47:04 -06:00
Drew T 07a167516d docs(s60): the Fable frontier audit + corrections it forced to my own checkpoint
The audit's headline, measured: THE WALL IS AN INTEGRATION WALL, NOT A CODEGEN WALL. Of the 292
functions the gate has refused 6+ times, 178 (61%) have ALREADY produced a closeness-0 draft —
match_one byte-equality, whole-binary gate rejection. The blocker is symbols/decls/TU plumbing,
and the fleet keeps re-drafting them: 10,049 reject rows over 574 distinct functions. Highest-EV
build is a zero-token integration-resolver lane, not more drafting.

CORRECTIONS TO MY OWN NUMBERS, verified against the tree before accepting:
* siblings are 1,334 behind 480 multi-member groups, NOT ~3,900. 1,292 groups are SINGLETONS
  carrying 57% of open instruction mass. I conflated the never-drafted stub count with the sibling
  count and overstated remap leverage ~3x, in this checkpoint and repeatedly in conversation.
* 'everything drawable is gen6+' holds only for the collapsed wave-eligible view; whole-pool
  generation is 53% gen0/1, 25% gen6+, and only 292 fns are 6+ GATE-refused.
* '30-67 min gates at 8% CPU' conflated wall_min (includes drafting/queue) with gate wall (12-31
  min healthy). Gate cost is proportional to FAILURES, not drafts: ~3 whole-binary builds per
  failing draft, so banks/gate-min fell 17.5 -> 0.10 as conversion fell.
* the 5,388 closeness<=2 rows de-dupe to ~543 open functions; my own 19:40 re-measure found 290
  still open, down from its 470 — the re-gate and grinder are draining that pool now.
* campaign_status's 'banked today' undercounts: the stub invariant says ~2,644 net, because the
  A-prop lane's 357 rode in a chore commit its regex cannot see.

One documented counterexample to 'model quality is not a bottleneck': func_80181714, where
ox-alpha plateaued at closeness 4 while Opus/GLM/DeepSeek each reached reloc-verified MATCH —
argues for a small escalation tier AFTER the resolver drains the fake walls.

Taken on trust and flagged as such: the A-prop residual split (169 STRUCT / 121 no-seed-decl /
73 IMM / 12 void) — the refusal mechanisms exist in aprop_autodraft.py but no file carries those
counts; re-derive before building the decl-inference tool.
2026-08-25 19:36:51 -06:00
Drew T 8541238e60 docs(phase-31): S60 FINAL checkpoint — three registry wipes, six harness defects, and the endgame reframed
~2,200 banked today. Throughput went 65 -> 554 req/min peak by removing harness defects, not by
changing models. The registry was wiped THREE times by four non-atomic truncating writes, now
routed through tools/mk_write.py; each wipe made every gate reject every draft.

The strategic picture for next session: 3,062 open crackable collapse to ~334 drawable skeletons,
~308 of them generation 6+, with ~3,900 siblings behind them that bank by remap. Wide waves
convert at 1-5% and the GATE is the bottleneck (30-67 min at 8% CPU). Optimise banks per gate
minute. The reasoning budget is NOT the cause of the decline — truncation is inversely correlated
with bank rate.

A Fable analyst is writing docs/tool-designs/frontier-analysis-s60.md, briefed that we are not
married to the ox-wave model; that document is the first thing to read next session.

Five rule candidates (R51-R55), each earned by a defect that fired today.
2026-08-25 19:27:56 -06:00
Drew T 9e45486282 docs(phase-31): S60 checkpoint — five harness defects, the registry rescue, and the endgame reframed
1,947 banked today. Throughput went 65 -> 341 req/min peak and gates 63 -> 39 min, all by
removing harness defects rather than changing models. The registry incident (config/overlays.mk
committed EMPTY, taking main and every overlay gate down) is written up with its blast radius
and the config_sane guard that now prevents it.

The strategic finding is the part that matters for planning: 3,652 open crackable functions
collapse to 334 DRAWABLE skeletons, of which 308 are gen6+ walls — the ~2,900 untouched
functions sit behind those skeletons and bank by mechanical remap, not by drafting. Wide
drafting now converts at 5%. main is 327 crackable, not 1,288.

Four rule candidates for PhaseEnd (R51-R54), each earned by a defect that fired today.
2026-08-25 15:34:50 -06:00
Drew T 9696f8fb14 docs(phase-31): S59 FINAL checkpoint — six harness defects, six lanes, and the rules they earned
1,342 banked, 140 commits, open stubs 6,575 (main 1,493 / overlay-md 5,082).

The session's one lesson, measured six times: every lane that looked like the models
underperforming was a harness defect — an -O0 oracle nothing ever passed, a lane
retired on a card-size verdict, carve machinery nothing fed, a poisoned main baseline
that made 737 drafts read as bad, a soft 429 killing 44-72% of shards at turn 1, and a
stager consuming one bit of one verdict.

Records what landed (jtbl island split + gate automation, the -O0 census and unlock,
the new main and distill lanes, two RED binaries fixed, the throughput settings with
their probe evidence, the portable-workflow doc), seven rule candidates for PhaseEnd
approval, the ranked open threads with the A-prop residual named and sized, and a
resume procedure that starts from campaign_status.py and verifies from the process
rather than the file.
2026-08-25 00:25:46 -06:00
Drew T a0dc0d9b09 docs(phase-31): S59 second checkpoint — agents returned, every fix landed and verified
Records the per-type answer to 'can the waves draw and bank this now', the four
commits that landed after the agents returned, three rule candidates for PhaseEnd
(a card may not name a lever the knowledge base lacks; draw-time bankability; a
budget is part of the harness), and the ranked open work from the agents' docs.
2026-08-24 12:50:29 -06:00
Drew T a664636ce2 docs(phase-31): S59 checkpoint — the three S58 tooling lanes, each worked to a verdict
jtbl: the 154-A island split is byte-proven (one config line + jr_isolate_all --only),
with the object-level sh_size control a green SHA cannot give; four md_*/main tool
blindnesses fixed; one delay-slot instruction short of the first bank, logged.
o0: the census (167 real -O0 of 14,400; 116/14,148 ins stranded), the handoff's
md_MAIN_003/011 refutation shown to be itself wrong, the never-wired -O0 oracle fixed,
and 79 unbankable card-draws across 19 waves stopped at the source.
tells: restored and pinned to band 5-80 — the lane gap was a card-size gap, and 235
is refuted as the cause by the recorded pre-gate verdicts.
Three Fable agents in flight; their briefs and every fact they were given are recorded
here so a crash costs a re-spawn, not the knowledge.
2026-08-24 11:52:58 -06:00
Drew T 52b9e84208 docs(phase-31): S58 handoff — the three open tooling lanes in working detail
jtbl (36,685 ins, but 150 of 177 groups are singletons): Fable's review says the fix is SMALLER
than proposed — ONE inserted .rodata carve line plus jr_isolate_all --only. No _pre piece (it
cannot build), no ld_interleave change (the native script is already rodata-first). Harden
parse_config FIRST: it corrupts md_*/main configs on disk before erroring, which is why
jtbl_carve now hard-refuses them. The ox study's negative control is misattributed — build a
fresh one.

o0/cc1 (6,564 + 6,511 ins): the study is half refuted, and the doc header says which half.
md_MAIN_003/011 do NOT carry the -O0 fingerprint; the 311 files that contain $fp are the real
population. The unanswered load-bearing question for both is whether the EXISTING gate can bank
them unchanged — a lane that drafts what the gate cannot accept has already cost two sessions.

tells (86,602 ins): removed from drafting on four waves of evidence. aprop_autodraft is NOT the
destination (4.2% overlap, checked after I asserted it three times). The live hypothesis is
cookbook §235, the phantom symbol — one wave with it in the brief tests it cheaply.
2026-08-24 10:29:18 -06:00
Drew T e05dc116be chore(phase-31): S58 crash-recovery checkpoint + autonomous lane architecture
CURRENT_PHASE.md gains a CRASH-RECOVERY checkpoint (not a fresh-session handoff): what is
running, restart order, the measured fleet/scaling facts, the fixes that must not regress,
and the ordered work queue.

Lanes: drafter (never stop it), gater (restartable), maintenance (free A-prop sibling lane),
stallguard (60s auto-repair). Drafting holds no lock; one narrow draw-vs-gate lock exists
because build_wave_atlas reads corpus.stubs and misreads substituted drafts mid-gate.

main is off the wave critical path — 157 drafts parked to .run/main_queue/ rather than
stalling the gater for another hour on a bisecting whole-EXE rebuild.

api_agent: 5xx retried like 429 (a 502 was abandoning functions at near-19), HTTP_TIMEOUT
420s not 1800 (a hung request parked an agent 30 min), EXTRA_READABLE for tooling briefs,
and bare-directory paths no longer refused against their own granted root.
2026-08-24 00:24:11 -06:00
Drew T e7ebb1cd45 docs+rules(S58): R42 commit-banked-work-immediately, R43 refuse-unsupported-input
R42: gate_main reverted 61 byte-proven overlay banks it could not distinguish from its own
substitution (sweep_parallel gates commit=False by design). Fixed by committing overlay banks
before the main batch, chunking main at 8 to bound bisect cost, and replacing every blind
'git checkout -- src/ config/' with commit-or-refuse in ox_campaign and idiom_serial.

R43: sweep_parallel had an explicit branch admitting main, which cannot be gated incrementally
— wave ab banked 0/105 main cards while its non-main cards banked 94/115 (82%), and the wave
read as a drafting failure. sweep_parallel now refuses main and names gate_main.py.

Also: validate_targets now prefers the card's own addr field (named symbols like SYS_OBJ_F00
were MALFORMED and discarded whole 220-card waves); ox_campaign deals model lanes by
smallest-ratio scheduling (a 73-card wave had put 73 shards on ox and 0 on deepseek);
docs/accelerators.md gains the four vacuous-check defects.
2026-08-23 12:59:59 -06:00