Commit Graph

31 Commits

Author SHA1 Message Date
Drew T 4f7c3b64a3 docs(phase-33): commit-map + citations resolved to the rewritten history (C4–C7 — the tip commit)
- docs/commit-map.tsv: 4,032 rows (ordinal of the ORIGINAL main -> rewritten hash, author/committer dates, subject);
  1 pruned row of zeros (ordinal 1712, "session archive update"); 0 old hashes asserted; ordinal 1 unchanged by the
  rewrite (byte-identical)
- resolve_tokens: 1,238 commit:NNNN tokens -> shortest-unique new hashes in 98 files (docs, phase-ends, logs, tool
  docstrings, 2 C comments, the A5 evidence logs); residue left as tokens: commit:1712 x4 (the pruned commit),
  commit:orphan-24 x2, commit:orphan-26, commit:orphan-35 (cited commits that exist in no lineage)
- the rewrite (C4): filter-repo 2.47.0 on a bare clone of the C2 tip, 311 s, exactly 1 pruned, main 4,032 -> 4,031;
  the pre-rewrite history is mirrored in the private archive repo and in the local bundle
- the proof (C5): verify_rewrite 4,031 pairs / 0 failures; absent_scan 0 offenders; gate_scan 0 offenders on the clone
- adoption (C6): 100 text files differ at the tip, 0 purge paths, 0 added/deleted; leftover refs dropped; no gc yet
- resolver skips tools/public_rewrite/ (its self-test fixtures are the token grammar, not citations); repo-local
  identity is the GitHub noreply address from here on; CURRENT_PHASE: C4–C7 logged, checkpoint -> NEXT = C8
2026-09-06 23:28:39 -06:00
Drew T 433c7cd8ef feat(resolver): the integration-resolver lane — zero-token re-judging of the ledgers' shape-correct stock at the real TU (P31 S61 T10+/S61-1)
frontier-analysis-s60 §4 measured that ~571 open functions had FINISHED drafting (closeness-0 backlog
rows / reloc shape-MATCH rejects) and were being re-drafted wave after wave. tools/integration_resolver.py
treats those ledgers as an index: still-open? -> rtu_match at the real split TU (CC1: the gate ladder's
draft-side transforms, one retry) -> reloc_identity as the disagreeing oracle (rtu masks reloc fields)
-> aprop_symfix on MISMATCH/shape-MATCH -> stage -> sweep_parallel (whole-binary SHA, sole arbiter)
-> commit at once (R42). Refuses main by name (gate_main owns it), //@EDIT drafts, dirty trees, collapsed
registries; every drop is counted (R32); a negative control over recently-banked functions must pass
N/N before any verdict is trusted (R35/R39 — its first form picked carve moves as banks, 9/12 FAIL,
and was fixed before a single stock verdict was read). Ledger .run/resolver/verdicts.jsonl keyed by
(binary, fn, draft-sha, split-TU-sha) so unchanged rejects are never re-judged.

First pass (commit:2991): 1,352 nominated -> 901 already banked, 27 main -> 424 judged in 41 s ->
245 staged (57.8%; 242 raw, 3 via transforms) -> 63 banked (net INCLUDE_ASM delta; that commit's
subject says 72 = gross incl. 9 carve moves), 182 gate-refused, zero model tokens, ~10 min total.
Lane wrapper tools/lanes/resolver_lane.sh (holds .run/auto/draw.lock for judge+gate: rtu reads the
TUs a gate splices into).
2026-08-25 23:50:26 -06:00
Drew T 094dfebef4 fix(campaign): measure what the campaign can see — banked-today from the INCLUDE_ASM invariant, gate_min beside wall_min, a lock-aware fleet R22 window (P31 S61 T10+/S61-8)
campaign_status: 'today: N banked' summed '— N banked' commit subjects and missed every bank that
rode in a chore/maint commit (S60: 2,185 reported vs 2,644 net stubs removed); it now derives the
number from INCLUDE_ASM stub counts at last-commit-before-midnight / HEAD / working tree (R33), and
alive() is anchored so pgrep no longer matches its own wrapper (every lane read ok with 0 processes).
ox_campaign gater: the ledger's wall_min counts drafting + queue wait since the ready marker's t0;
gate_min is the gate alone (the '30-67 min gates' picture was this conflation).
maintenance lane: the fleet R22 sweep skipped whenever any gate was in flight, i.e. always (last
real sweep 12:54 08-25); it now takes .run/auto/draw.lock and waits its turn, skipping only for gate_main.
2026-08-25 23:49:43 -06:00
Drew T 065e122c64 docs(phase-31): S60 shutdown — all lanes stopped, tree clean, sentinels set
Every campaign process stopped deliberately at session end (0 alive, verified after settling).
.run/ox_campaign.stop and .run/auto/STOP are SET — delete both before relaunching, or every lane
exits immediately.

One dirty overlay TU left by a killed gate was BUILD-VERIFIED as an abandoned substitution (the
binary failed to build with it) and reverted rather than committed — R42's distinction between a
proven bank and mid-gate residue, decided by the bytes.

Two shutdown hazards recorded: pkill on a lane's shell leaves its python running (hit the
drafter, gater and main lane tonight — kill by PID, verify with ps -o lstart), and a bash case
pattern 'src/[a-z0-9_]*.c' matches ACROSS SLASHES, which classified an overlay TU as a main TU
and nearly reverted the wrong file.

Also committing the two lanes built today: tools/lanes/elastic.sh (starts serial idiom lanes when
the API window is idle and the gate queue is deep — it scales the work that is NOT gate-bound,
because adding drafters to a full gate queue makes the backlog worse) and
tools/lanes/grinder_lane.sh (runs tools/grinder.py, the Phase-21 LLM-free permuter, which had
never been run this campaign against 5,388 near-miss rows).
2026-08-25 22:53:09 -06:00
Drew T f78ea55b3e fix(gate): log what the gate actually does; --gate-jobs 24 -> 32; TAIL_DONE_FRAC 0.80
THE GATE WAS A BLACK BOX. sweep_parallel's stdout was captured and dropped, so a gate logged
"reloc_identity -> gating 216" and then THIRTY MINUTES OF SILENCE before its bank line — no
worker count, no per-binary progress, no phase-A/phase-B split. Gate times went 31 -> 37 ->
50 -> 67 min across ej/ek/en/eo with nothing to diagnose from, and I twice asserted things
about phase B that the log could not support (its absence measured LOG CAPTURE, not
behaviour). A lane that must run unattended has to leave evidence.

MEASURED WHILE DIAGNOSING, and it rules out the obvious suspects: load average 2.6 on 32
cores with 1-3 concurrent builds during a gate — the gate is NOT CPU-bound and is not
saturating its own -j 24. Raising to 32 is cheap given ~8% utilisation, but the real answer
will come from the log this change adds.

TAIL_DONE_FRAC 0.85 -> 0.80. 0.85 overcorrected: the fleet fell to 15 agents / 11 req/min
because the drafter parks between waves while the gater drains a deep queue. 0.75 was too
deep (26% 429s, draft completion sliding 94->91->73->47% across eq/er/es/et). Neither number
is really the lever: the drafter cannot start a wave the gater has no room for, so the gate
throughput is what bounds the campaign now.

Generational tiering confirmed already correct: the top-off orders by generation at both
assembly levels (group ranking and within-group) without FILTERING any tier out, so every
generation stays eligible and the scarce never-drafted work simply goes first.
2026-08-25 19:00:02 -06:00
Drew T 84706c9860 feat(draw): draft EVERY instance, siblings included (ONE_PER_GID=0) + K attempts per card
Two ways to spend a free drafting window on a tail that is 92% walls.

ONE_PER_GID=0 — the sibling collapse exists because a same-gid sibling banks by mechanical
remap once its exemplar cracks, so drafting it pays for what the remap does free. That prices
AGENT TOKENS as the scarce resource. On the free ox window they are not, and the collapse is
what makes 3,271 open crackable functions look like 334 drawable skeletons — of which 308 are
gen6+ walls whose exemplars have already refused six waves each. A sibling drafted directly
can crack on its OWN terms instead of waiting on an exemplar that never will.

Measured on a live draw rather than argued:
    uncollapsed  627 cards / 44,403 ins / 160 binaries / 215 gate groups = 2.9 drafts per rebuild
    collapsed    334 cards / 30,926 ins / 104 binaries / 127 gate groups = 2.6 drafts per rebuild
The gate cost is per (binary, TU) group and chunked, so siblings landing in binaries the wave
already touches are close to free at the gate — the card count nearly doubles and the gate gets
MORE efficient per build, not less.

ATTEMPTS=K — K independent shots at each card. The gate cost does not multiply: reloc_filter
keys by fn and staging writes <binary>/<fn>.c, so a function still gets exactly one whole-binary
build per wave; the attempts compete to BE that build, ranked by match_one, which is local and
needs no build. Alternates stay on disk for a later recovery pass. Default 1 (no-op).

A BUG I CAUGHT IN MY OWN SELECTOR before it shipped: it passed binof[fn] (a BINARY NAME) where
match_one wants --asm-subdir (an asm DIRECTORY). Every attempt would have scored identically at
infinity and the picker would have silently degraded to first-seen while appearing to rank —
the same "true number about the wrong thing" class as the day's other defects. Fixed with a
subof map, and a missing subdir now returns neutral instead of a fake score.
2026-08-25 16:18:43 -06:00
Drew T b0587c573e fix(idiom_serial): main is never a target, and a contended index is not a dirty tree
The serial lane's first-ever run on tells died 0 seconds in, on both counts:

  [1/6] func_8017D1C0 @ ov_SC06_027 · 118 ins  -> drafting
  [2/6] func_8001382C @ main · 103 ins         -> could not commit — refusing on a dirty tree

MAIN IS NOT A SERIAL-LANE TARGET. main gates through gate_main on the main lane's own cadence
(a clean whole-EXE rebuild with bisection), so this lane cannot bank it however good the draft
is — a main target burns a slot to learn that. md_* was already excluded for jtbl; main never
was, because the lane predates main being in the atlas at all.

A CONTENDED INDEX IS NOT A DIRTY TREE. Six campaign lanes plus the re-gate runner commit
continuously, so `git commit` here loses the index.lock race routinely — twice today it was a
STALE lock blocking every lane for 8 and 21 minutes. The lane refused over a tree that was
fine. It now retries the commit six times at 5 s before concluding, and the refusal is
reserved for FOREIGN DIRT it must not adopt — which is the property that rule was protecting.

Verified: pick_targets('extend-tell', 8, 80) now returns 8 overlay targets, no main.

Also in this commit: MAX_BINS 160 -> 50 in the drafter shell. Wave size was the right lever at
50% conversion (dd: 217 banked of 422 gated in 39 min); at 5% it is dead weight — dq banked 8
of 167 gated and took 98 MINUTES, while four drafted waves queued and the free-ox fleet sat at
14 agents / 10 req/min. At this conversion a 150-card wave banks what a 430-card wave banks,
in a third of the gate.
2026-08-25 14:57:13 -06:00
Drew T baabdd0478 fix(jtbl): end the stale-pad-spec RED class, and check the fleet every pass
THE CLASS. JTBL_PADS is a per-object spec written by jtbl_carve at CARVE time — one entry
per rodata `.align 3`, each 0 or 4 — describing how many jump tables the object emits. That
is a DERIVED property of the current source stored as static config, so any bank carrying a
`switch` (or any bank being reverted) invalidates it and nothing re-derives it. Three of the
five REDs on 08-25 were this one design choice: ov_SC02_005 and ov_SC07_006 from wave dd
banking switch-bearing functions, ov_SC04_018 from the identical symptom with the opposite
cause — a reverted bank taking its table with it.

WHY NOT DERIVE IT. The COUNT is derivable from the assembly stream; the VALUES are not — a
pad records where the ORIGINAL image has an inter-table pad, which lives in the retail
layout, not in our source. Guessing shifts every downstream data symbol: silent corruption,
the worst outcome available. So tools/jtbl_pads_fix.py does not derive. It ENUMERATES the
2^(N-1) candidate specs (first entry 0, rest in {0,4}) and accepts one ONLY if it is the
UNIQUE candidate that rebuilds the binary byte-identical to config/check.<bin>.sha; zero or
two matches restore the original and refuse. R39 negative control: on a healthy binary it
reports "no pad-count drift" and changes nothing.

TWO INSTRUMENT BUGS THIS TOOL FOUND IN ITSELF:
  * JTBL_PADS is a target-specific MAKE VARIABLE, so changing it does NOT make the .o out of
    date. The first run reported "no drift" against a spec I had deliberately broken. It now
    deletes the armed objects before every build — R22's incremental trap in config costume.
  * A failed object build leaves the PREVIOUS binary in build/<bin>/<bin>, so
    `make build; sha1sum build/<bin>/<bin>` reports the OLD artifact as if it were this
    build's — a FALSE GREEN over a build that never linked, which briefly convinced me two
    binaries were fixed. build_sha now deletes the output too and requires make to exit 0.
    Same family as R49: an error inside something shaped like success.

CADENCE: the fleet sweep runs EVERY maintenance pass, not every 4th. A RED fails at BUILD,
so every draft gated against it is rejected regardless of quality and the wave reads as a
drafting failure — detection latency is the whole cost. Gates now finish in ~35 min rather
than 60, so the sweep is affordable each pass. It still FIXES NOTHING by design, with this
single exception, admissible only because it proves itself against the byte gate first.
2026-08-25 12:04:14 -06:00
Drew T 188314f081 perf(throughput): unblock the main lane, both gate logs, deeper wave overlap, atomic atlas
1. MAIN LANE — the largest single block of unfinished work was drawing 32 cards a wave.
   main_lane.draw() never passed --max-bins, so it inherited build_wave_atlas's default of
   12 gate groups — a cap that exists because each group costs a whole-binary rebuild, and
   main's own --only-bins docstring says the opposite applies to it: "main is gated ONCE per
   SLATE, so main has no per-TU gate cost and --max-bins can be large". Nobody passed it.
   Measured cost: main banked ~19 stubs/hour against 1,291 remaining while the overlay lane
   ran 650-card waves beside it. Now --max-bins 400 (MAIN_MAX_BINS overrides), and the lane
   shell draws 600 cards with 600 workers instead of 200/150.

2. TWO LANES GATE, SO READ BOTH LOGS — a defect I introduced this session. The in-flight
   exclusion derived "this wave has been gated" from .run/gater.log only, but the main lane
   gates its own waves into .run/main_lane.log. Every m## wave therefore looked permanently
   in flight and main's draw lost 425 cards to an exclusion meant for work in progress.

3. TAIL_DONE_FRAC 0.80 -> 0.65. At 0.80 the fleet runs 2-3 overlapping waves at ~250
   req/min; the residual troughs are the gap between one wave draining and the next ramping.
   65% keeps 3-4 waves overlapping. Stragglers keep their full 700s grace in the finisher
   thread — this changes when the NEXT wave starts, never what lands.

4. ATOMIC ATLAS WRITE. The lanes read .run/atlas.json at every draw and atlas.py dumped
   straight onto it, leaving a truncated file readable for the length of the write. Now
   written to .tmp and os.replace'd.

Context for 1-3: the atlas both lanes draw from is dated 08-23 01:13 — two days stale,
predating ~4,600 banks — and its regen chain is running now (its own R32 assertion caught a
stale family map first and named the fix).
2026-08-25 11:28:51 -06:00
Drew T d421f1bbae perf(draw): every second wave was re-drafting the wave still in flight
MEASURED over 18 consecutive waves. Consecutive card sets: ck->cl 239/239 shared, co->cp
238/238, cv->cw 222/222, db->dc 208/209 — and the "different" pairs still shared 50-90%.
Yield alternated in lockstep: 47.6% / 3.8% / 35.3% / 3.6% / 29.9% / 3.7% / 43.4% / 14.6%,
because the duplicate wave gates AFTER the original banked its cards. Half of all drafting
went to work already in flight, and it read as campaign decay.

ROOT CAUSE: --retry-unbanked returns "previously waved but still an OPEN STUB" cards to the
pool — right in principle, unfinished work is not spent work. But the pre-draw for wave N+1
runs WHILE wave N drafts, when none of wave N's cards have been gated, so every one of them
is still an open stub and the filter hands the whole wave back. The ranking then rebuilds it
card for card. The filter knew about "banked" and "not banked" and had no notion of "in
flight".

FIX: a wave is finished when its GATE has run, and the gater already says so in its own log
(R33 — derive from the artifact that exists). Tags with no GATE line stay excluded; a tag
with no gate line whose cards are older than 6 h was killed, and is released so nothing is
locked out forever.

THROUGHPUT, same commit — the draw was setting the campaign's request rate:
* --max-bins 24 -> 160. Concentrating a wave into 24 gate groups was a CPU-economy choice
  made when CPU was scarce. It is not: a live gate runs at load 2.7 of 32 cores (8%), one
  harvest_verify at --chunk 1. Meanwhile the drafting fleet — the resource actually bounded
  by the free-model clock — got 196 cards out of 699 available. Re-drawn with 160 bins:
  644 drafts / 46,590 ins across 136 binaries, 3.3x the wave for the same gate economics.
* TAIL_DONE_FRAC 0.95 -> 0.80. Overlapping at 95% still left 25% of minutes under 20 req/min,
  because a wave's last 5% is its SLOWEST 5% and 12 stragglers cannot fill a fleet. Handing
  off at 80% starts the next ramp with ~40 agents still working. Stragglers keep their full
  700s grace in the finisher thread; nothing is cut short.
* --queue-depth 2 -> 4, so a bigger wave's longer gate never parks the drafter.

Arithmetic this is aimed at: req/min = agents-in-flight x ~0.8 (a 16k-token turn at ~30
tok/s emits few requests). 644 cards x two overlapping waves puts the fleet where the
endpoint has already been measured to sustain it — 764 req/min for 15 min at 8% 429s, peak
2,755 in one minute.
2026-08-25 10:19:17 -06:00
Drew T 80800911fd feat(decomp): main lane m11aab — 3 banked
One clean whole-EXE rebuild verified the batch (gate_main), and main re-checked
BYTE-IDENTICAL against config/check.us.sha before anything was credited.

  SYS_OBJ_2DD8
  _clr
  func_8005EA68
2026-08-25 00:55:41 -06:00
Drew T 75bfb04226 fix(maintenance): restore the S59 lane logic a stale second copy had reverted
.run/maintenance.sh (what runs) and tools/lanes/maintenance.sh (a pre-S59 copy) had
diverged. The 150->50 threshold tune landed on the stale copy and was then copied over
the live one, silently reverting five S59 fixes:

* the R47 shape filter (staging fell back to status=='AGREE' alone — the exact defect
  that staged 82 hopeless drafts every 45 minutes)
* the R48 (binary, fn) keying (bare-fn keys collide across overlays)
* reloc --fix MISMATCH auto-repair (measured 4/4 repaired to AGREE)
* rtu_second_chance (re-judges standalone COMPILE-FAILs against the real TU)
* fix_tu_ret_decls (the return-type half of the stale-decl wall)

Rebuilt from the S59 lineage with the 150->50 threshold and the periodic fleet R22
re-applied, both paths now byte-identical, `bash -n` clean, and the two-path hazard
documented in the header so the next edit cannot repeat it.

Also: relaunch_drafter_shell.sh 30s -> 5s ready-marker poll; regenerated backlog and
fleet progress artifacts.
2026-08-25 00:35:27 -06:00
Drew T 619f4dfaee tune(maintenance): trigger threshold 150 -> 50 banked (Drew)
150 was tuned for a lane that only re-swept an unchanged sibling pool and banked
nothing. The lane now consumes every verdict layer in the A-prop pipeline, carries the
free pre-gate reject recovery, and runs the periodic fleet R22 — a pass is worth
running on a much smaller refill.
2026-08-25 00:21:16 -06:00
Drew T 85f6b79471 feat(maintenance): a periodic fleet R22 — nothing was watching the binaries no lane touches
Two binaries sat RED for hours today and nothing noticed: ov_SC07_010 from a
maintenance commit whose final tree state was provably never built, and ov_SC07_002
from a stale 2-table jtbl pad spec written at wave bp. Every lane only ever checks the
binary it is currently touching, so a byte-gate — a correctness oracle — was silent
about everything it did not build. They were found by accident, by an agent's scoped
R22 sweeping 141 binaries.

Every 4th maintenance pass (~3h), skipped while any gate is in flight (check-all
rebuilds stale objects and must not race a gate), it runs the fleet check and writes
any REDs to .run/fleet_red.txt with a loud log line. It FIXES NOTHING: a wrong repair
to a pad spec or a config is exactly how a silent byte shift gets committed, and the
two we fixed today each needed a different, evidence-led remedy.
2026-08-25 00:03:45 -06:00
Drew T e40fe9c116 fix(main-lane): the baseline was RED — guard every gate with a no-draft control (S59)
PROVEN: from 14:57:01 to 18:43:23 today HEAD built main to 307aa45d… against the
expected 143dbb89…, with NO draft substituted (measured under gate.main.lock, no
gate_main alive). Auto-commit commit:2693 had adopted a mid-flight gate_main
substitution — its carve-out reverted main's TUs, gate_main re-wrote them, and
`git add -A src/` swept the unverified bodies in (a TOCTOU race, 14 s after a
bisect chunk banked). Every main batch after it was doomed before its first
draft was judged: m00–m03 card cycles drafted ~737, slated 160, banked 0, and
burned ~50 clean rebuilds bisecting innocent slates. commit:2712 restored the
green content by accident (it swept this investigation's diagnostic checkout).

gate_main: on any batch failure, ONE try_batch([]) control runs first — if HEAD
itself is red it prints BASELINE RED, leaves the slate reusable, exits 3 (R40).
clean_build no longer reports a linked-but-mismatched build as "no binary" (the
build target embeds the SHA check), the compile-conflict shortcut fires only on
error-shaped lines naming a symbol some draft in the slate actually uses (the
baseline's own func_800143AC implicit-decl WARNING was matching — every m04
chunk died with "drafts declaring it: []"), reverts narrow to top-level src/*.c
(main_tus) so a main gate can never destroy overlay lanes' in-flight work, and
--assert-baseline is a first-class mode.

main_lane: every cycle opens with gate_main --assert-baseline and REFUSES to
draft or gate against a red baseline (R43) — BaselineRed parks nothing, burns
no tries, writes .run/main_lane.BASELINE_RED, re-checks every 30 min.

Adopters (ox_campaign ×3, maintenance.sh, gate_stage, gate_lane, idiom_serial):
main's TUs (top-level src/*.c) are never staged and never reverted by an
overlay/maintenance lane — one writer (gate_main), one committer (main_lane,
after the whole-EXE SHA re-checks green). Unstage-after-add is race-free where
the old revert-then-add was the losing half of the TOCTOU.

Diagnosis, evidence and the full timeline: docs/tool-designs/main-lane-fix-s59.md
2026-08-24 19:08:47 -06:00
Drew T 86ebc95a81 feat(ops): one status view for ALL lanes; straggler grace = one full turn; distill dedupe
STATUS BLINDNESS (Drew): status checks kept reporting the overlay drafter and the
gater — the lanes whose logs scroll — while the main, maintenance and distill lanes
went unmentioned for hours. A lane you do not report is a lane you do not notice
failing: the main lane spent an afternoon on an old config and bisected a whole batch
to zero banks without that ever reaching a status line. tools/campaign_status.py
prints every lane with ITS OWN metrics, read from artefacts rather than memory.

STRAGGLER GRACE 120 -> 700, tied to HTTP_TIMEOUT so they cannot drift. collect_drafts
queues a wave once 95% of shards finish, then waits this long for the rest — and 120s
is shorter than a single turn (~530s for a 16k generation at ~30 tok/s). So raising
the token budget converted truncated turns into agents guillotined mid-thought with NO
draft: overlay draft completion fell from 84-89% at 8k to 41% (bt) and 69% (bu).

DISTILL DEDUPE: each pass re-offers everything unmined, so the lane wrote a fresh
overlapping marker every five minutes — eight queued, each a superset of the last, and
a reviewer cannot tell which one is the work. One pending marker at a time.
2026-08-24 18:27:05 -06:00
Drew T 6a19755969 feat(lanes): restart the main lane at its idle boundary, not mid-work
Same problem as the drafter: an env/arg change (MAXTOK, HTTP_TIMEOUT) only reaches a
fresh shell, and the main lane is usually either drafting or gating. This waits for
the one safe window — no main-lane agents alive and no gate_main running, i.e.
between the gate and the next draw — then restarts. Mid-draft would discard drafted
work; mid-gate would abort a batch (safe, since gate_main reverts its own
substitution, but wasteful).
2026-08-24 16:39:42 -06:00
Drew T db3fe3490b fix(lanes): raise HTTP_TIMEOUT with MAXTOK — they are one setting, not two
Probed ox-alpha directly on a real MIPS derivation:

  no reasoning cap      265.2s  finish=stop  completion=8,067  reasoning=0   30 tok/s
  reasoning cap 2000     22.3s  finish=stop  completion=  672  reasoning=0
  reasoning cap 6000     41.4s  finish=stop  completion=  618  reasoning=0

Three findings. (1) ox reports reasoning_tokens=0 — its thinking is IN the content
stream, so the output cap was capping the reasoning; that is exactly why turns ended
in 'no tool call (finish=length)'. (2) The uncapped hard prompt wanted 8,067 tokens —
it was finishing precisely where the old 8k cap cut it off. (3) It generates at ~30
tok/s, not the ~54 I estimated from turn gaps, so a full 16k generation needs ~530s
and the 420s socket would have killed the very turns the bigger budget exists to
allow. A timeout wastes the whole turn; truncation at least leaves a partial.

HTTP_TIMEOUT=700 on both drafting lanes. The ordering that must hold is generation <
HTTP_TIMEOUT (700) < stallguard's wedged-agent kill (1200s). 420 was itself deliberate
— 1800 once parked a hung agent for thirty minutes — and 700 keeps a hang under 12
minutes without strangling legitimate deep reasoning.

Also recorded: a reasoning cap DOES work on ox, but it shortens the ANSWER too (618-672
total tokens), so it is a quality knob, not a fix for truncation.
2026-08-24 16:36:47 -06:00
Drew T 99e71bfa83 perf(waves): MAXTOK 8000 -> 16000, and free recovery of pre-gate rejects
THE OUTPUT CAP WAS EATING THE TURN BUDGET. Wave bk's shard logs: 240 of 244
turn-finishes were 'no tool call (finish=length) — NUDGE n/6'. The model was
exhausting its 8,000-token output budget BEFORE emitting a tool call, so the turn did
no work; an agent gets six nudges before giving up. That is why MATCHes average 2.8
oracle calls against a 24-turn budget — the turns are going to truncation, not
iteration. ox is free, so a bigger output budget costs latency and nothing else.

Measured alongside it, and worth recording because it redirects the obvious fix: turn
caps are NOT binding on the default lane. Across 1,166 agent completions, non-MATCH
runs used a median of 4 oracle calls and a p90 of 12, and exactly 1 of 194 reached 20
of the 24 available. Agents are not running out of turns; they are giving up early
after truncated turns. (The tells lane WAS cap-bound — 98 of 270 — which is why it
already has 40 turns.)

Plus tools/recover_rejects.py, wired into the maintenance lane: rebase the pre-gate
rejects whose body already matches and only the symbols are wrong (§171), stage them
for the lane's existing free gate. Zero model tokens; it only stages, so a bad
recovery can waste a build but never a bank.
2026-08-24 16:23:37 -06:00
Drew T 801a32179e perf(lanes): one lane, one band — the rotation is obsolete and its large slot is the worst wave we run
The tells slot became redundant when build_wave_atlas started reserving 60 tell-lever
cards inside every ordinary wave: a dedicated tells wave draws 70-87 cards, a quarter
of a default wave, for a full 40-minute slot.

The 120-2000 slot is worse than redundant. Bank rate by size, measured: 57% under 50
instructions, 30% at 50-80, 22% at 80-120, 3% at 120-200, 6% above. Wave br drew 69
cards on that band — roughly 3 banks for a slot that a full-band wave turns into ~150.
Large functions are not abandoned: the full band contains them and the draw takes
mass-first within each gate group.

Takes effect at the next wave boundary via relaunch_drafter_shell.sh.
2026-08-24 16:08:34 -06:00
Drew T 86b96b2125 perf(lanes): gate-jobs 12 -> 24, and tells become a quota inside default waves
TELLS QUOTA (Drew approved): a dedicated tells wave drew only 70-87 cards — a full
40-minute drafting slot at a quarter of a default wave — because the 5-80 size cap and
the tells pool cannot fill more. Tells now ride inside ordinary waves with a 60-card
quota, same as jtbl. The size cap moved into the draw itself: tell-lever members above
--tells-max-ins (80) are not drawn at all, because the measured bank rate is 27-40%
at 5-80, 10% at 81-120, 1% at 121-200 and 0% above — those 383 members / 51,941 ins
are idiom_serial's work, and the skip counter names it (R45).

A QUOTA IS A FLOOR UNLESS IT IS ALSO A CEILING. First test: putting the tell levers in
the default list let them win the ranked fill too, and a 300-card wave came back 122
tells (41%). The size cap held; the mix did not. Tells now enter through the quota or
not at all.

GATE JOBS 24. The quotas deliberately pull cards from binaries outside the ranked gate
groups, so a measured draw went from ~24 groups to 63 — 63 whole-binary rebuilds per
wave, five serial batches at 12 jobs. The box is 32 cores at ~6% (load 3.1) with 39 GB
free. Lands via restart_gater_when_idle.sh so no sweep is killed mid-flight.

Correction to the record: the live draw already passed --max-bins 24 (plus
--one-per-gid and --exclude-bins main). An earlier measurement of mine used the tool's
default of 12 without those flags and read as 'max-bins is the cap' — it is not;
--one-per-gid is, and deliberately: it defers same-skeleton siblings to the free
deterministic remap instead of paying an agent twice.
2026-08-24 14:34:09 -06:00
Drew T c4a7a45bf5 fix(lanes): drop the paid lane before the credit floor kills drafting
Balance $2.56 and falling ~$1.43/h over the last three waves ($4.56 at 11:54 ->
$2.56 at 13:18) — about 23 minutes from --credit-floor 2.0. That floor does NOT pause
the paid lane: it breaks the whole drafting loop, and the shell then restarts a python
that breaks again, so the clock-limited resource dies on a check about money.

ox-alpha is free for the rest of this window, so drafting continues on ox alone at zero
burn, and the floor drops to 0.25 because with a free model the balance stops being a
proxy for 'can we draft'. deepseek was 280 of 2,000 workers — its value was an
independent 429 ceiling, not throughput.

Takes effect at the next wave boundary via relaunch_drafter_shell.sh, so wave bp's
in-flight drafts are not lost. Restoring it after a top-up is two edits, named in the
script's header.
2026-08-24 14:17:52 -06:00
Drew T 6c580a28a0 docs(distill): name the review tier — Opus/Sonnet, not Fable
Drew, 2026-08-24: distillation is judgement over an existing corpus (read harvested
notes, decide covered / addendum / new against 760+ sections), not a new wall class.
Fable is for the walls — an unsolved tooling problem, an adversarial design review, a
residual no documented lever reaches. I routed a distill batch to Fable; recorded here
so the next session reads the tier off the lane rather than guessing it.
2026-08-24 13:57:31 -06:00
Drew T 333a169458 fix(distill): state is {tag: novel-count}, never a done-list — it nearly buried 82 candidates
A re-gated wave rewrites its candidate file with NEW rows under the SAME tag, so
"have I seen this tag" answers the wrong question. Two instances in one hour:

  * wave `at` was re-gated hours after its first harvest, so an mtime-keyed seed
    called it new and 52 mostly-re-derived rows went to a reviewer;
  * my own hand-edit of the state folded "queued for review" into "reviewed", which
    marked `ax` and `bm` — 82 candidates, gated minutes earlier — as mined by nobody.
    Caught only because their files were newer than the edit.

The scan now compares COUNTS: a tag re-opens the moment its file grows past what was
mined from it. Extracted to tools/distill_scan.py so the logic is testable rather than
living inside a heredoc inside a lane loop (the heredoc-in-heredoc edit is also what
produced a syntax-broken lane script a minute earlier).

Verified: the lane now raises exactly the true pending batch — ax + bm, 82 novel.
2026-08-24 13:38:52 -06:00
Drew T 90fa74ba49 feat(lanes): wait_distill — block until the next distill batch, no polling
The harness re-invokes the main loop when a background command exits, so a blocking
wait IS the self-reminder; a sleep-then-tail poll loop would just burn tokens for the
same information.
2026-08-24 13:02:30 -06:00
Drew T fc24404452 feat(lanes): the distill lane — the second half of the flywheel, beside drafting
The gater already harvests every wave before the next draw (Drew's 2026-08-23 rule),
but that is EXTRACTION: it writes .run/idiom_candidates.<tag>.md and stops. What
changes the next wave's behaviour is the COOKBOOK, because that is what the drafting
agents grep — and distillation was batched per session, so the ore piled up: 165
novel candidates across 4 waves within three hours of the last cookbook update.

The lane does the zero-token half — watch, count novel rows, and raise a READY marker
naming the waves when a batch is worth a reviewer's turn (>=30 candidates or >=2
waves). It writes no cookbook, no src/, no config/: a bad harvest cannot pollute the
knowledge base on its own, and landing stays a reviewed step (us + a subagent).

It never blocks a draw. Wave N's ore is distilled while wave N+1 drafts, so wave N+2
is the first that can grep it — stopping the drafter to think cost 139 of 162 idle
minutes on 2026-08-23.

State seeded so it does not re-mine what S58 already landed as §233-§259: the "done"
list holds the 17 candidate files written before commit commit:2637.
2026-08-24 13:02:23 -06:00
Drew T 2dd00be292 fix(lanes): the relaunch watcher must watch marker NAMES, not the queue count
The ready queue is two-sided — the gater consumes markers while the drafter produces
them — so 'the count went up' is not 'a wave queued'. The first version latched
BEFORE=1, the gater consumed that marker, and the next wave queuing took the count
back to 1, which is not greater than 1: it would have waited out its full 90-minute
deadline while the event it waits for happened. True about a number, false about the
world.
2026-08-24 11:19:27 -06:00
Drew T d980540095 fix(lanes): tells draws a SMALL band — the lane gap was a card-SIZE gap
Joined the campaign ledger to each wave's cards and to the functions its own commit
banked (removed INCLUDE_ASM lines), pooled over bb/bg (tells) vs bc/bf (default):

  nins      default        tells
  0-50      303/528  57%   27/ 67  40%
  50-80      43/145  30%   20/ 73  27%
  80-120      9/ 41  22%   10/100  10%
  120-200     1/ 30   3%    1/ 68   1%
  200+        2/ 35   6%    0/ 30   0%

At equal size the lanes are close below 80 instructions and both collapse above it.
What separated them is the card size MIX: default's cards are median 37-39 ins, the
tells pool median 89-95 (2.4x), so 'the tells lane is broken' measured the population,
not the lever. Tells now draws 5-80.

This also refutes the S58 handoff's one live hypothesis for tells (235, the phantom
symbol). Checked the recorded reloc_identity verdicts first (R38): among MISMATCH?
rows, the fraction whose instruction SHAPE already matched — the symbol-only class 235
describes — is 25/87, 16/79, 17/66, 13/57 on default waves but 6/66, 4/81, 5/47 on
tells. Tells drafts fail because the BODY is wrong, not the symbols, which is what a
2.4x larger median predicts.

Also adds tools/lanes/relaunch_drafter_shell.sh. bash parses a while...done body in
full before running it, so a lane-ARG change is invisible to the running shell and a
python bounce re-runs the OLD command line — measured at 11:02, when the bounced
python came back on the pre-S59 lane list 18 minutes after the file changed.
2026-08-24 11:09:38 -06:00
Drew T 6752653e6f feat(lanes): one-shot watcher that bounces the drafter python once the in-flight wave queues
Ships a lane-arg change without losing drafts: the running python still carries the
pre-S59 args, and killing it mid-wave discards everything the fleet has drafted for
that wave. Waits for .run/ready/<wave>.json, then pkills only the python — the
drafter SHELL, the lane that must never stop, relaunches it from .run/drafter.sh.
90-minute deadline so it never lingers.
2026-08-24 10:45:46 -06:00
Drew T 2257be920c fix(lanes): tells was removed for the wrong reason — pin it to the full band
The S58 removal cited four waves (as/aw/az/bd) and blamed the lane. All four ran
at band 120-2000. The campaign ledger splits the population by band instead:

  tells @ 120-2000   228 drafts ->   18 banked =  7.9%
  tells @ full band  655 drafts ->  161 banked = 24.6%
  default @ full     2,996 drafts -> 1,335 banked = 44.6%

Of drafts that actually reach a gate the lanes are indistinguishable (tells 54.4%,
default 56.4%) — the whole loss is reloc_identity discarding phantom-symbol drafts,
which is cookbook 235 and a BRIEF fix, not a lane deletion (R40).

Rotation now pins lane index%4 against band index%4 so slot 1 (tells) always draws
the full band and slot 3 (the large band) is always default; asserted over 200 waves.
Takes effect on the next drafter-python bounce; the running process still carries the
old args and the drafter SHELL is never stopped to ship a change.
2026-08-24 10:37:27 -06:00
Drew T d53c4bf796 fix(R42): gate_main self-reverts on abort; campaign refuses to adopt main sources
Closes the gap that made main red for nine hours. R42 ('commit a dirty tree rather than
revert it') is correct for a per-binary gate that leaves PROVEN banks uncommitted, and WRONG
for gate_main, whose substitution is unverified by construction until the SHA matches.

Two guards, defense in depth:
1. gate_main installs atexit + SIGTERM/SIGINT/SIGHUP handlers that revert its own substitution
   unless a bank actually succeeded. Killed mid-run, it now cleans up after itself.
2. ox_campaign's dirty-tree commit REFUSES top-level src/*.c (main's sources), reverting those
   and committing the rest. Verified: src/800c.c and src/800.c refused, src/ov_*/... and
   src/shared/engine_core.h still commit.

Also versions the autonomous lane scripts under tools/lanes/ — they lived only in gitignored
.run/, so a fresh clone had no drafter, gater, maintenance or stallguard at all.
2026-08-24 10:27:34 -06:00