Files
BFM-decomp/tools/lanes/drafter.sh
T
Drew T f78ea55b3e fix(gate): log what the gate actually does; --gate-jobs 24 -> 32; TAIL_DONE_FRAC 0.80
THE GATE WAS A BLACK BOX. sweep_parallel's stdout was captured and dropped, so a gate logged
"reloc_identity -> gating 216" and then THIRTY MINUTES OF SILENCE before its bank line — no
worker count, no per-binary progress, no phase-A/phase-B split. Gate times went 31 -> 37 ->
50 -> 67 min across ej/ek/en/eo with nothing to diagnose from, and I twice asserted things
about phase B that the log could not support (its absence measured LOG CAPTURE, not
behaviour). A lane that must run unattended has to leave evidence.

MEASURED WHILE DIAGNOSING, and it rules out the obvious suspects: load average 2.6 on 32
cores with 1-3 concurrent builds during a gate — the gate is NOT CPU-bound and is not
saturating its own -j 24. Raising to 32 is cheap given ~8% utilisation, but the real answer
will come from the log this change adds.

TAIL_DONE_FRAC 0.85 -> 0.80. 0.85 overcorrected: the fleet fell to 15 agents / 11 req/min
because the drafter parks between waves while the gater drains a deep queue. 0.75 was too
deep (26% 429s, draft completion sliding 94->91->73->47% across eq/er/es/et). Neither number
is really the lever: the drafter cannot start a wave the gater has no room for, so the gate
throughput is what bounds the campaign now.

Generational tiering confirmed already correct: the top-off orders by generation at both
assembly levels (group ranking and within-group) without FILTERING any tier out, so every
generation stays eligible and the scarce never-drafted work simply goes first.
2026-08-25 19:00:02 -06:00

147 lines
11 KiB
Bash

#!/usr/bin/env bash
# THE LANE THAT MUST NEVER STOP. Drafting is the only thing gated by the free-ox clock; it touches
# nothing but .run/, so no code change to the gater, the tooling or the rules ever needs to stop it.
# Measured 2026-08-23: 139 of 162 idle minutes were this lane being killed to ship a fix.
#
# 1720 ox = the 10x Drew asked for. Both caps I previously set were measurement artifacts:
# * "ox saturates at 9.8% 429s" — 961 of 964 429s landed in the FIRST 5-MINUTE BUCKET, the
# thundering herd of 818 shards starting at once. Every later bucket was 0.0%. Fixed by
# staggering shard startup, not by capping concurrency.
# * "39 MB per agent" — a STARTUP snapshot with the card file freshly loaded. Steady state is
# ~10 MB, so 45 GB carries thousands, not hundreds.
# TELLS: RESTORED, BUT PINNED TO THE FULL BAND (P31 S59, 2026-08-24). S58 removed the lane on four
# waves — as 5/9 gated of 57 · aw 6/12 of 60 · az 4/10 of 55 · bd 3/9 of 56 — and attributed the
# failure to the LANE. The campaign ledger says it was the BAND: every one of those four ran at
# 120-2000, and the whole tells population splits cleanly by band:
# tells @ 120-2000 (as/aw/az/bd): 228 drafts -> 18 banked = 7.9%
# tells @ full band (ao/au/bb/bg): 655 drafts -> 161 banked = 24.6%
# default @ full band : 2,996 drafts -> 1,335 banked = 44.6%
# So tells is ~2x worse per draft than default, not dead, and it is the only lane that touches
# 1,040 members / 86,602 instructions. Of what actually reaches a gate the two lanes are the SAME
# (tells 54.4% of gated, default 56.4%) — the entire loss is reloc_identity discarding drafts that
# name symbols the target .s never references, i.e. cookbook §235 (the phantom symbol), which is a
# BRIEF fix, not a lane deletion. R40: exonerate the instrument before blaming the subject.
#
# S59 REFINEMENT, from the same ledger joined to the wave cards and to the BANKED FUNCTIONS the
# wave's own commit names (bank rate by size, cards->banked, pooled over waves bb/bg vs bc/bf):
# nins default tells
# 0-50 303/528 57% 27/ 67 40%
# 50-80 43/145 30% 20/ 73 27%
# 80-120 9/ 41 22% 10/100 10%
# 120-200 1/ 30 3% 1/ 68 1%
# 200+ 2/ 35 6% 0/ 30 0%
# At EQUAL SIZE the two lanes are close below 80 instructions and both collapse above it. What
# actually separated them is the card SIZE MIX: default's cards are median 37-39 ins, the tells
# pool is median 89-95 — 2.4x larger — so "the tells lane is broken" was measuring the population,
# not the lever. Tells therefore draws a SMALL band (5-80), where its yield is within a few points
# of default's; widen it only when that stratum is worked out.
#
# The rotation pins the arithmetic: lane = index%4, band = index%4, so slot 1 (tells) always draws
# the small band and slot 3 (the large band) is always default. Changing the length of either list
# breaks that alignment — change both together.
#
# NOTE: do NOT assume aprop_autodraft is the answer for tells. Its input population overlaps the
# 1,040 tells member functions by only 44 (4.2%) — checked, after asserting the opposite three
# times from the failure signature alone.
#
# The remaining real constraint is CARD SUPPLY: a fleet is only as busy as the wave is large.
#
# BANDS ARE NOW MOSTLY FULL-RANGE. Narrow bands were right when each held thousands of
# candidates; they now FRAGMENT a shrinking pool — measured P31 S58: the 400-2000 band drew NINE
# cards for a 2,000-worker fleet (0.45% utilisation) because that band has 37 groups total and
# most are banked. One targeted 120-2000 slot is kept so large functions still get drawn
# deliberately; the rest draw from everything.
# CREDIT STOP, 2026-08-24 ~14:15 (P31 S59). The paid deepseek lane was burning the OpenRouter
# balance at ~$1.43/h over the last three waves ($4.56 -> $2.56 between 11:54 and 13:18) and the
# balance was ~23 minutes from `--credit-floor 2.0`, which does not pause the paid lane — it BREAKS
# the whole drafting loop, and the shell then restarts a python that breaks again. ox-alpha is free
# for the rest of this window, so drafting continues on ox alone at zero burn; the floor drops to
# 0.25 because with a free model the balance is no longer a proxy for "can we draft".
#
# TO RESTORE the second provider pool after a top-up: put the deepseek lane back in --models and
# raise --credit-floor to 2.0. Its value is a pool with an independent 429 ceiling (it has never
# returned one), not throughput: it was 280 of 2,000 workers.
# ONE LANE, ONE BAND (P31 S59, superseding the rotation above). Two measurements collapsed it:
#
# * THE TELLS SLOT IS NOW A QUOTA. build_wave_atlas reserves 60 tell-lever cards (<=80 ins) inside
# every ordinary wave. A dedicated tells WAVE draws only 70-87 cards — a whole 40-minute slot at
# a quarter of a default wave — and with the quota in place the slot is also redundant.
# * THE LARGE BAND IS THE WORST WAVE WE RUN. Bank rate by size, measured: 57% under 50 ins, 30% at
# 50-80, 22% at 80-120, 3% at 120-200, 6% above. The 120-2000 slot drew 69 cards at ~4% — about
# 3 banks for a 40-minute slot, against ~150 for a full-band wave. Big functions are still drawn:
# the full band includes them, and the draw takes mass-first inside each gate group.
#
# MAXTOK 16000 (P31 S59, measured). The output cap was 8,000 and the shard logs say the model was
# hitting it BEFORE emitting a tool call: in wave bk, 240 of 244 turn-finishes were
# `no tool call (finish=length) — NUDGE n/6`. A turn that ends in a nudge did no work at all, and an
# agent gets six of them before it gives up, which is why the mean is 2.8 oracle calls per MATCH on
# a 24-turn budget: the turns are being spent on truncation, not on iteration. ox is free, so a
# larger output budget costs latency and nothing else.
set -u
cd /home/musashi/bfm-decomp
# HTTP_TIMEOUT 700 (P31 S59) — MUST be raised WITH MAXTOK; they are one setting, not two.
# Measured on ox-alpha with a real MIPS derivation: 30 tok/s, and an UNCAPPED hard prompt ran
# 265 s for 8,067 completion tokens — i.e. it wanted to finish exactly where the old 8k cap cut it
# off, which is the nudge storm we were seeing. At 30 tok/s a full 16k generation needs ~530 s, so
# leaving the socket at 420 s would have killed the very turns the bigger budget exists to allow —
# and a timeout wastes the whole turn, where truncation at least leaves a partial.
# The ordering that must hold: generation < HTTP_TIMEOUT (700) < stallguard's wedged-agent kill
# (1200 s). 420 was itself a deliberate choice after 1800 parked a hung agent for THIRTY minutes;
# this keeps that concern (a hang costs <12 min) without strangling legitimate deep reasoning.
export HTTP_TIMEOUT=700
# STRAGGLER_GRACE = ONE FULL TURN (P31 S59). collect_drafts queues a wave once 95% of shards are
# done, then waits this long for the rest. At 120 s that was shorter than a single turn — at ~30
# tok/s a 16k generation runs ~530 s — so raising the token budget converted truncated turns into
# agents guillotined mid-thought with NO draft at all: overlay draft completion fell from 84-89%
# (8k) to 41% on wave bt and 69% on bu. A grace shorter than one turn guarantees that loss, so it
# is tied to HTTP_TIMEOUT rather than set independently.
# ...and since P31 S60 the grace no longer COSTS anything: wait_for_tail() hands the tail to a
# finisher thread, so the next wave draws and ramps while the stragglers run. Blocking on this
# grace was 11m40s of 2-5% fleet utilisation at the end of every wave, four waves for four.
export STRAGGLER_GRACE=700
# TAIL_DONE_FRAC 0.80 -> 0.65 (P31 S60, Drew: "get the throughput up"). At 0.80 the fleet ran
# 2-3 overlapping waves and averaged ~250 req/min; the remaining troughs are the gap between
# one wave draining and the next ramping. Handing off at 65% keeps 3-4 waves overlapping, so
# the fleet is always carrying a full ramp somewhere. Stragglers still keep the full 700s
# grace in the finisher thread — this changes WHEN THE NEXT WAVE STARTS, never what lands.
# TAIL_DONE_FRAC 0.75 -> 0.85 (P31 S60, Drew). Deeper overlap bought concurrency and then
# started spending it on retries: eight simultaneous waves means near-continuous ramping, and
# ramps are where throttling bites. Measured at 0.75 with 877 agents: 429s at 26% over the
# hour and draft completion sliding 94% -> 91% -> 73% -> 47% across eq/er/es/et, against
# 97-99% completion earlier today at sub-10% 429s. Half of et's cards were being spent for
# nothing. Fewer waves in flight, same ~500-card uncollapsed draws: trade peak req/min for
# the number that actually converts.
# 0.85 overcorrected: the fleet fell to 15 agents / 11 req/min because the drafter parks
# between waves while the gater drains a deep queue. 0.75 was too deep (26% 429s, completion
# sliding to 47%), 0.85 too shallow. 0.80 splits it — and the real lever is the GATE, not the
# overlap: the drafter cannot start a wave the gater has no room for.
export TAIL_DONE_FRAC=0.80
# MAX_BINS 160 -> 50 (P31 S60). Wave size was the right lever at 50% conversion (wave dd:
# 217 banked of 422 gated in 39 min). It is dead weight at 5%: dq banked 8 of 167 gated and
# took 98 MINUTES of gate to do it, while four drafted waves queued behind it and the free-ox
# fleet sat at 14 agents / 10 req/min. At this conversion a 150-card wave banks what a
# 430-card wave banks, in a third of the gate. Raise it again when the tail starts converting.
# ONE_PER_GID=0 + MAX_BINS 250 (P31 S60, Drew: "draft all remaining funcs, sib family be damned").
# The sibling collapse was priced when agent tokens were scarce: a same-gid sibling banks by
# mechanical remap once its exemplar cracks, so drafting it pays for what the remap does free. On a
# free ox window that reasoning inverts — and the collapse is what makes 3,271 open crackable
# functions look like 334 drawable skeletons, 92% of which are gen6+ walls whose exemplars have
# already refused six waves. A sibling drafted directly can crack on its OWN terms.
# Measured on a live draw: 627 cards / 44,403 ins / 215 gate groups uncollapsed, vs 334 / 30,926 /
# 127 collapsed — and the gate got MORE efficient per build, 2.9 drafts per rebuild vs 2.6, because
# siblings land in binaries the wave already touches and the gate is per (binary, TU), chunked.
# MAX_BINS goes back up because the whole pool is now the target; revert both if the gate backs up.
export ONE_PER_GID=0
export MAX_BINS=250
export MAX_429=10
while [ ! -e .run/ox_campaign.stop ]; do
.venv/bin/python tools/ox_campaign.py --drafter \
--workers 2000 \
--models 'stealth/ox-alpha:2000' \
--bands '5-2000' \
--maxtok 16000 \
--cards-per-wave 3000 --queue-depth 4 --credit-floor 0.25 2>&1
echo "[$(date +%H:%M:%S)] [drafter] exited; restarting in 20s"
sleep 20
done