mirror of
https://github.com/Druthulu/BFM-decomp
synced 2026-10-03 08:07:25 -04:00
f78ea55b3e
THE GATE WAS A BLACK BOX. sweep_parallel's stdout was captured and dropped, so a gate logged "reloc_identity -> gating 216" and then THIRTY MINUTES OF SILENCE before its bank line — no worker count, no per-binary progress, no phase-A/phase-B split. Gate times went 31 -> 37 -> 50 -> 67 min across ej/ek/en/eo with nothing to diagnose from, and I twice asserted things about phase B that the log could not support (its absence measured LOG CAPTURE, not behaviour). A lane that must run unattended has to leave evidence. MEASURED WHILE DIAGNOSING, and it rules out the obvious suspects: load average 2.6 on 32 cores with 1-3 concurrent builds during a gate — the gate is NOT CPU-bound and is not saturating its own -j 24. Raising to 32 is cheap given ~8% utilisation, but the real answer will come from the log this change adds. TAIL_DONE_FRAC 0.85 -> 0.80. 0.85 overcorrected: the fleet fell to 15 agents / 11 req/min because the drafter parks between waves while the gater drains a deep queue. 0.75 was too deep (26% 429s, draft completion sliding 94->91->73->47% across eq/er/es/et). Neither number is really the lever: the drafter cannot start a wave the gater has no room for, so the gate throughput is what bounds the campaign now. Generational tiering confirmed already correct: the top-off orders by generation at both assembly levels (group ranking and within-group) without FILTERING any tier out, so every generation stays eligible and the scarce never-drafted work simply goes first.
147 lines
11 KiB
Bash
147 lines
11 KiB
Bash
#!/usr/bin/env bash
|
|
# THE LANE THAT MUST NEVER STOP. Drafting is the only thing gated by the free-ox clock; it touches
|
|
# nothing but .run/, so no code change to the gater, the tooling or the rules ever needs to stop it.
|
|
# Measured 2026-08-23: 139 of 162 idle minutes were this lane being killed to ship a fix.
|
|
#
|
|
# 1720 ox = the 10x Drew asked for. Both caps I previously set were measurement artifacts:
|
|
# * "ox saturates at 9.8% 429s" — 961 of 964 429s landed in the FIRST 5-MINUTE BUCKET, the
|
|
# thundering herd of 818 shards starting at once. Every later bucket was 0.0%. Fixed by
|
|
# staggering shard startup, not by capping concurrency.
|
|
# * "39 MB per agent" — a STARTUP snapshot with the card file freshly loaded. Steady state is
|
|
# ~10 MB, so 45 GB carries thousands, not hundreds.
|
|
# TELLS: RESTORED, BUT PINNED TO THE FULL BAND (P31 S59, 2026-08-24). S58 removed the lane on four
|
|
# waves — as 5/9 gated of 57 · aw 6/12 of 60 · az 4/10 of 55 · bd 3/9 of 56 — and attributed the
|
|
# failure to the LANE. The campaign ledger says it was the BAND: every one of those four ran at
|
|
# 120-2000, and the whole tells population splits cleanly by band:
|
|
# tells @ 120-2000 (as/aw/az/bd): 228 drafts -> 18 banked = 7.9%
|
|
# tells @ full band (ao/au/bb/bg): 655 drafts -> 161 banked = 24.6%
|
|
# default @ full band : 2,996 drafts -> 1,335 banked = 44.6%
|
|
# So tells is ~2x worse per draft than default, not dead, and it is the only lane that touches
|
|
# 1,040 members / 86,602 instructions. Of what actually reaches a gate the two lanes are the SAME
|
|
# (tells 54.4% of gated, default 56.4%) — the entire loss is reloc_identity discarding drafts that
|
|
# name symbols the target .s never references, i.e. cookbook §235 (the phantom symbol), which is a
|
|
# BRIEF fix, not a lane deletion. R40: exonerate the instrument before blaming the subject.
|
|
#
|
|
# S59 REFINEMENT, from the same ledger joined to the wave cards and to the BANKED FUNCTIONS the
|
|
# wave's own commit names (bank rate by size, cards->banked, pooled over waves bb/bg vs bc/bf):
|
|
# nins default tells
|
|
# 0-50 303/528 57% 27/ 67 40%
|
|
# 50-80 43/145 30% 20/ 73 27%
|
|
# 80-120 9/ 41 22% 10/100 10%
|
|
# 120-200 1/ 30 3% 1/ 68 1%
|
|
# 200+ 2/ 35 6% 0/ 30 0%
|
|
# At EQUAL SIZE the two lanes are close below 80 instructions and both collapse above it. What
|
|
# actually separated them is the card SIZE MIX: default's cards are median 37-39 ins, the tells
|
|
# pool is median 89-95 — 2.4x larger — so "the tells lane is broken" was measuring the population,
|
|
# not the lever. Tells therefore draws a SMALL band (5-80), where its yield is within a few points
|
|
# of default's; widen it only when that stratum is worked out.
|
|
#
|
|
# The rotation pins the arithmetic: lane = index%4, band = index%4, so slot 1 (tells) always draws
|
|
# the small band and slot 3 (the large band) is always default. Changing the length of either list
|
|
# breaks that alignment — change both together.
|
|
#
|
|
# NOTE: do NOT assume aprop_autodraft is the answer for tells. Its input population overlaps the
|
|
# 1,040 tells member functions by only 44 (4.2%) — checked, after asserting the opposite three
|
|
# times from the failure signature alone.
|
|
#
|
|
# The remaining real constraint is CARD SUPPLY: a fleet is only as busy as the wave is large.
|
|
#
|
|
# BANDS ARE NOW MOSTLY FULL-RANGE. Narrow bands were right when each held thousands of
|
|
# candidates; they now FRAGMENT a shrinking pool — measured P31 S58: the 400-2000 band drew NINE
|
|
# cards for a 2,000-worker fleet (0.45% utilisation) because that band has 37 groups total and
|
|
# most are banked. One targeted 120-2000 slot is kept so large functions still get drawn
|
|
# deliberately; the rest draw from everything.
|
|
# CREDIT STOP, 2026-08-24 ~14:15 (P31 S59). The paid deepseek lane was burning the OpenRouter
|
|
# balance at ~$1.43/h over the last three waves ($4.56 -> $2.56 between 11:54 and 13:18) and the
|
|
# balance was ~23 minutes from `--credit-floor 2.0`, which does not pause the paid lane — it BREAKS
|
|
# the whole drafting loop, and the shell then restarts a python that breaks again. ox-alpha is free
|
|
# for the rest of this window, so drafting continues on ox alone at zero burn; the floor drops to
|
|
# 0.25 because with a free model the balance is no longer a proxy for "can we draft".
|
|
#
|
|
# TO RESTORE the second provider pool after a top-up: put the deepseek lane back in --models and
|
|
# raise --credit-floor to 2.0. Its value is a pool with an independent 429 ceiling (it has never
|
|
# returned one), not throughput: it was 280 of 2,000 workers.
|
|
# ONE LANE, ONE BAND (P31 S59, superseding the rotation above). Two measurements collapsed it:
|
|
#
|
|
# * THE TELLS SLOT IS NOW A QUOTA. build_wave_atlas reserves 60 tell-lever cards (<=80 ins) inside
|
|
# every ordinary wave. A dedicated tells WAVE draws only 70-87 cards — a whole 40-minute slot at
|
|
# a quarter of a default wave — and with the quota in place the slot is also redundant.
|
|
# * THE LARGE BAND IS THE WORST WAVE WE RUN. Bank rate by size, measured: 57% under 50 ins, 30% at
|
|
# 50-80, 22% at 80-120, 3% at 120-200, 6% above. The 120-2000 slot drew 69 cards at ~4% — about
|
|
# 3 banks for a 40-minute slot, against ~150 for a full-band wave. Big functions are still drawn:
|
|
# the full band includes them, and the draw takes mass-first inside each gate group.
|
|
#
|
|
# MAXTOK 16000 (P31 S59, measured). The output cap was 8,000 and the shard logs say the model was
|
|
# hitting it BEFORE emitting a tool call: in wave bk, 240 of 244 turn-finishes were
|
|
# `no tool call (finish=length) — NUDGE n/6`. A turn that ends in a nudge did no work at all, and an
|
|
# agent gets six of them before it gives up, which is why the mean is 2.8 oracle calls per MATCH on
|
|
# a 24-turn budget: the turns are being spent on truncation, not on iteration. ox is free, so a
|
|
# larger output budget costs latency and nothing else.
|
|
set -u
|
|
cd /home/musashi/bfm-decomp
|
|
# HTTP_TIMEOUT 700 (P31 S59) — MUST be raised WITH MAXTOK; they are one setting, not two.
|
|
# Measured on ox-alpha with a real MIPS derivation: 30 tok/s, and an UNCAPPED hard prompt ran
|
|
# 265 s for 8,067 completion tokens — i.e. it wanted to finish exactly where the old 8k cap cut it
|
|
# off, which is the nudge storm we were seeing. At 30 tok/s a full 16k generation needs ~530 s, so
|
|
# leaving the socket at 420 s would have killed the very turns the bigger budget exists to allow —
|
|
# and a timeout wastes the whole turn, where truncation at least leaves a partial.
|
|
# The ordering that must hold: generation < HTTP_TIMEOUT (700) < stallguard's wedged-agent kill
|
|
# (1200 s). 420 was itself a deliberate choice after 1800 parked a hung agent for THIRTY minutes;
|
|
# this keeps that concern (a hang costs <12 min) without strangling legitimate deep reasoning.
|
|
export HTTP_TIMEOUT=700
|
|
# STRAGGLER_GRACE = ONE FULL TURN (P31 S59). collect_drafts queues a wave once 95% of shards are
|
|
# done, then waits this long for the rest. At 120 s that was shorter than a single turn — at ~30
|
|
# tok/s a 16k generation runs ~530 s — so raising the token budget converted truncated turns into
|
|
# agents guillotined mid-thought with NO draft at all: overlay draft completion fell from 84-89%
|
|
# (8k) to 41% on wave bt and 69% on bu. A grace shorter than one turn guarantees that loss, so it
|
|
# is tied to HTTP_TIMEOUT rather than set independently.
|
|
# ...and since P31 S60 the grace no longer COSTS anything: wait_for_tail() hands the tail to a
|
|
# finisher thread, so the next wave draws and ramps while the stragglers run. Blocking on this
|
|
# grace was 11m40s of 2-5% fleet utilisation at the end of every wave, four waves for four.
|
|
export STRAGGLER_GRACE=700
|
|
# TAIL_DONE_FRAC 0.80 -> 0.65 (P31 S60, Drew: "get the throughput up"). At 0.80 the fleet ran
|
|
# 2-3 overlapping waves and averaged ~250 req/min; the remaining troughs are the gap between
|
|
# one wave draining and the next ramping. Handing off at 65% keeps 3-4 waves overlapping, so
|
|
# the fleet is always carrying a full ramp somewhere. Stragglers still keep the full 700s
|
|
# grace in the finisher thread — this changes WHEN THE NEXT WAVE STARTS, never what lands.
|
|
# TAIL_DONE_FRAC 0.75 -> 0.85 (P31 S60, Drew). Deeper overlap bought concurrency and then
|
|
# started spending it on retries: eight simultaneous waves means near-continuous ramping, and
|
|
# ramps are where throttling bites. Measured at 0.75 with 877 agents: 429s at 26% over the
|
|
# hour and draft completion sliding 94% -> 91% -> 73% -> 47% across eq/er/es/et, against
|
|
# 97-99% completion earlier today at sub-10% 429s. Half of et's cards were being spent for
|
|
# nothing. Fewer waves in flight, same ~500-card uncollapsed draws: trade peak req/min for
|
|
# the number that actually converts.
|
|
# 0.85 overcorrected: the fleet fell to 15 agents / 11 req/min because the drafter parks
|
|
# between waves while the gater drains a deep queue. 0.75 was too deep (26% 429s, completion
|
|
# sliding to 47%), 0.85 too shallow. 0.80 splits it — and the real lever is the GATE, not the
|
|
# overlap: the drafter cannot start a wave the gater has no room for.
|
|
export TAIL_DONE_FRAC=0.80
|
|
# MAX_BINS 160 -> 50 (P31 S60). Wave size was the right lever at 50% conversion (wave dd:
|
|
# 217 banked of 422 gated in 39 min). It is dead weight at 5%: dq banked 8 of 167 gated and
|
|
# took 98 MINUTES of gate to do it, while four drafted waves queued behind it and the free-ox
|
|
# fleet sat at 14 agents / 10 req/min. At this conversion a 150-card wave banks what a
|
|
# 430-card wave banks, in a third of the gate. Raise it again when the tail starts converting.
|
|
# ONE_PER_GID=0 + MAX_BINS 250 (P31 S60, Drew: "draft all remaining funcs, sib family be damned").
|
|
# The sibling collapse was priced when agent tokens were scarce: a same-gid sibling banks by
|
|
# mechanical remap once its exemplar cracks, so drafting it pays for what the remap does free. On a
|
|
# free ox window that reasoning inverts — and the collapse is what makes 3,271 open crackable
|
|
# functions look like 334 drawable skeletons, 92% of which are gen6+ walls whose exemplars have
|
|
# already refused six waves. A sibling drafted directly can crack on its OWN terms.
|
|
# Measured on a live draw: 627 cards / 44,403 ins / 215 gate groups uncollapsed, vs 334 / 30,926 /
|
|
# 127 collapsed — and the gate got MORE efficient per build, 2.9 drafts per rebuild vs 2.6, because
|
|
# siblings land in binaries the wave already touches and the gate is per (binary, TU), chunked.
|
|
# MAX_BINS goes back up because the whole pool is now the target; revert both if the gate backs up.
|
|
export ONE_PER_GID=0
|
|
export MAX_BINS=250
|
|
export MAX_429=10
|
|
while [ ! -e .run/ox_campaign.stop ]; do
|
|
.venv/bin/python tools/ox_campaign.py --drafter \
|
|
--workers 2000 \
|
|
--models 'stealth/ox-alpha:2000' \
|
|
--bands '5-2000' \
|
|
--maxtok 16000 \
|
|
--cards-per-wave 3000 --queue-depth 4 --credit-floor 0.25 2>&1
|
|
echo "[$(date +%H:%M:%S)] [drafter] exited; restarting in 20s"
|
|
sleep 20
|
|
done
|