mirror of
https://github.com/Druthulu/BFM-decomp
synced 2026-09-30 23:37:39 -04:00
188314f081
1. MAIN LANE — the largest single block of unfinished work was drawing 32 cards a wave. main_lane.draw() never passed --max-bins, so it inherited build_wave_atlas's default of 12 gate groups — a cap that exists because each group costs a whole-binary rebuild, and main's own --only-bins docstring says the opposite applies to it: "main is gated ONCE per SLATE, so main has no per-TU gate cost and --max-bins can be large". Nobody passed it. Measured cost: main banked ~19 stubs/hour against 1,291 remaining while the overlay lane ran 650-card waves beside it. Now --max-bins 400 (MAIN_MAX_BINS overrides), and the lane shell draws 600 cards with 600 workers instead of 200/150. 2. TWO LANES GATE, SO READ BOTH LOGS — a defect I introduced this session. The in-flight exclusion derived "this wave has been gated" from .run/gater.log only, but the main lane gates its own waves into .run/main_lane.log. Every m## wave therefore looked permanently in flight and main's draw lost 425 cards to an exclusion meant for work in progress. 3. TAIL_DONE_FRAC 0.80 -> 0.65. At 0.80 the fleet runs 2-3 overlapping waves at ~250 req/min; the residual troughs are the gap between one wave draining and the next ramping. 65% keeps 3-4 waves overlapping. Stragglers keep their full 700s grace in the finisher thread — this changes when the NEXT wave starts, never what lands. 4. ATOMIC ATLAS WRITE. The lanes read .run/atlas.json at every draw and atlas.py dumped straight onto it, leaving a truncated file readable for the length of the write. Now written to .tmp and os.replace'd. Context for 1-3: the atlas both lanes draw from is dated 08-23 01:13 — two days stale, predating ~4,600 banks — and its regen chain is running now (its own R32 assertion caught a stale family map first and named the fix).
38 lines
2.6 KiB
Bash
38 lines
2.6 KiB
Bash
#!/usr/bin/env bash
|
|
# THE MAIN LANE — the EXE's own drafting+gating cadence, beside the overlay lanes.
|
|
#
|
|
# main is excluded from every overlay wave draw because its gate is a clean whole-EXE rebuild that
|
|
# bisects on failure, and on the overlay critical path that cost three measured stalls. The result
|
|
# was that main had no lane at all: 1,713 open stubs and no cadence, while the overlay lane ran at
|
|
# roughly a quarter of the API ceiling because CARD SUPPLY is its constraint, not throughput.
|
|
#
|
|
# Safe to kill and restart, like the gater: gate_main holds .run/auto/gate.main.lock (blocking) and
|
|
# reverts its own substitution if it dies, so the worst a restart costs is one batch's gate time.
|
|
#
|
|
# It works the .run/main_queue backlog FIRST (drafts the overlay waves produced before main was
|
|
# excluded — already drafted, never gated, free), then draws and drafts its own main cards.
|
|
set -u
|
|
cd /home/musashi/bfm-decomp
|
|
# HTTP_TIMEOUT 700 (P31 S59) — MUST be raised WITH MAXTOK; they are one setting, not two.
|
|
# Measured on ox-alpha with a real MIPS derivation: 30 tok/s, and an UNCAPPED hard prompt ran
|
|
# 265 s for 8,067 completion tokens — i.e. it wanted to finish exactly where the old 8k cap cut it
|
|
# off, which is the nudge storm we were seeing. At 30 tok/s a full 16k generation needs ~530 s, so
|
|
# leaving the socket at 420 s would have killed the very turns the bigger budget exists to allow —
|
|
# and a timeout wastes the whole turn, where truncation at least leaves a partial.
|
|
# The ordering that must hold: generation < HTTP_TIMEOUT (700) < stallguard's wedged-agent kill
|
|
# (1200 s). 420 was itself a deliberate choice after 1800 parked a hung agent for THIRTY minutes;
|
|
# this keeps that concern (a hang costs <12 min) without strangling legitimate deep reasoning.
|
|
export HTTP_TIMEOUT=700
|
|
# STRAGGLER_GRACE = ONE FULL TURN (P31 S59). collect_drafts queues a wave once 95% of shards are
|
|
# done, then waits this long for the rest. At 120 s that was shorter than a single turn — at ~30
|
|
# tok/s a 16k generation runs ~530 s — so raising the token budget converted truncated turns into
|
|
# agents guillotined mid-thought with NO draft at all: overlay draft completion fell from 84-89%
|
|
# (8k) to 41% on wave bt and 69% on bu. A grace shorter than one turn guarantees that loss, so it
|
|
# is tied to HTTP_TIMEOUT rather than set independently.
|
|
export STRAGGLER_GRACE=700
|
|
while [ ! -e .run/ox_campaign.stop ]; do
|
|
.venv/bin/python tools/main_lane.py --workers 600 --batch 40 --cards 600 --max-ins 200 --maxtok 16000 2>&1
|
|
echo "[$(date +%H:%M:%S)] [main-lane] exited; restarting in 20s"
|
|
sleep 20
|
|
done
|