perf(throughput): unblock the main lane, both gate logs, deeper wave overlap, atomic atlas

1. MAIN LANE — the largest single block of unfinished work was drawing 32 cards a wave.
   main_lane.draw() never passed --max-bins, so it inherited build_wave_atlas's default of
   12 gate groups — a cap that exists because each group costs a whole-binary rebuild, and
   main's own --only-bins docstring says the opposite applies to it: "main is gated ONCE per
   SLATE, so main has no per-TU gate cost and --max-bins can be large". Nobody passed it.
   Measured cost: main banked ~19 stubs/hour against 1,291 remaining while the overlay lane
   ran 650-card waves beside it. Now --max-bins 400 (MAIN_MAX_BINS overrides), and the lane
   shell draws 600 cards with 600 workers instead of 200/150.

2. TWO LANES GATE, SO READ BOTH LOGS — a defect I introduced this session. The in-flight
   exclusion derived "this wave has been gated" from .run/gater.log only, but the main lane
   gates its own waves into .run/main_lane.log. Every m## wave therefore looked permanently
   in flight and main's draw lost 425 cards to an exclusion meant for work in progress.

3. TAIL_DONE_FRAC 0.80 -> 0.65. At 0.80 the fleet runs 2-3 overlapping waves at ~250
   req/min; the residual troughs are the gap between one wave draining and the next ramping.
   65% keeps 3-4 waves overlapping. Stragglers keep their full 700s grace in the finisher
   thread — this changes when the NEXT wave starts, never what lands.

4. ATOMIC ATLAS WRITE. The lanes read .run/atlas.json at every draw and atlas.py dumped
   straight onto it, leaving a truncated file readable for the length of the write. Now
   written to .tmp and os.replace'd.

Context for 1-3: the atlas both lanes draw from is dated 08-23 01:13 — two days stale,
predating ~4,600 banks — and its regen chain is running now (its own R32 assertion caught a
stale family map first and named the fix).
This commit is contained in:
Drew T
2026-08-25 11:28:51 -06:00
parent 7e1f554a05
commit 188314f081
5 changed files with 33 additions and 9 deletions
+4 -1
View File
@@ -742,7 +742,10 @@ def survey(args):
"groups": groups,
"knn": knn,
}
json.dump(out, open(".run/atlas.json", "w"))
# ATOMIC (P31 S60): the lanes read .run/atlas.json at every draw, and a plain dump leaves a
# truncated file readable for the length of the write. Rename is atomic on the same fs.
json.dump(out, open(".run/atlas.json.tmp", "w"))
os.replace(".run/atlas.json.tmp", ".run/atlas.json")
render(out)
print(f"atlas: {len(groups)} groups / {len(inst)} instances / "
f"{sum(g['ins'] for g in groups)} ins; t1.5 merges={t15_merges} warm merges={warm_merges}; "
+14 -7
View File
@@ -219,14 +219,21 @@ if a.retry_unbanked:
# flight and stay excluded; a tag whose card file is older than STALE_H with no gate line never
# got one (a killed wave) and is released, so nothing is locked out forever.
STALE_H = 6
# TWO LANES GATE, SO READ BOTH LOGS (P31 S60, caught by its own probe). The overlay gater
# writes "GATE <tag>: banked N/M"; the MAIN lane gates its own waves and writes
# "[main-lane] <tag><batch>: N banked" to a different file. Reading only the gater's log made
# every m## wave look permanently in flight, and main's draw — the lane with 1,291 stubs left —
# lost 425 cards to an exclusion meant for work in progress.
gated_tags = set()
try:
for line in open('.run/gater.log', errors='replace'):
m = re.search(r'GATE (\w+): banked ', line)
if m:
gated_tags.add(m.group(1))
except OSError:
pass
for path, pat in (('.run/gater.log', r'GATE (\w+): banked '),
('.run/main_lane.log', r'\[main-lane\] (m\d+)[a-z]*:.*\bbanked\b')):
try:
for line in open(path, errors='replace'):
m = re.search(pat, line)
if m:
gated_tags.add(m.group(1))
except OSError:
pass
_inflight = 0
_still_open = set()
for p in PRIORS:
+6
View File
@@ -99,6 +99,12 @@ export HTTP_TIMEOUT=700
# finisher thread, so the next wave draws and ramps while the stragglers run. Blocking on this
# grace was 11m40s of 2-5% fleet utilisation at the end of every wave, four waves for four.
export STRAGGLER_GRACE=700
# TAIL_DONE_FRAC 0.80 -> 0.65 (P31 S60, Drew: "get the throughput up"). At 0.80 the fleet ran
# 2-3 overlapping waves and averaged ~250 req/min; the remaining troughs are the gap between
# one wave draining and the next ramping. Handing off at 65% keeps 3-4 waves overlapping, so
# the fleet is always carrying a full ramp somewhere. Stragglers still keep the full 700s
# grace in the finisher thread — this changes WHEN THE NEXT WAVE STARTS, never what lands.
export TAIL_DONE_FRAC=0.65
export MAX_429=10
while [ ! -e .run/ox_campaign.stop ]; do
.venv/bin/python tools/ox_campaign.py --drafter \
+1 -1
View File
@@ -31,7 +31,7 @@ export HTTP_TIMEOUT=700
# is tied to HTTP_TIMEOUT rather than set independently.
export STRAGGLER_GRACE=700
while [ ! -e .run/ox_campaign.stop ]; do
.venv/bin/python tools/main_lane.py --workers 150 --batch 40 --cards 200 --max-ins 200 --maxtok 16000 2>&1
.venv/bin/python tools/main_lane.py --workers 600 --batch 40 --cards 600 --max-ins 200 --maxtok 16000 2>&1
echo "[$(date +%H:%M:%S)] [main-lane] exited; restarting in 20s"
sleep 20
done
+8
View File
@@ -113,7 +113,15 @@ def draw(tag, n, lo, hi):
if os.path.exists(cards):
log(f"{tag}: cards already drawn, reusing")
return cards
# --max-bins MUST be passed here (P31 S60). build_wave_atlas defaults to 12 gate groups, a cap
# that exists because each group costs one whole-binary rebuild — and main's own --only-bins
# docstring says the opposite applies to it: "main is gated ONCE per SLATE, so main has no
# per-TU gate cost and --max-bins can be large." Nobody passed it, so this lane asked for 200
# cards and drew 32 (12 of main's TUs), every wave. Measured cost: main banked ~19 stubs/hour
# against 1,291 remaining while the overlay lane ran 650-card waves beside it.
bins = os.environ.get("MAIN_MAX_BINS", "400")
r = OX.sh(f"{PY} tools/build_wave_atlas.py {cards} {n} --min-ins {lo} --max-ins {hi} "
f"--max-bins {bins} "
f"--only-bins main --one-per-gid --retry-unbanked --tells-quota 0 --jtbl-quota 0",
timeout=3600, quiet=False)
if not os.path.exists(cards):