From f78ea55b3e67e6cf2f8ac4d7032930c726a3fc4a Mon Sep 17 00:00:00 2001 From: Drew T <50529377+Druthulu@users.noreply.github.com> Date: Tue, 25 Aug 2026 19:00:02 -0600 Subject: [PATCH] fix(gate): log what the gate actually does; --gate-jobs 24 -> 32; TAIL_DONE_FRAC 0.80 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit THE GATE WAS A BLACK BOX. sweep_parallel's stdout was captured and dropped, so a gate logged "reloc_identity -> gating 216" and then THIRTY MINUTES OF SILENCE before its bank line — no worker count, no per-binary progress, no phase-A/phase-B split. Gate times went 31 -> 37 -> 50 -> 67 min across ej/ek/en/eo with nothing to diagnose from, and I twice asserted things about phase B that the log could not support (its absence measured LOG CAPTURE, not behaviour). A lane that must run unattended has to leave evidence. MEASURED WHILE DIAGNOSING, and it rules out the obvious suspects: load average 2.6 on 32 cores with 1-3 concurrent builds during a gate — the gate is NOT CPU-bound and is not saturating its own -j 24. Raising to 32 is cheap given ~8% utilisation, but the real answer will come from the log this change adds. TAIL_DONE_FRAC 0.85 -> 0.80. 0.85 overcorrected: the fleet fell to 15 agents / 11 req/min because the drafter parks between waves while the gater drains a deep queue. 0.75 was too deep (26% 429s, draft completion sliding 94->91->73->47% across eq/er/es/et). Neither number is really the lever: the drafter cannot start a wave the gater has no room for, so the gate throughput is what bounds the campaign now. Generational tiering confirmed already correct: the top-off orders by generation at both assembly levels (group ranking and within-group) without FILTERING any tier out, so every generation stays eligible and the scarce never-drafted work simply goes first. --- tools/lanes/drafter.sh | 13 ++++++++++++- tools/lanes/gater.sh | 2 +- tools/ox_campaign.py | 8 ++++++++ 3 files changed, 21 insertions(+), 2 deletions(-) diff --git a/tools/lanes/drafter.sh b/tools/lanes/drafter.sh index e60f5b091e..20d005d425 100644 --- a/tools/lanes/drafter.sh +++ b/tools/lanes/drafter.sh @@ -104,7 +104,18 @@ export STRAGGLER_GRACE=700 # one wave draining and the next ramping. Handing off at 65% keeps 3-4 waves overlapping, so # the fleet is always carrying a full ramp somewhere. Stragglers still keep the full 700s # grace in the finisher thread — this changes WHEN THE NEXT WAVE STARTS, never what lands. -export TAIL_DONE_FRAC=0.75 +# TAIL_DONE_FRAC 0.75 -> 0.85 (P31 S60, Drew). Deeper overlap bought concurrency and then +# started spending it on retries: eight simultaneous waves means near-continuous ramping, and +# ramps are where throttling bites. Measured at 0.75 with 877 agents: 429s at 26% over the +# hour and draft completion sliding 94% -> 91% -> 73% -> 47% across eq/er/es/et, against +# 97-99% completion earlier today at sub-10% 429s. Half of et's cards were being spent for +# nothing. Fewer waves in flight, same ~500-card uncollapsed draws: trade peak req/min for +# the number that actually converts. +# 0.85 overcorrected: the fleet fell to 15 agents / 11 req/min because the drafter parks +# between waves while the gater drains a deep queue. 0.75 was too deep (26% 429s, completion +# sliding to 47%), 0.85 too shallow. 0.80 splits it — and the real lever is the GATE, not the +# overlap: the drafter cannot start a wave the gater has no room for. +export TAIL_DONE_FRAC=0.80 # MAX_BINS 160 -> 50 (P31 S60). Wave size was the right lever at 50% conversion (wave dd: # 217 banked of 422 gated in 39 min). It is dead weight at 5%: dq banked 8 of 167 gated and # took 98 MINUTES of gate to do it, while four drafted waves queued behind it and the free-ox diff --git a/tools/lanes/gater.sh b/tools/lanes/gater.sh index 00c6ea3adf..bb20b3fd51 100644 --- a/tools/lanes/gater.sh +++ b/tools/lanes/gater.sh @@ -8,7 +8,7 @@ while [ ! -e .run/ox_campaign.stop ]; do # idle. It matters now because the tells and jtbl QUOTAS deliberately pull cards from binaries # outside the ranked gate groups: a measured draw went from ~24 groups to 63, i.e. 63 rebuilds per # wave, and at 12 jobs that is five serial batches. Raise this with the quota sizes, not on its own. - .venv/bin/python tools/ox_campaign.py --gater --gate-jobs 24 2>&1 + .venv/bin/python tools/ox_campaign.py --gater --gate-jobs 32 2>&1 echo "[$(date +%H:%M:%S)] [gater] exited; restarting in 20s" sleep 20 done diff --git a/tools/ox_campaign.py b/tools/ox_campaign.py index 7c165f2d25..19fbd55e69 100644 --- a/tools/ox_campaign.py +++ b/tools/ox_campaign.py @@ -665,6 +665,14 @@ def gate(tag, keep, jobs, run_id=None): return 0, [] t0 = time.time() r = sh(f"{PY} tools/sweep_parallel.py --drafts {d} -j {jobs}", timeout=28800, quiet=False) + # LOG WHAT THE GATE ACTUALLY DID (P31 S60). sweep_parallel's stdout was captured and dropped, + # so a gate showed `reloc_identity -> gating 216` and then THIRTY MINUTES OF SILENCE before its + # bank line — no per-binary progress, no phase-A/phase-B split, no worker count. When gates went + # from 32 to 67 minutes there was nothing to diagnose from, and I twice drew conclusions about + # phase B that the log could not support. A lane that runs unattended must leave evidence. + for _ln in (r.stdout or "").splitlines(): + if _ln.strip() and not _ln.startswith(" ["): # skip the per-binary chatter + log(f" sweep: {_ln.strip()[:150]}") banked = [] for f in glob.glob(".run/auto/bulk/*.verified.txt"): if os.path.getmtime(f) >= t0: