mirror of
https://github.com/Druthulu/BFM-decomp
synced 2026-09-29 15:18:24 -04:00
188314f081
1. MAIN LANE — the largest single block of unfinished work was drawing 32 cards a wave. main_lane.draw() never passed --max-bins, so it inherited build_wave_atlas's default of 12 gate groups — a cap that exists because each group costs a whole-binary rebuild, and main's own --only-bins docstring says the opposite applies to it: "main is gated ONCE per SLATE, so main has no per-TU gate cost and --max-bins can be large". Nobody passed it. Measured cost: main banked ~19 stubs/hour against 1,291 remaining while the overlay lane ran 650-card waves beside it. Now --max-bins 400 (MAIN_MAX_BINS overrides), and the lane shell draws 600 cards with 600 workers instead of 200/150. 2. TWO LANES GATE, SO READ BOTH LOGS — a defect I introduced this session. The in-flight exclusion derived "this wave has been gated" from .run/gater.log only, but the main lane gates its own waves into .run/main_lane.log. Every m## wave therefore looked permanently in flight and main's draw lost 425 cards to an exclusion meant for work in progress. 3. TAIL_DONE_FRAC 0.80 -> 0.65. At 0.80 the fleet runs 2-3 overlapping waves at ~250 req/min; the residual troughs are the gap between one wave draining and the next ramping. 65% keeps 3-4 waves overlapping. Stragglers keep their full 700s grace in the finisher thread — this changes when the NEXT wave starts, never what lands. 4. ATOMIC ATLAS WRITE. The lanes read .run/atlas.json at every draw and atlas.py dumped straight onto it, leaving a truncated file readable for the length of the write. Now written to .tmp and os.replace'd. Context for 1-3: the atlas both lanes draw from is dated 08-23 01:13 — two days stale, predating ~4,600 banks — and its regen chain is running now (its own R32 assertion caught a stale family map first and named the fix).