Files
BFM-decomp/cookbook/C0509.md
T
2026-09-29 18:59:02 -06:00

2.6 KiB
Raw Blame History

C0509 — fleet-tool-parallelism-defaults

tags: legacy,memory · date: 2026-09-29 · phase/task: - · origin: legacy memory fleet-tool-parallelism-defaults.md

Drew (2026-08-07): "if these work successfully I want a new memory and defaults created so we forever use these speedup techniques." Measured on dedup_propagate (S46): 24 min → 11.4 min, same 29 functions, +62 MORE member instances (R22 213/213 both ways).

The four defaults for any fleet-wide tool in this repo:

  1. Return EVERY verdict a sweep already computed. gate_all byte-gated all 141 overlays and returned only the first failure, so the recovery loop paid a full sweep to rediscover each of the next 137. Batching them took convergence from ~138 rounds to 1–3. Same builds, same determinism.
  2. PROCESSES for CPU-bound work; threads ONLY for subprocess waits. A ThreadPoolExecutor over 138 "independent" searches kept 0–4 builds alive at load 3 — the work was regex over 15k-line files, so every thread queued on the GIL. The same code in a ProcessPoolExecutor: 14–29 builds, load 34.75 on 32 cores. Threads are right for the byte-gate (each is a subprocess.run), wrong for anything that parses or rewrites source.
  3. Longest-first scheduling. ex.map starts work in list order, so the giant overlays landing last left 31 cores watching one build for ~25 s of every 56 s sweep. Sort by source size descending; re-sort results into the caller's order so the verdict stays bit-identical.
  4. Per-item search beats lock-step sweeps when items are independent — and say why they are. Here: the shared header carries every macro regardless, so writing it once up front leaves each overlay owning only its own .c files and build/<bin>/. That cuts BUILDS, not just overlap.

Two traps this exposed, both worth checking in any tool about to run parallel:

  • A fixed temp path (.run/dpcc/t.c) is a correctness bug the day something runs concurrently — the same fake-isolation class as match_one's shared --work dir. Make it per-call.
  • A pool submitted all at once shares no learning. Every worker got an empty suspect list and paid a full bisection. Seed with one item in-process first, then fan out with the result.

Prove it, don't assume it: the acceptance test was a regression, not a stopwatch — revert to the pre-run state, re-run the identical command, and require the same functions and a byte-verified fleet (R22). It came back faster AND with 62 more instances banked, which is how the over-exclusion in the old path was found at all. See decomp-accelerator-ledger (A8) and docs/accelerators.md.