Files
BFM-decomp/cookbook/C0194.md
T

3.7 KiB
Raw Blame History

§176j — STOPPING A WAVE MID-FLIGHT COSTS THE IN-FLIGHT TAIL (and how much is recoverable)

Wave Q was stopped early to save tokens. Measured consequence: 51/90 drafts verified MATCH (3,631 of 6,249 ins) against the 96–97% the same pipeline produced when allowed to finish.

But the loss is suspended, not destroyed — every draft persists on disk, and the stopped agents' partial work is closer than it looks:

closeness fns ins
≤10 15 833
11–30 10 729
31–60 10 699
>60 4 357

Do NOT resume the workflow to recover this. resumeFromRunId replays cached agents and re-runs the unfinished ones from scratch with the original prompt — full cost, no memory of their partial work. The cheap path is a repair-only pass: feed the existing draft plus its diff to the repair prompt (which is written to start from a draft, not from the .s), scoped to the ≤30-closeness band. Most of those need a statement moved, not a decompile.

The decision rule worth keeping: before killing a long agent run, price the tail. If the median in-flight draft is near-matching, the tokens are already spent and stopping converts them from "nearly banked" into "needs a second, cheaper pass" — which is fine, but it is a deferral, not a saving.

§176j-2 — THE REPAIR PASS, MEASURED (do this instead of resuming)

Wave Q's 39 unfinished drafts were run through a repair-only workflow: one stage, no draft phase, each agent handed its own on-disk draft plus that draft's measured closeness, with the prompt opening THIS IS A REPAIR, NOT A REWRITE. Model routing deliberately cheap (6 haiku / 29 sonnet / 4 opus — opus only for the four >130-ins functions).

Result: 12 of 39 recovered, 579 instructions, taking wave Q from 51 verified matches (3,631 ins) to 64 (4,245 ins). Roughly a quarter of a stopped wave's tail comes back for a fraction of a fresh wave's cost.

Two calibration notes for next time:

  • Closeness must be counted, not read off the first differing index. My first measurement sorted by the index of the first mismatch and reported six drafts at "closeness 0"; they were truncated drafts (agent stopped mid-write) that matched to instruction 35–48 and then simply ended. Count the differing instructions.
  • The yield concentrates in the small-residual band. Of the 12 recovered, most came from the ≤15-differing-instruction band; the 35–40 band mostly stayed stuck (and much of what remained turned out to be §177's epilogue rule, not per-function work at all).

§176k — TWO SELECTOR BUGS THAT SILENTLY SHRINK A WAVE

Both found while building wave Q, both silent, both would have quietly cost instructions forever:

  1. Ranking gate groups by MEMBER COUNT collapses a wide band to the smallest functions. build_wave_atlas ranked (binary,TU) groups by how many candidates they held — correct for overlays, where each group costs its own rebuild. For main the gate cost is per SLATE, so that ranking filled the wave from the biggest-by-count group, which is the smallest-by- instruction one: measured 60 cards / 2,604 ins selected when 46 cards / 4,829 ins were available. Fixed with --rank mass. Whenever a selector ranks by a proxy, check the proxy still means what it meant when the cost model was written.
  2. A selector that globs its own output poisons itself. Deriving the already-waved set from glob('.run/wave_*_cards.json') matched the file the run was about to write, so re-running with identical filters counted the previous attempt's cards as spent: candidate pool 106 → 46. Fixed by excluding the output path. Any derive-from-disk rule (R33) must exclude the artifact it is about to produce.