The x1 parallel_gate merged 23 files then ABORTED on 'check-all: 211 passed, 2 failed'
(ov_SC03_124, ov_SC03_001), correctly leaving the merge in the tree for inspection (R42).
That red was a STALE INCREMENTAL ARTIFACT, not a bad draft: after reverting the two named
binaries' files, a bare 'make check-all' still reported the same 211/213, but rebuilding
ov_SC03_124 from a cleaned build dir produced 0b3991853c39d8c3d8d12e145e3312dd333f2189 —
exactly config/check.ov_SC03_124.sha. The authoritative sweep (make clean && make extract-all
&& make check-all) then returned 213 passed, 0 failed of 213.
Lesson re-earned (R22's own rationale): a bare check-all is an incremental read and must never
be used to diagnose a red fleet — it nearly cost 21 byte-verified banks. The two reverted
drafts (func_8018009C, func_80182BB8) go back to the drafting queue.
Measured against corpus.stubs at gate time on this session's own waves:
w2 80 drawn / 12 still open -> 85% of its agents re-derived already-banked functions
w3 80 drawn / 39 still open -> 51%
x1 80 drawn / 71 still open -> 11%
~109 of 240 agents in w2+w3 spent full budgets on work that was already banked, by the concurrent
parallel_gate commit or by sibling propagation from a twin remap. The agents DID notice ('stale
pack', 'target .s no longer exists') but only after reading the pack, and several reconstructed the
banked body just to have something to submit.
draw_waves filters against the DRAWN ledger, which stops drawing a target twice but says nothing
about whether it is still OPEN. In a campaign where gates land continuously, a wave drawn at T and
launched at T+2h is stale by construction. Openness is the same class of assertion wave_args already
makes about the .s and the pack, so it belongs here, where it costs nothing.
Negative control (both extremes): x1 drops exactly its 9 closed targets and keeps 71; w2 drops 68
and keeps 12 — matching the independent corpus.stubs measurement.
Auto-reconcile loop dropped 1 batch-aborting draft (func_8005E8E8), then bisected 40 candidates in
24 full-EXE rebuilds: 7 explicitly rejected, 5 chunks unresolved when the bisect hit its own bounded
step limit. gate_main reported 'BANKED 3'; corpus.stubs('main') went 132 -> 130, so 2. SECOND
instance of gate_main over-reporting its bank count by one (R14) — trust the oracle, not the tool.
gate_main.py clean-rebuild gate over the 41 compatible drafts of the 54 drawn (11 dropped for
in-TU decl conflicts, 3 more dropped as batch-aborting compile conflicts by the auto-reconcile
loop). Bisect stopped itself after 24 full-EXE rebuilds with 3 chunks unresolved -- a bounded
failure by design. Final: 143dbb89f34491258bbc27810d0a12ec8b43a8dd BYTE-IDENTICAL.
Explicitly rejected: func_80013154, func_8001AAD0, func_8001BC6C, func_80021174, func_800241C0,
func_80031988.
- match_one.py: a CPP/CC1/MASPSX/AS failure printed bare text and exited 1 even under --json, so
every programmatic caller got json.loads of a non-JSON line. claude_wave_packs swallowed 4 of 19
prior drafts as 'residual not measured: Expecting value' — the most actionable datum a pack can
carry (the draft does not COMPILE, here is the error) was the one it dropped. Negative control:
a near draft still measures identically (closeness 2, same residual rows); human mode unchanged.
- claude_wave_packs.py: renders that verdict, pointing the agent at the card's decl_prior block.
Coverage on wave r1 went 15/19 -> 19/19 packs carrying a measured verdict.
- wave_args.py (new): emits the claude_wave_draft.js args from <wave>/targets.json, asserting the
.s exists, that sub is exactly its parent dir, and that the pack exists. Written because I
hand-typed sub as 'ov_SC03_112/jr_80181D08' when the truth is
'asm/ov_SC03_112/nonmatchings/ov_SC03_112_jr_801817E0' (a stub's asm dir is named for its jr-carve
block, not itself) — all 19 agent oracles would have failed identically and read as a model
failure. Negative control: that exact string is REFUSED.
- draw_waves.py (new): draws waves off corpus.stubs cheapest-first, ledger-filtered, with
name-collision deferral (packs are name-keyed and refuse a duplicate). Re-measured the frontier:
the S65 tier map's '~557 cheap singletons (3-17 ins)' conflated one-member FAMILIES with small
functions — only 28 undrawn non-main stubs are <=17 ins; the bulk is 51-120 (258) and >120 (245).
CURRENT_PHASE.md now holds 27 checkpoint blocks and several older ones say 'supersedes every earlier
block' — true when written, false now. The S64 FINAL block sits ~370 lines above the live S65 FINAL-4
one and makes the same claim, so a fresh session reading top-down could anchor on a state the tree has
moved past by 647 banked functions. Banner at the top states the rule: the LAST block is the live one.
Step 6 of the wave-closing sequence had not run since t5t. This is that backlog: every wave from t5u
to t8b, 139 transcripts, 14 carrying a novelty signal, each distilled and then adversarially verified
against this book before anything was written.
Verdicts: 7 COVERED, 6 ADDENDUM, 1 NEW (section 314).
The COVERED half is the point, not a loss. The verifiers did real work: the "3 new levers" claimed on
func_80183DA4 were traced to 164-71 + 30 stating the identical composite law (scalar decl places the
frame slot at &-time, COMPONENT_REF keeps MEM_IN_STRUCT_P) with a byte-evidenced worked example, and
the claimed-new delay-slot polarity on func_800CAEC0 turned out to REFINE the 220 S64 addendum rather
than contradict it — that addendum's "polarity means aliasing, never arity" holds only when the draft
has a pin or named alias, which this one did not. That bound is now stated.
Four entries are marked UNPROVEN: they come from drafts the whole-binary gate REFUSED, so the lever
moved match_one's closeness but no byte-equality is claimed (R14/G3). Section 314 is one of them — the
abs-range if/else-if wall, with four structurally different C rewrites measured byte-identical.
Index 921 -> 922 sections, green.
wave-harvest-is-a-pipeline-step step 2 (RECOVER the failure set: near-misses, gate-drops, errored
cards) ran for NO wave this session, nor for the earlier t5b-t5r waves. This is that backlog, built
and classified but deliberately NOT run (Drew, S65: queue it for next session).
69 unbanked wave targets: 47 with a draft on disk, 22 errored with none.
Lane A 14 GATE-DROPS — match_one MATCH, gate refused: integration/JTBL_PADS, probe first ($0).
Includes ov_MAIN_012:func_80144B9C at 770 ins, the largest recoverable item.
Lane B 33 NEAR-MISSES — closest are closeness 1, 2, 2, 2, 2, 2, 3, 4, 4, 5. The S65 pack builder
now embeds the measured residual, so a redraft starts from the diff.
Lane C 22 ERRORED — no draft was ever written (rate limits); these are not failures at all.
The lock I added earlier today guards a real hazard (the driver mutates the shared tree and is not
parallel-safe), but I scoped it to the whole tool instead of the mutating path. --probe-only execs
blocker_probe, which compiles in its own scratch dir and touches nothing — excluding it buys no
safety and costs a free diagnostic.
Measured cost: a t7b drafting agent (func_801832E0) tried the probe TWICE, was refused both times by
a concurrent sweep holding the lock, and submitted with its blocker unconfirmed — exactly the $0
diagnostic the pack tells agents to run first.
Control: with the lock held, --probe-only now returns rc 0 and its verdict table; the mutating path
still returns rc 1 REFUSED.
The seam decays hard as it is worked out. Measured across four rounds today: 139/157 = 88.5%, then
178/234 = 76%, then 71/165 = 43%, then 1/95 = 1%. By the last round almost every candidate was one a
previous round had already tried and the gate had refused, so the sweep spent ~7 minutes of builds to
bank one function.
A refusal is deterministic for a given (target, EXEMPLAR) pair — the same exemplar remaps to the same
text — so the ledger keys on both and skips those by default. A NEW exemplar for the same target is a
different question and is retried automatically, which matters because the pool refills as banking
mints exemplars. --retry-refused overrides.
The ledger starts empty (today's rounds predate it) and populates from .run/pgate_results.json.
The distill prompt asserted 'the whole-binary byte-gate ACCEPTED the final draft, so the final body is
ground truth' for EVERY target. With --with-unbanked now feeding it drafts the gate REFUSED, that
sentence would have laundered an unproven body into a byte-proven cookbook entry (R14/G3). A target
with banked=False now gets an explicit PROVENANCE WARNING telling the distiller to extract the lever
anyway (cookbook 52: a model that FAILS still distils the idiom that cracks its siblings) but to mark
every claim UNPROVEN and never assert byte-equality; the verifier is told the same so its
entry_markdown carries the label. Exercised this session: the one surviving ADDENDUM came from an
unbanked draft and is filed UNPROVEN.
Also regenerates the scoreboard after the day's banks.
parallel_gate.py — the per-binary byte gate was serial BY HARNESS, not by nature. Each binary already
compiles into its own build/<bin>/, links its own .ld and checks its own SHA; what serialized it was
shared mutable state in the ONE checkout (the splice, assert_write_set's GLOBAL git status, and the
deliberately-broad `git add -u src/` that must stay broad). Measured: a 109-binary sweep ran ~1
min/binary on a 32-core box at load 1.4 (~4% utilisation), and an xargs -P 4 attempt over the shared
tree CORRUPTED it earlier today.
Fix is ISOLATION, not locking: one git worktree per worker (own index, own src/, own build/). Workers
gate and NEVER commit; the orchestrator adopts only drafts the gate ACCEPTED, and only where the main
tree still matches the pinned baseline (otherwise REFUSED, never clobbered), then ONE commit and ONE
R22 clean-fleet sweep verifies the merged whole.
Measured on 85 binaries / 234 candidates: 178 banked in 12m19s wall for 127m40s CPU = 10.4x
parallelism, ~7x end-to-end vs serial, 99 files merged, 0 refused, check-all 213/213.
Five things a fresh worktree does NOT have, each found by measurement and each first appearing as
"the draft failed": splat-generated include/*.inc, the EMPTY tools/maspsx submodule, gitignored
tools/bin (cc1) + tools/psyq, build/{<bin>,assets/<bin>} outputs, and extracted/retail. Dirs mixing
tracked and untracked content cannot be symlinked wholesale (ln -s nests INSIDE them) — hence
link_missing(). The negative control that catches all of it: an UNMODIFIED binary must build
BYTE-IDENTICAL in the worktree. Before that control, the first parallel run reported a clean,
plausible "0 banked across 4 binaries" that was pure environment artifact.
twin_sweep.py — enumerate every open stub that has an ALREADY-BANKED structural twin, remap it
mechanically, gate it. Yields measured today: h_exact 139/157 = 88.5%, h_norm 36/45 then 178/234
= 76-80%. h_seq stays refused (Phase-26) and is not used. THE POOL REFILLS: each bank becomes an
exemplar for its siblings — 165 fresh candidates existed immediately after banking 178. Run it
BEFORE drawing any wave; t5_cards does not build seed_ref, so cards assert "no banked twin" for
these and agents redraft answers we already hold (a t5u Opus slot ground ov_SC03_023:func_8017BEBC
to closeness 45 while ov_SC02_004 held a byte-identical banked copy).