I shipped a guard that used drafting-scratch mtimes as a liveness proxy. It failed
in both possible directions within minutes:
* FALSE PASS: the find included '.run/*wave*', which expanded past ARG_MAX
('Argument list too long'). find then matched nothing, the guard PASSED, and I
ran clean: removed build/, expected/, and the regenerated splat tree (asm/, assets/, include macros, undefined_*_auto.txt). on a live lane — deleting asm/ under five drafting agents. I
restored it immediately (extract-all 212/212) but that is damage control, not a
design.
* FALSE PASS, structurally: even with the glob fixed, an agent that THINKS longer
than the window is indistinguishable from a finished one — the exact flaw I had
already written into gater_lane's docstring for the verdicts file ('a quiet file
mtime is deliberately NOT accepted as one') and then rebuilt here anyway.
tools/lane_inflight.py is the fix: liveness is RECORDED, not inferred. The
orchestrator adds a target when it launches the workflow and removes it when the
verdict returns — both actions it already performs, so the ledger cannot drift
without skipping a step that is taken anyway. exits non-zero when any agent
is live, which IS the guard, and both r22_verify.sh and parallel_gate --r22 now use
it instead of touching the filesystem.
Negative-controlled both directions: refuses with 5 live agents named and their
start times; passes when the ledger is drained.
The lesson worth more than the fix: I had already identified 'a quiet mtime is not
a completion signal' as a defect class, documented it, and then re-implemented it
in a different file. Writing a rule down does not stop you applying its opposite
somewhere else.
A scripted patch I applied inserted three lines that (a) re-ran the whole
check-all inside a process substitution and (b) grepped /dev/null. Caught by
reading the file back instead of trusting the edit reported success.
The rewrite does what was intended: capture check-all's output ONCE, clear
.run/R22_DEBT only when the summary line says '0 failed' AND the exit code is 0
(R53 -- a failed build leaves the previous binary on disk and sha1sum reads green,
so the exit code alone is not enough), and leave the debt standing otherwise.
tools/r22_verify.sh (NEW, promoted from .run so it survives the session):
'make clean' deletes asm/ AND build/, and THREE times this session that raced a
live lane -- a subagent authorised to splice src/800.c produced a FALSE
'212 passed, 1 failed' red, and two drafting agents reported their target's asm/
tree MISSING mid-draft (one survived only by finding an old snapshot). Drafting
agents never WRITE src/, which is exactly why 'check for a dirty tree' does not
catch them: they DEPEND on state this operation destroys. The guard refuses when
any wave scratch dir was touched in the last 6 minutes, names the live agents, and
offers R22_FORCE for a drained lane. R54 -- a guard that is not running is not a
guard, so this refuses instead of relying on me remembering.
Negative-controlled BOTH directions: refuses with 5 live agents named; passes on an
idle lane AND on a lane whose scratch is 30 minutes stale (no false positives).
fix(gater): the in-tree main commit message said '0 fn(s)' for a commit that
contained a real bank. corpus memoizes, so querying corpus.stubs immediately after
the bank returns the STALE pre-bank set. Derive the list from harvest_verify's own
verified-out file instead (R33: derive from the invariant the tool already wrote).
§372 ★★★ THE COPY-CAPTURE PAIR. Tell: a REGALLOC-PERM residual whose wrong-register
rows READ the destination of a nearby MATCHING copy insn. Two passes re-base uses
onto a copy's destination -- cse.c make_regs_eqv (canonical-reg rewrite of later
same-EBB uses) and local-alloc.c optimize_reg_copy_1 (forward-substitution when the
copy's src does not die in it) -- and BOTH die to one zero-byte edit: spell the copy
'P = X + zr' so SET_SRC is a PLUS, which is not a reg-reg copy and records no reg
equivalence, while emitting the byte-identical 'addu $rd,$rs,$zero'.
Notably the escalation was told to CHECK whether §368's tell applied rather than
assume it; it reported that it did NOT (pure shift/slti rows, no commutative
operands) and found the real cause from RTL dumps. That is §361's procedure working.