docs: close the S71 documentation gaps - SETUP tools table, playbook steps, decision log, gate ledger

Audit found real gaps rather than assumed coverage:
* SETUP.md (R21) had NONE of the five tools written this session. Added a table for
  journal_notes / launch_check / gate_triage / restage_matching / weave_sweep, each with
  when you need it, plus the two gating rules now enforced in code (parallel_gate refuses
  main; gate_main refuses a no-op draft and counts banks from the source).
* wave-playbook: launch_check as step 4c (payloads go stale while gates run - 3 of 27
  wave-2 targets were already banked) and gate_triage as step 6b with the measured
  blocker census.
* decision-log (R31) held only the §406 pivot. Added the two strategic entries this
  session actually turned on: gating main with a tool documented as unable to gate it
  (false PASS, caught only by R22), and the drafting pool running dry while the lever
  was an exclude list nobody re-probed after a tool fix.
* CURRENT_PHASE: the per-gate ledger for all 14 cycles plus the carve/rebase/main gates.
* Two memories: gate-main-only-with-gate-main, reprobe-exclude-lists-after-tool-fixes.
This commit is contained in:
Drew T
2026-09-02 11:19:49 -06:00
parent 9631949339
commit 6c904ebd0c
4 changed files with 125 additions and 0 deletions
+18
View File
@@ -931,3 +931,21 @@ fills fast). Nothing is leaking — but the host does not get the memory back on
* **Automatic, until the config applies:** `.run/memkeeper.sh` (nohup'd) drops the cache whenever it
exceeds 12 GB, every 5 minutes, logging to `.run/memkeeper.log`.
* `autoMemoryReclaim=dropcache` is the aggressive variant if `gradual` proves too slow.
### Tools added 2026-09-02 (S71) — R21 record
| tool | what it does | when you need it |
|---|---|---|
| `tools/journal_notes.py` | mines the agent journals for a `(binary, fn)`'s PAST ATTEMPTS and appends them to its pack; also reads `.run/journal_notes_local.jsonl` for hand-recorded evidence | **automatic** — `claude_wave_packs.py` calls it at the end of pack generation. Run `--wave <dir>` to back-fill a wave built another way (idempotent), `--fn F --binary B` to read what we hold on one function |
| `tools/launch_check.py` | refuses to launch an agent at an **already-banked** target; `--payload p.json` filters a `{wave,targets}` payload in place | before every launch. `wave_args` asserts open-ness at DRAW time, and payloads sit on disk while gates run — S71 launched one stale card and burned a full agent run |
| `tools/gate_triage.py` | routes a `parallel_gate` result set to the repair lane each verdict names (CARVE / UNDEF / CONFLICT / ARITY / PARSE / NO-DIAG / DIFF), asserting the staged-draft denominator | after any gate that banked less than it staged — tells you which lane the failures belong to instead of guessing |
| `tools/restage_matching.py` | rebuilds a gate plan from `recover_integration --probe-only` verdicts, keeping only drafts that compile-and-MATCH in their REAL TU | when a binary banks 0 and you suspect one bad draft is failing its siblings' shared build. **Caveat measured S71:** that probe compiles but never LINKS or CARVES, so its MATCH is not a bank prediction |
| `tools/weave_sweep.py` | the §406/§408 derived-selector sweep: scores each open stub's stored drafts, classifies the `sw $ra` disagreement from the residual, applies the clobber only to WEAVE-SUNK, `--lever-all` is the ablation control | as the template for "price a class by its RESIDUAL, not its SHAPE" — the sweep itself is a measured null (§408) |
**Two gating rules that are now enforced in code, not remembered:**
* `parallel_gate` **REFUSES `main`** — main's extract rewrites the linker script, so an incremental
gate is a false PASS (§414). Use `tools/gate_main.py`: baseline assert → one clean rebuild per
slate → bisect on failure.
* `gate_main` **REFUSES a draft containing its own `INCLUDE_ASM`** (substituting it restores the stub,
so the build passes for free and the function counts as banked), and counts banks from the SOURCE.
+76
View File
@@ -2712,3 +2712,79 @@ member must be priced by the residuals of the others before it is written down a
`--json` field on scores we were already running would have said "15, not 134" on the night it was
claimed. Cost of learning it here: about one hour of deterministic compute and no agent tokens, which
is exactly what a probe-first rule is supposed to buy.
---
## 2026-09-02 (P31 S71) — I gated `main` with a tool documented as unable to gate it, and only R22 caught it
**Context and belief.** The session's integration lane was running well: `parallel_gate` in isolated
worktrees had banked cleanly across 30-odd overlays all night. When 15 of the 64 standalone-match
bodies turned out to be `main`'s, I put them through the same tool. It reported **11 banked**, the
merge committed them, and every signal I was watching — worker exit codes, the bank oracle (a stub
disappeared), the summary line — agreed.
**What failed.** The R22 clean-fleet verify returned **212/213**. `main` did not compile from clean
(two `conflicting types` errors). Reconciling both declarations made it build — and it was **still not
byte-identical**. Re-gated one function at a time against a clean tree: **11 of 11 REJECTED.** The
commit was reverted and `main` was verified byte-identical again before anything else proceeded.
**The rule already existed, three files away.** `ox_campaign.gate_main_batch`'s docstring:
*"main is gated by ONE CLEAN REBUILD of the whole EXE, never incrementally … main's extract rewrites
the linker script, so an incremental main gate returns a FALSE DIFF. Measured P31 S58: wave `ab` drew
105 main cards and banked 0 of them."* `parallel_gate`'s worker **is** `gate_stage`, so it inherits
that constraint exactly. I had read that docstring earlier the same session, while looking at
something else.
**Why the failure direction was worse than the one on record.** S58 recorded the false-DIFF direction:
competent drafts thrown away, loud and wasteful. This was the false-PASS direction: wrong bytes
committed, reading green until the next clean fleet check. R53 names the mechanism — *a failed build
leaves the previous object on disk, so a SHA1 check downstream of it reads green* — and R53 was
written for a different tool and never applied here.
**The pivot.** Fixed as a **refusal in the wrapper**, not a note in the callee: `parallel_gate` now
returns REFUSED for `binary == 'main'` and names `tools/gate_main.py`. The main lane was then reopened
properly the next morning and banked 5 (4 after the source-truth correction below), with a bisect
isolating the one bad draft in 7 rebuilds.
**Hindsight / better path.** Three things would each have caught it earlier, in increasing order of
generality: (a) run R22 **before** committing a gate against a binary the lane has not gated before,
not at session close; (b) `gate_main`'s own bank count was also derived rather than measured — it
printed "BANKED 5 of 6" when 4 had applied, because `len(good)` is *what we decided to keep*, not
*what was substituted* — so **count from the source in every gating tool**; (c) the general rule this
session kept re-teaching: **a tool that wraps another tool inherits its refusals**, and the place to
encode that is a refusal in the wrapper. Every constraint documented on `gate_stage` binds
`parallel_gate`, `harvest_verify`, and anything else that shells it.
**Cost of learning it here:** one bad commit, ~40 minutes of revert-and-bisect, and an inflated bank
count I had already reported to Drew and had to correct. Cheap only because R22 exists and was run.
---
## 2026-09-02 (P31 S71) — the drafting pool ran dry, and the lever was an exclude list nobody re-probed
**Context and belief.** With ~40 agents landing at near-100% MATCH, the working assumption was that
drafting capacity was the constraint and the campaign would continue as draw → draft → gate until the
frontier was gone.
**What failed.** Wave 3 drew **1 target** and reported *"0 left in pool"*. Measured at that moment:
174 open, of which `main` 64, and of the 110 non-main — **41 drafted this session, 68 on the exclude
list, 2 proven walls, ZERO genuinely undrawn**. More agents would have had nothing to work on.
**The pivot.** The 68 excluded functions were excluded because the TOOLING could not carve them —
`jtbl_carve` refused their plans with *"subseg would host NON-CONTIGUOUS `.rodata` carves"*. But
tooling had changed **that same session**: `jr_isolate_all` had been fixed twice (file-local `static`
placement, and §323's `__attribute__`-blind regex). Re-probing all 68 found **17 now reporting `tail`
— a standard §8a carve**. Every one already had drafts on disk; scoring them put **10 at closeness 0
for zero drafting**, and the gate banked 9 — four of them in **57 seconds**.
**The grounded why.** An exclude list is a snapshot of *what the tooling could not do at the moment it
was written*. It is treated thereafter as a property of the FUNCTIONS. Nothing in the pipeline
re-examines it, so every tool improvement leaves behind a population that is now tractable and still
marked impossible — invisible, because the draw filters it out before anything measures it.
**Hindsight / better path.** **Re-probe the exclude list after every tool fix, as part of the fix.**
The probe is deterministic, costs no agents, and here it was worth more than the entire drafting lane
at that moment. Generalised: *any list that records a tool's limitation must be regenerated when the
tool changes, or it silently becomes a list of work you have decided not to do.* The same reasoning
applies to `.run/S71_walls_found.txt` — a wall proven against today's compiler knowledge is not a wall
forever, and each entry should carry the refutation list that would have to be beaten.
+24
View File
@@ -292,6 +292,19 @@ Notably variant 3 (`conflicting types`, RETURN type only, decl already `()`, sym
is fixed by `--sync-decls` ALONE: both other levers no-op, and the sync is safe precisely because an
address-taken site has no arguments to convert.
### 4c. LAUNCH-TIME OPEN CHECK — `wave_args` asserts at DRAW time, and payloads go stale
```
python3 tools/launch_check.py --payload .run/<wave>/wf_args.json # filters in place
python3 tools/launch_check.py <binary> <fn> # exit 2 = already banked
```
A wave's payload sits on disk while gates run, so by launch time some of its targets are banked. An
agent handed one burns a full run to report "STALE CARD — already banked today", with no `.s` left to
score against. Measured S71: `ov_SC01_006/func_8017F9F8` did exactly that, and filtering the wave-2
payload found **3 of 27** already banked. Also skip any target that already has a FRESH draft from
this session — it needs a gate, not another agent.
## 5. Draft
`tools/workflows/claude_wave_draft.js`, `args = {wave, targets}`. One agent per target,
@@ -360,6 +373,17 @@ python3 tools/gate_wave.py --drafts <dir> --workers 8 --commit [--r22]
about an hour for what should have taken minutes. **`gate_wave.py`'s split is now an optimisation
(same-binary drafts share a build), NOT a safety requirement.**
### 6b. READ THE VERDICTS — `tools/gate_triage.py`
```
python3 tools/gate_triage.py --plan <gate_plan.json>
```
Routes every verdict to the lane it names and asserts the staged denominator: CARVE (probe it —
`jr_isolate_all` is usually the unblock) · UNDEF-D (§171 `aprop_symfix`) · CONFLICT/ARITY (§376/§378)
· PARSE · NO-DIAG · DIFF (real codegen). S71's census over 37 verdicts: DIFF 18 · CARVE 7 · PARSE 3 ·
NO-DIAG 3 · CONFLICT 2 · ARITY 2 · UNDEF 2 — which corrected an impression that carve dominated.
## 7. After ANY bank
```
+7
View File
@@ -5569,6 +5569,13 @@ Every one is fuel keyed by BARE NAME or asserted without a freshness check (R48/
time last night) — §376 in its purest form. Do not re-slate them without a TU-level fix.
**No new waves from here (Drew).**
- 2026-09-02 — **S71 gate ledger, all 14 cycles** (each `parallel_gate --commit` unless noted; every
bank R22-verified at close): g1 12 (integration pile, main's 11 later REVERTED — §414) · g2 0 ·
g3 9 · g4 6 · g5 5 · g6 2 · g7 4 · g8 3 · g9 2 · g10 2 · g11 3 · g12 2 · g13 2 · g14 4 · g15 2 ·
**carve gate 5** (the five jr-isolated overlays) · **§420 rebase gate 4** (one body, four overlays,
57 s) · **`gate_main` 4 + 1** (the only tool that may gate main).
Net: **210 → 147 = 63 banked**, `check-all: 213 passed, 0 failed of 213`.
## 🛑 SESSION CHECKPOINT — S71 CLOSE (2026-09-02 11:12). SUPERSEDES every earlier block in this file. Phase 31 T10 CONTINUES.
**FLEET VERIFIED GREEN FROM A CLEAN REBUILD — `check-all: 213 passed, 0 failed of 213`**