From 6c904ebd0c514c7d47bcded97e1e8110c1e006be Mon Sep 17 00:00:00 2001 From: Drew T <50529377+Druthulu@users.noreply.github.com> Date: Wed, 2 Sep 2026 11:19:49 -0600 Subject: [PATCH] docs: close the S71 documentation gaps - SETUP tools table, playbook steps, decision log, gate ledger MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Audit found real gaps rather than assumed coverage: * SETUP.md (R21) had NONE of the five tools written this session. Added a table for journal_notes / launch_check / gate_triage / restage_matching / weave_sweep, each with when you need it, plus the two gating rules now enforced in code (parallel_gate refuses main; gate_main refuses a no-op draft and counts banks from the source). * wave-playbook: launch_check as step 4c (payloads go stale while gates run - 3 of 27 wave-2 targets were already banked) and gate_triage as step 6b with the measured blocker census. * decision-log (R31) held only the §406 pivot. Added the two strategic entries this session actually turned on: gating main with a tool documented as unable to gate it (false PASS, caught only by R22), and the drafting pool running dry while the lever was an exclude list nobody re-probed after a tool fix. * CURRENT_PHASE: the per-gate ledger for all 14 cycles plus the carve/rebase/main gates. * Two memories: gate-main-only-with-gate-main, reprobe-exclude-lists-after-tool-fixes. --- docs/SETUP.md | 18 +++++++++ docs/decision-log.md | 76 +++++++++++++++++++++++++++++++++++++ docs/wave-playbook.md | 24 ++++++++++++ phase-ends/CURRENT_PHASE.md | 7 ++++ 4 files changed, 125 insertions(+) diff --git a/docs/SETUP.md b/docs/SETUP.md index 0e32b1c38a..1da6fca5a7 100644 --- a/docs/SETUP.md +++ b/docs/SETUP.md @@ -931,3 +931,21 @@ fills fast). Nothing is leaking — but the host does not get the memory back on * **Automatic, until the config applies:** `.run/memkeeper.sh` (nohup'd) drops the cache whenever it exceeds 12 GB, every 5 minutes, logging to `.run/memkeeper.log`. * `autoMemoryReclaim=dropcache` is the aggressive variant if `gradual` proves too slow. + + +### Tools added 2026-09-02 (S71) — R21 record + +| tool | what it does | when you need it | +|---|---|---| +| `tools/journal_notes.py` | mines the agent journals for a `(binary, fn)`'s PAST ATTEMPTS and appends them to its pack; also reads `.run/journal_notes_local.jsonl` for hand-recorded evidence | **automatic** — `claude_wave_packs.py` calls it at the end of pack generation. Run `--wave ` to back-fill a wave built another way (idempotent), `--fn F --binary B` to read what we hold on one function | +| `tools/launch_check.py` | refuses to launch an agent at an **already-banked** target; `--payload p.json` filters a `{wave,targets}` payload in place | before every launch. `wave_args` asserts open-ness at DRAW time, and payloads sit on disk while gates run — S71 launched one stale card and burned a full agent run | +| `tools/gate_triage.py` | routes a `parallel_gate` result set to the repair lane each verdict names (CARVE / UNDEF / CONFLICT / ARITY / PARSE / NO-DIAG / DIFF), asserting the staged-draft denominator | after any gate that banked less than it staged — tells you which lane the failures belong to instead of guessing | +| `tools/restage_matching.py` | rebuilds a gate plan from `recover_integration --probe-only` verdicts, keeping only drafts that compile-and-MATCH in their REAL TU | when a binary banks 0 and you suspect one bad draft is failing its siblings' shared build. **Caveat measured S71:** that probe compiles but never LINKS or CARVES, so its MATCH is not a bank prediction | +| `tools/weave_sweep.py` | the §406/§408 derived-selector sweep: scores each open stub's stored drafts, classifies the `sw $ra` disagreement from the residual, applies the clobber only to WEAVE-SUNK, `--lever-all` is the ablation control | as the template for "price a class by its RESIDUAL, not its SHAPE" — the sweep itself is a measured null (§408) | + +**Two gating rules that are now enforced in code, not remembered:** +* `parallel_gate` **REFUSES `main`** — main's extract rewrites the linker script, so an incremental + gate is a false PASS (§414). Use `tools/gate_main.py`: baseline assert → one clean rebuild per + slate → bisect on failure. +* `gate_main` **REFUSES a draft containing its own `INCLUDE_ASM`** (substituting it restores the stub, + so the build passes for free and the function counts as banked), and counts banks from the SOURCE. diff --git a/docs/decision-log.md b/docs/decision-log.md index 2252e80927..528a65c45b 100644 --- a/docs/decision-log.md +++ b/docs/decision-log.md @@ -2712,3 +2712,79 @@ member must be priced by the residuals of the others before it is written down a `--json` field on scores we were already running would have said "15, not 134" on the night it was claimed. Cost of learning it here: about one hour of deterministic compute and no agent tokens, which is exactly what a probe-first rule is supposed to buy. + +--- + +## 2026-09-02 (P31 S71) — I gated `main` with a tool documented as unable to gate it, and only R22 caught it + +**Context and belief.** The session's integration lane was running well: `parallel_gate` in isolated +worktrees had banked cleanly across 30-odd overlays all night. When 15 of the 64 standalone-match +bodies turned out to be `main`'s, I put them through the same tool. It reported **11 banked**, the +merge committed them, and every signal I was watching — worker exit codes, the bank oracle (a stub +disappeared), the summary line — agreed. + +**What failed.** The R22 clean-fleet verify returned **212/213**. `main` did not compile from clean +(two `conflicting types` errors). Reconciling both declarations made it build — and it was **still not +byte-identical**. Re-gated one function at a time against a clean tree: **11 of 11 REJECTED.** The +commit was reverted and `main` was verified byte-identical again before anything else proceeded. + +**The rule already existed, three files away.** `ox_campaign.gate_main_batch`'s docstring: +*"main is gated by ONE CLEAN REBUILD of the whole EXE, never incrementally … main's extract rewrites +the linker script, so an incremental main gate returns a FALSE DIFF. Measured P31 S58: wave `ab` drew +105 main cards and banked 0 of them."* `parallel_gate`'s worker **is** `gate_stage`, so it inherits +that constraint exactly. I had read that docstring earlier the same session, while looking at +something else. + +**Why the failure direction was worse than the one on record.** S58 recorded the false-DIFF direction: +competent drafts thrown away, loud and wasteful. This was the false-PASS direction: wrong bytes +committed, reading green until the next clean fleet check. R53 names the mechanism — *a failed build +leaves the previous object on disk, so a SHA1 check downstream of it reads green* — and R53 was +written for a different tool and never applied here. + +**The pivot.** Fixed as a **refusal in the wrapper**, not a note in the callee: `parallel_gate` now +returns REFUSED for `binary == 'main'` and names `tools/gate_main.py`. The main lane was then reopened +properly the next morning and banked 5 (4 after the source-truth correction below), with a bisect +isolating the one bad draft in 7 rebuilds. + +**Hindsight / better path.** Three things would each have caught it earlier, in increasing order of +generality: (a) run R22 **before** committing a gate against a binary the lane has not gated before, +not at session close; (b) `gate_main`'s own bank count was also derived rather than measured — it +printed "BANKED 5 of 6" when 4 had applied, because `len(good)` is *what we decided to keep*, not +*what was substituted* — so **count from the source in every gating tool**; (c) the general rule this +session kept re-teaching: **a tool that wraps another tool inherits its refusals**, and the place to +encode that is a refusal in the wrapper. Every constraint documented on `gate_stage` binds +`parallel_gate`, `harvest_verify`, and anything else that shells it. + +**Cost of learning it here:** one bad commit, ~40 minutes of revert-and-bisect, and an inflated bank +count I had already reported to Drew and had to correct. Cheap only because R22 exists and was run. + +--- + +## 2026-09-02 (P31 S71) — the drafting pool ran dry, and the lever was an exclude list nobody re-probed + +**Context and belief.** With ~40 agents landing at near-100% MATCH, the working assumption was that +drafting capacity was the constraint and the campaign would continue as draw → draft → gate until the +frontier was gone. + +**What failed.** Wave 3 drew **1 target** and reported *"0 left in pool"*. Measured at that moment: +174 open, of which `main` 64, and of the 110 non-main — **41 drafted this session, 68 on the exclude +list, 2 proven walls, ZERO genuinely undrawn**. More agents would have had nothing to work on. + +**The pivot.** The 68 excluded functions were excluded because the TOOLING could not carve them — +`jtbl_carve` refused their plans with *"subseg would host NON-CONTIGUOUS `.rodata` carves"*. But +tooling had changed **that same session**: `jr_isolate_all` had been fixed twice (file-local `static` +placement, and §323's `__attribute__`-blind regex). Re-probing all 68 found **17 now reporting `tail` +— a standard §8a carve**. Every one already had drafts on disk; scoring them put **10 at closeness 0 +for zero drafting**, and the gate banked 9 — four of them in **57 seconds**. + +**The grounded why.** An exclude list is a snapshot of *what the tooling could not do at the moment it +was written*. It is treated thereafter as a property of the FUNCTIONS. Nothing in the pipeline +re-examines it, so every tool improvement leaves behind a population that is now tractable and still +marked impossible — invisible, because the draw filters it out before anything measures it. + +**Hindsight / better path.** **Re-probe the exclude list after every tool fix, as part of the fix.** +The probe is deterministic, costs no agents, and here it was worth more than the entire drafting lane +at that moment. Generalised: *any list that records a tool's limitation must be regenerated when the +tool changes, or it silently becomes a list of work you have decided not to do.* The same reasoning +applies to `.run/S71_walls_found.txt` — a wall proven against today's compiler knowledge is not a wall +forever, and each entry should carry the refutation list that would have to be beaten. diff --git a/docs/wave-playbook.md b/docs/wave-playbook.md index edd8577def..4bec4f5bf7 100644 --- a/docs/wave-playbook.md +++ b/docs/wave-playbook.md @@ -292,6 +292,19 @@ Notably variant 3 (`conflicting types`, RETURN type only, decl already `()`, sym is fixed by `--sync-decls` ALONE: both other levers no-op, and the sync is safe precisely because an address-taken site has no arguments to convert. +### 4c. LAUNCH-TIME OPEN CHECK — `wave_args` asserts at DRAW time, and payloads go stale + +``` +python3 tools/launch_check.py --payload .run//wf_args.json # filters in place +python3 tools/launch_check.py # exit 2 = already banked +``` + +A wave's payload sits on disk while gates run, so by launch time some of its targets are banked. An +agent handed one burns a full run to report "STALE CARD — already banked today", with no `.s` left to +score against. Measured S71: `ov_SC01_006/func_8017F9F8` did exactly that, and filtering the wave-2 +payload found **3 of 27** already banked. Also skip any target that already has a FRESH draft from +this session — it needs a gate, not another agent. + ## 5. Draft `tools/workflows/claude_wave_draft.js`, `args = {wave, targets}`. One agent per target, @@ -360,6 +373,17 @@ python3 tools/gate_wave.py --drafts --workers 8 --commit [--r22] about an hour for what should have taken minutes. **`gate_wave.py`'s split is now an optimisation (same-binary drafts share a build), NOT a safety requirement.** +### 6b. READ THE VERDICTS — `tools/gate_triage.py` + +``` +python3 tools/gate_triage.py --plan +``` + +Routes every verdict to the lane it names and asserts the staged denominator: CARVE (probe it — +`jr_isolate_all` is usually the unblock) · UNDEF-D (§171 `aprop_symfix`) · CONFLICT/ARITY (§376/§378) +· PARSE · NO-DIAG · DIFF (real codegen). S71's census over 37 verdicts: DIFF 18 · CARVE 7 · PARSE 3 · +NO-DIAG 3 · CONFLICT 2 · ARITY 2 · UNDEF 2 — which corrected an impression that carve dominated. + ## 7. After ANY bank ``` diff --git a/phase-ends/CURRENT_PHASE.md b/phase-ends/CURRENT_PHASE.md index 140863ab4e..b7d2283541 100644 --- a/phase-ends/CURRENT_PHASE.md +++ b/phase-ends/CURRENT_PHASE.md @@ -5569,6 +5569,13 @@ Every one is fuel keyed by BARE NAME or asserted without a freshness check (R48/ time last night) — §376 in its purest form. Do not re-slate them without a TU-level fix. **No new waves from here (Drew).** +- 2026-09-02 — **S71 gate ledger, all 14 cycles** (each `parallel_gate --commit` unless noted; every + bank R22-verified at close): g1 12 (integration pile, main's 11 later REVERTED — §414) · g2 0 · + g3 9 · g4 6 · g5 5 · g6 2 · g7 4 · g8 3 · g9 2 · g10 2 · g11 3 · g12 2 · g13 2 · g14 4 · g15 2 · + **carve gate 5** (the five jr-isolated overlays) · **§420 rebase gate 4** (one body, four overlays, + 57 s) · **`gate_main` 4 + 1** (the only tool that may gate main). + Net: **210 → 147 = 63 banked**, `check-all: 213 passed, 0 failed of 213`. + ## 🛑 SESSION CHECKPOINT — S71 CLOSE (2026-09-02 11:12). SUPERSEDES every earlier block in this file. Phase 31 T10 CONTINUES. **FLEET VERIFIED GREEN FROM A CLEAN REBUILD — `check-all: 213 passed, 0 failed of 213`**