mirror of
https://github.com/Druthulu/BFM-decomp
synced 2026-10-03 00:05:11 -04:00
docs: close the S71 documentation gaps - SETUP tools table, playbook steps, decision log, gate ledger
Audit found real gaps rather than assumed coverage: * SETUP.md (R21) had NONE of the five tools written this session. Added a table for journal_notes / launch_check / gate_triage / restage_matching / weave_sweep, each with when you need it, plus the two gating rules now enforced in code (parallel_gate refuses main; gate_main refuses a no-op draft and counts banks from the source). * wave-playbook: launch_check as step 4c (payloads go stale while gates run - 3 of 27 wave-2 targets were already banked) and gate_triage as step 6b with the measured blocker census. * decision-log (R31) held only the §406 pivot. Added the two strategic entries this session actually turned on: gating main with a tool documented as unable to gate it (false PASS, caught only by R22), and the drafting pool running dry while the lever was an exclude list nobody re-probed after a tool fix. * CURRENT_PHASE: the per-gate ledger for all 14 cycles plus the carve/rebase/main gates. * Two memories: gate-main-only-with-gate-main, reprobe-exclude-lists-after-tool-fixes.
This commit is contained in:
@@ -931,3 +931,21 @@ fills fast). Nothing is leaking — but the host does not get the memory back on
|
||||
* **Automatic, until the config applies:** `.run/memkeeper.sh` (nohup'd) drops the cache whenever it
|
||||
exceeds 12 GB, every 5 minutes, logging to `.run/memkeeper.log`.
|
||||
* `autoMemoryReclaim=dropcache` is the aggressive variant if `gradual` proves too slow.
|
||||
|
||||
|
||||
### Tools added 2026-09-02 (S71) — R21 record
|
||||
|
||||
| tool | what it does | when you need it |
|
||||
|---|---|---|
|
||||
| `tools/journal_notes.py` | mines the agent journals for a `(binary, fn)`'s PAST ATTEMPTS and appends them to its pack; also reads `.run/journal_notes_local.jsonl` for hand-recorded evidence | **automatic** — `claude_wave_packs.py` calls it at the end of pack generation. Run `--wave <dir>` to back-fill a wave built another way (idempotent), `--fn F --binary B` to read what we hold on one function |
|
||||
| `tools/launch_check.py` | refuses to launch an agent at an **already-banked** target; `--payload p.json` filters a `{wave,targets}` payload in place | before every launch. `wave_args` asserts open-ness at DRAW time, and payloads sit on disk while gates run — S71 launched one stale card and burned a full agent run |
|
||||
| `tools/gate_triage.py` | routes a `parallel_gate` result set to the repair lane each verdict names (CARVE / UNDEF / CONFLICT / ARITY / PARSE / NO-DIAG / DIFF), asserting the staged-draft denominator | after any gate that banked less than it staged — tells you which lane the failures belong to instead of guessing |
|
||||
| `tools/restage_matching.py` | rebuilds a gate plan from `recover_integration --probe-only` verdicts, keeping only drafts that compile-and-MATCH in their REAL TU | when a binary banks 0 and you suspect one bad draft is failing its siblings' shared build. **Caveat measured S71:** that probe compiles but never LINKS or CARVES, so its MATCH is not a bank prediction |
|
||||
| `tools/weave_sweep.py` | the §406/§408 derived-selector sweep: scores each open stub's stored drafts, classifies the `sw $ra` disagreement from the residual, applies the clobber only to WEAVE-SUNK, `--lever-all` is the ablation control | as the template for "price a class by its RESIDUAL, not its SHAPE" — the sweep itself is a measured null (§408) |
|
||||
|
||||
**Two gating rules that are now enforced in code, not remembered:**
|
||||
* `parallel_gate` **REFUSES `main`** — main's extract rewrites the linker script, so an incremental
|
||||
gate is a false PASS (§414). Use `tools/gate_main.py`: baseline assert → one clean rebuild per
|
||||
slate → bisect on failure.
|
||||
* `gate_main` **REFUSES a draft containing its own `INCLUDE_ASM`** (substituting it restores the stub,
|
||||
so the build passes for free and the function counts as banked), and counts banks from the SOURCE.
|
||||
|
||||
@@ -2712,3 +2712,79 @@ member must be priced by the residuals of the others before it is written down a
|
||||
`--json` field on scores we were already running would have said "15, not 134" on the night it was
|
||||
claimed. Cost of learning it here: about one hour of deterministic compute and no agent tokens, which
|
||||
is exactly what a probe-first rule is supposed to buy.
|
||||
|
||||
---
|
||||
|
||||
## 2026-09-02 (P31 S71) — I gated `main` with a tool documented as unable to gate it, and only R22 caught it
|
||||
|
||||
**Context and belief.** The session's integration lane was running well: `parallel_gate` in isolated
|
||||
worktrees had banked cleanly across 30-odd overlays all night. When 15 of the 64 standalone-match
|
||||
bodies turned out to be `main`'s, I put them through the same tool. It reported **11 banked**, the
|
||||
merge committed them, and every signal I was watching — worker exit codes, the bank oracle (a stub
|
||||
disappeared), the summary line — agreed.
|
||||
|
||||
**What failed.** The R22 clean-fleet verify returned **212/213**. `main` did not compile from clean
|
||||
(two `conflicting types` errors). Reconciling both declarations made it build — and it was **still not
|
||||
byte-identical**. Re-gated one function at a time against a clean tree: **11 of 11 REJECTED.** The
|
||||
commit was reverted and `main` was verified byte-identical again before anything else proceeded.
|
||||
|
||||
**The rule already existed, three files away.** `ox_campaign.gate_main_batch`'s docstring:
|
||||
*"main is gated by ONE CLEAN REBUILD of the whole EXE, never incrementally … main's extract rewrites
|
||||
the linker script, so an incremental main gate returns a FALSE DIFF. Measured P31 S58: wave `ab` drew
|
||||
105 main cards and banked 0 of them."* `parallel_gate`'s worker **is** `gate_stage`, so it inherits
|
||||
that constraint exactly. I had read that docstring earlier the same session, while looking at
|
||||
something else.
|
||||
|
||||
**Why the failure direction was worse than the one on record.** S58 recorded the false-DIFF direction:
|
||||
competent drafts thrown away, loud and wasteful. This was the false-PASS direction: wrong bytes
|
||||
committed, reading green until the next clean fleet check. R53 names the mechanism — *a failed build
|
||||
leaves the previous object on disk, so a SHA1 check downstream of it reads green* — and R53 was
|
||||
written for a different tool and never applied here.
|
||||
|
||||
**The pivot.** Fixed as a **refusal in the wrapper**, not a note in the callee: `parallel_gate` now
|
||||
returns REFUSED for `binary == 'main'` and names `tools/gate_main.py`. The main lane was then reopened
|
||||
properly the next morning and banked 5 (4 after the source-truth correction below), with a bisect
|
||||
isolating the one bad draft in 7 rebuilds.
|
||||
|
||||
**Hindsight / better path.** Three things would each have caught it earlier, in increasing order of
|
||||
generality: (a) run R22 **before** committing a gate against a binary the lane has not gated before,
|
||||
not at session close; (b) `gate_main`'s own bank count was also derived rather than measured — it
|
||||
printed "BANKED 5 of 6" when 4 had applied, because `len(good)` is *what we decided to keep*, not
|
||||
*what was substituted* — so **count from the source in every gating tool**; (c) the general rule this
|
||||
session kept re-teaching: **a tool that wraps another tool inherits its refusals**, and the place to
|
||||
encode that is a refusal in the wrapper. Every constraint documented on `gate_stage` binds
|
||||
`parallel_gate`, `harvest_verify`, and anything else that shells it.
|
||||
|
||||
**Cost of learning it here:** one bad commit, ~40 minutes of revert-and-bisect, and an inflated bank
|
||||
count I had already reported to Drew and had to correct. Cheap only because R22 exists and was run.
|
||||
|
||||
---
|
||||
|
||||
## 2026-09-02 (P31 S71) — the drafting pool ran dry, and the lever was an exclude list nobody re-probed
|
||||
|
||||
**Context and belief.** With ~40 agents landing at near-100% MATCH, the working assumption was that
|
||||
drafting capacity was the constraint and the campaign would continue as draw → draft → gate until the
|
||||
frontier was gone.
|
||||
|
||||
**What failed.** Wave 3 drew **1 target** and reported *"0 left in pool"*. Measured at that moment:
|
||||
174 open, of which `main` 64, and of the 110 non-main — **41 drafted this session, 68 on the exclude
|
||||
list, 2 proven walls, ZERO genuinely undrawn**. More agents would have had nothing to work on.
|
||||
|
||||
**The pivot.** The 68 excluded functions were excluded because the TOOLING could not carve them —
|
||||
`jtbl_carve` refused their plans with *"subseg would host NON-CONTIGUOUS `.rodata` carves"*. But
|
||||
tooling had changed **that same session**: `jr_isolate_all` had been fixed twice (file-local `static`
|
||||
placement, and §323's `__attribute__`-blind regex). Re-probing all 68 found **17 now reporting `tail`
|
||||
— a standard §8a carve**. Every one already had drafts on disk; scoring them put **10 at closeness 0
|
||||
for zero drafting**, and the gate banked 9 — four of them in **57 seconds**.
|
||||
|
||||
**The grounded why.** An exclude list is a snapshot of *what the tooling could not do at the moment it
|
||||
was written*. It is treated thereafter as a property of the FUNCTIONS. Nothing in the pipeline
|
||||
re-examines it, so every tool improvement leaves behind a population that is now tractable and still
|
||||
marked impossible — invisible, because the draw filters it out before anything measures it.
|
||||
|
||||
**Hindsight / better path.** **Re-probe the exclude list after every tool fix, as part of the fix.**
|
||||
The probe is deterministic, costs no agents, and here it was worth more than the entire drafting lane
|
||||
at that moment. Generalised: *any list that records a tool's limitation must be regenerated when the
|
||||
tool changes, or it silently becomes a list of work you have decided not to do.* The same reasoning
|
||||
applies to `.run/S71_walls_found.txt` — a wall proven against today's compiler knowledge is not a wall
|
||||
forever, and each entry should carry the refutation list that would have to be beaten.
|
||||
|
||||
@@ -292,6 +292,19 @@ Notably variant 3 (`conflicting types`, RETURN type only, decl already `()`, sym
|
||||
is fixed by `--sync-decls` ALONE: both other levers no-op, and the sync is safe precisely because an
|
||||
address-taken site has no arguments to convert.
|
||||
|
||||
### 4c. LAUNCH-TIME OPEN CHECK — `wave_args` asserts at DRAW time, and payloads go stale
|
||||
|
||||
```
|
||||
python3 tools/launch_check.py --payload .run/<wave>/wf_args.json # filters in place
|
||||
python3 tools/launch_check.py <binary> <fn> # exit 2 = already banked
|
||||
```
|
||||
|
||||
A wave's payload sits on disk while gates run, so by launch time some of its targets are banked. An
|
||||
agent handed one burns a full run to report "STALE CARD — already banked today", with no `.s` left to
|
||||
score against. Measured S71: `ov_SC01_006/func_8017F9F8` did exactly that, and filtering the wave-2
|
||||
payload found **3 of 27** already banked. Also skip any target that already has a FRESH draft from
|
||||
this session — it needs a gate, not another agent.
|
||||
|
||||
## 5. Draft
|
||||
|
||||
`tools/workflows/claude_wave_draft.js`, `args = {wave, targets}`. One agent per target,
|
||||
@@ -360,6 +373,17 @@ python3 tools/gate_wave.py --drafts <dir> --workers 8 --commit [--r22]
|
||||
about an hour for what should have taken minutes. **`gate_wave.py`'s split is now an optimisation
|
||||
(same-binary drafts share a build), NOT a safety requirement.**
|
||||
|
||||
### 6b. READ THE VERDICTS — `tools/gate_triage.py`
|
||||
|
||||
```
|
||||
python3 tools/gate_triage.py --plan <gate_plan.json>
|
||||
```
|
||||
|
||||
Routes every verdict to the lane it names and asserts the staged denominator: CARVE (probe it —
|
||||
`jr_isolate_all` is usually the unblock) · UNDEF-D (§171 `aprop_symfix`) · CONFLICT/ARITY (§376/§378)
|
||||
· PARSE · NO-DIAG · DIFF (real codegen). S71's census over 37 verdicts: DIFF 18 · CARVE 7 · PARSE 3 ·
|
||||
NO-DIAG 3 · CONFLICT 2 · ARITY 2 · UNDEF 2 — which corrected an impression that carve dominated.
|
||||
|
||||
## 7. After ANY bank
|
||||
|
||||
```
|
||||
|
||||
@@ -5569,6 +5569,13 @@ Every one is fuel keyed by BARE NAME or asserted without a freshness check (R48/
|
||||
time last night) — §376 in its purest form. Do not re-slate them without a TU-level fix.
|
||||
**No new waves from here (Drew).**
|
||||
|
||||
- 2026-09-02 — **S71 gate ledger, all 14 cycles** (each `parallel_gate --commit` unless noted; every
|
||||
bank R22-verified at close): g1 12 (integration pile, main's 11 later REVERTED — §414) · g2 0 ·
|
||||
g3 9 · g4 6 · g5 5 · g6 2 · g7 4 · g8 3 · g9 2 · g10 2 · g11 3 · g12 2 · g13 2 · g14 4 · g15 2 ·
|
||||
**carve gate 5** (the five jr-isolated overlays) · **§420 rebase gate 4** (one body, four overlays,
|
||||
57 s) · **`gate_main` 4 + 1** (the only tool that may gate main).
|
||||
Net: **210 → 147 = 63 banked**, `check-all: 213 passed, 0 failed of 213`.
|
||||
|
||||
## 🛑 SESSION CHECKPOINT — S71 CLOSE (2026-09-02 11:12). SUPERSEDES every earlier block in this file. Phase 31 T10 CONTINUES.
|
||||
|
||||
**FLEET VERIFIED GREEN FROM A CLEAN REBUILD — `check-all: 213 passed, 0 failed of 213`**
|
||||
|
||||
Reference in New Issue
Block a user