From 8a698e1b79d72573f8692bf2b152df6302cb85b0 Mon Sep 17 00:00:00 2001 From: Drew T <50529377+Druthulu@users.noreply.github.com> Date: Mon, 7 Sep 2026 23:24:14 -0600 Subject: [PATCH] =?UTF-8?q?kit+docs(phase-33.5):=20task=2014.5=20part=202?= =?UTF-8?q?=20=E2=80=94=20the=20phase-log=20mining=20pass=20banked:=20DK-6?= =?UTF-8?q?9=E2=80=93DK-80=20from=20143=20worklog=20lessons;=20the=20S92?= =?UTF-8?q?=20accelerators=20+=20decision-log=20entries?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - 21 read-only Opus slices over the 26 worklogs (30,540 lines read; the three giants by line range), briefed by .run/P33.5/log-mining/BRIEF.md with the "already banked?" grep protocol: 777 candidates, 634 already banked, 143 NEW; every cited log line verified to exist by harvest.py and read; 142 banked + 1 dropped (the miner's own low-value verdict) - decomp-kernels.md: twelve kernels DK-69–DK-80 (instrument blind spots; verdict staleness + the health suite; what earns belief; denominators/units/labels; leverage vs tractability + campaign scoping; models and prompts; the unattended run; agents and the tree; edits that keep proofs; the search harness + the compiler as evidence; maintaining the knowledge base; hosts and services), each provenance line generated from the slices' cited worklog lines (bank.py); section 8 retitled; Coverage "In all: DK-1 … DK-80"; every count mention → 80 (kit README, methodology, SETUP row, wiki page) - docs/accelerators.md "P33.5 S92" (two accelerators: a distillation ships with a coverage check; the end-of-project worklog pass recovers the fixed-but-never-generalised lessons — 1 in 5 here) → cited by the twelve kernels (kit_coverage: 59 entries, 42 cited + 15 dispositioned, 0 UNCOVERED; the matcher now requires an un-numbered entry's FULL group text); docs/decision-log.md "P33.5 S92" (R31: the question, the measurement, the pivot, the 82%/18% why, the hindsight path) - evidence tracked under .run/P33.5/log-mining/ (BRIEF, 21 slice reports, HARVEST.md, HARVEST_TABLE.md, harvest.py, bank.py; a dated .gitignore block); expected-manifest +1 (docs/inherited-record.md); make kit-corpus regenerated the record copies - verify: tool_census --check OK (358 copies); kit_coverage OK; kit_lint OK (0 leaks over 436 files); doc_links --strict OK; wiki_render --selftest 32/0; audit_public OK on the new files; 80 DK ids cited, 0 dangling --- .gitignore | 7 + .run/P33.5/kit-dryrun/expected-manifest.txt | 1 + .run/P33.5/log-mining/BRIEF.md | 69 ++++ .run/P33.5/log-mining/HARVEST.md | 330 ++++++++++++++++++ .run/P33.5/log-mining/HARVEST_TABLE.md | 147 ++++++++ .run/P33.5/log-mining/Phase15-16.md | 104 ++++++ .run/P33.5/log-mining/Phase17-18.md | 102 ++++++ .run/P33.5/log-mining/Phase19-20-22.md | 63 ++++ .run/P33.5/log-mining/Phase21.md | 84 +++++ .run/P33.5/log-mining/Phase23-27.md | 80 +++++ .run/P33.5/log-mining/Phase24.md | 66 ++++ .run/P33.5/log-mining/Phase25.md | 100 ++++++ .run/P33.5/log-mining/Phase26.md | 79 +++++ .run/P33.5/log-mining/Phase28-32.md | 86 +++++ .run/P33.5/log-mining/Phase29-1of4.md | 159 +++++++++ .run/P33.5/log-mining/Phase29-2of4.md | 194 ++++++++++ .run/P33.5/log-mining/Phase29-3of4.md | 111 ++++++ .run/P33.5/log-mining/Phase29-4of4.md | 88 +++++ .run/P33.5/log-mining/Phase30-1of2.md | 138 ++++++++ .run/P33.5/log-mining/Phase30-2of2.md | 90 +++++ .run/P33.5/log-mining/Phase31-1of3.md | 135 +++++++ .run/P33.5/log-mining/Phase31-2of3.md | 166 +++++++++ .run/P33.5/log-mining/Phase31-3of3.md | 109 ++++++ .run/P33.5/log-mining/Phase33.md | 280 +++++++++++++++ .run/P33.5/log-mining/Phase7.md | 70 ++++ .run/P33.5/log-mining/Phase8-13.md | 158 +++++++++ .run/P33.5/log-mining/bank.py | 319 +++++++++++++++++ .run/P33.5/log-mining/harvest.py | 40 +++ decomp-architect/README.md | 2 +- decomp-architect/corpus/decomp-kernels.md | 248 ++++++++++++- .../corpus/record/docs/accelerators.md | 26 ++ .../corpus/record/docs/decision-log.md | 28 ++ .../corpus/tools/P9/kit_coverage.py | 2 +- decomp-architect/decomp-architect.md | 2 +- docs/SETUP.md | 2 +- docs/accelerators.md | 26 ++ docs/decision-log.md | 28 ++ docs/wiki/Start-a-new-decomp-project.md | 7 +- tools/kit_coverage.py | 2 +- 39 files changed, 3737 insertions(+), 11 deletions(-) create mode 100644 .run/P33.5/log-mining/BRIEF.md create mode 100644 .run/P33.5/log-mining/HARVEST.md create mode 100644 .run/P33.5/log-mining/HARVEST_TABLE.md create mode 100644 .run/P33.5/log-mining/Phase15-16.md create mode 100644 .run/P33.5/log-mining/Phase17-18.md create mode 100644 .run/P33.5/log-mining/Phase19-20-22.md create mode 100644 .run/P33.5/log-mining/Phase21.md create mode 100644 .run/P33.5/log-mining/Phase23-27.md create mode 100644 .run/P33.5/log-mining/Phase24.md create mode 100644 .run/P33.5/log-mining/Phase25.md create mode 100644 .run/P33.5/log-mining/Phase26.md create mode 100644 .run/P33.5/log-mining/Phase28-32.md create mode 100644 .run/P33.5/log-mining/Phase29-1of4.md create mode 100644 .run/P33.5/log-mining/Phase29-2of4.md create mode 100644 .run/P33.5/log-mining/Phase29-3of4.md create mode 100644 .run/P33.5/log-mining/Phase29-4of4.md create mode 100644 .run/P33.5/log-mining/Phase30-1of2.md create mode 100644 .run/P33.5/log-mining/Phase30-2of2.md create mode 100644 .run/P33.5/log-mining/Phase31-1of3.md create mode 100644 .run/P33.5/log-mining/Phase31-2of3.md create mode 100644 .run/P33.5/log-mining/Phase31-3of3.md create mode 100644 .run/P33.5/log-mining/Phase33.md create mode 100644 .run/P33.5/log-mining/Phase7.md create mode 100644 .run/P33.5/log-mining/Phase8-13.md create mode 100644 .run/P33.5/log-mining/bank.py create mode 100644 .run/P33.5/log-mining/harvest.py diff --git a/.gitignore b/.gitignore index 130cad0ce2..b6de55fcd4 100644 --- a/.gitignore +++ b/.gitignore @@ -385,6 +385,13 @@ unsloth_compiled_cache/ !/.run/P33.5/kit-dryrun/*.json !/.run/P33.5/kit-dryrun/*.py +# P33.5 task 14.5 (S92, 2026-09-07): the phase-log mining pass — the brief, the 21 slice reports, the harvest and its table, +# the two scripts. Evidence of what the kit's kernels DK-69–DK-80 were distilled from; quotes of our own worklogs, nothing ROM-derived. +!/.run/P33.5/log-mining/ +/.run/P33.5/log-mining/* +!/.run/P33.5/log-mining/*.md +!/.run/P33.5/log-mining/*.py + # P33.5 task 13.5 (S91, 2026-09-07): the tools-audit evidence — the survey inputs, the review verdicts, the need-keys, the layout # probe and the run-4 judgment. Evidence, tracked. !/.run/P33.5/tools_audit/ diff --git a/.run/P33.5/kit-dryrun/expected-manifest.txt b/.run/P33.5/kit-dryrun/expected-manifest.txt index d0cde73f01..090aa2a152 100644 --- a/.run/P33.5/kit-dryrun/expected-manifest.txt +++ b/.run/P33.5/kit-dryrun/expected-manifest.txt @@ -42,6 +42,7 @@ CREATED config/mcp.json.template CREATED docs/README.md CREATED docs/decomp-architect.md CREATED docs/decomp-kernels.md +CREATED docs/inherited-record.md CREATED docs/knowledge-corpus.md CREATED docs/tools-manifest.md CREATED docs/wave-playbook.md diff --git a/.run/P33.5/log-mining/BRIEF.md b/.run/P33.5/log-mining/BRIEF.md new file mode 100644 index 0000000000..169f44de9e --- /dev/null +++ b/.run/P33.5/log-mining/BRIEF.md @@ -0,0 +1,69 @@ +# BRIEF — the phase-log mining pass (Phase 33.5, task 14.5, item 3) + +You are a READ-ONLY subagent on one bounded task. Do NOT run the repository's session-start protocol; do NOT read +PROJECT_CONTEXT.md, phase-ends/DIGEST.md or the PhaseEnd files as a preamble. Your whole job is below. + +## The question + +The project's day-one kit (`decomp-architect/`) claims to hold the whole of the project's hindsight. Its lessons were taken from +the distilled records — the cookbook, the decision log, the accelerators ledger, the retrospective, the how-to chapters — and +NOT from the raw phase worklogs under `phase-ends/logs/`. Your slice is one of those worklogs (or a line range of one, or a few +small ones). Find every lesson in it that a FUTURE decompilation project should know on day one **and that is banked NOWHERE +in the distilled records**. Only those. A lesson that is already recorded anywhere is a duplicate, not a finding. + +## What counts as a lesson + +A generalisable "do this sooner / never do this / the instrument was the wall / this is the mechanism" — the kind of thing the +accelerators ledger (`docs/accelerators.md`) and the kernels (`decomp-architect/corpus/decomp-kernels.md`) record. Not: a +function-specific matching idiom (those belong in the cookbook and almost all are there), a status line, a number, a task +ticked, a pivot already in the decision log. Prefer lessons with a COST attached (what it cost to learn late) and a +portable shape (true for another console / another compiler / another agent harness). + +## The "already banked?" test — mandatory, recorded per candidate + +For every candidate, run 2–3 greps with DISTINCT phrasings of its key terms (case-insensitive, `grep -n -i`) over ALL of: +``` +docs/matching-cookbook.md docs/cookbook-index.md docs/decision-log.md docs/accelerators.md docs/retrospective.md +docs/wave-playbook.md docs/how-to-ai-decomp/*.md decomp-architect/corpus/decomp-kernels.md +decomp-architect/templates/registry-E.decomp.md phase-ends/DIGEST.md +``` +(`docs/matching-cookbook.md` is 3.5 MB — grep it, never read it.) Record the exact grep commands and their hit counts. If any +hit records the same lesson (read the hit's surrounding lines to judge), the candidate is ALREADY-BANKED: list it in one line +with `file:line` of where it lives, and move on. If the greps find nothing that records it, it is NEW. + +## Reading your slice + +Read your file(s) with the Read tool in pages (offset/limit of ~400 lines); very long lines are normal. Read ALL of your +assigned range — your report states the lines read as a denominator. Do not skim; the lessons hide in the middle of task +entries ("what went wrong", "deviation", "finding", "lesson", "sooner", "cost", "wrong", "false", "instrument", "wall"). + +## The deliverable — write it EARLY, append as you go + +Create your output file FIRST (header + empty sections), then append each candidate as you find it. Path and schema: + +``` +# Log mining — +Files/ranges: :- … · Lines read: N of N +Candidates considered: N · NEW: N · ALREADY-BANKED: N + +## NEW +### C1 — +- **Evidence:** `:` — a quotation of at most 3 lines +- **What happened / what it cost:** +- **Not banked — greps:** `grep -n -i '' ` → 0; `grep -n -i '' …` → 0; `grep -n -i '' …` → N (hits are about ) +- **Proposed home:** DK (a kernel) | G (a rule) | accelerator | cookbook | decision-log | NOT-PORTABLE (this project only — say why) +- **Portable because:** + +## ALREADY-BANKED (one line each) +- — lives at `:` +``` + +## Hard constraints + +- Read-only except your ONE output file under `.run/P33.5/log-mining/`. No git commands. Nothing under `/tmp` or `~/.claude`. + Do not edit any repository file. Do not create other files. +- Quote line numbers from the Read tool; a lesson without an evidence line is not a finding. +- Do not invent lessons the log does not contain, and do not "improve" a banked lesson into a new one. +- Your FINAL message is exactly one JSON line: `{"slice": "...", "file": ".run/P33.5/log-mining/.md", "lines_read": N, "considered": N, "new": N, "banked": N}` +- Budget: finish within your context; if you are running out, write what you have (the file is the deliverable) and end + with the JSON line marking `"partial": true` and the last line read. diff --git a/.run/P33.5/log-mining/HARVEST.md b/.run/P33.5/log-mining/HARVEST.md new file mode 100644 index 0000000000..89d5ef44ec --- /dev/null +++ b/.run/P33.5/log-mining/HARVEST.md @@ -0,0 +1,330 @@ + +=================== Phase15-16: considered 28 · NEW 8 (found 8) · banked 19 · lines read 167 of 167 + C1 — Measure a pipeline's yield on the RESIDUAL, not on already-solved functions: build a known-answer ladder (revert a matched function to a stub and make the pipeline re-derive it), then read the ga + LOG phase-ends/logs/Phase16.md:36: - **Known-answer ladder (Drew's method), `tools/p16_known_answer.py`:** on 12 already-matched fns (known-reachable answers), **m2c-DIRECT re-derivation = 8/12 = 67%** (macro-only, NO permuter, NO struct types). Remainder: 2 near-misses @15 mismatch (permuter), + C2 — A search harness (permuter) must compile in the SAME declaration context as the real build; an isolated context does not merely fail to verify, it makes the search converge on the wrong answer + LOG phase-ends/logs/Phase16.md:71: ## CRITICAL FINDING (Fri 2026-06-19) — the extern-context bug (byte-gate caught a false 42%) - Overnight permuter "closed" 17/40 near-misses (42%) — BUT **0/17 whole-binary-gated.** Root cause: the permuter's `base.c` STRIPPED callee externs → compiled with im + C3 — Before concluding a drafting pipeline is weak, histogram the compiler's error TEXT: over half of ours were one missing declaration in the shared header, fixed byte-neutrally in one line + LOG phase-ends/logs/Phase16.md:40: - **393 m2c-target characterization (macro-only, match_one):** 70 direct-MATCH / 106 CC1-fail / 217 near-miss. **CC1-fail breakdown: 58 = `NULL` undeclared (TRIVIAL fix — add to common.h), ~15 stack-struct (sp* vars → --stack-structs), ~10 m2c-incomplete.** + C4 — Re-run the deterministic declaration-canonicaliser over OLD quarantined drafts after every large bank: a draft's recovery odds rise as the banked corpus grows, and the pile costs nothing to keep + LOG phase-ends/logs/Phase15.md:57: - 2026-06-17 — **T6 tail: quarantine recovery — fleet 44.64% → 46.27% (+1.63).** `canon_draft_decls` on the combined 505 quarantined (now canonicalizing against the much-larger banked set: 268 decls rewritten) → **47 verified** (1175→1128 stubs) → **42 propaga + C5 — Write numeric kill-criteria into the plan, before the data exists, for any expensive campaign — and let them fire + LOG phase-ends/logs/Phase16.md:24: - [ ] **S3** 10-medium validation (Max) → **GATE-B:** ≥6/10 byte-identical; record permuter yield (<4/10 → STOP run) + C6 — Design an unattended run for a human with no agent session: a STOP file honoured at a safe boundary, a supervisor that tells a clean exit from a crash, a status one-liner — and prove crash-resume + LOG phase-ends/logs/Phase16.md:89: - **Safe-exit mechanism (REQUIRED).** Driver checks `.run/auto/STOP` at every function boundary; if present → finish current fn's gate+propagate+commit → final heartbeat "stopped safely" → exit 0; supervisor sees STOP + clean exit → does NOT relaunch. **Trigge + C7 — The permuter's search unit is the C expression: it cannot freeze the instructions you already have right, because register allocation couples them + LOG phase-ends/logs/Phase16.md:43: - **Drew Q&A (permuter):** not one-shot-from-nothing (m2c draft = the info/jumping-off point); function-by-function not whole-file; permuter hill-climbs (additive) but can't freeze individual instructions (regalloc couples them) — compositional fixes happen at + C8 — An idempotency guard keyed on PRESENCE freezes every record created before the system matured; key it on COMPLETENESS, or re-audit records written before the last capability jump + LOG phase-ends/logs/Phase15.md:51: - Minor: the 4 Phase-13/T2 proof groups (E_func_80128EA8/8012A568/80132EC4/80138C30) are at 16 members not 134 (registered before the bulk, so --auto-from skips them) — negligible; top up later if desired. + +=================== Phase17-18: considered 32 · NEW 6 (found 6) · banked 26 · lines read 538 of 538 + C1 — Leverage and tractability are ANTI-correlated: the most-duplicated functions are systematically the hardest, so a leverage-first queue front-loads hand-tier work and its early bank rate is not a + LOG phase-ends/logs/Phase17.md:247: goto). **Selection lesson:** low-m2c-mismatch struct = the quirk tail; clean closes = fnptr/relocs-low/ m2c-mis-structured. **Sizing (§6):** tractable easy classes are LOW-reach; ×134 yield is in quirk-heavy STRUCTURAL_MISS(7.6%)/PERMUTER_CLASS(3.6 + C2 — Another decomp project's `INCLUDE_ASM` is a record of what they did not crack, never proof that a class is uncrackable; cross-project corroboration multiplies confidence in a WRONG verdict as rea + LOG phase-ends/logs/Phase18.md:52: - [x] **T2 — Xenogears mine ✓ 2026-06-20** (2 bg agents; full synthesis `.run/p18/T2_T3_synthesis.md`). DECISIVE: Xenogears (independent decomp, IDENTICAL gcc-2.7.2-psx -O2) has **NO C lever** for the call-crossing $s0/$s1 ORDER class — no `register`, no a + C3 — Your own corpus of byte-matches is an experiment you have already run on the toolchain: settle "is my rebuilt compiler faithful to the original?" from it before installing the original vendor too + LOG phase-ends/logs/Phase18.md:64: invoked: Wine is a heavy install on this WSL (106 pkgs + i386 arch not enabled + wineprefix) and the divergence question is already closed by stronger evidence: (1) **gcc-2.7.2-psx byte-matches ~700 functions**, many with call-crossing callee-saved value + C4 — Accelerate the stage that is actually the bottleneck: a byte-exact search loop costs a compiler invocation plus a whole-binary gate per candidate, so GPU/ML brute force buys nothing + LOG phase-ends/logs/Phase17.md:96: - **CUDA/ML brute-force:** CUDA can't run gcc-2.7.2; the bottleneck is the gcc compile + the whole-binary gate, NOT search speed, so raw GPU permutation doesn't apply. The real ML angle = a *learned gcc-2.7.2 codegen predictor/ranker* (Drew has CUDA + Ligh + C5 — A contradiction between two entries of your own knowledge base is a work item, not noise: replay the levers you have already written down, under the correct oracle, before commissioning any new r + LOG phase-ends/logs/Phase18.md:26: ~443–571). Phase 17 used the **wrong oracle** (the permuter's floor-polluted score) — §10:518 says use the object-level metric (`match_one.py`), which was **never applied** to the exemplars. §16 ("not source-steerable") contradicts §10 — Phase 18 reconciles it + C6 — Scope a research phase by a PER-CLASS VERDICT, not by a percentage: the deliverable is a validated lever or an honest, falsifiable wall verdict for each class + LOG phase-ends/logs/Phase18.md:93: ## Milestone (knowledge-gated, NOT a fleet-% target) Per-class byte-gated verdict (validated C idiom proven on a known-answer exemplar via `p16_known_answer --gate`, OR honest "unsteerable" verdict naming the exact gcc pass); §10-vs-§16 + +=================== Phase19-20-22: considered 36 · NEW 2 (found 2) · banked 34 · lines read 187 of 187 + C1 — An open-ended grind phase's milestone is invariants held plus a clean checkpoint, never a percentage target + LOG phase-ends/logs/Phase19.md:24: -O0 lever banked + recovery tooling proven + toolkit scaled across 2–3 batches of 50, fleet up materially (target +3–5%, ~56.6% → ~60%), 136/136 byte-identical from clean (R22), 0 NON_MATCHING (G4), closed at a clean checkpoint with a Phase-20 backlog. + C2 — Choose the exemplar for cracking a codegen class by the SIZE of its residual: the rows at a one-instruction decompiler-vs-target mismatch are the cleanest real-function isolates of the class + LOG phase-ends/logs/Phase20.md:14: `tools/exemplar_miner.py` → `docs/exemplar_curriculum.md` + `.run/exemplar_routing.json` (reach computed from sigs like dedup_propagate). **835 residual stubs routed:** WAVE 472 (218 reach-134, T6 fuel) · STRUCT 159 · **PINS 114 (45 reach-134)** · STUB 9 + +=================== Phase21: considered 27 · NEW 3 (found 3) · banked 24 · lines read 799 of 799 + C1 — An unattended agent run that looks throttled is usually blocked on an interactive approval prompt; check the pending prompt before diagnosing the provider + LOG phase-ends/logs/Phase21.md:731: - **Wave-2 "6.5h" was NOT throttling** — it was idle on a CC **permission prompt** (Drew approved on check-in); waves 1 & 3 ran in ~25–32 min. (R14: corrected my earlier rate-limit read.) + C2 — A class-distribution assessor only sees the population that has already been attempted; "analyse ALL remaining work" is a cheap triage pass, not a static analysis + LOG phase-ends/logs/Phase21.md:63: **THE ONE REAL GAP (the refinement worth making): `--assess` clusters the BACKLOG (already-DRAFTED near-misses), not ALL remaining stubs.** A function's "decomp issue" is only known AFTER a draft attempt (the residual = the class). So "analyze ALL remaining" = + C3 — Say which currency a wave buys — percentage or idioms — before launching it, and judge it in that currency + LOG phase-ends/logs/Phase21.md:475: 1. **Waves 17–29 — reach-1 (×1) smallest-first harvest:** ~183 banks but fleet only **+0.06%** — reach-1 fns are overlay-unique → ×1, and the fleet metric counts functions across all 136 binaries. Band climbed 14→75 ins, close-rate fell 0.9→0.33. **LESSON: rea + +=================== Phase23-27: considered 22 · NEW 8 (found 8) · banked 14 · lines read 172 of 172 + C1 — Mine new idioms from FRESH cracks, never from the failed backlog: your failure pile only re-teaches you what you already know + LOG phase-ends/logs/Phase23.md:93: - 2026-07-01 (**evening — T10.6 ≤15 campaign + T10.7/8/9 the OpenRouter/GLM5.2 frontier-model exploration**): CONTEXT CHECKPOINT (85% ctx, wind-down). **T10.6 DONE:** `bulk_harvest.py` ≤15 saturation campaign ran 23 cycles → **1,297 banks, fleet 64.6%, ~92% of + C2 — Give a drafting model a TARGETED slice of the knowledge base, never the whole thing: full context measurably made the model worse + LOG phase-ends/logs/Phase23.md:24: - [x] **T2 — Stock-local floor** — Qwen3.6-35B-A3B (LM Studio): structurally smart but **0 reliable byte-matches** (can't refine to byte-exact); format-robust; full-cookbook context made it *worse* (dilution). The floor to beat. + C3 — Name the drafter's degenerate output in the prompt: an empty body compiles, so "translate EVERY instruction, never an empty body" is a required instruction + LOG phase-ends/logs/Phase23.md:91: - 2026-06-30 (**8-hour autonomous run, Drew away**): prompt-fix + local serving + corpus-v3 + v3 + big batch. **LM Studio ejected** → built **`tools/serve_local.py`** (Unsloth GPU serving as an OpenAI endpoint; the prebuilt llama-cpp-python CUDA wheel SIGILLs + C4 — The output-token cap is a two-sided knob, and both failure modes read as "the model is bad" + LOG phase-ends/logs/Phase23.md:54: **FIXES ALREADY APPLIED THIS SESSION (don't redo):** `api_draft` max_tokens 4096→**512** (killed the no-stop-token ramble, ~80–130s→~10s on the bad cases); `lora_grind` runs `progress.py --fleet` **only on propagate sweeps** (was every batch = ~14s overhead); + C5 — Build the unattended campaign so a crash is a pause: probe the dependency at the top of each cycle, commit per cycle, persist the tried-set + LOG phase-ends/logs/Phase23.md:92: - 2026-07-01 (**T10 — the phase-separated + parallel-gate harvester, BUILT + MEASURED**): plan-mode Tier-1, harvester-first (Drew picked it over vLLM-first — the exploration showed parallel-gate is the bigger single lever, no install, and unblocks the measurem + C6 — Prove the aggregate check target is fail-closed before adding checks to it, and give every audit oracle a dependent + LOG phase-ends/logs/Phase27.md:17: 2. **`make report` is NOT fail-closed — roadmap §5 asserts it is.** `Makefile:9-10` sets `.ONESHELL` with **no `-e`** in `.SHELLFLAGS` (verified via `make -p`: `.SHELLFLAGS := -c`), so the recipe is one `bash -c` and only the **last** command's exit survives. + C7 — A document that cites a repository path is an untested claim about the repository; lint it + LOG phase-ends/logs/Phase27.md:71: - **2026-07-15 · Task 3 (addendum) — the allowlist was still too narrow; cookbook §45 cites untracked files.** While reading the seeds for Task 1 I hit a real defect: **cookbook §45 names `.run/giants/func_80133CD4.fable.c` as its worked example and `.run/gian + C8 — Assert that the work ledger PARTITIONS the live work, and treat a row the invariant refutes as a lie a fresh session will act on + LOG phase-ends/logs/Phase27.md:45: - [x] **Task 8 — The byte-gate-honest re-scan + partition + ledger rebuild** `[Max]` — **DONE (deterministic core; the 1,670-triage scoped to P29 — see below).** Built **`worklist --assert-partition`** (R32, the audit's literal prescription — enumerate live st + +=================== Phase24: considered 35 · NEW 4 (found 4) · banked 31 · lines read 138 of 138 + C1 — Escalations to the expensive tier run STRICTLY SERIAL with idiom-banking between them; only the tier that cannot learn is run in parallel + LOG phase-ends/logs/Phase24.md:57: - **T7 (resumed, Drew's choice) — sibling-giant harvest via parallel Opus-Max agents (2026-07-03, IN PROGRESS):** launched 5 Opus agents (one per remaining region-a giant), each applying §32/§34, escalating to Fable5 only on stall. **KEY FINDING (answers the f + C2 — Inside one leverage class, schedule by MEASURED remaining effort — and pull the payoff-dominating outlier out of that queue for an immediate cheap triage + LOG phase-ends/logs/Phase24.md:29: > The remaining **~9 reach-134 giants (≥150 ins)** to harvest next session (12 reach-134 giants were stubbed post-T7; **this session's batch = func_801392FC/8013A530/8013AF20**, banked HERE first). **Strategy: closest-to-completion FIRST + fast-track the whale + C3 — The wall-breaker tier is a MATCH tier, not a PLUMBING tier: work that a deterministic arbiter can judge does not need the expensive model + LOG phase-ends/logs/Phase24.md:61: - **T7 — giant `func_80129CF8` CRACKED (Fable5) + §32 idiom + banked ×1; ×134 WALL found → Drew chose to build the canonical-decl tool (2026-07-03, `8b48f482f`):** Fable5Max cracked the 191-ins region-a camera giant (match_one MATCH) — struct-base hoisting via + C4 — When a byte-proven body will not integrate, the repair moves the CALLER's declaration to the definition's signature — never the definition to the caller's — and only where the change is width-com + LOG phase-ends/logs/Phase24.md:119: - **Straggler/caller reconcile is byte-neutral ONLY on the caller's extern** (width-compatible types; the call site casts/passes-wide). NEVER touch the matched def. + +=================== Phase25: considered 34 · NEW 3 (found 3) · banked 31 · lines read 561 of 561 + C1 — Leverage ordering does not predict tractability: carry a MEASURED closeness read per target, and never let (reach × size) stand in for "crackable" + LOG phase-ends/logs/Phase25.md:473: - **Key findings (frontier data → T6):** (1) **byte-weight ≠ tractability** — `func_8014D3E0` (22 ins × 1997) looked like the mega-ROI freebie but is a `$sp` stack-switcher, matched only by porting an already-matched sibling (`func_8014D04C`); the 369-gian + C2 — A fan-out script generated by the orchestrator runs sandboxed with NO access to the repo: it must be self-contained, so target selection belongs to the wave generator, not the workers + LOG phase-ends/logs/Phase25.md:487: - **Generated a §12-robust scale-up wave** (`tools/workflows/t5_scaleup.js`: worker_wave's drafter prompt + sequential waves-of-10 + retry×2, targets embedded — sandbox can't read files). Launched over the 83 ≤149-ins targets. + C3 — Run reference-compiler dump/inspection passes from a scratch CWD (or set an explicit dump base); a `-da` run with CWD at the repo root silently deposits RTL dumps that later read as committed art + LOG phase-ends/logs/Phase25.md:337: - Housekeeping: `gccdump.lreg` (repo root) = gcc's DEFAULT RTL dump (dump-base "gccdump", `.lreg` = local-reg pass; `toplev.c:1973/2077`), left by a one-off `cc1 -da` RTL-inspection run with CWD=root — **not** any committed tool (grep hits only the gcc sou + +=================== Phase26: considered 39 · NEW 3 (found 3) · banked 36 · lines read 1106 of 1106 + C1 — A round-trip selftest is a serialisation check, not a coverage check: it passes by construction when a parser's missed item is absorbed into its neighbour's span. Every partition/rewrite tool nee + LOG phase-ends/logs/Phase26.md:764: FIRST token and returned at the first depth-0 `;`, so the def after it was never anchored — absorbed into the next anchor's preamble. **The selftest was structurally blind** (round-trip = `"\n".join(item_texts)` stays exact by construction when a miss la + C2 — A set that gates work must be reconstructible from committed artifacts. A roster kept in gitignored scratch is an unversioned oracle: it is silently wrong for anything it was not named after, and + LOG phase-ends/logs/Phase26.md:742: R33 applied to "the purest R33 case in the group" (audit). `jr_inventory`'s `banked` set was filtered by an **EPHEMERAL, gitignored `.run/banked_func_*.json` roster** — `rm -rf .run`/a fresh clone would blind ALL banked jr at once, cross-address siblings + C3 — A guard that is allowed to sit RED and UNWIRED does not exist. A detector's value is zero until it is green on HEAD and called by the standing report — and a docstring claiming it is wired is not + LOG phase-ends/logs/Phase26.md:101: - [ ] **A9 — `lint_symbol_refs`: a guard allowed to sit red does not exist** `[xHigh]` *(blocked on A2)* — currently **RED** (43 false positives) and **UNWIRED** (`make report` never calls it, though its docstring claims it does). No `__asm__("label")` model; + +=================== Phase28-32: considered 32 · NEW 7 (found 7) · banked 25 · lines read 480 of 480 + C1 — Never round-trip a curated config through a serializer: every oracle you own measures BYTES, so a formatting-destructive write is invisible to all of them + LOG phase-ends/logs/Phase28.md:181: - **❌ SELF-INFLICTED #1 — I DESTROYED the registry's documentation (H5), and every gate called it green.** My first cut wrote the registry with `yaml.safe_dump`, which round-tripped the whole file: **47 comment lines → 0** (including the curated Phase-11 hea + C2 — The file a function lives in is not evidence of its class: read the recorded attribute, never the hosting split + LOG phase-ends/logs/Phase28.md:182: - **❌ SELF-INFLICTED #2 — I mis-reported the DIFFs, twice (R14).** (a) I claimed ov_SC07_006's 71 non-banks were *"ALL PLUMBING, ZERO DIFF"* — from reading `head -6` of the classified file and generalizing. It has the same 4 DIFFs; the claim is **false and i + C3 — Re-verify a task's premise in the code at EXECUTION time: roadmap lines, audit findings, and even an audit's own correction footer go stale — usually the document that named a defect is the first + LOG phase-ends/logs/Phase28.md:143: - **R14 on the premise first (twice):** (1) the roadmap's two named grinder bugs are **already fixed** (`e91859fb4`, `4927f38c4`) — stale line. (2) `docs/tooling-audit.md:933-937` **downgraded its own finding** with three corrections: *the prescribed fix is + C4 — A hardcoded assumption fixed in one tool survives in its siblings: grep the tree for the literal and fix every twin in the same change + LOG phase-ends/logs/Phase28.md:36: - `family_remap.img_path` (`:32-35`) hardcodes `0.4.dec` → `None` for the 4 SC07 overlays → `stream_words`→`None` → `classify_member` → `("LEN",[])` → **member silently dropped as not-templatable**. Twin of the P27 T7 `new_overlay.sh` bug, left in a second t + C5 — A batch gate that bisects on failure re-runs the singleton against an unchanged baseline: special-case n==1 or pay a duplicate build on the hot path + LOG phase-ends/logs/Phase28.md:58: - **Unconditional:** the `--chunk 1` double-build (`harvest_verify.py:199-213`) — the bisect re-runs `attempt()` on the same single element vs an unchanged baseline = a duplicate build. 1.35→1.0 builds/draft = **~26% fewer builds on the hot path**, one line. + C6 — The generated disassembly tree is shared mutable state: a fleet verify/clean chain and the per-function instruments cannot run at the same time + LOG phase-ends/logs/Phase32.md:255: - **Gotchas (live):** `make clean` deletes `asm/` — never run rtu_match/twin_rescan/frontier_classify while an R22 chain runs · `gate_main` rewrites main's `asm/` too · the standalone `cc1_dumps.sh` is not faithful on main TUs (use `cc1_dumps_tu.sh`) · the c + C7 — Write the phase synthesis in a FRESH session that re-reads the committed state cold; the cold re-read is what catches stale artifacts + LOG phase-ends/logs/Phase28.md:86: **PhaseEnd_Phase28 is the only remaining work** (Tier-1 Max synthesis — write it in a fresh session with headroom; it re-reads the committed state cold, which is the discipline that caught tonight's stale-map/stale-.md traps). Everything it needs is in this fi + +=================== Phase29-1of4: considered 46 · NEW 6 (found 6) · banked 40 · lines read 2580 of 2580 + C1 — A stop/continue instrument must aggregate at exactly the unit the decision is made in; one that averages a finer unit manufactures a false "we are at the floor" + LOG phase-ends/logs/Phase29.md:1283: **⚠️ THE FLOOR VERDICT WAS AN ARTIFACT — CORRECTED.** `burndown.py` averaged the last 3 INTER-COMMIT deltas, but the ROI criterion is per-SESSION yield. Three mid-session snapshots of a **+0.7pp** session averaged to **+0.23** and printed **"AT THE FLOOR + C2 — Batch size is a RISK lever, not a token lever: isolated agents cost ~N× one agent whether concurrent or serial, so size a batch by the unverified spend you are willing to lose before the next mea + LOG phase-ends/logs/Phase29.md:2387: > Batch size is a RISK lever, not a token lever: each agent drafts one fn in its own context, so N > agents cost ~N× one agent whether concurrent or serial (~110k tok/drafted fn, measured s14+s15). > Concurrency buys wall-clock only. Therefore **batches of ~5– + C3 — A repair/`--recover` mode whose cost is (exceptions × population) must be gated on a MEASURED exception count; for a broadly divergent set, drop rather than recover + LOG phase-ends/logs/Phase29.md:1635: > **⚠️ INCIDENT (recovered, 0 work lost):** `dedup_propagate --recover` on func_80169228 (many byte-divergent > stragglers) THRASHED >1hr (re-gates the fleet per excluded straggler = quadratic). Killed + reverted 326 > half-mutated files to the committed basel + C4 — A defect reasoned into a sibling tool is LATENT until a run shows its signature; do not patch it on theory right after that tool produced a clean run + LOG phase-ends/logs/Phase29.md:1191: were computable, so they are in the tool, not in prose). **(2) LATENT, evidence-gated:** `jtbl_family_bank` isolates AROUND THE STUB (the body is spliced later, after `remap_hseq`) — the same ordering defect just fixed in `harvest_verify`. It did **not** + C5 — Sweep for orphaned worker processes at every session boundary; a dead-pipe compiler or a self-matching wait loop holds a core forever and nothing reports it + LOG phase-ends/logs/Phase29.md:1029: **Housekeeping:** killed an orphaned `cc1` from the Jul-21 session that had been burning a full core for **13h23m** (pid 104350, dead pipe); committed the 4 wave-4 `.o0` drafts left untracked (R20). + C6 — Land a pure rename and a semantic/layout change as separate gated edits, so a gate failure attributes itself + LOG phase-ends/logs/Phase29.md:1959: **`VECTOR` left UNCHANGED deliberately:** ours is 12B vs PsyQ's 16B (missing the trailing `pad`), and I first called it dead — **WRONG, it has 1 live use** (`engine_core.h` `gte_ldlv0((VECTOR*)sp)`). vx/vy/vz offsets already agree so a fix is likely byte + +=================== Phase29-2of4: considered 43 · NEW 7 (found 7) · banked 36 · lines read 2580 of 2580 + C1 — A verification flag that short-circuits the tool's WRITE path leaves the stale artefact in place and still exits 0 + LOG phase-ends/logs/Phase29.md:2762: **⚠️ The sharp edge that hid it:** `worklist.py --assert-partition` **exits at the assertion and never rewrites the doc** — so "regenerating" with that flag leaves the stale file in place and still exits 0. (Its own assertion printed "160 live stubs, 160 + C2 — A diagnosis earns belief when it predicts its own RESIDUAL membership, not when it explains the failures already seen + LOG phase-ends/logs/Phase29.md:3690: binaries** — `func_80174CB0` VERIFIED in **132**, failed in **3**. The class-A census (`func_80012ABC`: **73 `s32` vs 7 `s16`**) had predicted the class-B fix would leave exactly the class-A overlays behind, and it left **3** — the same 3 the original sw + C3 — Make the DRAFTER run the pre-gate guard and return the NAMED banking prerequisite; a batch of diagnosed candidates is worth far more than a batch of opaque MATCHes + LOG phase-ends/logs/Phase29.md:5093: **⚠️ These are CANDIDATES, not banks (§58/G3/P9).** `match_one` MATCH is a proxy; the whole-binary byte-gate decides, and an earlier 11-core wave had 9/9 match_one MATCHes all gate-fail on integration. What makes this wave different is that the agents were *to + C4 — When two blockers are orthogonal, a classifier's if-chain ORDER silently becomes the label — cross-tabulate, never bucket + LOG phase-ends/logs/Phase29.md:4958: The T2 table above ordered its if-chain with `jr` FIRST, so any family carrying a mid-jr was bucketed as "jr → §81 carve chain" **regardless of whether its exemplar needed a crack at all**. That conflated two orthogonal axes and under-reported the zero-crack p + C5 — Sequence a phase so the cheapest thing that can invalidate everything below it runs FIRST + LOG phase-ends/logs/Phase29.md:4322: ### Why T0 leads Not caution — **T0 is the cheapest thing that can invalidate everything below it.** If the families template, T2.2 becomes "crack N exemplars and stamp" and most of that bucket evaporates. If they do + C6 — A reach-weighted gain figure (size × copies) is not a size; every number must say which of the two it is + LOG phase-ends/logs/Phase29.md:4775: **⚠️ SIZING TRAP (Drew caught me on this):** `nins × members` is **reach-weighted gain-ins**, NOT a function size. `func_80144090` is **154 ins × 136 copies**, not a 20,944-ins monster. **Always label which one you are quoting.** + C7 — A repair ladder must probe whether each stage is NEEDED before applying it, or it silently escalates a binary-local bank into a fleet-shared one + LOG phase-ends/logs/Phase29.md:3377: **⚠️ HYPOTHESIS TO TEST AFTER R22 — the recovery tool may have taken a FLEET-TIER edit it did not need.** The bank rewrote `src/shared/engine_core.h` (2 lines, `DEFINE_func_80174C60` + `DEFINE_func_80174C80`) relaxing `extern s32 func_80174CB0(s32, s32); + +=================== Phase29-3of4: considered 36 · NEW 6 (found 6) · banked 30 · lines read 2580 of 2580 + C1 — GNU C sources carry form-feed page separators, and `str.splitlines()` splits on them while `grep`/`sed` do not — so any Python line-number checker over compiler source silently drifts, and blames + LOG phase-ends/logs/Phase29.md:6915: ### The bug: `splitlines()` vs FORM FEEDS GNU C sources use **form-feed (`\f`) page separators** — `loop.c` has 47, `cse.c` 36, `reload1.c` 27, `local-alloc.c` 21. **Python's `str.splitlines()` splits on `\f`; `grep`/`sed`/editors do not.** So + C2 — Never restructure a proven tool with blind string replaces at the end of a long session; specify it as the next session's first task instead + LOG phase-ends/logs/Phase29.md:6134: **⚠️ AND I BROKE `family_sweep` TWICE TRYING TO WIRE THE PARALLEL DEFAULT** (missed import, then a closure-scope error) — on the tool that banked 543 members today. **Reverted, not committed.** Restructuring a proven tool with blind string replaces at the end + C3 — An oracle that can always be RUN is not always APPLICABLE: state the applicability precondition beside the recipe, or a coarse run returns a large number that reads as a verdict + LOG phase-ends/logs/Phase29.md:7359: ### 🔧 MAP REFINEMENT OWED — §H's oracle has an unstated PRECONDITION `regalloc.md` §H presents the swap oracle as the way to "discriminate RC-6 (allocation) from S3 (scheduling) in ONE gdb run". It only works when **the contested registers are held by PSEUDOS* + C4 — Before a fleet-wide mechanical edit, census the whole population for the exact preconditions the edit assumes; perfect uniformity is the licence to apply it, and non-uniformity is the design inpu + LOG phase-ends/logs/Phase29.md:7520: ### MEASURED BEFORE BUILDING (R35) Ran the blocker census over all 132 still-stubbed siblings before writing a line. It is **perfectly uniform**, which is the strongest possible signal that one mechanical edit fixes all of them: + C5 — A status line a script prints unconditionally is not a measurement; derive every conclusion the script emits from the command's own output + LOG phase-ends/logs/Phase29.md:7496: - **Three unconditional `echo` conclusions** (`[shared clean]`, `[none = ...]`) that asserted things + C6 — Measure what fraction of a cycle a parallelism knob can actually touch before adopting it: `make -j16` bought 12%, because the build was 5 s of a 16 s per-item cycle and the real cost was a four- + LOG phase-ends/logs/Phase29.md:6115: **The `-j` theory was WRONG, and measuring said so** (baseline ~18 s/sibling): | | | |---|--:| + +=================== Phase29-4of4: considered 36 · NEW 6 (found 6) · banked 30 · lines read 2576 of 2576 + C1 — A clean-looking verdict that appears immediately AFTER your own repair transform is a SUSPECT, not a result: re-measure the artefact the transform produced before routing the residual + LOG phase-ends/logs/Phase29.md:8207: ### ⚠️ A DEFECT I INTRODUCED, CAUGHT BY MEASURING `func_8016163C` read as a clean **DIFF** after T60 and I reported it as "genuine codegen". It is not. `match_one` says **`SIZE-MISMATCH`: draft 58 ins vs target 78** (Δ−20, ratio 0.74, bucket `redraft`). + C2 — Classify a harness fix as a LOGIC defect or a PATH-REACHABILITY gap and price it accordingly: only the logic defect generalises + LOG phase-ends/logs/Phase29.md:9957: ### THE BLAST-RADIUS PATTERN, NOW FOUR DATA POINTS | lever | members | kind | |---|---|---| + C3 — A wrong prescription left in the knowledge base is worse than no entry: when evidence refutes an entry you wrote, correct THAT entry in place, in the same session, carrying the refutation + LOG phase-ends/logs/Phase29.md:9029: **Reverted** (`git checkout -- src/`, tree clean, nothing committed). Cookbook **§116 corrected in place** — it now carries the refutation and the corollary, because a wrong prescription left in the cookbook is worse than no entry: the next session would have + C4 — A name grep is not a "defined / banked here" oracle: a DECLARATION carrying the name reads as a definition; use the structured stub oracle, and believe the pipeline's map over your own check + LOG phase-ends/logs/Phase29.md:8080: ### FIRST, A CORRECTION TO MY OWN T58 REPORT (R14) I said "7 remaining families all have banked exemplars". **Wrong — there were 5.** `0x80175820` (276 members) and `0x8016ec0c` (138) have **no matched exemplar anywhere**: both are INCLUDE_ASM + C5 — A yield estimator that counts "unclaimed at the moment it runs" over-projects: it ranks correctly and overstates absolutely; never plan off its absolute numbers + LOG phase-ends/logs/Phase29.md:9350: ### ⚠️ MY `new_distinct` ESTIMATOR OVER-PROJECTS ~2× (R14) I priced these two families at 130 + 129 = **259** new distinct; the measured gain is **125**. The estimator counts a family's h_exact classes that have no matched instance *at the time it runs*, so + C6 — Declare a mechanical lever SPENT only on a positive, three-part measurement — every built lever applied and returning zero, the residue split by structure, and the decay curve priced against what + LOG phase-ends/logs/Phase29.md:10300: The final four days decayed **+2.7 → +2.2 → +0.6 → +0.3pp/day with every lever this phase built applied** — and T98 characterised the residue as **80 families / 960 members / 173 distinct** (29 all-STRUCT refused by design + 51 gate-failing at ~3 distinct per + +=================== Phase30-1of2: considered 50 · NEW 13 (found 13) · banked 37 · lines read 2684 of 2684 + C1 — An instrument's refusal is a FINDING, not an obstacle: overriding it means explaining why the instrument is wrong, never finding another route + LOG phase-ends/logs/Phase30.md:1369: 4. **VALIDATE:** `tools/validate_targets.py --targets --out clean.json`. Schema needs `name` + `binary`. It fails closed; **an instrument's refusal is a finding, not an obstacle** — S46 routed around it and lost 87 of 119 agents. + C2 — Prose in an agent prompt is not enforcement: snapshot `git status` around every agent and name the offender + LOG phase-ends/logs/Phase30.md:449: > **3. An agent wrote a TRACKED file** (`src/shared/engine_types.h`, added a typedef) despite the > prompt forbidding it twice. `gate_lane`'s entry guard refused to gate on a tree it did not own — > caught before any commit, cost one gate cycle. **Prose is not + C3 — Read the FIRST ten results of a long run before trusting the other 190; and negative-control any NEW REFUSAL against everything that already succeeded + LOG phase-ends/logs/Phase30.md:250: ## 🔑 THE ONE THING TO CARRY FORWARD **Three runs launched at scale, three stopped early — and every stop was right.** The verdicts the tools already write (`harvest_failed..classified.txt`) named each defect within the first + C4 — When a tool is repaired, every verdict it produced becomes a hypothesis again — but re-gate only the drafts the repair's BLAST RADIUS plausibly touched, not the whole ledger + LOG phase-ends/logs/Phase30.md:2481: claim is suspect: **when a tool is repaired, every verdict it produced becomes a HYPOTHESIS again.** The backlog's `closeness` values and residual classes were produced by tooling that has changed materially this session (§143 alone). **Re-measure before respe + C5 — A ledger row with no draft artifact is a rumour, not a result + LOG phase-ends/logs/Phase30.md:2352: **Treat the 14 as UNVERIFIED; start from the committed 63/47 drafts; purge the row if it cannot be reproduced.** General rule worth adopting: *a backlog row with no draft artifact is a rumour, not a result* — `backlog.py log` should require a draft path or mar + C6 — Check group identity from the signature files BEFORE probing a family: if the members are structurally identical, a 0% result is a compile-error CERTAINTY, not evidence about codegen + LOG phase-ends/logs/Phase30.md:168: **🚨 STANDING PRE-PROBE RULE:** check h_norm identity across a family's members from the sig files BEFORE probing. If members are h_norm-identical, a 0% is a **compile-error certainty**, not evidence about codegen. And **read `.run/hseq_failed + C7 — Never test a helper by importing its module: a tool with no `__main__` guard runs its whole pipeline on import + LOG phase-ends/logs/Phase30.md:2163: ## 🧰 HAZARD INTRODUCED-AND-DOCUMENTED THIS SESSION `tools/harvest_verify.py` has **no `if __name__ == '__main__'` guard**: `import harvest_verify` runs a full build, splices drafts, and overwrites `.run/harvest_*.txt`. I tripped it unit-testing + C8 — When A/B-ing a harness knob (model, effort, reasoning level), ship a POSITIVE CONTROL that the knob actually moved + LOG phase-ends/logs/Phase30.md:1541: 3. **The effort/model experiment Drew raised** (design in the S46 log below): a 3-way on ONE target list in the ≥50-ins band — Sonnet-default vs Fable-low vs Opus-default — with a **positive control that the effort knob actually moved** (compare per-agen + C9 — Keep the wave harness IN THE REPO with its contracts; a harness rebuilt from memory each run silently goes stale + LOG phase-ends/logs/Phase30.md:688: **THE HARNESS IS NOW IN THE REPO** (`tools/wave/crack_wave.js` + README, `54de1bf24`). It had lived only in the workflow scratch dir, so each wave rebuilt it from memory — which is how its cookbook citation list went stale at §162 while §163 (5) and §164 (82) + C10 — Do not cross-price two economies: a conversion rate measured on the RESIDUE QUEUE does not price a FRESH wave + LOG phase-ends/logs/Phase30.md:899: it rather than my framing. **From here the mover is volume with multipliers, not more plumbing.** Corollary it also corrected: **do not cross-price the two economies** — 5:1 is a property of the RESIDUE QUEUE, while a FRESH wave converted 81% (and stored MATCH + C11 — Derive the target pool from the BUILD'S OWN INVARIANT, not from a reach/similarity metric — reach ranked the most-DONE work first + LOG phase-ends/logs/Phase30.md:1810: ## 🎯 THE CORRECTED FRONTIER DEFINITION (the session's most useful output) **Reach-141 identifies the most-DONE work, not the most valuable** — those are the shared engine functions banked over 29 phases, present in each overlay as `DEFINE_func_*` macros (~1,61 + C12 — A metric that re-parses source is blind to a body banked through an `#include`; trust the metrics derived from the stub oracle + LOG phase-ends/logs/Phase30.md:2589: ## 📌 A METRIC SHAPE WORTH KNOWING (R30) A body banked by `#include`-ing a shared header is **invisible to fn-count's NUMERATOR** (the definition is not in the `.c`) while its stub leaves the denominator — the whale bank moved fn-count + C13 — Keep a glossary line for any term two documents use in OPPOSITE senses + LOG phase-ends/logs/Phase30.md:1343: 4. **"ZERO-CRACK" MEANS OPPOSITE THINGS** in roadmap §3 T3 ("61 zero-crack = propagation-only") and in the current map/this file (zero-crack = needs its FIRST crack). **30× mis-scope risk.** One glossary line in `family-hseq.md` fixes it. Roadmap v2 has + +=================== Phase30-2of2: considered 44 · NEW 6 (found 6) · banked 38 · lines read 2684 of 2684 + C1 — Drafting agents must never be able to write the build tree; every agent artifact lands in a scratch directory, so a killed or racing campaign costs build cycles and zero work + LOG phase-ends/logs/Phase30.md:4017: - Corollary proven twice: **agents must only ever write `.run/`** — that is why both incidents cost build cycles and zero work. + C2 — A stored verdict can be stale because the INSTRUMENT changed, not the draft: re-gate the drafts a tool repair plausibly touched, scoped by the repair's blast radius — never the whole ledger + LOG phase-ends/logs/Phase30.md:4660: 2. **Whole-payload averaging DILUTES code.** A real overlay is code + a large data tail, so its + C3 — Count agent completions from the run journal's result records, never from artifact existence: an agent writes its deliverable early and then iterates, so the file proves nothing + LOG phase-ends/logs/Phase30.md:4001: *(Counting note, R14: my first count said "2 remaining" because I measured draft-FILE existence. An agent writes its draft early and then iterates, so a file proves nothing about completion — the journal's `result` records are the truth. Drew's "12" w + C4 — A derived claim outranks a heuristic verdict; when two heuristics disagree, take the UNION and queue the disagreements — under-reporting hides work, over-reporting only costs review + LOG phase-ends/logs/Phase30.md:4657: **L2, the SECOND DISAGREEING ORACLE (R34), earned its keep immediately — it found two L1 defects:** 1. **A claim outranks a heuristic.** I let the statistical verdict override a SHA match, so claimed binaries were being filed as `classified-data`. Onboarded + C5 — A failure that will not reproduce earns a negative-control-proven DETECTOR, not a speculative fix; and every abort path must PROVE its revert by diffing the worktree against a baseline captured a + LOG phase-ends/logs/Phase30.md:4064: checked out. **I did not "fix the bug"; I made the next occurrence name itself.** + C6 — Never let model-authored prose reach the shell inside double quotes: a backticked command in a `git commit -m` message EXECUTED + LOG phase-ends/logs/Phase30.md:2893: 4. **A backticked `` `make extract` `` in a `-m` commit message EXECUTED** — corrupted the message and ran a real extract. Use quoted heredocs. (No damage; re-committed.) + +=================== Phase31-1of3: considered 52 · NEW 12 (found 12) · banked 40 · lines read 2902 of 2902 + C1 — More output-token budget is NOT more quality; measure it as a paired A/B and treat truncation as recoverable, not as a defect signal + LOG phase-ends/logs/Phase31.md:2292: snapshot-restore per draft under per-binary locks. **BANKABLE: 8k 10/10 · 16k 10/10 · 24k 8/10 · 32k 8/10** (big arms: two budget-never-converged NO-DRAFTs plus the 55-ins fn BYTE-DIFF in both). Truncated turns 6/1/0/0 — truncation recovers across the turn loo + C2 — In a pipelined drafter/gater, "still open" is not "not yet attempted": consecutive waves re-drafted the wave still in flight + LOG phase-ends/logs/Phase31.md:258: 3. **Every second wave re-drafted the wave still in flight** — `--retry-unbanked` returns still-open cards and the pre-draw runs WHILE a wave drafts, so ck->cl were 239/239 identical, co->cp 238/238. Yield alternated 47.6% / 3.8% / 35.3% / 3.6%. Fixed by + C3 — Snapshot every target's disassembly BEFORE gating: a successful bank prunes it, and every downstream step needs both sides + LOG phase-ends/logs/Phase31.md:1612: * Snapshot every target's `.s` BEFORE gating — banking prunes it and recovery/harvest need both sides. + C4 — Draft first, then carve: a build-unit split is only safe when the new unit is immediately populated with proven bodies + LOG phase-ends/logs/Phase31.md:922: 6,414 ins, zero drafting). **Draft-first ordering is what makes carves safe: populated-at-carve is 137/137; stubs-in-new-object was 1/4 and 2/2.** + C5 — A cracked idiom transfers WITHIN its family and not across it: price a lane by families, not by class size + LOG phase-ends/logs/Phase31.md:1300: **11 turns/3 oracle MATCH** -> cross-family **56 turns/0 compiles, FAILED**. So jtbl costs ~40 turns of learning **per family** (191 families), not per class. The jtbl quest is a project, not a lane — defer it. + C6 — Sibling count and never-drafted count are different denominators; conflating them overstated the free-remap leverage ~3× and hid that the remaining mass was singletons + LOG phase-ends/logs/Phase31.md:298: measures: 3,106 open instances · 2,245 open skeletons · 1,772 groups, of which **480 are multi-member holding 1,334 siblings** and **1,292 are SINGLETONS carrying 57% of the open instruction mass**. My "~3,900 siblings behind ~334 skeletons" conflated tw + C7 — "Never write the tree" in a drafting prompt is a request, not an enforcement + LOG phase-ends/logs/Phase31.md:2879: * **A wave agent WROTE to `src/` and then `git checkout`-reverted it** (SYS_OBJ_2264, self-reported). It was harmless ONLY because R42 meant every bank was already committed. The prompt's "never modify src/" is a request, not an enforcement — candidate: *a + C8 — A free or preview model tier can be withdrawn mid-campaign without notice; a fleet-wide 404 is an epoch event, not N model failures + LOG phase-ends/logs/Phase31.md:2319: ### THE OX WINDOW CLOSED — 2026-08-26 07:55 (probed: HTTP 404 on stealth/ox-alpha; stealth/* gone from the model list) The free-drafting era ended mid-m0b (its 29 agents all 404'd at turn 0 — a harness-epoch event, not 29 model failures, R40). THE GOAL BEAT TH + C9 — A metered API key's own cap is a separate limit from the account's credit + LOG phase-ends/logs/Phase31.md:2380: **DeepSeek push paused by a KEY CAP, not the budget (10:4x, R40-corrected):** the OpenRouter key carries a $60 LIFETIME limit; usage hit $60.21 mid-ds2 and every request 403s ("Key limit exceeded (total limit)") while the ACCOUNT still holds ~$10.6 credit. ds1 + C10 — Incremental gates never exercise the extraction/regeneration step, so regeneration rot is undated and invisible for weeks + LOG phase-ends/logs/Phase31.md:2376: nothing — a fresh-eyes day-session surgery. The general lesson repeats §61c with a new edge: INCREMENTALLY-GREEN HIDES EXTRACT ROT — a periodic `make extract` sweep per binary (not just check-all builds) would have dated defect (a) precisely. + C11 — A probe that splices the tree in order to measure it must restore it, or the progress oracle counts the splices as banks + LOG phase-ends/logs/Phase31.md:2598: next pass / T6. Instrument lessons of the task: a stale classification file re-labels every fn "CARVE-REFUSED" until re-gated (verify from the gate, not the ledger); an autopsy script without its scratch dir leaves TUs spliced and `corpus` then reads them as b + C12 — Keep prompt/law text in data, not inside the launcher's source template + LOG phase-ends/logs/Phase31.md:1702: * A wave script's LAWS block is a JS template literal: **backticks inside the law text terminate it**. Wave U's first launch died on exactly that; write law text with single quotes. + +=================== Phase31-2of3: considered 52 · NEW 11 (found 11) · banked 41 · lines read 2902 of 2902 + C1 — The health suite must assert that a tool DID ITS WORK, not only that the data is intact: zero inputs, an impossible wall-clock and a missing persistent effect are each a DEFECT, not a result + LOG phase-ends/logs/Phase31.md:5122: **ITEMS 2+3 — `tools/work_evidence.py`, wired and negative-controlled.** One module, three assertions, all about OBSERVABLE CONSEQUENCE rather than internal state: * `assert_inputs` — zero readable inputs is a DEFECT, not a zero-yield result. **"0 of 0" is a f + C2 — A health check that cannot finish is not a check: keep the health target sampled and fast, and put the exhaustive form behind its own name + LOG phase-ends/logs/Phase31.md:5175: - **S70 — `make tools-health` WAS UNRUNNABLE AND IS NOW 333s GREEN.** Drew: *"this is a tools health test, not a full regression test."* Correct — `audit-cdecl` re-parsed **every declaration in all 4,168 TUs** and handed each to real gcc: **~787s of pure-P + C3 — "Independent" names the INSTRUMENT, not the input: two refusals of two separately-written drafts from one tool is one test repeated + LOG phase-ends/logs/Phase31.md:3589: **"Independent" means a DIFFERENT INSTRUMENT, not a different input.** Two runs of one tool on two drafts is one test repeated. R40, sharpened. + C4 — A drafting agent must run where it CANNOT write the source tree; "never modify src/" in a prompt is a request, not an enforcement + LOG phase-ends/logs/Phase31.md:2879: * **A wave agent WROTE to `src/` and then `git checkout`-reverted it** (SYS_OBJ_2264, self-reported). It was harmless ONLY because R42 meant every bank was already committed. The prompt's "never modify src/" is a request, not an enforcement — candidate: *a + C5 — A proven transform that is not a rung of the ladder the agents' drafts actually pass through does not exist for those drafts + LOG phase-ends/logs/Phase31.md:3810: `scope_data_externs.fix()` (§8d) has been byte-proven since Phase 26 and is used by `family_sweep` / `bank_exemplar` / `jtbl_family_bank` — but **nothing in `gate_stage`'s ladder ever called it**, so a draft written by a wave agent had never seen it. Built as + C6 — The knowledge-harvest selector must be able to see FAILED attempts; a filter that can only read banked work learns from the easy half + LOG phase-ends/logs/Phase31.md:3236: 5. **The distill novelty selector was INVERTED** (`91e624c04`): `'no cookbook lever'` matched "no cookbook lever *needed*" (a TRIVIAL note) and was the only pick of 24, while three multi-lever notes went unseen — **and it structurally could not see UNBAN + C7 — A refusal names the branch the caller entered, not the subject — make the applier consult the classifier it already has + LOG phase-ends/logs/Phase31.md:3346: 4. `--probe-only` crashed for `--funcs`/`--auto` (exec'd before staging). 5. The distill novelty selector was INVERTED (its only pick of 24 was a "nothing to learn here" note) and could not see UNBANKED fns at all — where the hardest functions write their r + C8 — Every status claim in an agent's context must be expiry-checked against live state, or agents will report it back to you as an observation + LOG phase-ends/logs/Phase31.md:5394: 5. **stale BASELINE-RED verdict** — OPEN. THREE agents reported a red baseline on binaries I verified BYTE-IDENTICAL; the third revealed the source: *"the pack's last gate verdict was BASELINE-RED"* — they read it off the card. Cost me two phantom-regres + C9 — A validity stamp must be honoured by every downstream consumer; and a wrong write-side label cannot be repaired by a correct read key + LOG phase-ends/logs/Phase31.md:5437: * **#4 SYMBOL MISMATCHES — sound=True, SHIPPED.** `gate_feedback` gated on `shape=='MATCH'` when reloc_identity's binding condition is **`aligned`** (shape AND equal reloc-stream lengths). Below that bar reloc_identity itself downgrades status to `MI + C10 — Yield is clustered by binary, not spread over the fleet: draw per-binary once two independent lanes concentrate in the same place + LOG phase-ends/logs/Phase31.md:5040: **20 of 52 twin remaps banked (38%)**, well above §398's ~15% straight-through — and **all 20 in `ov_SC06_011`**, the same binary that carried 15 of the 21 standalone banks. Two lanes, same concentration: `ov_SC06_011` is simply a binary whose open tail is hig + C11 — (minor) A build-system conditional that expands at PARSE time makes its own negative control vacuous + LOG phase-ends/logs/Phase31.md:5155: *Make gotcha worth keeping: my first patch used `ifeq ($(filter $*,...))`, which make evaluates at PARSE time when `$*` is empty — it would have silently always taken the maspsx branch and the "byte-inert" result would have been vacuous. `$(if ...)` expa + +=================== Phase31-3of3: considered 30 · NEW 7 (found 7) · banked 23 · lines read 2931 of 2931 + C1 — Never wrap a project tool in a `timeout` shorter than its own internal budget; you pre-empt its documented recovery handler and lose its buffered output + LOG phase-ends/logs/Phase31.md:7024: * **Never wrap a project tool in a shorter `timeout` than its own budget.** My `timeout 2400` beat `gate_stage`'s 3600 s budget, SIGTERM'd the tree mid-propagation, and Python lost its buffered stdout — three gates with NO verdict and three half-applied, n + C2 — A programmatic edit to a long-lived knowledge document silently truncates or duplicates it; verify the sections, never the commit + LOG phase-ends/logs/Phase31.md:6529: 1. **§429 had been SILENTLY DELETED from the cookbook.** My §428a rewrite wrote `t[:start] + new` instead of `t[:start] + new + t[end:]`, truncating everything below it. §429 ("every held pointer needs its own local") was gone for the rest of the session + C3 — The live hand-off block must be strictly APPENDED at the end of its file: file order is the only recency signal a fresh session has + LOG phase-ends/logs/Phase31.md:6126: > **The "last block is the live one" rule was BROKEN when this session started** — S70 FINAL-5 sat > below S71 CLOSE in file order while being a day older. This block is appended at the END, which > restores the rule. Keep appending. + C4 — A long stateless batch must persist each confirmed result the moment it is confirmed + LOG phase-ends/logs/Phase31.md:7121: 3. **`gate_main` writes banks only at the END.** `try_batch` is stateless — every attempt is `git checkout` main's TUs → `make extract` → substitute → build — so a 34-minute bisection holds its result in memory and a kill loses all of it. Writing each co + C5 — "Idempotent by SKIPPING" seals an artifact against later evidence; make regenerated artifacts idempotent by REPLACEMENT + LOG phase-ends/logs/Phase31.md:6171: * **`journal_notes.py`** — was idempotent by SKIPPING, so a pack with one old note could never receive a newer one, and `claude_wave_packs` calls it at build time, so **every pack with any history was sealed against later evidence**. Now idempotent by repl + C6 — A verification target too slow to complete is not a check: sample it by default and keep the exhaustive form as a separate target + LOG phase-ends/logs/Phase31.md:6027: * `make tools-health` — was UNRUNNABLE (>15 min, never once completed). `audit-cdecl` was a full-corpus regression test in a health target (~787s of pure-Python collection before the first cc1 call). Now sampled (`CDECL_AUDIT_TUS ?= 60`, 61s); `audit-cdecl + C7 — Never adopt a subagent's worktree wholesale: it is a snapshot of an older tree and may predate a bank; re-gate its artifacts against HEAD + LOG phase-ends/logs/Phase31.md:6808: * **Never adopt an agent's worktree wholesale.** One predated a bank of `func_8017F9C0`; copying its TU would have destroyed it. Re-gate against HEAD with the fixed tools instead. + +=================== Phase33: considered 37 · NEW 16 (found 16) · banked 21 · lines read 1372 of 1372 + C1 — A Ghidra script directory compiles as ONE bundle: a single non-compiling script disables every script in it + LOG phase-ends/logs/Phase33.md:186: with '.'; `resident_funcs.txt` from the ELF = 1,304 entries). **Blocker found:** in the resident rebuild the runs after the import failed with `Failed to get OSGi bundle containing script: …/tools/ghidra_scripts/ApplySymbols.java` (same for ExportAnnotat + C2 — Ghidra refuses a project path containing a component that starts with `.` — a scratch project cannot live under `.run/` + LOG phase-ends/logs/Phase33.md:185: step works: the scratch project must live under `build/ghidra_rebuild/proj` — Ghidra refuses a path component starting with '.'; `resident_funcs.txt` from the ELF = 1,304 entries). **Blocker found:** in the resident rebuild the runs after the + C3 — `git check-ignore` is SILENT for tracked paths: an ignore-coverage audit run before the untracking passes vacuously + LOG phase-ends/logs/Phase33.md:236: project-authored inside a purged dir — CHECKSUMS had moved in B4); ignore coverage proven with `git check-ignore --no-index` on every path (the plain form is BLIND to tracked files — it reported nothing); **control: main byte-identical + C4 — A coverage instrument that infers its denominator from OPEN work INVERTS at 100% — carry the scanned denominator in the artifact, and test the instrument at both endpoints + LOG phase-ends/logs/Phase33.md:126: - **2026-09-06 (S86) — A4 (family map regen → instrument fix, R35).** Regenerating `.run/family_hseq.json` made the `audit-binaries` warning WORSE (6 → 217 "missing"): at 100% the map's `families` list is empty and CHECK 4 inferred coverage from family mem + C5 — An annotator that writes into the text it reads must never treat its own output as evidence + LOG phase-ends/logs/Phase33.md:499: lines — the input to an override). **Three instrument defects, found by its own controls/verify before any tag was written (R39/R57):** (1) span pairing inside a ±160-char window inverted the backtick pairs and dropped the identifiers right next to a cit + C6 — Regex-extracted evidence needs a structural marker, or prose becomes data + LOG phase-ends/logs/Phase33.md:503: caught by `--verify`. Bare ALL-CAPS prose words (`NOT`, `AND`, `DEST`) had also passed as evidence: identifiers now need an underscore, as every real gcc macro/function cited has. **Final census:** `[2.7.2]` 79 · `[2.8.1 pm]` 55 · `[repo]` 1 · + C7 — `objdump -dr` interleaves relocation records only for OBJECT files; a linked ELF lists them separately and shifts the address column + LOG phase-ends/logs/Phase33.md:526: **8/8 OK**. Two gotchas, recorded: `objdump -dr` interleaves relocation records only for OBJECT files — a linked ELF lists them separately (`-r`, section-relative offsets), so the generator merges them into the listing in the object-listing format; and a + C8 — A derive-then-apply pipeline over a live repository needs a freshness guard, and a stated sequencing law + LOG phase-ends/logs/Phase33.md:285: added to `run_filter.py`: it refuses a dictionary whose main count/HEAD differ from the clone's (a stale dictionary would drop rows from the public map; trial #2 matched by construction). **Sequencing law for C4:** the dictionary + ids are rebuilt from t + C9 — A content-hash "no ROM bytes" audit collides on zero-length files + LOG phase-ends/logs/Phase33.md:205: TUs from `_SRC_DIR` with nested pruning, skips DERIVED: 70 LINKED + 47 INCLUDE_ASM, -O0 TUs compiled at -O0; measured PR scope 54/54 in 1.6 s, **fleet 4,170/4,170 in 123 s at -j32, failed 0**; unknown alias refused); `.github/workflows/no-rom.yml` + C10 — `git push --mirror` does not push `refs/stash` + LOG phase-ends/logs/Phase33.md:295: **private, not a fork, no parent**; `git ls-remote archive` == local refs except **`refs/stash`, which `--mirror` does not push** (the bundle holds it; Drew may push it as a branch). C4c: FINAL dictionary on the committed tip (4,422 commit objects, + C11 — Route a host purge request through the flow that actually exists; the obvious form is a trap + LOG phase-ends/logs/Phase33.md:635: - **2026-09-07 (S89) — C10: the GitHub Support ticket is FILED — #4736982** (`https://support.github.com/ticket/personal/0/4736982`, Drew, via the Support portal's **Virtual Agent "Clear cached views"** flow — the route that actually works: Repositories form + C12 — A host feature can be gated on the very flip it was meant to precede: read the settings page, don't infer + LOG phase-ends/logs/Phase33.md:692: ~~The wiki can be pushed NOW~~ — **WRONG (R14, corrected minutes later):** GitHub's settings page reads "Upgrade or make this repository public to enable Wikis"; on the free plan wikis exist only on public repos, so the wiki waits for the flip. `wiki_sync. + C13 — A public scratch service's compiler image is NOT your pinned toolchain; rebuild it locally and prove byte-identity before asking for a preset + LOG phase-ends/logs/Phase33.md:645: is old-gcc **0.13** + maspsx **`86ccd7d8`** (not our 0.17 + `874855c5` — SETUP's "same pinned commit we use" was stale; three rows corrected, R14) with `as` = a maspsx `--run-assembler` wrapper (so `-Wa,--aspsx-version=2.56,--expand-div` reaches maspsx); + C14 — A miner over your own records finds only what its pattern anticipates: measure the widened pattern's yield + LOG phase-ends/logs/Phase33.md:429: - **2026-09-07 (S87, Max) — F2 the retrospective.** `tools/mine_hindsight.py` (stdlib; the decision-log's `Hindsight` bullets AND `### Hindsight` sections — 19 over 79 entries after widening the pattern, the first cut found 11 —, the two "What we believed" + C15 — A harness's low-memory guard silently kills a long BACKGROUND job — and workers from closed phases survive for days + LOG phase-ends/logs/Phase33.md:508: `doc_links --strict` PASS (41 documents, 301 links, 0 broken). Gotcha, recorded: the first tools-health run was KILLED by the harness's low-memory guard during the report step (a transient spike; 29 GB available afterwards) — and the process table held * + C16 — A `cd` in one agent shell call persists into the next + LOG phase-ends/logs/Phase33.md:660: SETUP §6.5 rewritten + 2 inventory rows + the P33 E1 section + ledger row 14. Gotcha, recorded: a `cd` in one Bash call persists into the next — the first download landed inside `tools/maspsx/` (moved out; submodule clean). Commit: see below. + +=================== Phase7: considered 28 · NEW 5 (found 5) · banked 23 · lines read 201 of 201 + C1 — A decompiler's `unaff_` output means it decompiled ONE entry path of a multi-entry function and handed you a fragment; that is an instrument limit, never evidence the function is hard + LOG phase-ends/logs/Phase7.md:109: **(session G cont.)** Then the 4 jtbl loaders: **CdReadStateMachine / CdReadSectorReadyCB / StreamLoadStateMachine NON_MATCHING-drafted** (faithful translations of the Ghidra decompile; all valid C under `-DNON_MATCHING`, verified via a cpp→cc1 syntax check; + C2 — A milestone counted in "matched functions" must exclude the splitter's auto-generated empty bodies; define the bar on substantive matches at the moment you set it + LOG phase-ends/logs/Phase7.md:12: - **≥25 = real substantive matches** (the 42 splat-auto empties do NOT count). 14 real now → need ≥11 more. + C3 — Duplicate/family census run before the vendor library is linked out is contaminated: the duplicate groups are library fragments and epilogues, not game functions + LOG phase-ends/logs/Phase7.md:26: - [x] **Task 4 — Harvest to ≥25 real** ✓ DONE. **36 real matches** (14 + 22 trivial accessor leaves: getters/setters of globals, one mask, one two-store), build BYTE-IDENTICAL → all 22 byte-perfect. **≥25 bar PASS** (margin 36/25). `.run/harvest.py` = the appl + C4 — A byte-identical baseline is only proven when the whole gate is green from clean across several INDEPENDENT sessions; one green run hides a nondeterministic extract + LOG phase-ends/logs/Phase7.md:75: NOTE: Gen1 exit needs ≥3 SESSIONS of green `make check` — **satisfied** (A, B, C, D, E, F, +G); Tasks 6/7 pending. + C5 — When a foundation task hits a structural wall mid-phase, reorder: bank the tractable wins first and re-scope the wall as its own focused sub-project + LOG phase-ends/logs/Phase7.md:3: **Status:** Plan APPROVED 2026-06-14. **REORDERED 2026-06-14** (Drew): bank non-switch wins first; the rodata-island foundation + LZSS match are DEFERRED to a focused sub-project after the easier tasks. + +=================== Phase8-13: considered 38 · NEW 4 (found 4) · banked 34 · lines read 400 of 400 + C1 — A toolchain/SDK-version detector's verdict is a HYPOTHESIS until a placement COUNT backs it; when a version stamp and a byte-probe disagree, the probe wins — and the refuted stamp must be un-bank + LOG phase-ends/logs/Phase10.md:44: ## Key finding (T4, 2026-06-15) — resident is PsyQ **4.7.0**, not 4.0.0 DetectPsyQ on the imported resident program reports **PsyQ Version = 470** (the EXE is 4.0.0); the import associated a `psyq470` source archive alongside `psyq400`; the one PsyQ-signature + C2 — Histogram a secondary binary's `jal` targets BY ADDRESS RANGE before assuming it carries its own copy of anything + LOG phase-ends/logs/Phase12.md:45: - **Ghidra verification (Drew-requested, G1 — CONFIRMS the pivot):** switched the headless MCP to the `resident` program (296 funcs); decompiled a sample. `FUN_800cf854` = game-global accessor (`return DAT_800ae6bf != 0`); `DsMix` = `{FUN_800d1bf8(); return + C3 — A raw blob's load address is a hypothesis; the free confirmation is arithmetic against the NEXT known segment's base + LOG phase-ends/logs/Phase10.md:13: - Load **vram 0x800CEDF8** (Phase-3 T6b proven); `VRAM_BASE`=0x800CEDF8 (fileoff 0→vram). Computed **end vram 0x80128154** (4 B under overlay slot 0x80128158 — boundary corroboration). + C4 — ⚠ LOW VALUE, recommend dropping: during a parameterization refactor whose oracle is a byte-locked binary, generalize only what a second target actually needs + LOG phase-ends/logs/Phase9.md:26: - [x] **T1** — Makefile data block (BINARIES/main_*/aliases; wrapped 9 integrate blocks in `ifeq ($(BINARY),main)`). DONE: aliases resolve byte-identically; `BINARY=bogus` errors; clean build ⇒ `143dbb89` + report 52/959/7/50.24%. **Refinement vs plan:** SDK-r + +TOTAL over 21 slices: considered 777 · NEW 143 · already-banked 633 · lines read 30540 diff --git a/.run/P33.5/log-mining/HARVEST_TABLE.md b/.run/P33.5/log-mining/HARVEST_TABLE.md new file mode 100644 index 0000000000..f7ec709dfa --- /dev/null +++ b/.run/P33.5/log-mining/HARVEST_TABLE.md @@ -0,0 +1,147 @@ +# Harvest table — every NEW candidate of the log-mining pass, its verdict and its home + +| Slice | C | Lesson | Log line | Verdict | +|---|---|---|---|---| +| Phase15-16 | C1 | Measure a pipeline's yield on the RESIDUAL, not on already-solved functions: build a known-answer ladder (reve | `phase-ends/logs/Phase16.md:36` | DK-73 | +| Phase15-16 | C2 | A search harness (permuter) must compile in the SAME declaration context as the real build; an isolated contex | `phase-ends/logs/Phase16.md:71` | DK-78 | +| Phase15-16 | C3 | Before concluding a drafting pipeline is weak, histogram the compiler's error TEXT: over half of ours were one | `phase-ends/logs/Phase16.md:40` | DK-74 | +| Phase15-16 | C4 | Re-run the deterministic declaration-canonicaliser over OLD quarantined drafts after every large bank: a draft | `phase-ends/logs/Phase15.md:57` | DK-77 | +| Phase15-16 | C5 | Write numeric kill-criteria into the plan, before the data exists, for any expensive campaign — and let them f | `phase-ends/logs/Phase16.md:24` | DK-73 | +| Phase15-16 | C6 | Design an unattended run for a human with no agent session: a STOP file honoured at a safe boundary, a supervi | `phase-ends/logs/Phase16.md:89` | DK-75 | +| Phase15-16 | C7 | The permuter's search unit is the C expression: it cannot freeze the instructions you already have right, beca | `phase-ends/logs/Phase16.md:43` | DK-78 | +| Phase15-16 | C8 | An idempotency guard keyed on PRESENCE freezes every record created before the system matured; key it on COMPL | `phase-ends/logs/Phase15.md:51` | DK-77 | +| Phase17-18 | C1 | Leverage and tractability are ANTI-correlated: the most-duplicated functions are systematically the hardest, s | `phase-ends/logs/Phase17.md:247` | DK-73 | +| Phase17-18 | C2 | Another decomp project's `INCLUDE_ASM` is a record of what they did not crack, never proof that a class is unc | `phase-ends/logs/Phase18.md:52` | DK-78 | +| Phase17-18 | C3 | Your own corpus of byte-matches is an experiment you have already run on the toolchain: settle "is my rebuilt | `phase-ends/logs/Phase18.md:64` | DK-78 | +| Phase17-18 | C4 | Accelerate the stage that is actually the bottleneck: a byte-exact search loop costs a compiler invocation plu | `phase-ends/logs/Phase17.md:96` | DK-74 | +| Phase17-18 | C5 | A contradiction between two entries of your own knowledge base is a work item, not noise: replay the levers yo | `phase-ends/logs/Phase18.md:26` | DK-79 | +| Phase17-18 | C6 | Scope a research phase by a PER-CLASS VERDICT, not by a percentage: the deliverable is a validated lever or an | `phase-ends/logs/Phase18.md:93` | DK-73 | +| Phase19-20-22 | C1 | An open-ended grind phase's milestone is invariants held plus a clean checkpoint, never a percentage target | `phase-ends/logs/Phase19.md:24` | DK-73 | +| Phase19-20-22 | C2 | Choose the exemplar for cracking a codegen class by the SIZE of its residual: the rows at a one-instruction de | `phase-ends/logs/Phase20.md:14` | DK-73 | +| Phase21 | C1 | An unattended agent run that looks throttled is usually blocked on an interactive approval prompt; check the p | `phase-ends/logs/Phase21.md:731` | DK-75 | +| Phase21 | C2 | A class-distribution assessor only sees the population that has already been attempted; "analyse ALL remaining | `phase-ends/logs/Phase21.md:63` | DK-72 | +| Phase21 | C3 | Say which currency a wave buys — percentage or idioms — before launching it, and judge it in that currency | `phase-ends/logs/Phase21.md:475` | DK-72 | +| Phase23-27 | C1 | Mine new idioms from FRESH cracks, never from the failed backlog: your failure pile only re-teaches you what y | `phase-ends/logs/Phase23.md:93` | DK-74 | +| Phase23-27 | C2 | Give a drafting model a TARGETED slice of the knowledge base, never the whole thing: full context measurably m | `phase-ends/logs/Phase23.md:24` | DK-74 | +| Phase23-27 | C3 | Name the drafter's degenerate output in the prompt: an empty body compiles, so "translate EVERY instruction, n | `phase-ends/logs/Phase23.md:91` | DK-74 | +| Phase23-27 | C4 | The output-token cap is a two-sided knob, and both failure modes read as "the model is bad" | `phase-ends/logs/Phase23.md:54` | DK-74 | +| Phase23-27 | C5 | Build the unattended campaign so a crash is a pause: probe the dependency at the top of each cycle, commit per | `phase-ends/logs/Phase23.md:92` | DK-75 | +| Phase23-27 | C6 | Prove the aggregate check target is fail-closed before adding checks to it, and give every audit oracle a depe | `phase-ends/logs/Phase27.md:17` | DK-70 | +| Phase23-27 | C7 | A document that cites a repository path is an untested claim about the repository; lint it | `phase-ends/logs/Phase27.md:71` | DK-79 | +| Phase23-27 | C8 | Assert that the work ledger PARTITIONS the live work, and treat a row the invariant refutes as a lie a fresh s | `phase-ends/logs/Phase27.md:45` | DK-79 | +| Phase24 | C1 | Escalations to the expensive tier run STRICTLY SERIAL with idiom-banking between them; only the tier that cann | `phase-ends/logs/Phase24.md:57` | DK-74 | +| Phase24 | C2 | Inside one leverage class, schedule by MEASURED remaining effort — and pull the payoff-dominating outlier out | `phase-ends/logs/Phase24.md:29` | DK-73 | +| Phase24 | C3 | The wall-breaker tier is a MATCH tier, not a PLUMBING tier: work that a deterministic arbiter can judge does n | `phase-ends/logs/Phase24.md:61` | DK-74 | +| Phase24 | C4 | When a byte-proven body will not integrate, the repair moves the CALLER's declaration to the definition's sign | `phase-ends/logs/Phase24.md:119` | DK-77 | +| Phase25 | C1 | Leverage ordering does not predict tractability: carry a MEASURED closeness read per target, and never let (re | `phase-ends/logs/Phase25.md:473` | DK-73 | +| Phase25 | C2 | A fan-out script generated by the orchestrator runs sandboxed with NO access to the repo: it must be self-cont | `phase-ends/logs/Phase25.md:487` | DK-75 | +| Phase25 | C3 | Run reference-compiler dump/inspection passes from a scratch CWD (or set an explicit dump base); a `-da` run w | `phase-ends/logs/Phase25.md:337` | DK-80 | +| Phase26 | C1 | A round-trip selftest is a serialisation check, not a coverage check: it passes by construction when a parser' | `phase-ends/logs/Phase26.md:764` | DK-69 | +| Phase26 | C2 | A set that gates work must be reconstructible from committed artifacts. A roster kept in gitignored scratch is | `phase-ends/logs/Phase26.md:742` | DK-69 | +| Phase26 | C3 | A guard that is allowed to sit RED and UNWIRED does not exist. A detector's value is zero until it is green on | `phase-ends/logs/Phase26.md:101` | DK-69 | +| Phase28-32 | C1 | Never round-trip a curated config through a serializer: every oracle you own measures BYTES, so a formatting-d | `phase-ends/logs/Phase28.md:181` | DK-77 | +| Phase28-32 | C2 | The file a function lives in is not evidence of its class: read the recorded attribute, never the hosting spli | `phase-ends/logs/Phase28.md:182` | DK-72 | +| Phase28-32 | C3 | Re-verify a task's premise in the code at EXECUTION time: roadmap lines, audit findings, and even an audit's o | `phase-ends/logs/Phase28.md:143` | DK-71 | +| Phase28-32 | C4 | A hardcoded assumption fixed in one tool survives in its siblings: grep the tree for the literal and fix every | `phase-ends/logs/Phase28.md:36` | DK-77 | +| Phase28-32 | C5 | A batch gate that bisects on failure re-runs the singleton against an unchanged baseline: special-case n==1 or | `phase-ends/logs/Phase28.md:58` | DK-77 | +| Phase28-32 | C6 | The generated disassembly tree is shared mutable state: a fleet verify/clean chain and the per-function instru | `phase-ends/logs/Phase32.md:255` | DK-76 | +| Phase28-32 | C7 | Write the phase synthesis in a FRESH session that re-reads the committed state cold; the cold re-read is what | `phase-ends/logs/Phase28.md:86` | DK-79 | +| Phase29-1of4 | C1 | A stop/continue instrument must aggregate at exactly the unit the decision is made in; one that averages a fin | `phase-ends/logs/Phase29.md:1283` | DK-72 | +| Phase29-1of4 | C2 | Batch size is a RISK lever, not a token lever: isolated agents cost ~N× one agent whether concurrent or serial | `phase-ends/logs/Phase29.md:2387` | DK-75 | +| Phase29-1of4 | C3 | A repair/`--recover` mode whose cost is (exceptions × population) must be gated on a MEASURED exception count; | `phase-ends/logs/Phase29.md:1635` | DK-75 | +| Phase29-1of4 | C4 | A defect reasoned into a sibling tool is LATENT until a run shows its signature; do not patch it on theory rig | `phase-ends/logs/Phase29.md:1191` | DK-71 | +| Phase29-1of4 | C5 | Sweep for orphaned worker processes at every session boundary; a dead-pipe compiler or a self-matching wait lo | `phase-ends/logs/Phase29.md:1029` | DK-75 | +| Phase29-1of4 | C6 | Land a pure rename and a semantic/layout change as separate gated edits, so a gate failure attributes itself | `phase-ends/logs/Phase29.md:1959` | DK-77 | +| Phase29-2of4 | C1 | A verification flag that short-circuits the tool's WRITE path leaves the stale artefact in place and still exi | `phase-ends/logs/Phase29.md:2762` | DK-69 | +| Phase29-2of4 | C2 | A diagnosis earns belief when it predicts its own RESIDUAL membership, not when it explains the failures alrea | `phase-ends/logs/Phase29.md:3690` | DK-71 | +| Phase29-2of4 | C3 | Make the DRAFTER run the pre-gate guard and return the NAMED banking prerequisite; a batch of diagnosed candid | `phase-ends/logs/Phase29.md:5093` | DK-74 | +| Phase29-2of4 | C4 | When two blockers are orthogonal, a classifier's if-chain ORDER silently becomes the label — cross-tabulate, n | `phase-ends/logs/Phase29.md:4958` | DK-72 | +| Phase29-2of4 | C5 | Sequence a phase so the cheapest thing that can invalidate everything below it runs FIRST | `phase-ends/logs/Phase29.md:4322` | DK-73 | +| Phase29-2of4 | C6 | A reach-weighted gain figure (size × copies) is not a size; every number must say which of the two it is | `phase-ends/logs/Phase29.md:4775` | DK-72 | +| Phase29-2of4 | C7 | A repair ladder must probe whether each stage is NEEDED before applying it, or it silently escalates a binary- | `phase-ends/logs/Phase29.md:3377` | DK-77 | +| Phase29-3of4 | C1 | GNU C sources carry form-feed page separators, and `str.splitlines()` splits on them while `grep`/`sed` do not | `phase-ends/logs/Phase29.md:6915` | DK-78 | +| Phase29-3of4 | C2 | Never restructure a proven tool with blind string replaces at the end of a long session; specify it as the nex | `phase-ends/logs/Phase29.md:6134` | DK-79 | +| Phase29-3of4 | C3 | An oracle that can always be RUN is not always APPLICABLE: state the applicability precondition beside the rec | `phase-ends/logs/Phase29.md:7359` | DK-71 | +| Phase29-3of4 | C4 | Before a fleet-wide mechanical edit, census the whole population for the exact preconditions the edit assumes; | `phase-ends/logs/Phase29.md:7520` | DK-77 | +| Phase29-3of4 | C5 | A status line a script prints unconditionally is not a measurement; derive every conclusion the script emits f | `phase-ends/logs/Phase29.md:7496` | DK-69 | +| Phase29-3of4 | C6 | Measure what fraction of a cycle a parallelism knob can actually touch before adopting it: `make -j16` bought | `phase-ends/logs/Phase29.md:6115` | DK-72 | +| Phase29-4of4 | C1 | A clean-looking verdict that appears immediately AFTER your own repair transform is a SUSPECT, not a result: r | `phase-ends/logs/Phase29.md:8207` | DK-70 | +| Phase29-4of4 | C2 | Classify a harness fix as a LOGIC defect or a PATH-REACHABILITY gap and price it accordingly: only the logic d | `phase-ends/logs/Phase29.md:9957` | DK-71 | +| Phase29-4of4 | C3 | A wrong prescription left in the knowledge base is worse than no entry: when evidence refutes an entry you wro | `phase-ends/logs/Phase29.md:9029` | DK-79 | +| Phase29-4of4 | C4 | A name grep is not a "defined / banked here" oracle: a DECLARATION carrying the name reads as a definition; us | `phase-ends/logs/Phase29.md:8080` | DK-69 | +| Phase29-4of4 | C5 | A yield estimator that counts "unclaimed at the moment it runs" over-projects: it ranks correctly and overstat | `phase-ends/logs/Phase29.md:9350` | DK-72 | +| Phase29-4of4 | C6 | Declare a mechanical lever SPENT only on a positive, three-part measurement — every built lever applied and re | `phase-ends/logs/Phase29.md:10300` | DK-73 | +| Phase30-1of2 | C1 | An instrument's refusal is a FINDING, not an obstacle: overriding it means explaining why the instrument is wr | `phase-ends/logs/Phase30.md:1369` | DK-71 | +| Phase30-1of2 | C2 | Prose in an agent prompt is not enforcement: snapshot `git status` around every agent and name the offender | `phase-ends/logs/Phase30.md:449` | DK-76 | +| Phase30-1of2 | C3 | Read the FIRST ten results of a long run before trusting the other 190; and negative-control any NEW REFUSAL a | `phase-ends/logs/Phase30.md:250` | DK-71 | +| Phase30-1of2 | C4 | When a tool is repaired, every verdict it produced becomes a hypothesis again — but re-gate only the drafts th | `phase-ends/logs/Phase30.md:2481` | DK-70 | +| Phase30-1of2 | C5 | A ledger row with no draft artifact is a rumour, not a result | `phase-ends/logs/Phase30.md:2352` | DK-70 | +| Phase30-1of2 | C6 | Check group identity from the signature files BEFORE probing a family: if the members are structurally identic | `phase-ends/logs/Phase30.md:168` | DK-78 | +| Phase30-1of2 | C7 | Never test a helper by importing its module: a tool with no `__main__` guard runs its whole pipeline on import | `phase-ends/logs/Phase30.md:2163` | DK-76 | +| Phase30-1of2 | C8 | When A/B-ing a harness knob (model, effort, reasoning level), ship a POSITIVE CONTROL that the knob actually m | `phase-ends/logs/Phase30.md:1541` | DK-74 | +| Phase30-1of2 | C9 | Keep the wave harness IN THE REPO with its contracts; a harness rebuilt from memory each run silently goes sta | `phase-ends/logs/Phase30.md:688` | DK-76 | +| Phase30-1of2 | C10 | Do not cross-price two economies: a conversion rate measured on the RESIDUE QUEUE does not price a FRESH wave | `phase-ends/logs/Phase30.md:899` | DK-72 | +| Phase30-1of2 | C11 | Derive the target pool from the BUILD'S OWN INVARIANT, not from a reach/similarity metric — reach ranked the m | `phase-ends/logs/Phase30.md:1810` | DK-73 | +| Phase30-1of2 | C12 | A metric that re-parses source is blind to a body banked through an `#include`; trust the metrics derived from | `phase-ends/logs/Phase30.md:2589` | DK-69 | +| Phase30-1of2 | C13 | Keep a glossary line for any term two documents use in OPPOSITE senses | `phase-ends/logs/Phase30.md:1343` | DK-72 | +| Phase30-2of2 | C1 | Drafting agents must never be able to write the build tree; every agent artifact lands in a scratch directory, | `phase-ends/logs/Phase30.md:4017` | DK-76 | +| Phase30-2of2 | C2 | A stored verdict can be stale because the INSTRUMENT changed, not the draft: re-gate the drafts a tool repair | `phase-ends/logs/Phase30.md:4660` | DK-70 | +| Phase30-2of2 | C3 | Count agent completions from the run journal's result records, never from artifact existence: an agent writes | `phase-ends/logs/Phase30.md:4001` | DK-70 | +| Phase30-2of2 | C4 | A derived claim outranks a heuristic verdict; when two heuristics disagree, take the UNION and queue the disag | `phase-ends/logs/Phase30.md:4657` | DK-71 | +| Phase30-2of2 | C5 | A failure that will not reproduce earns a negative-control-proven DETECTOR, not a speculative fix; and every a | `phase-ends/logs/Phase30.md:4064` | DK-71 | +| Phase30-2of2 | C6 | Never let model-authored prose reach the shell inside double quotes: a backticked command in a `git commit -m` | `phase-ends/logs/Phase30.md:2893` | DK-76 | +| Phase31-1of3 | C1 | More output-token budget is NOT more quality; measure it as a paired A/B and treat truncation as recoverable, | `phase-ends/logs/Phase31.md:2292` | DK-74 | +| Phase31-1of3 | C2 | In a pipelined drafter/gater, "still open" is not "not yet attempted": consecutive waves re-drafted the wave s | `phase-ends/logs/Phase31.md:258` | DK-75 | +| Phase31-1of3 | C3 | Snapshot every target's disassembly BEFORE gating: a successful bank prunes it, and every downstream step need | `phase-ends/logs/Phase31.md:1612` | DK-76 | +| Phase31-1of3 | C4 | Draft first, then carve: a build-unit split is only safe when the new unit is immediately populated with prove | `phase-ends/logs/Phase31.md:922` | DK-77 | +| Phase31-1of3 | C5 | A cracked idiom transfers WITHIN its family and not across it: price a lane by families, not by class size | `phase-ends/logs/Phase31.md:1300` | DK-73 | +| Phase31-1of3 | C6 | Sibling count and never-drafted count are different denominators; conflating them overstated the free-remap le | `phase-ends/logs/Phase31.md:298` | DK-72 | +| Phase31-1of3 | C7 | "Never write the tree" in a drafting prompt is a request, not an enforcement | `phase-ends/logs/Phase31.md:2879` | DK-76 | +| Phase31-1of3 | C8 | A free or preview model tier can be withdrawn mid-campaign without notice; a fleet-wide 404 is an epoch event, | `phase-ends/logs/Phase31.md:2319` | DK-75 | +| Phase31-1of3 | C9 | A metered API key's own cap is a separate limit from the account's credit | `phase-ends/logs/Phase31.md:2380` | DK-75 | +| Phase31-1of3 | C10 | Incremental gates never exercise the extraction/regeneration step, so regeneration rot is undated and invisibl | `phase-ends/logs/Phase31.md:2376` | DK-70 | +| Phase31-1of3 | C11 | A probe that splices the tree in order to measure it must restore it, or the progress oracle counts the splice | `phase-ends/logs/Phase31.md:2598` | DK-76 | +| Phase31-1of3 | C12 | Keep prompt/law text in data, not inside the launcher's source template | `phase-ends/logs/Phase31.md:1702` | DK-74 | +| Phase31-2of3 | C1 | The health suite must assert that a tool DID ITS WORK, not only that the data is intact: zero inputs, an impos | `phase-ends/logs/Phase31.md:5122` | DK-70 | +| Phase31-2of3 | C2 | A health check that cannot finish is not a check: keep the health target sampled and fast, and put the exhaust | `phase-ends/logs/Phase31.md:5175` | DK-70 | +| Phase31-2of3 | C3 | "Independent" names the INSTRUMENT, not the input: two refusals of two separately-written drafts from one tool | `phase-ends/logs/Phase31.md:3589` | DK-71 | +| Phase31-2of3 | C4 | A drafting agent must run where it CANNOT write the source tree; "never modify src/" in a prompt is a request, | `phase-ends/logs/Phase31.md:2879` | DK-76 | +| Phase31-2of3 | C5 | A proven transform that is not a rung of the ladder the agents' drafts actually pass through does not exist fo | `phase-ends/logs/Phase31.md:3810` | DK-77 | +| Phase31-2of3 | C6 | The knowledge-harvest selector must be able to see FAILED attempts; a filter that can only read banked work le | `phase-ends/logs/Phase31.md:3236` | DK-74 | +| Phase31-2of3 | C7 | A refusal names the branch the caller entered, not the subject — make the applier consult the classifier it al | `phase-ends/logs/Phase31.md:3346` | DK-71 | +| Phase31-2of3 | C8 | Every status claim in an agent's context must be expiry-checked against live state, or agents will report it b | `phase-ends/logs/Phase31.md:5394` | DK-70 | +| Phase31-2of3 | C9 | A validity stamp must be honoured by every downstream consumer; and a wrong write-side label cannot be repaire | `phase-ends/logs/Phase31.md:5437` | DK-70 | +| Phase31-2of3 | C10 | Yield is clustered by binary, not spread over the fleet: draw per-binary once two independent lanes concentrat | `phase-ends/logs/Phase31.md:5040` | DK-73 | +| Phase31-2of3 | C11 | (minor) A build-system conditional that expands at PARSE time makes its own negative control vacuous | `phase-ends/logs/Phase31.md:5155` | DK-77 | +| Phase31-3of3 | C1 | Never wrap a project tool in a `timeout` shorter than its own internal budget; you pre-empt its documented rec | `phase-ends/logs/Phase31.md:7024` | DK-75 | +| Phase31-3of3 | C2 | A programmatic edit to a long-lived knowledge document silently truncates or duplicates it; verify the section | `phase-ends/logs/Phase31.md:6529` | DK-79 | +| Phase31-3of3 | C3 | The live hand-off block must be strictly APPENDED at the end of its file: file order is the only recency signa | `phase-ends/logs/Phase31.md:6126` | DK-79 | +| Phase31-3of3 | C4 | A long stateless batch must persist each confirmed result the moment it is confirmed | `phase-ends/logs/Phase31.md:7121` | DK-75 | +| Phase31-3of3 | C5 | "Idempotent by SKIPPING" seals an artifact against later evidence; make regenerated artifacts idempotent by RE | `phase-ends/logs/Phase31.md:6171` | DK-77 | +| Phase31-3of3 | C6 | A verification target too slow to complete is not a check: sample it by default and keep the exhaustive form a | `phase-ends/logs/Phase31.md:6027` | DK-70 | +| Phase31-3of3 | C7 | Never adopt a subagent's worktree wholesale: it is a snapshot of an older tree and may predate a bank; re-gate | `phase-ends/logs/Phase31.md:6808` | DK-76 | +| Phase33 | C1 | A Ghidra script directory compiles as ONE bundle: a single non-compiling script disables every script in it | `phase-ends/logs/Phase33.md:186` | DK-80 | +| Phase33 | C2 | Ghidra refuses a project path containing a component that starts with `.` — a scratch project cannot live unde | `phase-ends/logs/Phase33.md:185` | DK-80 | +| Phase33 | C3 | `git check-ignore` is SILENT for tracked paths: an ignore-coverage audit run before the untracking passes vacu | `phase-ends/logs/Phase33.md:236` | DK-80 | +| Phase33 | C4 | A coverage instrument that infers its denominator from OPEN work INVERTS at 100% — carry the scanned denominat | `phase-ends/logs/Phase33.md:126` | DK-69 | +| Phase33 | C5 | An annotator that writes into the text it reads must never treat its own output as evidence | `phase-ends/logs/Phase33.md:499` | DK-69 | +| Phase33 | C6 | Regex-extracted evidence needs a structural marker, or prose becomes data | `phase-ends/logs/Phase33.md:503` | DK-69 | +| Phase33 | C7 | `objdump -dr` interleaves relocation records only for OBJECT files; a linked ELF lists them separately and shi | `phase-ends/logs/Phase33.md:526` | DK-78 | +| Phase33 | C8 | A derive-then-apply pipeline over a live repository needs a freshness guard, and a stated sequencing law | `phase-ends/logs/Phase33.md:285` | DK-79 | +| Phase33 | C9 | A content-hash "no ROM bytes" audit collides on zero-length files | `phase-ends/logs/Phase33.md:205` | DK-80 | +| Phase33 | C10 | `git push --mirror` does not push `refs/stash` | `phase-ends/logs/Phase33.md:295` | DK-80 | +| Phase33 | C11 | Route a host purge request through the flow that actually exists; the obvious form is a trap | `phase-ends/logs/Phase33.md:635` | DK-80 | +| Phase33 | C12 | A host feature can be gated on the very flip it was meant to precede: read the settings page, don't infer | `phase-ends/logs/Phase33.md:692` | DK-80 | +| Phase33 | C13 | A public scratch service's compiler image is NOT your pinned toolchain; rebuild it locally and prove byte-iden | `phase-ends/logs/Phase33.md:645` | DK-80 | +| Phase33 | C14 | A miner over your own records finds only what its pattern anticipates: measure the widened pattern's yield | `phase-ends/logs/Phase33.md:429` | DK-79 | +| Phase33 | C15 | A harness's low-memory guard silently kills a long BACKGROUND job — and workers from closed phases survive for | `phase-ends/logs/Phase33.md:508` | DK-75 | +| Phase33 | C16 | A `cd` in one agent shell call persists into the next | `phase-ends/logs/Phase33.md:660` | DK-76 | +| Phase7 | C1 | A decompiler's `unaff_` output means it decompiled ONE entry path of a multi-entry function and handed yo | `phase-ends/logs/Phase7.md:109` | DK-78 | +| Phase7 | C2 | A milestone counted in "matched functions" must exclude the splitter's auto-generated empty bodies; define the | `phase-ends/logs/Phase7.md:12` | DK-72 | +| Phase7 | C3 | Duplicate/family census run before the vendor library is linked out is contaminated: the duplicate groups are | `phase-ends/logs/Phase7.md:26` | DK-72 | +| Phase7 | C4 | A byte-identical baseline is only proven when the whole gate is green from clean across several INDEPENDENT se | `phase-ends/logs/Phase7.md:75` | DK-77 | +| Phase7 | C5 | When a foundation task hits a structural wall mid-phase, reorder: bank the tractable wins first and re-scope t | `phase-ends/logs/Phase7.md:3` | DK-73 | +| Phase8-13 | C1 | A toolchain/SDK-version detector's verdict is a HYPOTHESIS until a placement COUNT backs it; when a version st | `phase-ends/logs/Phase10.md:44` | DK-71 | +| Phase8-13 | C2 | Histogram a secondary binary's `jal` targets BY ADDRESS RANGE before assuming it carries its own copy of anyth | `phase-ends/logs/Phase12.md:45` | DK-78 | +| Phase8-13 | C3 | A raw blob's load address is a hypothesis; the free confirmation is arithmetic against the NEXT known segment' | `phase-ends/logs/Phase10.md:13` | DK-78 | +| Phase8-13 | C4 | ⚠ LOW VALUE, recommend dropping: during a parameterization refactor whose oracle is a byte-locked binary, gene | `phase-ends/logs/Phase9.md:26` | DROP: the miner's own verdict: low value — a refactoring generality with no measured cost | diff --git a/.run/P33.5/log-mining/Phase15-16.md b/.run/P33.5/log-mining/Phase15-16.md new file mode 100644 index 0000000000..6d41f39346 --- /dev/null +++ b/.run/P33.5/log-mining/Phase15-16.md @@ -0,0 +1,104 @@ +# Log mining — Phase15-16 +Files/ranges: `phase-ends/logs/Phase15.md`:1-71 · `phase-ends/logs/Phase16.md`:1-96 · Lines read: 167 of 167 +Candidates considered: 28 · NEW: 8 · ALREADY-BANKED: 19 (+1 dropped as not-a-lesson, listed at the end) + +Grep file set used for every "already banked?" test (abbreviated `$F` below): +`docs/matching-cookbook.md docs/cookbook-index.md docs/decision-log.md docs/accelerators.md docs/retrospective.md +docs/wave-playbook.md docs/how-to-ai-decomp/*.md decomp-architect/corpus/decomp-kernels.md +decomp-architect/templates/registry-E.decomp.md phase-ends/DIGEST.md` + +## NEW + +### C1 — Measure a pipeline's yield on the RESIDUAL, not on already-solved functions: build a known-answer ladder (revert a matched function to a stub and make the pipeline re-derive it), then read the gap between that number and the unmatched tail — the gap is the go/no-go +- **Evidence:** `phase-ends/logs/Phase16.md:36-37`, `:48` — + > "on 12 already-matched fns (known-reachable answers), **m2c-DIRECT re-derivation = 8/12 = 67%**" + > "**Unmatched smallest-80 macro-only-no-permuter:** 2/80 byte-gated … The gap vs 67% confirms the unmatched tail is self-selected hard" + > "**Known-answer WHOLE-BINARY capability (the honest number):** m2c+sig_unify = **4/16 = 25%** … Unmatched hard tail ≈ **1%**." +- **What happened / what it cost:** Phase 16 planned a multi-day unattended run against a computed "85-90% ceiling" (`:50`: "the '85-90% ceiling' was theoretical"). Drew's known-answer method — take an ALREADY-matched function, revert it to an `INCLUDE_ASM` stub, re-extract its `.s`, run the FULL pipeline and confirm it re-derives the byte-match already known correct (`:82`, `:87`) — gave 67% (per-function oracle) then 25% (whole-binary gate) then ~1-2% on the actual unmatched residual. Three numbers for the same pipeline, spanning 60 points; only the last one was the run's real yield, and it killed the run (`:63-69`). The measurement cost one weekend; the run would have cost five days. +- **Not banked — greps:** `grep -n -i 'known.answer' $F` → 2 (both about mining known answers from prose / a known-answer population for a build-path bug, not the revert-a-match method); `grep -n -i 'graduated' $F` → 0; `grep -n -i 'difficulty ramp' $F` → 0; `grep -n -i 'yield baseline' $F` → 0; `grep -n -i 'self.select' $F` → 0; `grep -n -i 'hard by selection' $F` → 0; `grep -n -i 'prove the pipeline' $F` → 0. The nearest banked items are DK-61 (`decomp-kernels.md:762`, a known-TRUE case before reading an *instrument's* output — about scans/censuses, not about measuring a matching pipeline's yield) and `docs/how-to-ai-decomp/04-oracles-and-instruments.md:82` (the same instrument rule). Neither states that a capability measured on solved work is an upper bound, nor prescribes the revert-to-stub ladder. +- **Proposed home:** DK (a kernel), sitting next to DK-8 (census the shape) and DK-15 (the permuter) — "before committing compute to an automated matching campaign, measure it three ways: on known answers, through the whole-binary gate, and on the actual residual; the residual number is the only one that funds the run." +- **Portable because:** every decomp reaches a point where a pipeline is proposed for a long unattended run, and every decomp's remaining tail is self-selected hard by everything that already succeeded — the arithmetic of survivorship is not project-specific. + +### C2 — A search harness (permuter) must compile in the SAME declaration context as the real build; an isolated context does not merely fail to verify, it makes the search converge on the wrong answer +- **Evidence:** `phase-ends/logs/Phase16.md:71-73` — + > "Overnight permuter 'closed' 17/40 near-misses (42%) — BUT **0/17 whole-binary-gated.** Root cause: the permuter's `base.c` STRIPPED callee externs → compiled with implicit-`int` callees → matched the target in the WRONG signature context." + > "**FIX:** `p16_permute.make_base_c` now KEEPS the canonical externs … `winner_to_draft` strips only the TYPEDEFS block." +- **What happened / what it cost:** an overnight 4.5-hour permuter run (40 near-misses × 7 min) reported a 42% close-rate — the single number the phase's go/no-go depended on. Every one of the 17 wins was an artifact: with callee externs stripped for self-containment, the callees became implicit-`int`, so the optimizer was hill-climbing toward a *different* target than the real build produces. The whole-binary gate caught it (0/17), but a whole night of compute and the phase's key measurement were spent on a number that meant nothing. +- **Not banked — greps:** `grep -n -i -E 'signature context' $F` → 0; `grep -n -i -E 'stripped.*extern|extern.*stripp' $F` → 1 (`matching-cookbook.md:1752`, a redundant inline *typedef*, different subject); `grep -n -i -E 'permuter.*extern' $F` → 1 (`how-to-ai-decomp/07-compiler-source.md:48`, "a residual on a call to an external callee is outside its search space" — a different claim); `grep -n -i 'illusory' $F` → 0; `grep -n -i '17/40' $F` → 0. **Worse than absent:** the banked `base.c` recipe actively invites the bug — `matching-cookbook.md:112` says *"`base.c` — the near-match C, **self-contained**"* and names only the typedefs as what to inline. The related banked rules cover verification, not search: `accelerators.md:365-378` / DK-20 ("a claim names the compilation it survived", the `match_one` flavor) and `matching-cookbook.md:176-177` ("a permuter `output-0-*` is a strong CANDIDATE to gate, never a bank"). +- **Proposed home:** cookbook §3 (amend the `base.c` recipe: keep the canonical callee/data externs, strip only what the TU already provides) **and** a DK addendum extending DK-20/DK-21 from "an isolated verdict describes its environment" to "an isolated *search* optimises for its environment — a false positive, not just an unverified one". +- **Portable because:** every decomp that runs decomp-permuter (or any automated mutation search) builds a reduced translation unit for it; the reduction is where the target silently changes. + +### C3 — Before concluding a drafting pipeline is weak, histogram the compiler's error TEXT: over half of ours were one missing declaration in the shared header, fixed byte-neutrally in one line +- **Evidence:** `phase-ends/logs/Phase16.md:40`, `:47` — + > "**CC1-fail breakdown: 58 = `NULL` undeclared (TRIVIAL fix — add to common.h), ~15 stack-struct (sp* vars → --stack-structs), ~10 m2c-incomplete.**" + > "**NULL fix DONE** (byte-neutral, committed c8c9fd20c): recovers **54/106** CC1-fails. Compiling drafts now 341/393." +- **What happened / what it cost:** of 393 characterized targets, 106 failed to compile at all — 27% of the population, read at first as pipeline capability. Grouping the failures by the compiler's message showed 58 of the 106 were the single token `NULL`, undeclared in the project's `common.h`. One byte-neutral header line recovered 54 and moved compiling drafts from 287 to 341. The pipeline was never as weak as its first pass-rate said. +- **Not banked — greps:** `grep -n -i -E 'NULL.*undeclared|undeclared.*NULL' $F` → 0; `grep -n -i -E 'histogram' $F` → 6 (all opcode/register histograms for similarity work); `grep -n -i -E 'compile failure.*classif|bucket.*compile error' $F` → 0; `grep -n -E '\bNULL\b' $F` → 26, none about a missing `NULL` declaration. The adjacent banked items are the *downstream* ladder (`accelerators.md:282` "CC1-FAIL → the recovery ladder"; `wave-playbook.md:563-566` "a `CC1-FAIL` says the declaration blocked COMPILATION") and one sibling instance (`matching-cookbook.md:1344`: `common.h` lacking `s64/u64/f64`) — but nothing says *group the failures by message text first*, and the `NULL` case itself is nowhere. +- **Proposed home:** accelerator (or a G-rule): "a compile-failure population is triaged by error text before it is triaged by function; fix the shared header once and re-run the gate before spending a single agent token." +- **Portable because:** every decomp hands generated C to an old compiler through one shared preamble; a single missing declaration in that preamble is indistinguishable from model failure until the messages are counted. + +### C4 — Re-run the deterministic declaration-canonicaliser over OLD quarantined drafts after every large bank: a draft's recovery odds rise as the banked corpus grows, and the pile costs nothing to keep +- **Evidence:** `phase-ends/logs/Phase15.md:57` — + > "`canon_draft_decls` on the combined 505 quarantined (now canonicalizing against the much-larger banked set: 268 decls rewritten) → **47 verified** (1175→1128 stubs) → **42 propagated** fleet-wide … Fleet 44.64% → 46.27% (+1.63)." +- **What happened / what it cost:** drafts that failed the whole-binary gate were quarantined to `.run/drafts-*-fail/` rather than discarded (`:58`, `:61`). A later deterministic pass over the *combined* pile — with a canonical set now far larger than when those drafts were first gated — rewrote 268 declarations and banked 47 more functions for ~0 agent tokens. The same tool over the same drafts had already run once and left them; what changed was the corpus, not the drafts. +- **Not banked — greps:** `grep -n -i 'banked set grows' $F` → 0; `grep -n -i 'as the corpus grows' $F` → 0; `grep -n -i -E 're-run.*recovery' $F` → 1 (`matching-cookbook.md:2231`, which asserts the *opposite* for one case: "close=0 + already-recovery-gated = a WALL … do NOT re-run recovery on it"); `grep -n -i 'retry the quarantine' $F` → 0; `grep -n -i 'quarantin' $F` → 1 (`matching-cookbook.md:1283`, quarantine only "so re-gates stay fast"). §14d (`matching-cookbook.md:1288-1300`) banks *that* deterministic recovery beats agent waves; it does not bank the time-dependence — that a re-run of the same recovery on the same drafts pays again later. +- **Proposed home:** cookbook §14d addendum + a DK line under DK-16 (the widening review): "which recovery just got a bigger canonical set?" alongside "which scanner's denominator just got wider?". +- **Portable because:** it is the same compounding property as the twin-graph law already banked (`registry-E.decomp.md:293`, "a bank changes the twin graph") applied to declarations: every bank enlarges the canonical vocabulary that a mechanical repair pass draws on. + +### C5 — Write numeric kill-criteria into the plan, before the data exists, for any expensive campaign — and let them fire +- **Evidence:** `phase-ends/logs/Phase16.md:24`, `:58` — + > "**S3** 10-medium validation (Max) → **GATE-B:** ≥6/10 byte-identical; record permuter yield (<4/10 → STOP run)" + > "K1 type-graph won't compile (GATE-A <8/10) → STOP … K2 permuter yield <4/10 (GATE-B) → STOP unattended plan. K3 driver corrupts state (GATE-D not 136/136) → no unattended run … K4 silent under-propagation → assert fleet-% rose per batch; flat-but-green → halt." +- **What happened / what it cost:** the phase declared four kill-criteria and three numbered gates with explicit thresholds at plan time, then measured against them over three days and stopped: "**PAUSE the brute-force approach (yields ~3%, bounded by the loose-typing wall, not a fixable bug)**" (`:63-65`). The thresholds existed before the disappointing numbers arrived, which is why a 3% yield read as a stop instead of as "nearly there". The harness work built for the run was kept and reused (`:66`). +- **Not banked — greps:** `grep -n -i 'kill.criteri' $F` → 0; `grep -n -i -E 'declare.*threshold|threshold.*before' $F` → 0 (1 unrelated hit on a variable named `threshold`); `grep -n -i 'stop condition' $F` → 2 (`how-to-ai-decomp/01-governance.md:47`, the harness stop conditions for a failing check; `decision-log.md:3497`, a permuter scorer's zero). The governance chapter banks *two human gates per phase* (`01-governance.md:31,45-52`) — direction approval, not a pre-committed number that can kill a plan the human already approved. +- **Proposed home:** G (a rule) in the kit's registry, adjacent to the two-gates protocol: "an expensive campaign's plan names the measurement, the threshold and the action at each gate before the measurement is taken." +- **Portable because:** it is a defence against motivated reasoning, not against a compiler; it applies to any project where a human approves a plan and an agent then produces the numbers that judge it. + +### C6 — Design an unattended run for a human with no agent session: a STOP file honoured at a safe boundary, a supervisor that tells a clean exit from a crash, a status one-liner — and prove crash-resume with a deliberate kill before trusting it +- **Evidence:** `phase-ends/logs/Phase16.md:89`, `:27` — + > "**Safe-exit mechanism (REQUIRED).** Driver checks `.run/auto/STOP` at every function boundary; if present → finish current fn's gate+propagate+commit → final heartbeat 'stopped safely' → exit 0; supervisor sees STOP + clean exit → does NOT relaunch. **Trigger:** … or Drew runs the one-liner himself (works with no Claude session)." + > "**S6** Hands-on trial (xHigh) → **GATE-D:** multi-hr pass + forced kill→auto-resume→`check-all` 136/136 + negative control" +- **What happened / what it cost:** the owner was leaving for several days; the run had to be stoppable and inspectable by a person who could not open an agent session. Four small tools were specified for that alone (`:39`: driver with STOP-sentinel safe-exit + heartbeat, a pure-bash supervisor that relaunches on crash and reaps stray permuters, `auto_stop.sh`, `auto_status.sh`), and the go/no-go gate included *deliberately killing the run* and proving it resumed to 136/136 green. Writing this before the run was the condition of the run existing at all. +- **Not banked — greps:** `grep -n -i 'safe.exit' $F` → 0; `grep -n -i 'heartbeat' $F` → 0; `grep -n -i -E 'forced kill|kill.*resume|auto-resume' $F` → 0; `grep -n -i 'sentinel' $F` → 14 (one relevant: `matching-cookbook.md:16770`, "stop the grinder (STOP sentinel)" as a concurrency note, no design or drill); `grep -n -i 'unattended' $F` → 15, all about R55/G30 (`DIGEST.md:240`, `registry-E.decomp.md:199-201`: an unattended lane must leave *evidence*). Evidence-of-liveness is banked; stoppability, clean-vs-crash exit and the kill drill are not. +- **Proposed home:** DK (a kernel) or a G-rule extending G30: "an unattended lane is stoppable and readable without the agent, and its resume path is proven by a deliberate kill before the first real run." +- **Portable because:** every long autonomous agent run eventually runs while its owner is away or its session is dead; the properties needed are independent of what the run is matching. + +### C7 — The permuter's search unit is the C expression: it cannot freeze the instructions you already have right, because register allocation couples them +- **Evidence:** `phase-ends/logs/Phase16.md:43` — + > "**Drew Q&A (permuter):** not one-shot-from-nothing (m2c draft = the info/jumping-off point); function-by-function not whole-file; permuter hill-climbs (additive) but can't freeze individual instructions (regalloc couples them) — compositional fixes happen at the C-expression level." +- **What happened / what it cost:** recorded as the answer to the owner's direct question about what the tool can and cannot do, at the moment the phase was deciding how much of the plan to hang on it. No cost is attached in the log — it is a calibration fact captured before the money was spent, which is exactly when a day-one kit wants it. +- **Not banked — greps:** `grep -n -i 'freeze' $F` → 0; `grep -n -i -E 'cannot pin.*instruction' $F` → 0; `grep -n -i -E 'permuter.*extern' $F` → 1 (unrelated). The banked statements are about the *search space* rather than the coupling: `matching-cookbook.md:123-125` ("NOT in the permuter's C-randomization search space; it needs a structural insight or `PERM_*` macros") and `how-to-ai-decomp/07-compiler-source.md:48` ("outside its search space entirely"). Neither says why partial correctness cannot be locked in. +- **Proposed home:** cookbook §3 (one line in the harness description) / DK-15 addendum. Lower-value than C1-C6; include only if the kernel has room. +- **Portable because:** it is a property of any allocator-coupled backend plus a C-level mutation search, and it sets the expectation — improvements are compositional at the source level or not at all — that decides whether the permuter is a closer or a strategy. + +### C8 — An idempotency guard keyed on PRESENCE freezes every record created before the system matured; key it on COMPLETENESS, or re-audit records written before the last capability jump +- **Evidence:** `phase-ends/logs/Phase15.md:51` — + > "Minor: the 4 Phase-13/T2 proof groups (E_func_80128EA8/8012A568/80132EC4/80138C30) are at 16 members not 134 (registered before the bulk, so --auto-from skips them) — negligible; top up later if desired." +- **What happened / what it cost:** `dedup_propagate --auto-from` skips addresses already in the registry — the property that makes it additive and resumable (`matching-cookbook.md:1205`). Four functions proven during the 16-binary era therefore stayed registered at 16 members after the fleet grew to 134, permanently and invisibly: the byte-gate is green either way, and the tool will never revisit them. The log's own disposition ("negligible; top up later if desired") is the lesson's shape — the class is cheap to note, easy to defer, and its cost is only ever paid later. +- **Not banked — greps:** `grep -n -i -E 'already registered|skip.*already.registered' $F` → 1 (`matching-cookbook.md:1205`, which documents the skip as a *feature* and says nothing about stale members); `grep -n -i 'under.propagat' $F` → 1 (`matching-cookbook.md:1303`, the `find_site` brace bug — a *detector-format* cause, plus "always sanity-check the OUTCOME metric"); `grep -n -i 'unpropagated' $F` → 1 (`matching-cookbook.md:10329`, a later symptom carried for three sessions); `grep -n -i 'member count' $F` → 4 (verify a handoff's ×1 claim; add new split suffixes) — the *presence-keyed skip* mechanism is in none of them. +- **Proposed home:** cookbook §14d addendum + a DK line: "a resumability guard is keyed on the record being COMPLETE (reach == population), never on it existing." +- **Portable because:** any decomp with a propagation or bank registry adds the same skip-if-present guard for resumability, and every such project's population grows after its first records are written. + +## ALREADY-BANKED (one line each) +- Header-dependency tracking is mandatory once shared headers are build inputs; without it a header-only edit gives a stale, falsely-passing incremental check — lives at `docs/matching-cookbook.md:1176-1179` +- Callee-signature-aware target manifests break the extern-type-conflict wall (naive 16% → 60-67% gate yield) — lives at `docs/matching-cookbook.md:1232-1253` (§14c; index `docs/cookbook-index.md:557`) +- Leaf-first / easy-first ordering (no `jal`/`%hi`/`%lo` ⇒ the per-function oracle is exact; ~69% vs ~16% on call-heavy) — lives at `docs/matching-cookbook.md:1021` and `:1224` +- Under a sustained throttle, throttle the FAN-OUT: sequential waves of ≤10 agents beat a 48-batch burst (~42 drafts/60 s and 0 dead batches vs ~1/45 s and 33 dead) — lives at `docs/matching-cookbook.md:1014-1021` +- A retry wave misses truncated-but-non-null results; reconcile produced artifacts against the expected work-list after every fan-out — lives at `docs/matching-cookbook.md:1024-1032` +- The whole-binary gate cannot see an un-attempted target (a leftover stub is itself byte-identical) — lives at `docs/matching-cookbook.md:1027-1030` +- Deterministic recovery beats agent waves on the hard tail (a 50-agent wave: 27 banks for ~4.1M tokens; a deterministic pass: +2.67% for ~0) — lives at `docs/matching-cookbook.md:1288-1300` (§14d) +- Drafters invent colliding generic type names inline; the cure is one canonical shared type file from the first bank, one definition per shape — lives at `docs/decision-log.md:3598-3615`, `decomp-architect/corpus/decomp-kernels.md:798-813` (DK-65), `decomp-architect/templates/registry-E.decomp.md:414-418` +- Match once → propagate ×N, and bulk-propagate what is ALREADY matched before harvesting anything new — lives at `decomp-architect/corpus/decomp-kernels.md:68-77` (DK-5), `docs/accelerators.md:47` (A3), `docs/matching-cookbook.md:1150-1151` +- A mechanical fleet-wide edit needs a compile pre-filter plus a fail-closed byte-gate (bodies using binary-local types are skipped honestly) — lives at `docs/matching-cookbook.md:1194-1197` +- Silent under-propagation: a green byte-gate proves what landed is correct, never that everything that should have propagated did — check the OUTCOME metric — lives at `docs/matching-cookbook.md:1303-1309` +- At fleet scale the reporting tool becomes the bottleneck; cache the registry parse (136 parses → 1) — lives at `docs/matching-cookbook.md:1208`; the general bar "a slow gate is a bug" at `docs/how-to-ai-decomp/02-byte-gate.md:48` and `decomp-architect/corpus/decomp-kernels.md:439` +- Workflow argument-passing gotcha (the `args` payload does not arrive as you assume — pass an array / hardcode, defensively parse) — lives at `docs/matching-cookbook.md:1700-1702`; the "never type a payload, generate it" practice at `docs/wave-playbook.md:319-328` +- Never `git checkout src/.c` during a harvest — it silently reverts banked matches; use a `.run/_bak.c` copy — lives at `docs/matching-cookbook.md:1276-1279` +- Struct typing is byte-neutral in gcc-2.7.2 (`p->field` ≡ the cast form); types are a banking/width lever, not a codegen lever — lives at `decomp-architect/corpus/decomp-kernels.md:803` (DK-65), `phase-ends/DIGEST.md:86` (P17), `docs/matching-cookbook.md:1336-1339` +- m2c `--valid-syntax` + a byte-faithful macro header compiles without the structs; the non-faithful macros (`M2C_ERROR`/`MULT_HI`/`CLZ`…) mean "defer, not matchable" — lives at `docs/matching-cookbook.md:1336-1338` (§15) +- A per-function match is a statement about the BODY; banking is a statement about the TU — the isolated oracle's output is a gate-first candidate, never a bank — lives at `docs/accelerators.md:362-378` (#16 / DK-20), `docs/matching-cookbook.md:1046` +- A permuter `output-0` is a strong candidate to gate, never a bank; the whole-binary byte-gate is the sole arbiter — lives at `docs/matching-cookbook.md:176-177` +- Census the corpus for byte-identical duplication (the free tier) before choosing a strategy — covers Phase 15's "134 overlays, 129 distinct by sha1, 5 free dup pairs" — lives at `docs/how-to-ai-decomp/03-bootstrap-order.md:33`, `decomp-architect/corpus/decomp-kernels.md:108-120` (DK-8) + +## Considered and dropped (not a lesson) +- "Ghidra-free by default; escalate to per-overlay → per-area disassembler import only when the loop stalls" (`phase-ends/logs/Phase15.md:6`, `:45`) — a plan-time owner policy with no recorded outcome or cost anywhere in the slice; the escalation is never shown to have been taken. Recording it would be inventing a lesson the log does not contain. (Greps run anyway: `grep -n -i -E 'ghidra.free|escalate to ghidra|ghidra.*only when' $F` → 3, all incidental "Ghidra-free tooling" mentions.) diff --git a/.run/P33.5/log-mining/Phase17-18.md b/.run/P33.5/log-mining/Phase17-18.md new file mode 100644 index 0000000000..169f8efa33 --- /dev/null +++ b/.run/P33.5/log-mining/Phase17-18.md @@ -0,0 +1,102 @@ +# Log mining — Phase17-18 +Files/ranges: `phase-ends/logs/Phase17.md`:1-301 · `phase-ends/logs/Phase18.md`:1-237 · Lines read: 538 of 538 +Candidates considered: 32 · NEW: 6 · ALREADY-BANKED: 26 + +Grep denominator for every "not banked" test below (the brief's file set, expanded): +`docs/matching-cookbook.md docs/cookbook-index.md docs/decision-log.md docs/accelerators.md docs/retrospective.md +docs/wave-playbook.md docs/how-to-ai-decomp/*.md decomp-architect/corpus/decomp-kernels.md +decomp-architect/templates/registry-E.decomp.md phase-ends/DIGEST.md` — referred to below as `$F`. + +## NEW + +### C1 — Leverage and tractability are ANTI-correlated: the most-duplicated functions are systematically the hardest, so a leverage-first queue front-loads hand-tier work and its early bank rate is not a harness fault +- **Evidence:** `phase-ends/logs/Phase17.md:247-249` — + > **Sizing (§6):** tractable easy classes are LOW-reach; ×134 yield is in quirk-heavy + > STRUCTURAL_MISS(7.6%)/PERMUTER_CLASS(3.6%). + + corroborated at `phase-ends/logs/Phase17.md:297-300`: *"the layer makes the high-reach circular targets declarable but they are the gcc-quirk tail … Wave yield is quirk-tail-limited (~2-4% fleet), NOT ~2×."* +- **What happened / what it cost:** Phase 17 hand-matched the 4 highest-reach (reach-134) targets after building the canonical-sig layer for them and banked **0 of 4** — every one was regalloc/hoist/layout-bound (`Phase17.md:286-296`). The same session's cheap closes were the *low*-reach classes (fn-ptr, low-reloc, badly-structured decompiles). Two sessions of top-of-queue effort produced one banked reach-134 function while the tractable set sat below it, and the phase's projected wave yield had to be re-cut from "~2×" to "quirk-tail-limited". +- **Not banked — greps:** `grep -n -i 'tractable.*low.reach' $F` → 0; `grep -n -i 'leverage.*anti' $F` → 0; `grep -n -i 'highest.reach.*hardest' $F` → 0; `grep -n -i 'easiest.*lowest.value' $F` → 0. Nearest hit is the *prescription* `docs/how-to-ai-decomp/03-bootstrap-order.md:93` ("Order by leverage, not difficulty … highest (reach × size) first") — it tells you to sort that way but nowhere records that the sort is also a difficulty sort, nor what to expect from it. +- **Proposed home:** DK (a kernel), appended as the empirical half of the leverage-first ordering already in chapter 03. +- **Portable because:** in any duplicated-binary project (overlays, statically linked libraries, multi-build ROMs) the most-shared code is the engine core — bigger, call-heavier, more optimized — so ROI-first and easy-first are opposite orders on every target, and a leverage-first lane will look like a failing lane for its first stretch. + +### C2 — Another decomp project's `INCLUDE_ASM` is a record of what they did not crack, never proof that a class is uncrackable; cross-project corroboration multiplies confidence in a WRONG verdict as readily as a right one +- **Evidence:** `phase-ends/logs/Phase18.md:52-54` — + > DECISIVE: Xenogears (independent decomp, IDENTICAL gcc-2.7.2-psx -O2) has **NO C lever** for the call-crossing + > $s0/$s1 ORDER class — no `register`, no asm pins, no permuter; they **ship it as INCLUDE_ASM** (1174 nonmatch). + + and the reversal at `phase-ends/logs/Phase18.md:136-138`: *"The T1-T5 'regalloc-order = UNSTEERABLE' verdict was **WRONG** (P9/R14 self-correction — I'd skipped the most direct lever). **Explicit register pinning works**."* +- **What happened / what it cost:** the sibling project's give-up point was treated as independent confirmation and hardened a wrong verdict through T2, T3, T4 and T5 — it was written into the cookbook (§17 and the §16 reconcile), and it re-scoped the *next* phase to "triage-and-stub the unsteerable" (`Phase18.md:80-82`). One directive to try pins reversed it in a session: `func_8012B8E4` went 21 → MATCH and propagated ×134 (`Phase18.md:139-143`). The cheapest lever in the toolkit was skipped because two independent sources agreed there was none. +- **Not banked — greps:** `grep -n -i 'include_asm.*proof' $F` → 0; `grep -n -i 'sibling project.*gave up' $F` → 0; `grep -n -i 'another project.*impossible' $F` → 0; `grep -n -i 'corroborat.*wrong verdict' $F` → 0. The corrected *fact* exists only as a parenthetical about this one class (`docs/matching-cookbook.md:1420`, "Xenogears ships this class as INCLUDE_ASM only because they hadn't found the pin lever"); the records elsewhere say only the positive half — mine sibling projects for idioms (`decomp-kernels.md:196`, `03-bootstrap-order.md:89`, `07-compiler-source.md:99`). +- **Proposed home:** DK (a kernel) or a G-rule attached to G51/G52 — the evidence rule for cross-project inputs: a sibling's matched code is positive evidence; a sibling's unmatched code is no evidence. +- **Portable because:** every mature decomp scene has neighbours on the same compiler, and their `NONMATCHING`/`INCLUDE_ASM` lists are the first thing an agent finds and the easiest thing to mistake for a proof of impossibility. + +### C3 — Your own corpus of byte-matches is an experiment you have already run on the toolchain: settle "is my rebuilt compiler faithful to the original?" from it before installing the original vendor toolchain +- **Evidence:** `phase-ends/logs/Phase18.md:64-70` — + > Plan escape clause invoked: Wine is a heavy install on this WSL (106 pkgs + i386 arch not enabled + wineprefix) + > and the divergence question is already closed by stronger evidence: (1) **gcc-2.7.2-psx byte-matches ~700 functions**, + > many with call-crossing callee-saved values → its global register allocation IS faithful to the original compiler +- **What happened / what it cost:** a whole planned task (W1) existed to run the real `CC1PSX.EXE`/`ASPSX.EXE` under Wine to test whether the compiler *build* explained the residual. It was dropped and answered for free from three existing sources: ~700 already-matching functions (many exercising the exact pass under suspicion — a divergent compiler would have broken them), a directly-tested sibling `cc1` build that diverged and was *worse* (32 vs 21 mismatches), and the sibling project's era toolchain. The install was never done and the verdict has held since. +- **Not banked — greps:** `grep -n -i 'build divergence' $F` → 0; `grep -n -i 'original compiler binary' $F` → 0; `grep -n -i 'existing matches.*compiler.*faithful' $F` → 0; `grep -n -i 'ruled out by proxy' $F` → 0; `grep -n -i 'wine\b' $F` → 0. The records cover choosing and pinning the compiler triple (`03-bootstrap-order.md:24`) but never say that the accumulating match corpus is itself the evidence about the toolchain. +- **Portable because:** every decomp runs on a rebuilt or substituted compiler and every one of them eventually asks "is my build the original build?"; the answer is a free query against work already banked, and the alternative is a day of emulation plumbing. The construction generalizes: **pick the suspect pass, count how many existing matches exercise it.** +- **Proposed home:** accelerator (a "do this instead" with a measured saving), cross-referenced from chapter 07. + +### C4 — Accelerate the stage that is actually the bottleneck: a byte-exact search loop costs a compiler invocation plus a whole-binary gate per candidate, so GPU/ML brute force buys nothing +- **Evidence:** `phase-ends/logs/Phase17.md:96-99` — + > **CUDA/ML brute-force:** CUDA can't run gcc-2.7.2; the bottleneck is the gcc compile + the whole-binary gate, + > NOT search speed, so raw GPU permutation doesn't apply. The real ML angle = a *learned gcc-2.7.2 codegen + > predictor/ranker* … a high-risk, months-long research wildcard … NOT the plan. +- **What happened / what it cost:** the project owner had the GPU and the ML skills, and the proposal was live enough to be written into the phase's deferred list. Naming the per-candidate cost (a serial cc1 run plus a link and hash of the whole binary) killed it in one paragraph and redirected the same instinct at the only place a learned model could sit — ranking candidate C shapes — which was correctly parked as research, not plan. It also records that no off-the-shelf byte-exact MIPS/gcc decompiler exists, so "just use an ML decompiler" is answered too. +- **Not banked — greps:** `grep -n -i 'cuda' $F` → 0; `grep -n -i 'search speed' $F` → 0; `grep -n -i 'bottleneck is the compile' $F` → 0; `grep -n -i 'gpu-hours\|gpu accelerat' $F` → 1 (`docs/decision-log.md:99`, about renting a bigger GPU for *fine-tuning a 7B drafter* — a different subject). +- **Proposed home:** DK (a kernel) — a one-line refutation to hand a future project the first time hardware is proposed as the answer. +- **Portable because:** the proposal is perennial (GPUs, ML decompilers, "just brute force it") and the refutation is architectural, not BFM-specific: the inner loop is a serial, CPU-bound, host-toolchain compile plus a link, and neither moves to a GPU. + +### C5 — A contradiction between two entries of your own knowledge base is a work item, not noise: replay the levers you have already written down, under the correct oracle, before commissioning any new research +- **Evidence:** `phase-ends/logs/Phase18.md:26-29` — + > Phase 17 used the **wrong oracle** (the permuter's floor-polluted score) — §10:518 says use the + > object-level metric (`match_one.py`), which was **never applied** to the exemplars. §16 ("not + > source-steerable") contradicts §10 — Phase 18 reconciles it. So: **existing-knowledge-first** +- **What happened / what it cost:** two sections of the cookbook had said opposite things about the same class for a whole phase — §10 documented C-level levers with a committed worked example, §16 declared the class not source-steerable — and nothing detected it, because the §16 verdict had been produced with an instrument §10 itself warns against. The first task of Phase 18 was simply to replay §10's documented lever under the right oracle: it closed 24 → 21 mismatches immediately (`Phase18.md:42-44`), a real win that had been sitting written down and unused. Phase 17's strategic conclusion ("the harness ceiling did NOT rise") had been drawn on top of it. +- **Not banked — greps:** `grep -n -i 'two sections contradict\|sections contradict' $F` → 0; `grep -n -i 'knowledge base.*contradict' $F` → 0; `grep -n -i 'replay.*documented lever' $F` → 0; `grep -n -i 'existing-knowledge-first' $F` → 0. `grep -n -i 'contradict' $F` → 11, all about *code* contradicting itself (a draft vs a header, a declaration vs a definition, a card vs the bytes), none about two knowledge-base entries disagreeing. +- **Proposed home:** DK (a kernel) for chapter 06 (the knowledge base) — the maintenance half of a compounding cookbook: entries are written by different sessions under different instruments, so a growing base acquires contradictions by construction, and each one is a lead. +- **Portable because:** any project whose knowledge base is written incrementally by agents will hold entries that disagree; the disagreement always encodes an instrument change or a missing lever, and it is the cheapest research a session can do. + +### C6 — Scope a research phase by a PER-CLASS VERDICT, not by a percentage: the deliverable is a validated lever or an honest, falsifiable wall verdict for each class +- **Evidence:** `phase-ends/logs/Phase18.md:93-96` — + > ## Milestone (knowledge-gated, NOT a fleet-% target) + > Per-class byte-gated verdict (validated C idiom proven on a known-answer exemplar via + > `p16_known_answer --gate`, OR honest "unsteerable" verdict naming the exact gcc pass) + + set at the gate as `phase-ends/logs/Phase18.md:18-19`: *"Knowledge-gated milestone (validated idiom OR honest verdict per class) + a bounded quirk-tail demonstration. The full tractable-247 harvest wave is **deferred to Phase 19**."* +- **What happened / what it cost:** Phase 17 was scoped as a percentage phase — five avenues, each a candidate match-% lever — and four of the five returned ~0 (`Phase17.md:47-50`: T2 = 0 functions, T4/T5 byte-neutral, T6 = 0 whole-binary), so a phase of real work read as a failure and ended in a NO-GO. Phase 18 was scoped knowledge-gated on the same wall and produced the transferable toolkit *and* the throughput: the wave close-rate went 33% → 56% → 90% match_one as the toolkit entered the prompt, and the session arc banked 55.58% → 56.64% (`Phase18.md:126-134`). Same wall, same models; the difference was what the phase was allowed to deliver. +- **Not banked — greps:** `grep -n -i 'knowledge-gated' $F` → 0; `grep -n -i 'milestone.*per-class' $F` → 0; `grep -n -i 'research phase.*milestone' $F` → 0; `grep -n -i 'not a fleet-% target\|not a percentage' $F` → 0. The governance chapter requires a milestone be observable and machine-checkable (`docs/how-to-ai-decomp/01-governance.md:50`, G3/P9 at `:27`) and G52 requires a wall verdict to name the pass (`decomp-architect/templates/registry-E.decomp.md:329`) — but nothing says a phase may be *scoped* on verdicts, which is what makes the honest negative a deliverable instead of a failure. +- **Proposed home:** DK (a kernel) or a G-rule for chapter 01 — a phase-scoping form, alongside the plan/milestone gates. +- **Portable because:** every matching decomp hits a stretch where the remaining work is compiler research; percentage milestones there either fail or divert the phase into easy wins, and a per-class verdict is both achievable and the thing that compounds (the toolkit that raises every later wave). + +## ALREADY-BANKED (one line each) +- Parallel drafters declare a shared callee two ways; the whole standalone-match-vs-bank gap was compile conflicts with zero codegen mismatches — lives at `docs/how-to-ai-decomp/02-byte-gate.md:54-57` (and `docs/matching-cookbook.md:1377`). +- Gate one draft at a time when the failure class is declaration conflicts, so a clash cannot mask good drafts (`--chunk 1`) — `docs/how-to-ai-decomp/02-byte-gate.md:56`. +- A relocation-masked isolated per-function oracle cannot certify a whole-binary match; a permuter driven by it banks ~0 — `docs/matching-cookbook.md:15062` and `:4059`. +- The canonical-sig layer's projected "~2× lever" shrank to 7% once the sample that produced it was banked (re-size a lever after the measurement's own work lands) — `docs/matching-cookbook.md:1386`. +- The wave pipeline is draft → mandatory declaration-normalisation → byte-gate → propagate (a recovery pass becomes mandatory once the shared TU carries a canonical block) — `docs/matching-cookbook.md:1384-1385`. +- Re-run the oracle AFTER the canonical retype — normalisation can itself change codegen — `docs/matching-cookbook.md:1476-1481`. +- Fix arity/void callee conflicts with codegen-neutral call-site casts, not redeclaration — `docs/matching-cookbook.md:1474-1481`. +- Teaching the current toolkit (and each callee's canonical signature) in the per-target wave prompt lifted the close rate 33% → 56% → 90% on the same models — `docs/matching-cookbook.md:1467-1472`. +- `register __asm__` pins + a scheduling barrier crack the call-crossing register-ORDER class; the "unsteerable" verdict was a missing lever — `docs/matching-cookbook.md:1396-1420`, `decomp-architect/templates/registry-E.decomp.md:322` (G51). +- Array-decay (pass `buf`, not `&buf`/`buf.w`/`*(T*)buf`) forces rematerialization instead of a callee-saved hoist — `docs/matching-cookbook.md:1423-1425`. +- The permuter cannot help a pinned function (pycparser rejects `register __asm__`) — `docs/matching-cookbook.md:1370`. +- Census the residual/corpus shape with validated instruments before choosing a strategy, and validate the classifier against the arbiter on a known-true case (the wall taxonomy, and its byte-refuted "LOOSE_TYPING_WALL" class) — `decomp-architect/corpus/decomp-kernels.md:108` (DK-8), `decomp-architect/templates/registry-E.decomp.md:211`. +- Struct/type recovery is byte-neutral for matching (0 better / 10 same / 2 worse) but essential for hand-writing and comprehension — `decomp-architect/corpus/decomp-kernels.md:808-813`. +- A blanket canonical header imposed over a correct draft is destructive — canonicalise surgically, per callee — `docs/decision-log.md:1356-1364`, `docs/how-to-ai-decomp/10-integration-and-propagation.md:16`. +- Probe before costing: ground a lever's estimate on one instance before scaling or committing a long run (the 5-day unattended run that would have banked ~0) — `decomp-architect/templates/registry-E.decomp.md:160` (G23), `docs/how-to-ai-decomp/04-oracles-and-instruments.md:56` (R37). +- Order by leverage (reach × size), not by difficulty — `docs/how-to-ai-decomp/03-bootstrap-order.md:93`. +- The mismatch count does not predict which draft banks — `docs/matching-cookbook.md:35838`. +- Bulk-cache the decompiler's output with a headless batch pre-pass instead of driving the interactive server per function — `docs/matching-cookbook.md:1372`. +- Per-file `-O0` modules are a real class and need their own split file — `phase-ends/DIGEST.md:51`, `docs/cookbook-index.md:857` (§18). +- A verdict produced in an isolated environment describes the environment; gitignored inputs are missing there (the background-worktree isolation that had to be disabled before Phase 17 could start) — `decomp-architect/corpus/decomp-kernels.md:291-297` (DK-21). +- A harness/argument channel that silently drops what you passed it (the wave `args` channel that would not transit arrays, so targets had to be embedded) — `decomp-architect/corpus/decomp-kernels.md:263` (DK-19) and `:306`. +- The emulator is the runtime oracle, with multi-datapoint verification, for recovering a live struct's fields — `decomp-architect/templates/registry-E.decomp.md:20`, `phase-ends/DIGEST.md:166` (R10). +- Seed the knowledge base from sibling projects on the same compiler (the positive half of C2) — `decomp-architect/corpus/decomp-kernels.md:196`, `docs/how-to-ai-decomp/03-bootstrap-order.md:89`. +- A wall verdict names the pass and quotes the dump line — `decomp-architect/templates/registry-E.decomp.md:329` (G52), `docs/how-to-ai-decomp/07-compiler-source.md:88`. +- Sanity-check the OUTCOME metric (fleet %), never only the gate — `docs/matching-cookbook.md:1307`. +- Widening a discarding caller's `extern void` → `s32` is byte-neutral and dissolves the return-type conflict — `docs/matching-cookbook.md:2408`, `:615`. diff --git a/.run/P33.5/log-mining/Phase19-20-22.md b/.run/P33.5/log-mining/Phase19-20-22.md new file mode 100644 index 0000000000..b530c22c30 --- /dev/null +++ b/.run/P33.5/log-mining/Phase19-20-22.md @@ -0,0 +1,63 @@ +# Log mining — Phase19-20-22 +Files/ranges: `phase-ends/logs/Phase19.md`:1-56 · `phase-ends/logs/Phase20.md`:1-57 · `phase-ends/logs/Phase22.md`:1-74 · Lines read: 187 of 187 +Candidates considered: 36 · NEW: 2 · ALREADY-BANKED: 34 + +> Grep corpus used for every "already banked?" test (abbreviated `` below): +> `docs/matching-cookbook.md docs/cookbook-index.md docs/decision-log.md docs/accelerators.md docs/retrospective.md docs/wave-playbook.md docs/how-to-ai-decomp/*.md decomp-architect/corpus/decomp-kernels.md decomp-architect/templates/registry-E.decomp.md phase-ends/DIGEST.md` + +## NEW + +### C1 — An open-ended grind phase's milestone is invariants held plus a clean checkpoint, never a percentage target +- **Evidence:** `phase-ends/logs/Phase19.md:24` — "Milestone (gate 2, Drew confirms) … toolkit scaled across 2–3 batches of 50, **fleet up materially (target +3–5%, ~56.6% → ~60%)**" + `phase-ends/logs/Phase19.md:42` — "Phase total: **56.64% → 58.00% (+1.36%)**; the +3–5% target was gated by (a) the cache-limited pool (~88 eligible, all run) and (b) this propagation gap." + `phase-ends/logs/Phase20.md:7` — "Open-ended phase (Phase 15/19 precedent): fleet % rises monotonically, 136/136 byte-identical throughout, 0 NON_MATCHING; **closed at a clean checkpoint, not a fixed %**." +- **What happened / what it cost:** Phase 19's gate-2 milestone was written as a number (+3–5%). The phase executed its plan (both wave batches ran, 88% and 92% match rates, every recovery lever built and banked) and still delivered +1.36% — because the two things that actually set the yield were a *consumable* seed-cache pool (~88 eligible targets, all of them run) and an integration gap discovered mid-phase, neither of which the plan controlled. The very next phase's plan rewrote the milestone form to "monotonic + invariants + close at a clean checkpoint, not a fixed %", and that form was then used for every later open-ended phase. The cost of the numeric form: a phase that executed correctly reads as a 2/3 miss, and the pressure at gate 2 is on the number rather than on the invariants. +- **Not banked — greps:** `grep -n -i -E "target (of )?\+?[0-9]+(–|-)?[0-9]*%|percentage target|numeric target" ` → 0; `grep -n -i -E "close(d)? (at|on) a (clean )?checkpoint" ` → 0; `grep -n -i -E "open-ended phase|not a fixed %|clean checkpoint" ` → 1 (`docs/matching-cookbook.md:1230`, a status note "open-ended Phase 15 continuation, not a milestone gate" — not the milestone-shape rule); `grep -n -i -E "milestone.*(percent|%)|redefined milestone" ` → 1 (`docs/how-to-ai-decomp/01-governance.md:6`, which *boasts* "33 phases … without a single redefined milestone" but never says what milestone SHAPE makes that possible). Chapter 01's phase-cadence block (`:42-59`) specifies gates and "the milestone demonstrated with its observable proof", not how a grind phase's milestone must be worded. +- **Proposed home:** G (a rule in the registry seed's governance group) — and one line in `docs/how-to-ai-decomp/01-governance.md` under "The phase cadence", since chapter 01's "no milestone was ever redefined" claim is *explained* by this rule. +- **Portable because:** any decomp phase whose yield depends on a consumable fuel pool (drawable targets, cached decompiler seeds, model budget) cannot honestly promise a percentage; the invariants (monotonic progress, every binary byte-identical, zero regressions, a replayable checkpoint) are entirely under the agent's control and are what gate 2 should actually be checking. + +### C2 — Choose the exemplar for cracking a codegen class by the SIZE of its residual: the rows at a one-instruction decompiler-vs-target mismatch are the cleanest real-function isolates of the class +- **Evidence:** `phase-ends/logs/Phase20.md:14` — "**Key T3 input: 17 reach-134 fns at m2c-mismatch=1 = the cleanest regalloc/schedule isolates** (likely §17-pins wins + the genuine class residuals)." +- **What happened / what it cost:** the exemplar-miner routed 835 residual stubs into recovery lanes, and its most useful by-product was not the lanes but this ranking: among real open functions, the ones whose automatic decompiler output differs from the target by a single instruction isolate one codegen decision with nothing else in the way, so they are the cheapest place to attack a class. The phase's own class-crack tasks (T3b/T3c) were instead worked on the two *named legacy* exemplars carried in from earlier phases, and both ended cited-irreducible; one of them (`func_8012C2D0`) then turned out to have been **mislabelled** all along (`Phase20.md:17`, strength-reduction, not operand-order) — a mis-diagnosis that a minimal-residual isolate would have made hard to sustain. +- **Not banked — greps:** `grep -n -i -E "m2c.?mismatch|mismatch=1" ` → 0; `grep -n -i -E "cleanest (isolate|target)|clean isolate" ` → 0; `grep -n -i -E "pick.*(smallest|cleanest).*residual|choose the crack target" ` → 0; `grep -n -i -E "one-instruction|single-instruction (diff|residual)|closeness (of )?1\b" ` → 27, all in `docs/matching-cookbook.md` and all describing *a particular function's* residual (e.g. `:17226`, `:19948`, `:20027`), never a rule for selecting which function to attack a class on. The adjacent banked rules route by a *different* key: `decomp-architect/corpus/decomp-kernels.md:657` (DK-51) and `docs/accelerators.md:786` route a row by its **blocker** ("an unread compiler pass"), and `DK-45:589` prescribes **synthetic** five-line reproducers; neither says "rank the real open rows by residual size and crack the smallest". +- **Proposed home:** accelerator (`docs/accelerators.md`), promotable to a kernel — it is a draw-time selection heuristic, not a matching idiom. +- **Portable because:** every project has a cheap automatic closeness number against the target (any decompiler output or draft, diffed); ranking the open set by it costs nothing and names the minimal real-world reproducer of whatever class you are about to spend a deep session on. +- **Caveat (stated honestly):** the log records the heuristic as a finding but never spends it — no cost or yield is attached, because T3b/T3c went to the legacy exemplars instead. + +## ALREADY-BANKED (one line each) + +- A saved best draft gets clobbered by a later, worse attempt; the directory is the inventory, never overwrite the top-level `.c` — `docs/matching-cookbook.md:15964-15966` +- A failed in-TU build leaves a STALE `.o`, so `objdump` shows a false byte-match; trust the whole-binary SHA — `docs/matching-cookbook.md:1675` (and `:3032`, `:3765`) +- Never run a second `make`-invoking job concurrently with the propagation; parallel make corrupted a partial `.o`; serialize all make jobs — `docs/matching-cookbook.md:2383` (§28b-6) +- The propagation is non-atomic (source + instantiations applied, registry unwritten) and must be run in the background and re-run idempotently; the registry-skip recovery — `docs/matching-cookbook.md:2362`, §28c at `:2372` +- `git checkout src/` does NOT revert `config/dedup.us.yaml` — reset BOTH on a redo — `docs/matching-cookbook.md:2362` +- "reach-134" conflated *function-present-at-this-vram* with *byte-identical*; the `-O0` ×134 rollout was built and reverted on an invalid premise — `docs/matching-cookbook.md:1600-1610` +- The grinder's lifetime ROI (7 all-time banks, 0 since Phase 21) — and the stronger successor finding that 99% of its queue lived in another TU so it *could not* bank — `docs/decision-log.md:804-808`, `:1429`; `docs/matching-cookbook.md:4477` +- A permuter "win" the whole-binary gate still rejects must be blacklisted or the daemon churns the same function forever — `docs/matching-cookbook.md:2045` (§22) +- A signature normalizer REGRESSES already-canonical agent drafts: canon-only FIRST, `sig_unify` as FALLBACK (worth ~5 banks/batch) — `docs/matching-cookbook.md:1622-1626` +- Garbled callee signatures in the agents' own context read as model failure; fixing the manifest regex took the close rate 88%→92% and recovered 2 of 3 retries — `docs/matching-cookbook.md:1637-1640` +- Gate in chunks of 1: one compile-erroring draft fails the whole chunk's build and mis-attributes innocent neighbours — `docs/matching-cookbook.md:1705` +- Do not build a recovery transform for a class with ~0 measured instances (the data-cast was code for 0 cases) — `docs/matching-cookbook.md:1747-1753` +- Classify EVERY gate-fail by build-error class before assuming one cause — "the first pass saw one callee example and generalized it; the bytes said otherwise" (covers both the §20 "~33" projection and the close=0 giants that were not uniformly near-free) — `docs/matching-cookbook.md:1761-1764`, `:2348-2353` +- `INCLUDE_ASM` declares nothing, so a shared header's guessed `extern` becomes load-bearing project state in every stub overlay and there is no caller-side escape — `docs/matching-cookbook.md:1761`, `:3848` +- A def-side signature wall cannot be fixed by a text transform; it needs RE-DRAFTING under the caller-canonical signature, with the canonical pinned into the draft prompt — `docs/matching-cookbook.md:1745`; `docs/decision-log.md:1799` +- A standalone per-function compile over-predicts: it uses the draft's own externs and masks `jal`/`%hi`/`%lo`; only the real TU and the whole-binary gate decide — `docs/how-to-ai-decomp/04-oracles-and-instruments.md:16`, `:23` +- Propagation/integration, not matching, caps the yield — `docs/how-to-ai-decomp/10-integration-and-propagation.md` (chapter), `00-README.md`, `03-bootstrap-order.md`; `docs/retrospective.md:29` +- An "irreducible" verdict usually names the wrong mechanism and must be re-derived, not honoured (Phase 20's loop-guard exemplar was strength-reduction, not operand-order) — `decomp-architect/corpus/decomp-kernels.md:654-657` (DK-51); the cookbook's "misattributed mechanism" attacks, e.g. `docs/matching-cookbook.md:19736`, `:20512` +- Community/wiki compiler patterns must be triaged against your exact ISA and source language before adoption (`bnel` is MIPS II+, `.lit4` is PS2, C++ `bool` is not C) — `docs/matching-cookbook.md:2369` (§28a); the general form is G67 at `decomp-architect/templates/registry-E.decomp.md:395` +- Function-scoped `register __asm__` pins cannot express a register-LIFETIME reuse; the lever is one reused variable — `docs/matching-cookbook.md:36868` +- Instruction/byte weight beats head count, and the work queue is ranked by byte-weight, not by which member banks easiest — `decomp-architect/corpus/decomp-kernels.md:541` (DK-42); `docs/how-to-ai-decomp/09-economics.md:82`; `docs/decision-log.md:408-412` +- Per-agent reasoning depth is not the last-mile lever (the controlled 80-function experiment: 13 vs 14 banks); the wall is compiler determinism — `docs/how-to-ai-decomp/08-models-and-budgets.md:60-62` +- Preserve a structurally-perfect but scheduler-walled draft as a head start; per-function past-attempt notes are pack fuel — `decomp-architect/corpus/decomp-kernels.md:420` (DK-32); `docs/how-to-ai-decomp/05-cards-lanes-waves.md:17`, `06-knowledge-base.md:15` +- The decompiler-seed cache is a CONSUMABLE created by prefetch, not found; re-measure the fuel before each wave (Phase 19's batch 2 was cache-limited to 38 targets and the pool was one of the two named reasons the phase missed its target) — `docs/matching-cookbook.md:5275-5277` +- A shared dedup macro cannot carry a body compiled at a different optimization level; that path needs a per-binary split plus a shared header — `docs/matching-cookbook.md:2552-2562`; and mixed optimization levels as a wall class: `docs/how-to-ai-decomp/12-failure-museum.md:14`, `:32`, `docs/how-to-ai-decomp/03-bootstrap-order.md:24` +- The 3-object (`before` / different-flags / `after`) split for carving a middle region out of one object — `docs/matching-cookbook.md:1610` (§18 split machinery), `:2552` +- A wave's DIFFs are free compiler-class diagnoses; the drafting agent records the residual's class — `docs/matching-cookbook.md:149` +- Calibrate a new loop on an easier member first, and gate the calibration whole-binary rather than standalone — `docs/decision-log.md:1577`, `:1596`, `:1625` +- Regenerate a derived census/map after banking; a stale exclude/target list is refused at draw time — `docs/wave-playbook.md:39-41`, `:166`; `decomp-architect/templates/registry-E.decomp.md:245` +- "This class is exhausted" is unsafe: re-probing the still-live close=0 set banked a 15-20% tail (Phase 22 banked 7 + 7 bonus) — `docs/matching-cookbook.md:2372` (§28c); `docs/decision-log.md:804` +- Already-matched-but-unpropagated bodies are free banks; harvest → widen the tool → bank — `decomp-architect/corpus/decomp-kernels.md:226`; `docs/how-to-ai-decomp/06-knowledge-base.md:51`, `03-bootstrap-order.md:99` +- Inject the declarations a lifted body needs instead of rejecting it as "not self-contained" (the macro-extern-injection lever for the 7 plumbing-capped functions) — `docs/matching-cookbook.md:2385` (§28d) +- Reproduce the actual near-misses through the gate to find the real blocker before scoping a recovery tool; baseline the existing ladder first — `docs/matching-cookbook.md:1492`, `:5125` (§65d) +- Canonicalize a draft's raw-address callee names to the curated names as the FIRST pipeline step (the resident-callee link-miss) — `docs/matching-cookbook.md` ×16 (`canon_resident_calls`); `phase-ends/DIGEST.md` +- The canonical-extern + block-scope-extern recovery that turns a standalone MATCH into a bankable, propagatable body — `docs/matching-cookbook.md:2348-2362` (§28), `:2372` (§28c) diff --git a/.run/P33.5/log-mining/Phase21.md b/.run/P33.5/log-mining/Phase21.md new file mode 100644 index 0000000000..f4d849bad8 --- /dev/null +++ b/.run/P33.5/log-mining/Phase21.md @@ -0,0 +1,84 @@ +# Log mining — Phase21 +Files/ranges: phase-ends/logs/Phase21.md:1-799 · Lines read: 799 of 799 +Candidates considered: 27 · NEW: 3 · ALREADY-BANKED: 24 + +Grep corpus used for every "already banked?" test (`$CORP`): +`docs/matching-cookbook.md docs/cookbook-index.md docs/decision-log.md docs/accelerators.md docs/retrospective.md +docs/wave-playbook.md docs/how-to-ai-decomp/*.md decomp-architect/corpus/decomp-kernels.md +decomp-architect/templates/registry-E.decomp.md phase-ends/DIGEST.md` + +## NEW + +### C1 — An unattended agent run that looks throttled is usually blocked on an interactive approval prompt; check the pending prompt before diagnosing the provider +- **Evidence:** `phase-ends/logs/Phase21.md:731-732` — + `**Wave-2 "6.5h" was NOT throttling** — it was idle on a CC **permission prompt** (Drew approved on check-in);` + `waves 1 & 3 ran in ~25–32 min. (R14: corrected my earlier rate-limit read.)` +- **What happened / what it cost:** The overnight worker loop was launched for an ~8h hands-off run. Wave 2 sat for + ~6.5 hours against a measured 25–32 min per wave, and I recorded it as provider throttling; the true cause was the + harness waiting on a permission prompt that only cleared when the human checked in. The cost is the whole delta — + roughly six wasted hours of a run whose entire premise was that it would keep working while nobody watched, plus a + false rate-limit entry in the session record that a later session would have planned around. +- **Not banked — greps:** `grep -n -i -E 'permission prompt' $CORP` → 0; `grep -n -i -E 'waiting for (a )?(user|human|approval)|blocked on (a )?prompt' $CORP` → 0; + `grep -n -i -E 'rate.?limit.{0,40}(wrong|not|actually)|not throttl|throttling' $CORP` → 2 (`docs/how-to-ai-decomp/09-economics.md:67` counts outage casualties, `docs/matching-cookbook.md:1283` calls a deterministic step "immune to throttling" — neither is the stall cause); + `grep -n -i -E 'idle' $CORP` → 11 (all about lane idleness from stopping a lane or from batch stragglers, none about an approval gate). Adjacent but not the same: R40 exonerate the instrument (`docs/how-to-ai-decomp/04-oracles-and-instruments.md:64`) lists rate limits as a cause to clear on the way to judging the *subject*, and R55 (`:68`) says an unattended lane must leave evidence — neither names the harness's own approval gate as the thing that stops an unattended run. +- **Proposed home:** accelerator (or a DK on unattended lanes) +- **Portable because:** every agent harness with a permission/approval gate can silently hold an autonomous run; the diagnosis order — "is a prompt pending?" before "is the provider throttling?" — costs one glance and is true on any harness and any project. + +### C2 — A class-distribution assessor only sees the population that has already been attempted; "analyse ALL remaining work" is a cheap triage pass, not a static analysis +- **Evidence:** `phase-ends/logs/Phase21.md:63-67` — + `**THE ONE REAL GAP … `--assess` clusters the BACKLOG (already-DRAFTED near-misses), not ALL remaining stubs.**` + `A function's "decomp issue" is only known AFTER a draft attempt (the residual = the class). So "analyze ALL` + `remaining" = a **TRIAGE wave**: point `worker_wave` at the UN-drafted pools → each draft self-reports `// @class`` +- **What happened / what it cost:** The phase's flywheel model was "when out of fuel, analyse all remaining functions, + group them by decomp issue, learn one idiom per group". The tool built for it (`idiom_loop.py --assess`) clustered + the *backlog* — i.e. only functions a wave had already drafted — so every "next idiom to learn" verdict was computed + over a biased sample and the un-drafted majority was invisible to planning. The resolution needed no new tool: send a + cheap triage wave at the un-drafted pools so each target self-reports a residual class, and only then assess. +- **Not banked — greps:** `grep -n -i -E 'triage wave' $CORP` → 0; `grep -n -i -E 'never attempted|not attempted|un-?drafted' $CORP` → 7 (all about the *gate/oracle* being blind to work never attempted, e.g. `docs/decision-log.md:704`, `docs/matching-cookbook.md:3689`, or about a per-run count — none about a class distribution computed over the attempted subset); + `grep -n -i -E 'class label|classifier' $CORP` → 14, nearest `docs/how-to-ai-decomp/03-bootstrap-order.md:86-87` ("Build the classifier before the backlog is large; 91% of BFM's open backlog carried no class label") — that says build the labeller early, not that the assessor's population is the attempted one and must be widened by attempting; + `grep -n -i -E '@class|self-report' $CORP` → 10 (self-reports as unreliable verdicts, R14 — the opposite concern). +- **Proposed home:** DK (a kernel, alongside the residual-classifier kernel) — or how-to chapter 03 +- **Portable because:** on any project, the label that routes work (a residual class, an error class, a failure mode) is *produced by an attempt*. Any planner that clusters a work ledger is therefore planning over the attempted subset, and the fix — one cheap attempt per un-attempted item, purely to generate labels — is compiler- and console-independent. + +### C3 — Say which currency a wave buys — percentage or idioms — before launching it, and judge it in that currency +- **Evidence:** `phase-ends/logs/Phase21.md:475` and `:18` — + `**Waves 17–29 — reach-1 (×1) smallest-first harvest:** ~183 banks but fleet only **+0.06%** … **LESSON: reach-1 is` + `poor fleet-ROI; its value was ov_SC01_077 completeness + idiom-mining.**` / `reach-1 is ×1, % negligible BY DESIGN;` + `value = ov_SC01_077 completeness + idiom-mining` +- **What happened / what it cost:** Thirteen waves banked ~183 functions and moved the fleet 61.16% → 61.22%. The very + next two waves, aimed at the high-reach pool, banked 26 and moved it +0.97% — about sixteen times the entire earlier + run, per wave. The reach-1 run was not worthless (it was the phase's idiom mine and it completed one overlay), but it + had been launched and reported as if it were percentage work, so thirteen waves' worth of head-count progress read as + progress it wasn't. When reach-1 was later run again deliberately, the log states the currency up front — "% negligible + BY DESIGN; value = … idiom-mining" — and the same pool became a defensible spend. +- **Not banked — greps:** `grep -n -i -E 'idiom-mining|idiom mining' $CORP` → 0; `grep -n -i -E 'not for (its|the) yield|value was the (idiom|knowledge)|training, not production' $CORP` → 0; + `grep -n -i -E 'curriculum' $CORP` → 11 (all the Phase-25 *exemplar-crack* curriculum, a different thing). Nearest adjacency, read and judged: `docs/matching-cookbook.md:5113` "for the 0-stubs completion contract that is real progress; for the decomp.dev display number it is not. **Say which one you are buying**" and `:6903` (byte-variant families move RE-completeness, high-reach families move the display number) — both are about choosing between two *progress metrics*; neither treats knowledge production as an output a wave can be bought for. DK-42 (`decomp-architect/corpus/decomp-kernels.md:541`) rules the opposite direction only (instruction weight over head count). +- **Proposed home:** accelerator, or a line in DK-42 +- **Portable because:** every campaign has a low-yield population that is nonetheless the cheapest source of new knowledge for its flywheel; the rule is to name the output (percent vs. lessons) at draw time so the wave is judged against what it was for, which holds for any target, compiler or harness. + +## ALREADY-BANKED (one line each) + +- Under a harness's `run_in_background`, do not also `nohup … &` inside — completion fires for the wrapper while the real job runs on (`:30`) — lives at `docs/matching-cookbook.md:9576-9577`. +- Delegate hand-matching and deep RE to isolated agents; doing them inline bloats the orchestrator's context (`:86-89`) — lives at `docs/how-to-ai-decomp/08-models-and-budgets.md:70-71`. +- The gate's own repair transforms regressed near-miss closeness and the tool logged the regressed value, hiding permuter-eligible drafts (`:20`, `:176-180`) — lives at `docs/matching-cookbook.md:2328` (re-log with true closeness + raw draft) and `docs/decision-log.md:1449-1455` (check what a ranking scalar measures before consuming it). +- A canonicalisation repair run unconditionally destroys an already-correct hand-pinned crack; gate first, repair only the failures (`:221-226`) — lives at `docs/matching-cookbook.md:2208-2218`. +- Batch propagation that is all-or-nothing lets one poisoned member revert ten clean banks and read as a wall (`:310-315`) — lives at `docs/how-to-ai-decomp/02-byte-gate.md:50-52` (bisect so one bad draft cannot sink the rest). +- A wrapper that swallows a subprocess's non-zero exit turns a failure into a false measurement (`:313`) — lives at `phase-ends/DIGEST.md:238` (R53), `docs/how-to-ai-decomp/02-byte-gate.md:30-32`, `docs/decision-log.md:2078-2084`. +- A crashed agent self-check is indistinguishable from a failed draft, so the wave drafts blind and nobody is told (`:319`) — lives at `docs/decision-log.md:1455` and `docs/matching-cookbook.md:4500-4503`. +- A per-function differ that builds without deleting the stale `.o`/`.elf` reports false byte-matches (`:407-410`) — lives at `docs/matching-cookbook.md:2059-2076` and `:1673-1676`. +- `objdump` elides runs of zero words, so cop2/GTE functions read as mismatched and a whole family was filed "unproducible"; verify on raw bytes (`:486`, `:514`) — lives at `docs/matching-cookbook.md:1980-1987`; the general "objdump's rendering elides repeated words" at `docs/decision-log.md:3043`. +- A drafting agent appended unverified idioms straight into the shared cookbook; that write path was blocked (`:530-531`) — lives at `docs/matching-cookbook.md:1997-1998` (the incident + the block) and `docs/how-to-ai-decomp/06-knowledge-base.md:32-33` ("harvest only from proven results"). +- `h_exact`-style relocation-masked reach over-counts real propagation; probe shareability on the cracked exemplar before waving a class (`:235-241`) — lives at `docs/matching-cookbook.md:2195-2205`. +- A recommender that counts already-banked work and cannot separate never-tried from tried-and-failed re-burns a dry lever every session (`:119-126`) — lives at `docs/matching-cookbook.md:2218-2246` (§26). +- A backlog closeness achieved BY THE PERMUTER is not source closeness; `match_one` the saved draft before assuming a pin crack (`:143-146`) — lives at `docs/matching-cookbook.md:2269-2274`. +- Giants do not auto-bank: a 1.16M-token six-giant wave banked 0; they are hand-finish/permuter fuel (`:168-175`) — lives at `docs/matching-cookbook.md:2310-2318` (§27). +- An autonomous grinder without a persistent gate-rejection blacklist re-permutes the same impossible functions forever (banked 0 in ~8h) (`:735-739`) — lives at `docs/matching-cookbook.md:2024-2048` (§22) and `:178-179`. +- Rank targets by duplication reach, not head count: 183 banks ≈ +0.06%, 26 high-reach banks ≈ +0.97% (`:475-476`) — lives at `decomp-architect/corpus/decomp-kernels.md:541-548` (DK-42) and `decomp-architect/templates/registry-E.decomp.md:57-60` (G7). +- Same-compiler sibling projects carry their own flag deltas, so an inherited idiom must be re-proven on your own bytes (`:38`) — lives at `decomp-architect/templates/registry-E.decomp.md:395-401` (G67) and `decomp-architect/corpus/decomp-kernels.md:45-53` (DK-3). +- Cross-project byte-identical code between same-compiler games exists only in the vendor SDK/BIOS — zero engine code (`:39`, `:42`) — lives at `docs/decision-log.md:2025-2027`; "negative results with evidence are the product" at `docs/decision-log.md:2317`. +- Don't build a probe until you have read what the existing instrument already computes (the resolved-reach probe was withdrawn, unbuilt) (`:138-141`) — lives at `decomp-architect/templates/registry-E.decomp.md:215` (G33, never re-implement a gate you have) and `docs/matching-cookbook.md:5125` (§65d, measure the existing ladder first). +- Objects built with different flags (`-O0`, overlay-local) must be excluded from cross-object propagation or they poison it (`:312`, `:317`, `:323`) — lives at `docs/matching-cookbook.md:2388`, `:1595`, `:2552`. +- Long foreground build/VCS commands are killed by the harness; detached jobs survive (`:653`, `:719-720`) — lives at `docs/wave-playbook.md:738-740` and `docs/matching-cookbook.md:34804-34805`. +- Wave arguments must be pasted from a derived manifest, never hand-transcribed (`:375`, `:590`, `:649`) — lives at `docs/matching-cookbook.md:8861-8863`. +- The draw must exclude self-MATCH-but-gate-rejected and repeatedly-redrafted walls, or waves re-draft churners; a router in auto-mode grinds a saturated class (`:488`, `:589`, `:759-763`) — lives at `docs/wave-playbook.md:37-49` + `:76-78`, `docs/how-to-ai-decomp/12-failure-museum.md:28`, `docs/how-to-ai-decomp/03-bootstrap-order.md:67` (R45). +- A per-function cache must be keyed by the entry address, not by the decompiler's default `FUN_` label (`:786`) — lives at `docs/matching-cookbook.md:12859` (§164-45) and `docs/wave-playbook.md:236-238`. diff --git a/.run/P33.5/log-mining/Phase23-27.md b/.run/P33.5/log-mining/Phase23-27.md new file mode 100644 index 0000000000..f69c86440d --- /dev/null +++ b/.run/P33.5/log-mining/Phase23-27.md @@ -0,0 +1,80 @@ +# Log mining — Phase23-27 +Files/ranges: `phase-ends/logs/Phase23.md`:1-101 · `phase-ends/logs/Phase27.md`:1-71 · Lines read: 172 of 172 +Candidates considered: 22 · NEW: 8 · ALREADY-BANKED: 14 + +> Grep file set (abbreviated `$F` below, run verbatim each time): +> `docs/matching-cookbook.md docs/cookbook-index.md docs/decision-log.md docs/accelerators.md docs/retrospective.md docs/wave-playbook.md docs/how-to-ai-decomp/*.md decomp-architect/corpus/decomp-kernels.md decomp-architect/templates/registry-E.decomp.md phase-ends/DIGEST.md` + +## NEW + +### C1 — Mine new idioms from FRESH cracks, never from the failed backlog: your failure pile only re-teaches you what you already know +- **Evidence:** `phase-ends/logs/Phase23.md:93` — "**T10.8 DONE (`idiom_hunt.py`, $0.51):** batch-clustered the FAILED backlog by class → GLM re-derives our OWN idioms (pins §17, array-of-struct §18) + confirms walls, **0 new bankable idioms** (the failed residual is the worst discovery material — Drew's catch)." +- **What happened / what it cost:** A frontier model was pointed at the accumulated near-miss/failed backlog to harvest unmapped compiler behaviour; it produced zero new idioms, because a function is in the failed pile precisely when every idiom you own has already been tried on it, so the model reconstructs your own catalogue. The reframe (Drew's) was that discovery material is *fresh, never-tried hard functions solved correctly* — and the log states the principle underneath it: "the idiom lives in the correct BODY, not the bank" (`Phase23.md:93`), i.e. a correct body that fails to integrate still teaches, and a bank of a trivial function teaches nothing. Cost: two paid campaigns ($0.51 + $2.00) and a false "the idiom well is dry" strategic verdict built on the wrong sample. +- **Not banked — greps:** `grep -n -i 'failed residual' $F` → 0; `grep -n -i 'residual is the worst' $F` → 0; `grep -n -i 'discovery material' $F` → 0; `grep -n -i 'mine the fail' $F` → 0; `grep -n -i 'idiom lives in' $F` → 0; `grep -n -i -E 'T10\.8' $F` → 1 (cookbook:2411, which banks the *model-relative* half of the verdict, not the sample-selection half). `grep -n -i 'discovering' $F` → 6 (cookbook:2414 banks discover-vs-apply COST, not where to point the discoverer). +- **Proposed home:** DK (a kernel), with a line in `docs/accelerators.md` +- **Portable because:** it is a statement about sampling for discovery under any verifier — the population that survived your whole toolkit is the population least likely to contain a *new* mechanism, and it is the population everyone reaches for first because it is already collected. + +### C2 — Give a drafting model a TARGETED slice of the knowledge base, never the whole thing: full context measurably made the model worse +- **Evidence:** `phase-ends/logs/Phase23.md:24` — "**T2 — Stock-local floor** — Qwen3.6-35B-A3B … structurally smart but **0 reliable byte-matches** … format-robust; **full-cookbook context made it *worse* (dilution)**. The floor to beat." +- **What happened / what it cost:** The obvious first move with a large knowledge base is to hand all of it to the drafter; measured against the same model with a minimal prompt, the full-cookbook context *lowered* output quality. The project's later per-function pack construction (the relevant sections only, plus that function's own history) is the corrected form, but this was measured at the very first local-model probe and never written down as a law. +- **Not banked — greps:** `grep -n -i 'dilution' $F` → 0; `grep -n -i 'full cookbook' $F` → 0; `grep -n -i 'entire cookbook' $F` → 0; `grep -n -i 'too much context' $F` → 0; `grep -n -i 'relevant sections' $F` → 0; `grep -n -i 'whole cookbook' $F` → 1 (cookbook:11054, about deduping new entries against the corpus, not about prompt context). +- **Proposed home:** DK (a kernel) — it belongs beside the pack-construction guidance in the wave playbook +- **Portable because:** any project that accumulates a large idiom base will be tempted to paste it into the prompt; the finding is that context size and context relevance trade off, and the trade is measurable per model on day one with two prompts. + +### C3 — Name the drafter's degenerate output in the prompt: an empty body compiles, so "translate EVERY instruction, never an empty body" is a required instruction +- **Evidence:** `phase-ends/logs/Phase23.md:91` — "**PROMPT FIX** (`api_draft.LEAN_SYS` + `format_finetune.SYS`, synced): \"translate EVERY instruction, never an empty body\" — the v2 empty-leaf overfit, small-leaf band **0/3→2/3**, banked 3 on a fresh ov_SC01_001 batch." +- **What happened / what it cost:** The model had learned that `void f(void){}` is a valid, compiling C function, and emitted it for real leaf functions — output that passes every cheap filter (it compiles, it parses, it looks like a decompilation) and can only ever fail at the byte gate. One sentence added to both the inference prompt and the training format string moved the small-leaf band from 0/3 to 2/3 and banked functions the previous model could not reach. Earlier the same class had been misread as a corpus problem. +- **Not banked — greps:** `grep -n -i 'never an empty' $F` → 0; `grep -n -i 'empty draft' $F` → 0; `grep -n -i 'stub body' $F` → 0; `grep -n -i 'empty body' $F` → 3 (all cookbook §-entries about gcc threading an empty switch arm onto the epilogue — a codegen fact, unrelated). +- **Proposed home:** DK (a kernel) or a `docs/accelerators.md` entry +- **Portable because:** every generate-and-verify harness has a degenerate output that satisfies the cheap checks and never the real one; the accelerator is to enumerate those and forbid them by name in the prompt, and to keep the inference prompt and the training/format prompt synchronised. + +### C4 — The output-token cap is a two-sided knob, and both failure modes read as "the model is bad" +- **Evidence:** `phase-ends/logs/Phase23.md:54` — "`api_draft` max_tokens 4096→**512** (killed the no-stop-token ramble, ~80–130s→~10s on the bad cases)"; `Phase23.md:93` — "GLM5.2 is a REASONING model (needs MAXTOK≥8k or it returns empty content — burns budget on reasoning tokens)"; `Phase23.md:100` — "MAXTOK-tuning: reasoning models starve low, 96k overkill, ~48k+iters=2 sweet spot; iter-2 rescues ~29% of near-misses". +- **What happened / what it cost:** With a generous cap, a non-reasoning drafter that failed to emit a stop token rambled to the limit — 80–130 s per call instead of ~10 s, i.e. an order of magnitude of campaign wall-clock spent on garbage. With a tight cap, a *reasoning* model spent the whole budget thinking and returned empty content, which reads exactly like a refusal or a model incapable of the task. The measured shape was per-model: a small cap for the local drafter, ~48k plus a second iteration for the frontier reasoner, with the second iteration recovering ~29% of near-misses. +- **Not banked — greps:** `grep -n -i 'max_tokens' $F` → 0; `grep -n -i 'token budget' $F` → 0; `grep -n -i 'reasoning token' $F` → 0; `grep -n -i 'output cap' $F` → 0; `grep -n -i 'MAXTOK' $F` → 8 (all naming the S61 `ab8/ab16/ab24/ab32` cookbook §294 wave arms as provenance; §294's content is harvested idiom addenda, and no hit states a budget conclusion). Adjacent but different: `registry-E.decomp.md:178` (truncation ruled out before blaming a draft — the diagnosis, not the calibration). +- **Proposed home:** DK (a kernel), or a row in `docs/how-to-ai-decomp/08-models-and-budgets.md` +- **Portable because:** it is a per-model calibration every harness must do once, in both directions, before any model verdict is trustworthy — and neither failure mode announces itself. + +### C5 — Build the unattended campaign so a crash is a pause: probe the dependency at the top of each cycle, commit per cycle, persist the tried-set +- **Evidence:** `phase-ends/logs/Phase23.md:92` — "**Ended gracefully on a serve_local crash** (bitsandbytes `ops.cu:81` after ~14h continuous serving … ) — bulk_harvest's per-cycle endpoint-check caught the dead server, committed the partial cycle 23, and exited clean." Same line: "23 cycles → 1,297 fns banked … tried 764→**4,214**". +- **What happened / what it cost:** A ~15-hour unattended campaign's model server died of its own memory bug after 14 hours. Because the driver re-probed the endpoint at the start of every cycle, committed each cycle's banks as it went, and kept a persisted `tried` set, the death cost one partial cycle instead of the run: 1,297 banks were already committed and the campaign was resumable by re-issuing the identical command. The log even records the one residue to clean up on resume (~132 functions marked `tried` in the crashed cycle that never ran). +- **Not banked — greps:** `grep -n -i 'resumable' $F` → 2 (both cookbook §-entries about the *propagation transform* being idempotent/resumable, not a driver); `grep -n -i 'crash-safe' $F` → 0; `grep -n -i 'heartbeat' $F` → 0; `grep -n -i 'endpoint check' $F` → 0; `grep -n -i 'health-check' $F` → 0. Adjacent: `decomp-kernels.md:312` (DK-22 — long-running-driver checks: banner, pgrep, telemetry) says nothing about surviving a dependency's death or per-cycle commit granularity. +- **Proposed home:** DK (a kernel) — an addendum to DK-22 +- **Portable because:** every long unattended lane depends on something that will die (a served model, an API key's budget, a disk); the three properties that turn that death into a pause are cheap and must be designed in before the first long run, not after. + +### C6 — Prove the aggregate check target is fail-closed before adding checks to it, and give every audit oracle a dependent +- **Evidence:** `phase-ends/logs/Phase27.md:17` — "**`make report` is NOT fail-closed** … `Makefile:9-10` sets `.ONESHELL` with **no `-e`** … so the recipe is one `bash -c` and only the **last** command's exit survives … `lint_symbol_refs`, `progress --audit`, `difficulty`, `dup_report` are **swallowed** … Neither audit target has any dependent." +- **What happened / what it cost:** The aggregate reporting target had looked green for the life of the project while silently discarding the exit status of every check but the last one — `dedup-check` was "gated" purely by being last in the list. As the log puts it (`Phase27.md:64`): "**This unblocks every downstream assertion**: until now, any gate added to a report-invoked tool was swallowed on arrival." Two further shapes came with it: `check-all`/`extract-all` asserted `fail == 0` rather than `pass == N`, so an empty pipeline was a vacuous pass; and the two audit oracles had no dependent target, so nothing ever ran them. The fix was proven by a negative control — the identical broken gate exits 0 under the old shell flags and 2 under the new (`Phase27.md:64`). +- **Not banked — greps:** `grep -n -i 'ONESHELL' $F` → 0; `grep -n -i 'SHELLFLAGS' $F` → 0; `grep -n -i 'pass == N' $F` → 0; `grep -n -i 'fail == 0' $F` → 0; `grep -n -i 'no dependent' $F` → 1 (cookbook:30304, a codegen sentence about a dependent RMW). Adjacent: R32 "assert your coverage" (`how-to-ai-decomp/02-byte-gate.md:30`) banks the coverage-assertion idea for *scanners*, and `decision-log.md:1141` lists "`make report` swallowing gates" as one instrument in a list — neither records the mechanism, the vacuous-pass shape, or the orphaned-oracle shape. +- **Proposed home:** G (a rule) — the natural pair to G30 ("a guard that is downstream, or not running, is not a guard") +- **Portable because:** it is generic build-system plumbing (Make `.ONESHELL`, CI aggregate steps, any `&&`-less script of checks): a suite of assertions is worth exactly what its aggregate target's exit status is worth, and the check is one deliberately-broken gate away. + +### C7 — A document that cites a repository path is an untested claim about the repository; lint it +- **Evidence:** `phase-ends/logs/Phase27.md:71` — "**cookbook §45 names `.run/giants/func_80133CD4.fable.c` as its worked example and `.run/giants/fable_cd4/` as the flagship's gdb oracle — and BOTH were untracked.** The documentation cites artifacts that were not in git. … **Lesson … a doc that cites a path is an untested claim about the repo** — the citation and the file were four days out of sync, and only reading the seed for an unrelated reason caught it. **Candidate for a lint**". +- **What happened / what it cost:** The knowledge base's flagship worked example pointed at evidence that lived only in an ignored scratch directory on one machine — so the entry was unverifiable for any other reader and one `git clean` from being unverifiable for its author. It was found by accident, four days late, while reading the seed for an unrelated task; the fix was to widen the preservation allowlist by *file type* rather than by directory (49 files / 460K of oracle harnesses, regression ladders and drafts). +- **Not banked — greps:** `grep -n -i 'cites a path' $F` → 0; `grep -n -i 'broken link' $F` → 0; `grep -n -i 'untracked' $F` → 8 (all about untracked *source region files* that `git checkout -- src/` won't remove, plus the P33 purge mechanics — none about doc citations); `grep -n -i 'doc_links' $F` → 4 (`DIGEST.md:296` R80 — the checker exists, but for links between *documents* and promised pages; `how-to-ai-decomp/11-publishing.md:100` describes the same relative-link checker). No hit extends the check to cited artifact paths or to trackedness. +- **Proposed home:** G (a rule), with the concrete lint named — every path cited by a knowledge-base document must resolve to a TRACKED file +- **Portable because:** any project whose knowledge base cites evidence produced in scratch will drift the same way; the property that makes a citation trustworthy is trackedness, and it is machine-checkable from day one. + +### C8 — Assert that the work ledger PARTITIONS the live work, and treat a row the invariant refutes as a lie a fresh session will act on +- **Evidence:** `phase-ends/logs/Phase27.md:45` — "Built **`worklist --assert-partition`** … enumerate live stubs from `corpus.stubs`, assert the manifest partitions its source overlay): **proven** by catching 5 stale rows (the pin-free cores Phase-26 banked but the manifest still listed)"; `Phase27.md:65` — "`func_80178004` carried a `close=0 \"MATCH\"` that a fresh session would read as *done*, when the invariant refutes it outright (a real match banks; it's still a stub)." +- **What happened / what it cost:** The frontier ledger (worklist/backlog) is regenerated from scans and had drifted against the source of truth in both directions: rows for functions already banked, and rows asserting a *match* for functions still stubbed — the latter a retracted claim that had survived in two files and that any fresh session would have read as finished work. The one-line invariant kills the whole class: a real match banks, so a "MATCH" row on a live stub is self-refuting. The partition assertion caught 5 stale rows the moment it was written. +- **Not banked — greps:** `grep -n -i 'assert-partition' $F` → 0; `grep -n -i 'residue-0' $F` → 1 (`how-to-ai-decomp/12-failure-museum.md:20` — the same instrument SHAPE applied to the *disc census*, "what is on the medium", not to the work ledger); `grep -n -i 'ledger' $F` → 142, of which the relevant `accelerators.md:449` records claim-TIER honesty ("a 'free win' ledger entry nobody re-checks") — the naming-your-evidence-tier lesson, not the partition assertion or the self-refuting-row invariant. +- **Proposed home:** G (a rule) or DK (a kernel) +- **Portable because:** every long project accumulates a derived work ledger that outlives the scan that produced it; asserting it partitions the live work set, and deriving one refutation invariant from the gate, is cheap and catches the drift that silently re-targets or retires real work. + +## ALREADY-BANKED (one line each) +- The corpus fix (capturing the declarations each body needs, so a training example is self-contained) was the unlock, not the model — "corpus quality > size" — lives at `docs/decision-log.md:88-89`. +- Don't chase band-extension by retraining a small local model; the reasoning must come from a frontier model, and the small deterministic RE-smartness already exists as rules, not weights — `docs/decision-log.md:92-99`. +- Model capacity, not truncation, is what a small model's 0/10 on the hard band measures (verified inside maxlen) — `docs/decision-log.md:79-83`. +- The "wall is intrinsic / the idiom well is dry" verdict was MODEL-RELATIVE, refuted by a frontier model reading the gcc source — `docs/matching-cookbook.md:2411`, `docs/retrospective.md:24` and `:105`, `docs/accelerators.md:62`. +- Standard agents APPLYING a documented idiom crack functions 3–5× cheaper than the frontier model DISCOVERING it, and the "unsteerable" backlog is largely mis-verdicted — `docs/matching-cookbook.md:2414`. +- Fix the measuring tool before you trust a measurement: "a 0% from a broken tool and a 0% from a working one are the same number and opposite facts" — `docs/decision-log.md:1139`. +- Red-team a forward plan against the committed artifacts of the same day, and make it self-expiring — `docs/decision-log.md:1127-1129`. +- The medium was never fully counted: hard-coded entry indices hid code-bearing payloads, and the byte gate is blind to what was never onboarded — `docs/decision-log.md:1135`, `docs/retrospective.md:30`, `docs/how-to-ai-decomp/12-failure-museum.md:20`, kernels G21. +- Surface the compiler's stderr and split a failure into DIFF / PLUMBING / CC1-FAIL, or a compile error is recorded as a byte miss — `docs/matching-cookbook.md:3223`, `docs/wave-playbook.md:563-564`, `docs/accelerators.md:282`. +- Validate a scanner's discriminator against positive AND negative controls before reporting a count (a permissive validity check called structured data "code") — G32 `decomp-architect/templates/registry-E.decomp.md` (Check against a known-true case) + `docs/decision-log.md:1135`. +- Replace a hand-maintained table with a derivation only after proving the derivation equals it, by running both paths and failing on disagreement — `docs/accelerators.md:297` (#15), `decomp-kernels.md:121` (DK-9), G31. +- A shared scratch directory is a shared blast radius, and irreplaceable RE artifacts must be committed rather than left ignored-but-present — `docs/accelerators.md:732`, R20 (`phase-ends/DIGEST.md:186`), G18. +- Bank COUNT and the instruction-weighted headline are different currencies — a pile of tiny banks barely moves the weighted metric — `docs/matching-cookbook.md:5111`, `:16996`, `:17021`; the bank-rate-by-size curve at `docs/how-to-ai-decomp/09-economics.md:22`. +- Separate the produce phase from the verify phase and run the verifier in parallel over distinct binaries so the expensive resource never idles — `docs/matching-cookbook.md:6780`, `decomp-kernels.md:515`, `docs/wave-playbook.md` (gating speed). diff --git a/.run/P33.5/log-mining/Phase24.md b/.run/P33.5/log-mining/Phase24.md new file mode 100644 index 0000000000..4548cda3ab --- /dev/null +++ b/.run/P33.5/log-mining/Phase24.md @@ -0,0 +1,66 @@ +# Log mining — Phase24 +Files/ranges: phase-ends/logs/Phase24.md:1-138 · Lines read: 138 of 138 +Candidates considered: 35 · NEW: 4 · ALREADY-BANKED: 31 + +## NEW + +### C1 — Escalations to the expensive tier run STRICTLY SERIAL with idiom-banking between them; only the tier that cannot learn is run in parallel +- **Evidence:** `phase-ends/logs/Phase24.md:57` — "**Two-pass verdict (Drew's Q):** Opus-first is right (structure cheap + flywheel; Fable5 only for genuine global-hoist/regalloc walls); parallel Opus is fine, sequence any Fable5 escalations with idiom-banking." Also `:56` ("all → **MATCH** via sequential Fable5 (idiom-banking between)") and `:29` ("**Bank this session's Fable5 idioms into §31 BEFORE starting** (compounding)"). +- **What happened / what it cost:** The phase ran two shapes side by side: 5–8 cheap draft-agents in parallel (each isolated, learning nothing from the others), and the expensive wall-breaker tier one function at a time with the cracked lever written into the codegen map between each run. The three residual monsters (close=2/10/15) all fell in one sequential chain; the note "next-session Opus starts richer" (`:56`) is the reason. Parallelising the expensive tier would have paid N times for the same lever — the exact waste the serial chain avoided at ~375k tokens per escalation. +- **Not banked — greps:** `grep -n -i 'sequential.*idiom' ` → 0; `grep -n -i 'idiom-banking' ` → 0; `grep -n -i 'starts richer' ` → 0; `grep -n -i 'between escalations' ` → 0; `grep -n -i 'in parallel.*cheap' ` → 0; `grep -n -i 'parallel.*escalat' ` → 0. The nearest banked rows are the WAVE-level form (`decomp-architect/templates/registry-E.decomp.md:285` G45 "Harvest before the next wave — a hard gate"; `decomp-architect/corpus/decomp-kernels.md:234`) and the routing rule (`decomp-kernels.md:555` DK-43) — neither says how to SEQUENCE escalations inside one batch, and G45's gate is between waves, not between agents. +- **Proposed home:** DK (a kernel), as an addendum to DK-43 / a sibling of the G45 harvest gate +- **Portable because:** it follows from a property every agentic decomp shares — cheap drafters are stateless with respect to each other, while the expensive tier's output is a reusable lever; so parallelism is free for the first and destroys compounding for the second. + +### C2 — Inside one leverage class, schedule by MEASURED remaining effort — and pull the payoff-dominating outlier out of that queue for an immediate cheap triage +- **Evidence:** `phase-ends/logs/Phase24.md:29,34` — "**Strategy: closest-to-completion FIRST + fast-track the whale via a cheap triage** — NOT smallest→largest (closeness = the effort; size = the payoff; a 204-ins giant can be close=10 while a 152-ins one is close=52)"; "Its byte-weight ≈ the next ~8 giants COMBINED … **Do NOT bury it at the end of the closeness queue.**" +- **What happened / what it cost:** The §G queue held ~12 equal-reach giants, so payoff was proportional to size and effort was proportional to (re-measured) closeness — two axes that the log shows are uncorrelated. The whale ranked worst on effort (770 ins, untouched) and best on payoff; the plan gave it a cheap parallel triage instead of its effort-rank slot, and it banked at ≈ +1.6% byte-weight (`:52`), more than the six §G cracks around it. +- **Not banked — greps:** `grep -n -i 'closest-first' ` → 0; `grep -n -i 'closeness = the effort' ` → 0; `grep -n -i 'size is the payoff' ` → 0; `grep -n -i 'order the queue' ` → 0; `grep -n -i 'payoff' ` → 9 (all about a different subject: leverage of a CLASS, DK-172's early-investment rule, measured harvest yields). The two adjacent banked rows each carry only one axis: `docs/how-to-ai-decomp/03-bootstrap-order.md:93` ("Order by leverage, not difficulty … highest (reach × size) first" — payoff only) and `docs/matching-cookbook.md:33473` §413 ("DIFFICULTY IS THE RESIDUAL CLASS, NOT `nins`" — difficulty only, and used for model routing, not ordering). Neither states the two-axis rule or the outlier exception. +- **Proposed home:** DK (a kernel), beside the leverage-ordering rule in the bootstrap chapter +- **Portable because:** every decomp queue has these same two independent axes (reach × size = payoff; measured closeness = remaining effort), and the payoff distribution is always long-tailed, so the outlier exception recurs. + +### C3 — The wall-breaker tier is a MATCH tier, not a PLUMBING tier: work that a deterministic arbiter can judge does not need the expensive model +- **Evidence:** `phase-ends/logs/Phase24.md:61` — "Build it **Opus, plan-mode, Max — NOT Fable5** (it's tooling; the byte-gate arbitrates; reserve Fable5 for stalled giant MATCHES). Fable5's giant-crack is a *match* tier; this is a *plumbing* tier." +- **What happened / what it cost:** At the decision point (build the canonical-decl reconcile tool) the obvious move was to spend the tier that had just cracked the giant. Drew split it on the criterion of whether an oracle exists: the byte gate can arbitrate every candidate transform a tooling task produces, so an ordinary tier plus the gate converges; a codegen wall has no oracle to iterate against, which is what the tier premium actually buys. The tool was then built by the ordinary tier and byte-proven (`:59`). +- **Not banked — greps:** `grep -n -i 'plumbing tier' ` → 0; `grep -n -i 'match tier' ` → 1 (`phase-ends/DIGEST.md:172`, R13 proto-provenance — a different sense of "tier"); `grep -n -i 'deterministic arbiter' ` → 0; `grep -n -i 'not for tooling' ` → 0. `decomp-kernels.md:555` (DK-43) and `docs/how-to-ai-decomp/08-models-and-budgets.md:19` bank the neighbouring rule — "the strongest model is for genuinely new wall classes … it is NOT for reviewing a corpus against an existing knowledge base" — as a list of examples; neither gives the CRITERION (does a deterministic arbiter exist for this work?) nor names tooling as an excluded class. +- **Proposed home:** DK — a sharpening of DK-43 (one added sentence: the criterion, and tooling as the second excluded class) +- **Portable because:** the criterion is stated in terms of "is there an automatic arbiter for this task", which any project can evaluate, rather than in terms of BFM's tier names or task list. + +### C4 — When a byte-proven body will not integrate, the repair moves the CALLER's declaration to the definition's signature — never the definition to the caller's — and only where the change is width-compatible +- **Evidence:** `phase-ends/logs/Phase24.md:119` — "**Straggler/caller reconcile is byte-neutral ONLY on the caller's extern** (width-compatible types; the call site casts/passes-wide). NEVER touch the matched def." Mechanism at `:94`: "rewrite the straggler overlay's `extern func_X(...)` to the def's canonical sig (width-compatible: `s32↔u32`, `s16↔u16`, ptr types — the call site casts or passes wide args, so codegen is identical) … **NEVER change the matched def; only the caller's extern.**" +- **What happened / what it cost:** The flagship (400 ins, 22 phases deferred) matched but capped at ×1 because ONE overlay declared a conflicting caller extern. The fix was a one-line edit on that caller, after which it propagated to 134 overlays byte-identical (`:64`) — +1.6% byte weight. The direction matters: the same `conflicting types` error can be silenced from the definition side, which mutates the one artifact the project actually paid for and is not byte-neutral in general. +- **Not banked — greps:** `grep -n -i 'NEVER change the matched def' ` → 0; `grep -n -i "never touch the matched def" ` → 0; `grep -n -i 'width-compatible' ` → 0; `grep -n -i "caller's extern" ` → 1 (`docs/matching-cookbook.md:27848`, about an extern going stale mid-campaign — a different problem); `grep -n -i 'declaration side' ` → 1 (`docs/matching-cookbook.md:7215`, an unrelated §43 note). `docs/how-to-ai-decomp/10-integration-and-propagation.md:17` banks the FAILURE CLASS ("Def-side loose typing … def-side canonical-signature reconciliation") but not the direction rule or the width-compatibility condition; `docs/accelerators.md:384` and `docs/wave-playbook.md:392` bank the no-proto mechanics and their call-site hazard, not this law. +- **Proposed home:** G (a rule) in the integration chapter / registry — the one-line discipline that guards a paid-for body +- **Portable because:** it is a property of C linkage, not of this compiler or console: declarations are the repairable side, definitions carry the bytes; and the byte-neutral subset (width-compatible scalars and pointer types under default argument promotion) is the same everywhere. + +## ALREADY-BANKED (one line each) +- The "×134 wall" was a stale-asm misdiagnosis — RUN the claimed-blocked operation before building a tool to unblock it — lives at `docs/matching-cookbook.md:2471` (and `:2451`) +- A stale handoff line is a claim, not ground truth; verify each "×1" claim against the bytes before finishing it (R14) — `docs/matching-cookbook.md:2560` +- The permuter reporting `no match (0s)` is a TOOLING failure signature, never a search result (the header-comment `cpp` no-op) — `docs/matching-cookbook.md:2544` +- Warm-restart iterated local search descends where a cold permuter run plateaus — `docs/matching-cookbook.md:2544` +- The stock permuter scorer's relocation floor destroys the gradient; a relocation-masked scorer is the fix (and why it was right for us and refused upstream) — `docs/decision-log.md:3486`-3500 +- `objdump -dr` without `-z` elides runs of identical instructions → a false closeness number — `docs/matching-cookbook.md:6809` +- A tool that mutates shared state must undo by SNAPSHOT-RESTORE, never by an inverse transform (the lossy `--revert`) — `docs/matching-cookbook.md:4863` (instance at `:4617`) +- A green byte-gate proves what landed is correct, not that everything that should have propagated did — check the outcome metric — `docs/matching-cookbook.md:1303`-1308 +- One straggler member drops the whole propagation group unless the tool can exclude it per-member (`--recover`) — `docs/matching-cookbook.md:4192` +- RE-MEASURE A WALL BEFORE YOU RESPECT IT; stored closeness goes stale/regresses — `docs/cookbook-index.md:1375` (§146, `matching-cookbook.md:10011`); `docs/decision-log.md:2354` +- Order by leverage, not difficulty; instruction count does not predict difficulty — `docs/how-to-ai-decomp/03-bootstrap-order.md:93`; `docs/matching-cookbook.md:33473` (§413) +- A giant's difficulty is global-array hoisting, not its `$s`-register count (retires the "8-$s = hardest" heuristic) — `docs/matching-cookbook.md:2493` (§35) +- "Unsteerable / intrinsic" means "not yet read"; reading the compiler's passes dissolves the class — `docs/retrospective.md:24`; `docs/how-to-ai-decomp/12-failure-museum.md:13` +- Reserve the strongest model for genuinely new wall classes; cheap tiers are honest filters, never the gate — `decomp-architect/corpus/decomp-kernels.md:555` (DK-43) +- Lift a function-local typedef under a UNIQUE name, never the bare name (a same-named different layout collides fleet-wide) — `docs/matching-cookbook.md:2568` +- The shipped cc1's RTL dumps are NOT stripped (`-dr -dj -dc -dl -dg`) — readable ground truth before gdb — `docs/matching-cookbook.md:2531` +- Bias the permuter's pass weights by the diagnosed residual class instead of searching undirected — `docs/matching-cookbook.md:147`-163 (§3b) +- A deterministic search worker must be re-opened only on changed input, never on a blind idle re-try — `docs/matching-cookbook.md:178`-181 +- A symbol rename must reach every reference form (shared-macro bodies, `INCLUDE_ASM` stubs, asm string bodies) and be lint-guarded — `docs/matching-cookbook.md:36195`, `:36230` +- Verify byte-matches from a CLEAN rebuild — incremental builds masked a broken tree for three phases (R22) — `phase-ends/DIGEST.md:190` +- Any new per-file split suffix must be added to the propagation tool's file list or its functions silently cannot propagate — `docs/matching-cookbook.md:2564` +- A `SIZE-MISMATCH/short` with an `addu $fp,$sp,$zero` prologue means the target is -O0; optimization level is a property of the FILE — `docs/cookbook-index.md:24`; `docs/matching-cookbook.md:274` +- splat rejects two symbols at one address — bind the second name with an `__asm__` label, never an alias pair — `docs/matching-cookbook.md:2554` +- The silent-skip class: a scanner extracts N items where the truth is M > N and nobody compares; the byte gate is blind to work never attempted — `docs/decision-log.md:608`-620 (names this phase's `overlay_files` bug) +- A saved "best draft" can be regressed or stale — re-derive it, don't trust the ledger's pointer — `docs/matching-cookbook.md:2351` +- A transform wired into a pipeline must be idempotent and a no-op on inputs that don't need it, so it cannot regress the batch — `docs/matching-cookbook.md:2467` +- Recover the tree first (revert + re-extract), THEN measure — a stale `asm/` tree fakes a byte-diverge everywhere — `docs/matching-cookbook.md:4186` +- The reference "gcc 2.7.2" tree in community circulation is gcc 2.8.1 — a behavioural difference, not just line drift — `docs/how-to-ai-decomp/07-compiler-source.md:54`; `decomp-architect/corpus/decomp-kernels.md:672` +- Prove a class exhausted by construction, not by anyone's memory (R33) — `docs/decision-log.md:930` +- Harvest before the next wave, as a hard gate (the wave-level form of C1) — `decomp-architect/templates/registry-E.decomp.md:285` (G45); `decomp-architect/corpus/decomp-kernels.md:234` +- The def-side loose-typing failure class and its canonical-signature repair (the class only — the direction rule is C4) — `docs/how-to-ai-decomp/10-integration-and-propagation.md:17` diff --git a/.run/P33.5/log-mining/Phase25.md b/.run/P33.5/log-mining/Phase25.md new file mode 100644 index 0000000000..649f048040 --- /dev/null +++ b/.run/P33.5/log-mining/Phase25.md @@ -0,0 +1,100 @@ +# Log mining — Phase25 +Files/ranges: phase-ends/logs/Phase25.md:1-561 · Lines read: 561 of 561 +Candidates considered: 34 · NEW: 3 · ALREADY-BANKED: 31 + +## NEW + +### C1 — Leverage ordering does not predict tractability: carry a MEASURED closeness read per target, and never let (reach × size) stand in for "crackable" +- **Evidence:** `phase-ends/logs/Phase25.md:473-475` — + `**Key findings (frontier data → T6):** (1) **byte-weight ≠ tractability** — `func_8014D3E0` (22 ins × 1997) looked like` + `the mega-ROI freebie but is a `$sp` stack-switcher, matched only by porting an already-matched sibling (`func_8014D04C`);` + `the 369-giant fell to the wave while a 304-giant (`func_8014D820`) is close=253 (Fable5).` +- **What happened / what it cost:** The T5 pilot ordered its exemplar queue by byte-weight (instances × size). Its single + highest-leverage target — a 22-instruction function replicated 1,997 times — was one of the hardest things in the phase + (a `$sp` stack-switcher, only matchable by porting an already-matched sibling), while a 369-instruction giant fell to a + cheap wave in the same batch and a 304-instruction one sat at closeness 253. The phase's answer was to stop predicting + and build a measured frontier map (`.run/t5_frontier.jsonl`, eventually 93 exemplars with per-target closeness + class) + before spending anything on the queue. +- **Not banked — greps:** `grep -n -i 'byte-weight ≠|byte-weight is not|size.*not.*tractab|tractability' ` → 0 relevant + (the only `tractability` hits are cookbook prose about h_exact inflation); + `grep -n -i 'ROI-order|ROI order|order.*by ROI|highest-ROI'` → 0; + `grep -n -i 'payoff.*difficulty|reach.*not.*difficulty|high-reach.*hard'` → 0; + `grep -n -i 'leverage ≠|reach ≠|weight ≠|not.*predict.*crack|hardest.*smallest'` → 1 hit + (`docs/how-to-ai-decomp/05-cards-lanes-waves.md:165`, about turn/cost caps, not target selection). + The two nearest banked doctrines say the OPPOSITE and this bounds both: `docs/how-to-ai-decomp/03-bootstrap-order.md:93` + "**Order by leverage, not difficulty** … the schedule is highest (reach × size) first", and + `decomp-architect/corpus/decomp-kernels.md:556` / `docs/how-to-ai-decomp/08-models-and-budgets.md:9` "measure bank rate + **by instruction count** — that is the routing cliff". Phase 25's data point is a counterexample to the size prior + (22 ins harder than 369 ins) and a caveat to the leverage prior, and neither caveat is recorded anywhere. +- **Proposed home:** DK (a kernel), as a companion clause on the leverage-ordering kernel; mirrored as a line in + `docs/how-to-ai-decomp/03-bootstrap-order.md` Phase-3 bullet list. +- **Portable because:** every decomp schedules a frontier by some payoff estimate (reach, size, byte weight); on any + console or compiler that estimate is computed from the binary and says nothing about how the codegen was reached, so + a cheap per-target difficulty probe must ride alongside the payoff sort or the campaign stalls on its own top target. + +### C2 — A fan-out script generated by the orchestrator runs sandboxed with NO access to the repo: it must be self-contained, so target selection belongs to the wave generator, not the workers +- **Evidence:** `phase-ends/logs/Phase25.md:487-488` — + `- **Generated a §12-robust scale-up wave** (`tools/workflows/t5_scaleup.js`: worker_wave's drafter prompt + sequential` + ` waves-of-10 + retry×2, targets embedded — sandbox can't read files). Launched over the 83 ≤149-ins targets.` + (restated at `:271` — "targets embedded — sandbox can't read files") +- **What happened / what it cost:** Both the batch-1 wave and its continuation had to embed their target lists and the + drafter prompt literally into the generated script because the workflow sandbox cannot read repository files. The log + records it twice as a construction constraint on every wave generator ("copy the `.run/t5_scaleup_cont.js` pattern"), + i.e. it was rediscovered rather than looked up. +- **Not banked — greps:** `grep -n -i "sandbox can.t read|embed.*targets|targets embedded" ` → 0; + `grep -n -i 'self-contained script|no filesystem|cannot read the repo|read files from|embed the target'` → 0; + `grep -n -i 'wave script|generated script|worker_wave|workflow script'` → 4 hits, all about wave YIELD + (`docs/matching-cookbook.md:2308`, `:4129`) or an unrelated linker script (`:24783`), none about the sandbox boundary. +- **Proposed home:** accelerator (a harness-mechanics row), with a pointer from the wave playbook's authoring step. +- **Portable because:** any orchestration harness that ships a script to an isolated runner has this boundary; the + consequence — the generator owns target selection and prompt assembly, workers receive data not paths — is a design + rule for the wave layer regardless of console, compiler or vendor. + +### C3 — Run reference-compiler dump/inspection passes from a scratch CWD (or set an explicit dump base); a `-da` run with CWD at the repo root silently deposits RTL dumps that later read as committed artifacts +- **Evidence:** `phase-ends/logs/Phase25.md:337-340` — + `- Housekeeping: `gccdump.lreg` (repo root) = gcc's DEFAULT RTL dump (dump-base "gccdump", `.lreg` = local-reg` + ` pass; `toplev.c:1973/2077`), left by a one-off `cc1 -da` RTL-inspection run with CWD=root — **not** any` + ` committed tool (grep hits only the gcc source under `tools/reference/`). Deleted; root clean;` +- **What happened / what it cost:** A stray `gccdump.lreg` sat in the repo root long enough that a later session had to + trace it — reading gcc's `toplev.c` for the default dump-base and grepping the whole tool tree — before it could safely + be deleted. The practice adopted afterwards: RTL-inspection runs use a `.run/` CWD or `-dumpbase .run/gccdump`. +- **Not banked — greps:** `grep -n -i 'dumpbase|gccdump|stray dump|RTL dump' ` → 8 hits, all about USING RTL + dumps for matching (`docs/how-to-ai-decomp/07-compiler-source.md:64`, cookbook `§36`/`§29` at `:2402`, `:2506`), none + about where the files land; `grep -n -i 'repo root|CWD|working directory.*dump|litter'` → hits are a + `--prefix` relative-path gotcha (`docs/wave-playbook.md:35`) and unrelated carve/`§167-19` text. +- **Proposed home:** accelerator (a one-line ops row alongside the compiler-source chapter's `-da` instructions). +- **Portable because:** every matching project eventually runs its reference compiler with dump flags; every compiler + writes dumps relative to the invocation directory, so the hygiene rule holds for any toolchain. + +## ALREADY-BANKED (one line each) +- The stale-object gate trap (a failed piped `make build` leaves a stale `.o`; asm-differ then reports a phantom score 0) — lives at `docs/matching-cookbook.md:3028-3040` (§42b addendum, "THE #1 METHODOLOGY BUG") +- Isolation-MATCH ≠ real-TU bank; the worker must verify against the reconciled real TU — `docs/matching-cookbook.md:2971`, `docs/cookbook-index.md:1020` (§42a), `docs/how-to-ai-decomp/10-integration-and-propagation.md:17` +- The object-only probe is not the gate — it is blind to rodata, link and in-TU codegen; never SIZE a mechanical tier from it — `docs/cookbook-index.md:785` (§41b), `docs/decision-log.md:185-186` +- Trace a tool's real exception, not its summary label ("remap-fail" was a swallowed reconcile throw) — `docs/matching-cookbook.md:3178-3179`, `docs/decision-log.md:233` +- A pre-classifier that compiles in isolation false-negatives every family whose types live in the real TU's headers — `docs/matching-cookbook.md:2628-2633` (`--no-preclassify`) +- Never `git checkout` a shared/split source file mid-harvest to remove one function (reverted functions rebuild byte-identical as stubs, so the loss is invisible) — `docs/matching-cookbook.md:1275-1279`, `:498` +- Hand register pins are TU-context-specific — a pinned body that compiles in its own TU can SIGABRT cc1 in a sibling TU (later REFUTED: it was `extract_unit` dropping macros) — `docs/decision-log.md:235-260` and `:1137`, `docs/matching-cookbook.md:3187-3199` +- Direct `register __asm__` pins often backfire on giants (they wreck prologue save-birthing); density levers beat them — `docs/matching-cookbook.md:2965`, `docs/cookbook-index.md:218` (§42) +- The "36k unique tail" was a strict-clustering-key artifact; regrouping by instruction skeleton collapsed ~90% into ~754/986 families — `docs/matching-cookbook.md:2655`, `docs/accelerators.md:89`, `docs/decision-log.md:266` +- Three completion metrics (fn-count is duplication-inflated; instruction-weighted is the display number; distinct-code is the RE truth) — say which number a piece of work buys — `docs/matching-cookbook.md:5113`, `:6903`, `:5237`; `docs/decision-log.md:661-662` +- The local fine-tuned drafter is capacity-bound; corpus quality > size; the v4 retrain measured worse and no more retrains — `docs/decision-log.md:74-98` +- Parallel isolated agents for breadth, never one batched agent or the serial main loop (context accumulation makes batching ~N× costlier) — `docs/how-to-ai-decomp/08-models-and-budgets.md:70` +- Prove the plan's load-bearing leverage assumption against the bytes on a sample BEFORE scaling on it — `docs/matching-cookbook.md:4367`, `:10242`, `docs/decision-log.md:1655`, `:1233` +- Duplicate members are byte-shattered only in relocation fields; crack one exemplar and symbol-remap the family at ~0 agent tokens — `docs/matching-cookbook.md:2890`, `:32429`, `docs/how-to-ai-decomp/10-integration-and-propagation.md` +- A blind "lift every local type into the shared header" is unsafe: sibling split TUs give one typedef different layouts and redefine SDK names — `docs/decision-log.md:60-63`, `:1770-1785` +- A fresh session misreading a compressed hand-off as "done" nearly closed an open phase — `docs/decision-log.md:38`, `docs/retrospective.md:68`, `decomp-architect/corpus/decomp-kernels.md:741` +- TU-local levers (frame pads, pins) do not survive propagation — check remap-ability of the exemplar body BEFORE counting the ×N leverage — `docs/matching-cookbook.md:2957-2958`, `docs/decision-log.md:209-210` +- A verifier defaulting to the wrong output path reports "final SHA None"/fail for every non-default binary even when byte-identical — `docs/matching-cookbook.md:2978` +- The periodic clean-fleet verification caught a bug in its OWN harness (an unexpanded make variable extracting 2 of 136 binaries) before a false "136/136" could be reported — `docs/decision-log.md:122-123` +- An idiom cracked on one large function does not universally transfer — each giant is its own class — `docs/matching-cookbook.md:3282`, `:3345`, `:5319` +- The expected checksum always comes from the generated contract file, never hand-typed — `docs/wave-playbook.md:316`, `docs/decision-log.md:150`, `docs/retrospective.md:53` +- The def-side loose-typing / caller-declaration wall is the dominant gate rejection at scale, and ~71% of it was tool-shaped — `docs/decision-log.md:102-134`, `docs/retrospective.md:26`, `docs/how-to-ai-decomp/10-integration-and-propagation.md:17` +- Spend the strongest model only on genuinely new wall classes, where the deliverable is a transferable idiom rather than one function — `decomp-architect/corpus/decomp-kernels.md:555-559`, `docs/how-to-ai-decomp/08-models-and-budgets.md:19` +- Independently drafted functions banking into one TU collide on invented type names and struct tags — uniquify draft-defined types — `docs/cookbook-index.md:579` (§120), `docs/decision-log.md:181` +- A tool that hardcodes one path/subdir silently skips a whole cohort and still exits zero — assert the denominator — `docs/how-to-ai-decomp/04-oracles-and-instruments.md:39`, `docs/how-to-ai-decomp/12-failure-museum.md:33`, `docs/cookbook-index.md:691` (§96) +- Gate a mechanically-derived (remapped) draft with the plain byte-gate — the repair pipeline's transforms perturb an already-correct body and bank 0 — `docs/matching-cookbook.md:2608-2609`, `:7866` +- Propagating a reconciled body verbatim banks 0; re-run the reconcile against EACH sibling's own TU (the per-sibling re-reconcile law, 94%) — `docs/cookbook-index.md:906` (§41c), `docs/matching-cookbook.md:2703-2706`, `docs/decision-log.md:182` +- When the normalizer chokes on a draft, gate the RAW draft first and treat the rewrite as the fallback — `docs/cookbook-index.md:1266` (§313), `:564` (§41d) +- A transform proven byte-neutral at one optimization level is not neutral at another (and the Phase-25 "-O0-specific reconcile" need was itself later refuted) — `docs/matching-cookbook.md:2834`, `docs/cookbook-index.md:564` (§41d) +- A propagation tool's join key bounds its reach — a same-address join misses the cross-address instances (~1,863 of 1,997) — `docs/decision-log.md:211`, `:317`, `:1160-1161`; `docs/wave-playbook.md:158` (§438 same-address lead) +- A tool's `--commit` firing on its own narrower gate while the fleet-wide verification is still running — governed by the banked commit-ordering rule (bank committed before the next command that can touch the tree; the periodic fleet check is the backstop) — `decomp-architect/templates/registry-E.decomp.md:239`, `docs/retrospective.md:53` (R50) diff --git a/.run/P33.5/log-mining/Phase26.md b/.run/P33.5/log-mining/Phase26.md new file mode 100644 index 0000000000..bb7290f523 --- /dev/null +++ b/.run/P33.5/log-mining/Phase26.md @@ -0,0 +1,79 @@ +# Log mining — Phase26 +Files/ranges: phase-ends/logs/Phase26.md:1-1106 · Lines read: 1106 of 1106 +Candidates considered: 39 · NEW: 3 · ALREADY-BANKED: 36 + +> Grep target set (used for every candidate below, referred to as `$F`): +> `docs/matching-cookbook.md docs/cookbook-index.md docs/decision-log.md docs/accelerators.md docs/retrospective.md +> docs/wave-playbook.md docs/how-to-ai-decomp/*.md decomp-architect/corpus/decomp-kernels.md +> decomp-architect/templates/registry-E.decomp.md phase-ends/DIGEST.md` +> +> Note on the slice: Phase 26 contains the inserted tooling-integrity half-phase (26-A), which was *itself* a +> distillation exercise — cookbook §51 (LAWS 1–11), R32/R33/R34 and ~15 decision-log entries were written live +> during it (R30). The banked-rate is therefore very high by design: nearly every "audit thesis" line in the log +> already has a home. The three below are the ones that were fixed but never generalised. + +## NEW + +### C1 — A round-trip selftest is a serialisation check, not a coverage check: it passes by construction when a parser's missed item is absorbed into its neighbour's span. Every partition/rewrite tool needs an independent detector of items it failed to anchor. +- **Evidence:** `phase-ends/logs/Phase26.md:764-773` — + > "**The selftest was structurally blind** (round-trip = `"\n".join(item_texts)` stays exact by construction when a miss lands in a preamble — a *serialisation* check, not *coverage*). … **R32 coverage oracle** `hidden_definitions()` wired into `selftest` (independent detector of `func_XXXX(...){` bodies not anchored) — the selftest is now a coverage check." +- **What happened / what it cost:** `overlay_src_split.py` — the parser that repartitions overlay `.c` files for TU isolation, i.e. a tool in the byte-changing path — shipped with a fleet-wide selftest that read **404/404 files, 341,902 items, round-trip exact** (`:564`). It was green while `scan_construct`'s `force_decl` latch silently swallowed **2 real function definitions** on lines shaped `extern A; extern B; void f(){...}`, absorbing them into the next anchor's preamble. Because the round-trip re-joins whole-line chunks, the join stayed byte-exact and the miss was undetectable by the test that existed. It survived from session 5 to session 13 (the audit) and was found only by writing a second, independent detector; the audit's own line refs were stale, so the real cases had to be re-found by the new oracle (`:774-777`). +- **Not banked — greps:** `grep -n -i 'selftest\|self-test' $F` → 3 (a render selftest's reachability assert; a `--selftest` known-true case; a fleet-wide self-test catch — none about round-trip blindness); `grep -n -i 'round-trip\|roundtrip' $F` → 12 (all about splat round-tripping bytes, cdecl declarations round-tripped through gcc, or wiki row round-trips); `grep -n -i 'is not a coverage\|coverage check' $F` → 2 (both the P31 S80 K&R backlog refusal, a different subject); `grep -n -i 'masquerad' $F` → 4 (build-integration walls, byte-gate optimism, wrong TYPE as codegen residual); `grep -n -i 'structurally blind' $F` → 7 (all R34: an *oracle* blind to an error class ⇒ add a second oracle — the prescription, not this test-design failure mode). +- **Proposed home:** DK (a kernel), sitting beside R32/R34 — closest existing text is R32, which says assert coverage against an over-approximating candidate set; this says *which* test cannot be that assertion, and why. +- **Portable because:** every decomp builds source-partitioning / draft-splicing / TU-rewriting tools, and the obvious test for all of them is "re-emit the input and diff". That test is green for the whole class of merge/absorption defects, which are exactly the ones that silently move bytes. + +### C2 — A set that gates work must be reconstructible from committed artifacts. A roster kept in gitignored scratch is an unversioned oracle: it is silently wrong for anything it was not named after, and a `rm -rf` scratch or a fresh clone blinds it completely. +- **Evidence:** `phase-ends/logs/Phase26.md:742-748` — + > "`jr_inventory`'s `banked` set was filtered by an **EPHEMERAL, gitignored `.run/banked_func_*.json` roster** — `rm -rf .run`/a fresh clone would blind ALL banked jr at once, cross-address siblings (roster named after the exemplar) were structurally invisible, and non-leader banked jr were missed. **FIX … `banked` is now DERIVED FROM THE IMAGE** … (config + image, both durable)." +- **What happened / what it cost:** the tool that decides which functions already own a jump-table carve — the input to every isolation, and therefore to the ×134 economic engine — trusted a scratch file. Two silent narrowings followed from the roster's *shape*, not from a parse bug: it was keyed on the exemplar's name, so every cross-address sibling was invisible, and non-leader banked functions were dropped. Deriving `banked` from the committed config plus the built image instead, and adding an R32 assertion that every committed carve resolve to exactly one owner, made the fleet run **134/134 OK, 0 false aborts, 1336 banked jr == 1336 carves** — and immediately found the cross-address sibling (`func_8017FCB0`) and three non-leader overlays the roster had missed (`:753-758`). +- **Not banked — greps:** `grep -n -i 'ephemeral' $F` → 0; `grep -n -i 'gitignored' $F` → 8 (worktree isolation missing a gitignored signature registry — DK/accelerator #19; publishing; the `.run/` policy line — none about a roster as an authoritative set in the main tree); `grep -n -i 'roster' $F` → 1 (`docs/matching-cookbook.md:24906`, a parenthetical "the removed `INCLUDE_ASM` lines are the ground truth — no roster, no ledger" inside §262's lane-yield size-matching argument; it uses the instinct, it does not record the lesson); `grep -n -i 'source of truth' $F` → 8 (wiki/docs governance, licence provenance, §436's "wrong source of truth" = a stale asm directory vs the config, a different mechanism). +- **Proposed home:** G (a rule) or DK — the operational half of R33: *derive from committed state*, stated about persistence rather than about re-parsing. +- **Portable because:** every agent-driven decomp accumulates a `.run/`-style gitignored scratch tree, and JSON rosters written there are the cheapest way to remember what has been done. The failure is invisible in the tree that wrote them and total in every other tree. + +### C3 — A guard that is allowed to sit RED and UNWIRED does not exist. A detector's value is zero until it is green on HEAD and called by the standing report — and a docstring claiming it is wired is not wiring. +- **Evidence:** `phase-ends/logs/Phase26.md:101` and `:851-853` — + > "**A9 — `lint_symbol_refs`: a guard allowed to sit red does not exist** … currently **RED** (43 false positives) and **UNWIRED** (`make report` never calls it, though its docstring claims it does)." + > "The ONLY detector for the R22 rename-drift failure mode (a symbols rename leaves a `func_` ref dangling; a clean build fails, an incremental masks it — **undetected Phase 21→23**). Was RED (262 FPs) + UNWIRED." +- **What happened / what it cost:** the project's only detector for symbol-rename drift existed, in tree, for the whole period in which that exact failure mode went undetected across three phases. It was never deleted and never fixed — it sat red, so its output was ignored, and it sat unwired, so nothing produced output to ignore. Closing it took fixing three blind spots (it never scanned `src/shared/*.h`, where `engine_core.h`'s 10k+ tokens are; it unioned two symbol files instead of using each binary's own splat stack; and **all 262 false positives were one modelled class**, the `__asm__("memcpy")` asm-label binding), a negative control proving it still fires, and then wiring it fail-closed into `make report` (`:853-858`). +- **Not banked — greps:** `grep -n -i 'sit red\|sits red\|allowed to be red\|RED and UNWIRED' $F` → 0; `grep -n -i 'unwired\|not wired' $F` → 5 (an unwired SDK object, an unwired h_exact pool, "the fix already existed and was simply not wired in" — about a *fix*, not a standing guard, and with no rule attached); `grep -n -i 'only detector\|docstring claim' $F` → 0; `grep -n -i 'false positives' $F` → 8 (R39 — *negative-control a NEW refusal-check*, which governs a check's introduction, not the standing state of one already red). +- **Proposed home:** G (a rule), next to R39 — R39 says a new check must be negative-controlled to zero FPs; this says an *existing* check that is red or uncalled must be fixed-and-wired or deleted, never left as a note. +- **Portable because:** every long decomp accumulates half-finished lint/consistency tools, and the standing report is the only place a check is actually read. The cost shape (the one detector for a live failure mode, dormant across three phases) transfers to any project with a CI-style report and a tool directory. + +## ALREADY-BANKED (one line each) + +- A tool that cannot bank a function is indistinguishable, in every log, from a function that cannot be banked — lives at `docs/matching-cookbook.md:3934` +- A loud failure that nobody counts is exactly as invisible as a silent one (R32's corrected form) — lives at `docs/matching-cookbook.md:3739`, `phase-ends/DIGEST.md:208` +- A fix is not landed until its caller stops overriding it (grep every call site after fixing a scanner) — lives at `docs/matching-cookbook.md:3916` (§51g LAW 11) +- A check applied outside its valid domain does not become more thorough — it becomes noise (the 914-vs-193 near miss) — lives at `docs/matching-cookbook.md:3758` +- A second, disagreeing oracle beats a better assertion when an oracle is structurally blind (R34) — lives at `phase-ends/DIGEST.md:211`, `docs/decision-log.md:878` +- A symbol whose address falls inside another binary's vram window must never enter that binary's symbol stack — lives at `docs/decision-log.md:896` +- A stale object gives a FALSE PASS: `.o ← .s` is not a dependency make sees, so an incremental build "verifies" a broken config change; a rule a human must remember is not a gate — lives at `docs/matching-cookbook.md:3761-3767` +- gcc-2.7.2 does not prefix errors with `error:`; grep the diagnostic text, not the word "error" — lives at `docs/matching-cookbook.md:545` +- Triage rule: "wrong BYTES" → read the compiler source; "won't COMPILE" → read your own tooling (R17) — lives at `docs/decision-log.md:528` +- Every recovery pass is a FALLBACK, never unconditional (gate raw first, reconciled only on failure) — lives at `docs/matching-cookbook.md:2928` +- `match_one closeness==0` (isolated, reloc-masked) systematically overstates whole-binary bankability — lives at `docs/accelerators.md:365-370`, `docs/decision-log.md:1017` +- Force a newly-built mechanism's un-built sub-case out on cheap targets first: de-risking and building-the-missing-piece are the same move — lives at `docs/decision-log.md:413-422` +- Don't grind the light tail because it "feels productive"; the cheap tier is a means (harden the pipeline), not the objective — lives at `docs/decision-log.md:419-421` +- `--no-propagate` on every per-group gate; commit the cheap verified banks before the expensive fleet-wide propagate, which mutates before it gates and dies mid-mutation on a timeout — lives at `docs/matching-cookbook.md:4155-4166` (§55b) +- A checker that checked nothing must never read as a pass; a tool written against one binary carries an untested hypothesis about every other — lives at `docs/matching-cookbook.md:18450` (§192b) +- The adjudicator must be the compiler that compiles your code — not the standard, not the gcc on PATH — lives at `docs/matching-cookbook.md:3873` (§51g LAW 9) +- Probe the compiler for FACTS (four three-line probes, 90 s) before accepting a documented dead-end's premise — lives at `docs/matching-cookbook.md:3894-3898` +- Hand-maintained models of the corpus layout DECAY measurably as the corpus is restructured, and an un-nominated target produces silence, not an error — lives at `docs/decision-log.md:788-798` +- Audit the SELECTION tools before fixing anything on top of them: a hole there makes work invisible to planning — lives at `docs/decision-log.md:715-721` +- A target-ranking document sorted on already-matched work is majority-fiction (62% of the advertised byte-weight was phantom) — lives at `docs/decision-log.md:812`, `docs/matching-cookbook.md:3683` +- Making a parser see more ARMS dormant downstream transforms; migrate consumers one at a time, byte-gated — lives at `docs/decision-log.md:958-961` +- Pair every auditor finding with an adversarial skeptic told to refute it (32 raised → 28 survived, 4 refuted, 16 downgraded) — lives at `docs/matching-cookbook.md:3708`, `docs/decision-log.md:783` +- Audit scope filter: "does it PARSE something, and does it GATE or SELECT work?" — not all 82 tools — lives at `docs/decision-log.md:719` +- The frontier model DISCOVERS a class; the cheap agents APPLY it — lives at `docs/matching-cookbook.md:3555`, `:3991`, `docs/decision-log.md:1058` +- Run deterministic parallel shell jobs, not agents, for a gate re-test (agents are wasted running a shell gate) — lives at `docs/decision-log.md:1020`, `:2403` +- Prefer a transformation correct BY CONSTRUCTION (never worse than raw, conflict-free forward-carry) over one that needs an oracle to dodge conflicts — lives at `docs/decision-log.md:559`, `docs/matching-cookbook.md:483`, `:533` +- A tool whose per-item revert restores from HEAD must refuse a dirty tree (commit each family before sweeping the next) — lives at `docs/matching-cookbook.md:499`, `:4524`; R42 in `docs/how-to-ai-decomp/01-governance.md:37` +- Profile and parallelise the fleet verification harness — a slow gate is a bug; it runs on every commit and compounds — lives at `decomp-architect/corpus/decomp-kernels.md:432-444` (DK-33), `docs/how-to-ai-decomp/02-byte-gate.md:48` +- The "reach-1 unique tail" was largely a reloc-tracker blind spot in our own normalisation, not unique code — lives at `docs/matching-cookbook.md:2653` (§40b), `docs/decision-log.md:264` +- The byte-gate happily banked a phantom function splat invented (`void listCdBuffer(void) {}` — byte-correct, gate-green, fictitious) — lives at `docs/decision-log.md:890` +- A dropped prototype is a silent byte-changer in C89 (implicit `int f()`, and return type drives delay-slot fill) — lives at `docs/matching-cookbook.md:6399-6400` +- A GO/NO-GO validation harvest on a zero-crack corpus (families whose exemplar is already matched) before scaling the engine; check every instrument against a known-true case — lives at `docs/decision-log.md:281`, `:292`, `:317`; `decomp-architect/corpus/decomp-kernels.md:762` (DK-61) +- An audit that is a prerequisite to the live milestone runs as an INSERTED half-phase, not as a new phase and not by closing the current one early — lives at `docs/decision-log.md:737-754` +- Don't trust a handoff's diagnosis — reproduce one case by hand (R14); the sharper picture was a smaller, safer fix — lives at `docs/decision-log.md:532` +- Agents die to usage limits; write deliverables early and treat the limit as a wall class (R67) — lives at `phase-ends/DIGEST.md:266`, `docs/how-to-ai-decomp/08-models-and-budgets.md:42` +- A fleet verification (`make clean` + extract) rewrites the shared `asm/` tree that running workers read — never run it under them — lives at `docs/wave-playbook.md:621`, `docs/accelerators.md:464-468` diff --git a/.run/P33.5/log-mining/Phase28-32.md b/.run/P33.5/log-mining/Phase28-32.md new file mode 100644 index 0000000000..20dea88cb3 --- /dev/null +++ b/.run/P33.5/log-mining/Phase28-32.md @@ -0,0 +1,86 @@ +# Log mining — Phase28-32 +Files/ranges: `phase-ends/logs/Phase28.md`:1-200 · `phase-ends/logs/Phase32.md`:1-280 · Lines read: 480 of 480 +Candidates considered: 32 · NEW: 7 · ALREADY-BANKED: 25 + +Grep corpus used for every "already banked?" test (abbreviated `` below): +`docs/matching-cookbook.md docs/cookbook-index.md docs/decision-log.md docs/accelerators.md docs/retrospective.md +docs/wave-playbook.md docs/how-to-ai-decomp/*.md decomp-architect/corpus/decomp-kernels.md +decomp-architect/templates/registry-E.decomp.md phase-ends/DIGEST.md` + +## NEW + +### C1 — Never round-trip a curated config through a serializer: every oracle you own measures BYTES, so a formatting-destructive write is invisible to all of them +- **Evidence:** `phase-ends/logs/Phase28.md:181` — "My first cut wrote the registry with `yaml.safe_dump`, which round-tripped the whole file: **47 comment lines → 0** … and **1,832 `vram: 0x80162FF4` → `vram: 2148937716`** … **It passed dedup-check 1840/0 AND check-all 140/140**" +- **What happened / what it cost:** A tool extended the dedup registry by loading and re-dumping its YAML. That destroyed the curated Phase-11 header explaining *why* the share is source-level, and converted 1,832 hex `vram` fields to decimal (PyYAML parses YAML-1.1 hex as int, dumps int as decimal) — 25,948 lines rewritten, landed in a commit. Every gate called it green because the loader accepts both forms: "the data was correct and the document was ruined". Recovery was a restore from `HEAD~1` plus a surgical text re-apply (1545 insertions / 1545 deletions, 0 non-`binaries:` lines changed, 47 comments + 1908 hex fields verified intact). +- **Not banked — greps:** `grep -n -i -E 'safe_dump|round-trip(ped)? the (whole )?file|yaml.*round.?trip' ` → 0; `grep -n -i -E 'formatting-destructive|destroyed the (registry|document)|comment lines|strip(ped)? the comments' ` → 1 (cookbook:7192, a call-site count — unrelated); `grep -n -i -E 'pyyaml|yaml-1\.1|hex to int|surgical (text )?edit' ` → 0; `grep -n -i -E 'never rewrite a curated|preserve.*comments|comment-preserving' ` → 0 +- **Proposed home:** DK (a kernel) — the byte gate's blind spot for *documents*; plus a G rule: "a tool that edits a curated file edits it surgically, never by serializer round-trip; its control is a diff limited to the intended key" +- **Portable because:** every decomp keeps hand-curated YAML/JSON beside a byte oracle that cannot see documentation loss — splat maps, symbol lists, dedup/registry files — and the same round-trip silently strips comments and reformats numbers on any console. + +### C2 — The file a function lives in is not evidence of its class: read the recorded attribute, never the hosting split +- **Evidence:** `phase-ends/logs/Phase28.md:182` — "I then built a jr guard on the assumption those 4 were the §53 jr class *because ov_SC01_077 hosts them in `_jr_8017A4AC.c`* — **`has_mid_jr` is False for all four** … They merely live in a carved jr-**region** split … **Hosting file ≠ function class.**" +- **What happened / what it cost:** A carve is named after the jump-table function that forced it and sweeps in every function in its address range. Twelve unexplained byte-DIFFs were attributed to the jump-table class purely because their bodies sat in a `_jr_.c` split, and a guard was written on that inference. The guard was kept (correct for the real class, skipping 0) with its docstring corrected; the 12 DIFFs went back to UNDIAGNOSED — the diagnosis had to be redone. +- **Not banked — greps:** `grep -n -i -E 'hosting file|file it lives in|the split it lives in' ` → 0; `grep -n -i -E 'carve sweeps in every|region split.*(class|not)' ` → 0; `grep -n -i -E "co-?locat|neighbou?r'?s? class|because it lives in|derived from the (file|path) name" ` → 3 (all about rodata co-locating in its function's object); `grep -n -i 'has_mid_jr' ` → 8 (all the flag's sweep-tool routing, never the file-name inference) +- **Proposed home:** DK (a kernel) or G — "classify from the recorded attribute, never from the artifact's path" +- **Portable because:** every decomp splits source by address range and names splits after a landmark symbol; the path is a tempting, wrong proxy for a function's properties on any console. + +### C3 — Re-verify a task's premise in the code at EXECUTION time: roadmap lines, audit findings, and even an audit's own correction footer go stale — usually the document that named a defect is the first thing to stop being true +- **Evidence:** `phase-ends/logs/Phase28.md:143` — "**R14 on the premise first (twice):** (1) the roadmap's two named grinder bugs are **already fixed** … (2) `docs/tooling-audit.md:933-937` **downgraded its own finding** with three corrections: *the prescribed fix is a NO-OP* … *nothing is being discarded now* … *the queue magnitude was inflated*"; `:144` — "**the audit's snapshot is itself now stale (verified in code):** `harvest_verify` is **fully multi-TU** today … **The gate that manufactured the blacklist no longer exists.**" +- **What happened / what it cost:** The phase plan's task was written from a tooling audit and a roadmap section. At execution the premise was checked in the code first: both named bugs were already fixed, the audit had already retracted itself in a footer nobody had read, the prescribed fix was a measured no-op (deleting the scanner banks zero functions), and the tool the audit described no longer existed in that form. The task's only real content turned out to be deleting a poisoned artifact. The same class had already cost the whole phase its charter: the Phase-Start finding (`:11`) is that "the evidence behind its presumed value does not survive contact with the bytes — the exact failure R35 was coined to prevent, by the phase that coined it", and two further roadmap premises were also stale (`:28`: the parallel gate farm already existed at ~75× serial). +- **Not banked — greps:** `grep -n -i -E "plan'?s? (own )?premise|premise.*(stale|verif)|re-verify (the|every) premise|stale (roadmap|plan) line" ` → 0; `grep -n -i -E 'downgrade(d|s)? its own|audit.*(stale|superseded)|prescribed fix is a no-?op' ` → 5 (all about the *exclude list* being stale, which is the drawn-work case, not the plan/audit case); `grep -n -i -E 'phase-?start (verification|finding)|verify the (plan|task).{0,20}before' ` → 0 +- **Proposed home:** G (a rule) — the execution-time premise check — and a how-to/governance line beside the phase cadence +- **Portable because:** any multi-month project executes plans written from documents authored weeks earlier; the tooling moves faster than the prose, and an audit's own retraction is the least-read paragraph in the repository. + +### C4 — A hardcoded assumption fixed in one tool survives in its siblings: grep the tree for the literal and fix every twin in the same change +- **Evidence:** `phase-ends/logs/Phase28.md:36` — "`family_remap.img_path` (`:32-35`) hardcodes `0.4.dec` → `None` for the 4 SC07 overlays → … **member silently dropped as not-templatable**. Twin of the P27 T7 `new_overlay.sh` bug, left in a second tool." +- **What happened / what it cost:** The identical hardcoded payload path had been found and fixed in `new_overlay.sh` one phase earlier; the copy in `family_remap` was never looked for. It silently classified every member of the 4 newly-onboarded overlays as `LEN` ("not templatable") *and* poisoned their families' class to MIXED — hiding a 1,255-family / 6,268-member / 230,612-instruction pool underneath the very number the phase was chartered to measure (`:141`). The fix derives the payload from the yaml the BUILD reads and raises on a miss; its negative control is a 4-row before/after table (`:116`) in which one row must NOT change. +- **Not banked — greps:** `grep -n -i -E 'sibling tool|same (bug|hardcode) in (a second|every) tool|grep for (the )?twins|fix.*siblings' ` → 0; `grep -n -i -E 'in the same change|updates its siblings|related tools wired|every tool that shares' ` → 4 (R21 docs-currency, an agent-report accelerator, and cookbook:7297 "fix a wall, grep for the guard erected against it" — the guard case, not the duplicated-defect case); `grep -n -i 'hardcod' ` → many, all single-tool instances (decision-log:1168 records THIS instance as a fact, with no rule drawn) +- **Proposed home:** accelerator + G — "a tool defect is a class until proven singular: grep the literal across `tools/` before closing the fix" +- **Portable because:** decomp tooling grows by copy-paste from a working script; the same magic path, extension, or address constant ends up in three tools, and only one of them gets fixed. + +### C5 — A batch gate that bisects on failure re-runs the singleton against an unchanged baseline: special-case n==1 or pay a duplicate build on the hot path +- **Evidence:** `phase-ends/logs/Phase28.md:58` — "the `--chunk 1` double-build (`harvest_verify.py:199-213`) — the bisect re-runs `attempt()` on the same single element vs an unchanged baseline = a duplicate build. 1.35→1.0 builds/draft = **~26% fewer builds on the hot path**, one line." +- **What happened / what it cost:** The gate's design is "apply the batch, build; if the SHA moves, bisect". With `--chunk 1` — the *recommended* setting for wave batches, so the common case — the bisect's first step is the same single draft against the same baseline, i.e. a build whose answer is already known. Measured overhead: 1.35 builds per draft where 1.0 suffices, on the project's single most-repeated operation. The fix was one line. +- **Not banked — greps:** `grep -n -i -E '26% fewer|builds per draft|fewer builds' ` → 0; `grep -n -i -E 'bisect.{0,60}(single|one element|degenerate)|re-runs? attempt|redundant (re-)?build' ` → 0; `grep -n -i 'bisect' ` → 20 (all about *when* to bisect and mis-attribution, never the degenerate single-element case) +- **Proposed home:** accelerator (the gating-speed playbook: "a slow gate is a bug") +- **Portable because:** every decomp's byte gate is batch-apply + bisect-on-failure, and the degenerate branch of a divide-and-conquer step is where a doubled cost hides on the hottest path in the project. + +### C6 — The generated disassembly tree is shared mutable state: a fleet verify/clean chain and the per-function instruments cannot run at the same time +- **Evidence:** `phase-ends/logs/Phase32.md:255-256` — "**Gotchas (live):** `make clean` deletes `asm/` — never run rtu_match/twin_rescan/frontier_classify while an R22 chain runs · `gate_main` rewrites main's `asm/` too" +- **What happened / what it cost:** The fleet verification is `make clean && make extract-all && make check-all`, and its first step removes the extracted `asm/` that every per-function instrument (the real-TU match probe, the twin rescan, the frontier census) reads as ground truth. The banking gate for main rewrites the same tree. The phase carried this as a standing live gotcha and scheduled around it — launching the verify chain only at a pause, and reading its logs in the next session (`:139`) rather than working alongside it. +- **Not banked — greps:** `grep -n -i -E 'make clean deletes|deletes asm/|while (an? )?(R22|gate|build) (chain )?runs|corrupts the shared' ` → 0; `grep -n -i -E 'corrupts? the shared|shared asm|asm/ subtree|concurrent extract' ` → 1 (wave-playbook:621 — a *parallel worker's* symlinked `asm/`, the worktree case, not the verify-chain-vs-session case); `grep -n -i -E 'do not run .{0,40}while|never run .{0,40}while' ` → 1 (about a permuter, unrelated) +- **Proposed home:** accelerator / how-to (oracles and instruments) — "name the tools that read the generated tree, and serialise them against the verify chain (or give the chain its own worktree)" +- **Portable because:** every decomp regenerates a large disassembly tree from the ROM and reads it from a dozen tools; the clean-rebuild verification is exactly the step that deletes it, and the failure mode is a silently wrong answer rather than an error. + +### C7 — Write the phase synthesis in a FRESH session that re-reads the committed state cold; the cold re-read is what catches stale artifacts +- **Evidence:** `phase-ends/logs/Phase28.md:86` — "**PhaseEnd_Phase28 is the only remaining work** (Tier-1 Max synthesis — write it in a fresh session with headroom; it re-reads the committed state cold, **which is the discipline that caught tonight's stale-map/stale-.md traps**)." +- **What happened / what it cost:** Two of the phase's findings were stale artifacts that only surfaced because something read the committed state without the authoring session's assumptions: a derived family map that had never been regenerated after four binaries were onboarded, and a per-function note whose prose header described a state its own `.c` had moved past. The phase's own resume note therefore prescribes the synthesis be written by a session that has read nothing but the committed files — the same property that made those catches possible. +- **Not banked — greps:** `grep -n -i -E 'fresh session.*(cold|re-read)|re-read(s)? the committed state|cold read' ` → 0; `grep -n -i -E 'PhaseEnd.*(fresh session|headroom)|write the (synthesis|phaseend) in a' ` → 0 (the governance chapter documents the cadence and the hard stop, but not who writes the synthesis or why) +- **Portable because:** any agent-run project ends units of work with a synthesis; a synthesis written by the session that did the work inherits that session's beliefs, and the cheapest audit available is a context that only knows what is on disk. +- **Proposed home:** G (a rule) / how-to governance, beside the phase cadence + +## ALREADY-BANKED (one line each) +- `match_one`'s "fully isolated" default shared `.run/match`, so concurrent agents compiled into one scratch and one read another agent's function — a confident wrong verdict, not a crash — lives at `docs/matching-cookbook.md:4504`, `:9062-9065`, `:2488` +- A blacklist/exclude entry is a verdict from a specific gate and expires when that gate changes; purge and re-derive, never inherit — lives at `docs/decision-log.md:2780-2785`, `docs/how-to-ai-decomp/12-failure-museum.md:28`, `decomp-architect/templates/registry-E.decomp.md:245-247`, `decomp-architect/corpus/decomp-kernels.md:860` +- A wrong fix derived from a true diagnosis (a compiler-generated switch table is never named in C, so the fix is placement, not symbol substitution) — lives at `docs/decision-log.md:1201-1202`, `docs/matching-cookbook.md:4108` +- A per-function note's prose header goes stale against its own `.c`; verify a claim against the bytes, never trust a stale note — lives at `docs/matching-cookbook.md:3274` +- One agent's tidy-up swept eleven sibling deliverables from a shared dir; per-function work dirs + a deliverable dir no agent cleans + transcript replay — lives at `docs/wave-playbook.md:789-793`, `docs/accelerators.md:733-734`, `docs/matching-cookbook.md:37026-37030`, `decomp-architect/corpus/decomp-kernels.md:462` (DK-34) +- The all-`INCLUDE_ASM` first build is a NULL oracle for fine load-base errors (+8 byte-identical, +0x1000 fails the link) — lives at `docs/matching-cookbook.md:36777` +- A base vote whose targets are shared-engine function starts is OUTWARD-EXPLAINED, not evidence — lives at `docs/matching-cookbook.md:36772` +- A residual gets a PRODUCER CENSUS before any spelling sweep, and "PROVED" names the list it was proved against — lives at `phase-ends/DIGEST.md:270` (R69), `decomp-architect/corpus/decomp-kernels.md:589` (DK-45), `docs/how-to-ai-decomp/07-compiler-source.md:77-85`, `docs/accelerators.md:827` +- Two "final verdicts" (PROVED / PLATEAU) fell to producers missing from the census; every register pin came off once the source shape was right — lives at `docs/retrospective.md:33`, `:137`, `docs/how-to-ai-decomp/12-failure-museum.md:25-26`, `docs/decision-log.md:3404` +- Every tool fix carries a negative control, and it corrupts a scratch copy, never the tracked file — lives at `docs/decision-log.md:2266`, `:2280`, `docs/matching-cookbook.md:8459` (§128a), `:6874` +- `disc_code_sweep` was structurally blind to compressed code, so its "type 4: 0 onboarded" row was vacuous for 138 known binaries — lives at `docs/decision-log.md:2160`; the general law (a second, disagreeing oracle) at `phase-ends/DIGEST.md:211` (R34) +- A newly-discovered binary is not real until every consumer knows it, asserted by a gate — lives at `phase-ends/DIGEST.md:214-215` (R36), `docs/how-to-ai-decomp/10-integration-and-propagation.md:62-63`, `decomp-architect/templates/registry-E.decomp.md:59` +- Fix the metric instrument and reconcile the denominator across independent oracles before planting a completion flag — lives at `docs/how-to-ai-decomp/00-README.md:37`, `docs/decision-log.md:3107-3112`, `:669`, `docs/accelerators.md:664` (R41) +- A milestone's "ledgered with cost" disposition nearly closed the phase 15 functions short; the directive to crack everything banked 13 of the 15 — lives at `docs/decision-log.md:3335-3350`, `:3404`, `decomp-architect/templates/registry-E.decomp.md:364` +- A concurrent tool's dotfile in `src/` broke the build's `find`-based `C_SRCS`; guard at the consumer with `-not -name '.*'` — lives at `docs/accelerators.md:765-766`, `decomp-architect/corpus/decomp-kernels.md:477`, `docs/matching-cookbook.md:37089-37094` +- The fleet's one atypical binary is where every majority-shaped tool assumption surfaces ("a tool written for the overlay class met the resident class and stayed silent") — lives at `docs/decision-log.md:3214-3220`, `:3238-3241` +- A helper must refuse an empty work list (it built the unchanged tree and exited 0); write "banked" only from the tool's printed line — lives at `phase-ends/DIGEST.md:268` (R68), `docs/how-to-ai-decomp/02-byte-gate.md:71-73`, `decomp-architect/templates/registry-E.decomp.md:189` +- Agents write deliverables early (draft first, verdict JSON last); budget for usage-limit outages and resume with context intact — lives at `phase-ends/DIGEST.md:266` (R67), `decomp-architect/corpus/decomp-kernels.md:446`, `docs/accelerators.md:787-788`, `docs/how-to-ai-decomp/05-cards-lanes-waves.md:159` +- A coordinator dies mid-wave and completions land in a dead session; recover verdicts and drafts from the harness transcripts by tool — lives at `docs/wave-playbook.md:789-795`, `decomp-architect/corpus/decomp-kernels.md:446`, `docs/decision-log.md:3251-3273` +- A masked scorer that normalises relocations is blind to relocation-class deltas and to `j`/`jal` destinations; the gate decides — lives at `docs/matching-cookbook.md:19925` (§195-D), `:1217`, `:5450`, `docs/decision-log.md:1331` +- A probe's failures must be classified (DIFF vs PLUMBING vs never-compiled); the gate number is not the rate — lives at `docs/matching-cookbook.md:7598`, `docs/how-to-ai-decomp/05-cards-lanes-waves.md:142`, `decomp-architect/templates/registry-E.decomp.md:316` +- Report a rate as "as-tooled, ceiling unknown" when the dominant failure mode keeps resolving to tooling — lives at `docs/decision-log.md:1209`, `:1218`, `:1239-1246` +- Propagation tooling needs an "extend an existing macro-backed group to a newly-onboarded binary" mode, distinct from crack→author→instantiate — lives at `docs/matching-cookbook.md:9852`, `:5953`, `:9556-9557`, `§75a` at `:5959` +- Refuting the evidence for a claim leaves it unmeasured, not refuted; ship every finding with its bounds — lives as the cookbook's pervasive ⚠ BOUNDS/UNPROVEN convention (e.g. `docs/matching-cookbook.md:13427`, `:16073`) and `decomp-architect/corpus/decomp-kernels.md:591-592` (DK-45) +- A 0% produced by the wrong tool for the class (a sweep with no carve step) and a real wall are the same number and opposite facts; route by the class flag — lives at `docs/decision-log.md:1262`, `:1277`, `docs/matching-cookbook.md:4098`, `:8180` diff --git a/.run/P33.5/log-mining/Phase29-1of4.md b/.run/P33.5/log-mining/Phase29-1of4.md new file mode 100644 index 0000000000..52706a3be7 --- /dev/null +++ b/.run/P33.5/log-mining/Phase29-1of4.md @@ -0,0 +1,159 @@ +# Log mining — Phase29-1of4 +Files/ranges: `phase-ends/logs/Phase29.md:1-2580` (read through :2619 for context; only :1-2580 reported) · Lines read: 2580 of 2580 +Candidates considered: 46 · NEW: 6 · ALREADY-BANKED: 40 + +Grep set used for every "already banked?" test (abbreviated `` below): +`docs/matching-cookbook.md docs/cookbook-index.md docs/decision-log.md docs/accelerators.md docs/retrospective.md docs/wave-playbook.md docs/how-to-ai-decomp/*.md decomp-architect/corpus/decomp-kernels.md decomp-architect/templates/registry-E.decomp.md phase-ends/DIGEST.md` +(the 3.5 MB cookbook was only ever grepped, never read). + +## NEW + +### C1 — A stop/continue instrument must aggregate at exactly the unit the decision is made in; one that averages a finer unit manufactures a false "we are at the floor" +- **Evidence:** `phase-ends/logs/Phase29.md:1283-1288` — + "`burndown.py` averaged the last 3 INTER-COMMIT deltas, but the ROI criterion is per-SESSION yield. Three mid-session + snapshots of a **+0.7pp** session averaged to **+0.23** and printed **"AT THE FLOOR — consider closing P29"**. + **I nearly closed the phase on it.**" +- **What happened / what it cost:** The phase's exit condition was "per-session yield floors out", but the instrument built + to decide it sampled at commit granularity, so one productive session read as three weak ones and printed a close-the-phase + verdict that was believed until re-derived. Fixed by an explicit `--session-close` boundary marker plus an honest + "0 SESSION-to-SESSION delta(s) logged — need >=3" output, and thereafter every checkpoint had to carry "DO NOT close + P29 on ROI — the floor is undetermined". The same granularity error recurred (`:1918-1921`, a `--session-close` logged + mid-session by mistake), i.e. it is a repeating shape, not a one-off. +- **Not banked — greps:** `grep -rn -i -c 'burndown' ` → 0; `grep -rn -i -c 'per-session yield' ` → 0; + `grep -rn -i -c 'stopping criterion' ` → 0; `grep -rn -i -c 'at the floor' ` → 1 (cookbook:30292, a scheduler + prologue "floor", unrelated); `grep -rn -i 'burn-down' ` → 2 (decision-log:1680 records the phase's *state*, + "the burn-down floor is still undetermined", never the lesson). +- **Proposed home:** DK (a kernel), with an accelerators cross-reference +- **Portable because:** every long decomp needs a "when do we stop this track" instrument, and the failure mode — sample at + commit/day granularity, decide at session/phase granularity — is independent of console, compiler and harness. + +### C2 — Batch size is a RISK lever, not a token lever: isolated agents cost ~N× one agent whether concurrent or serial, so size a batch by the unverified spend you are willing to lose before the next measured yield +- **Evidence:** `phase-ends/logs/Phase29.md:2387-2392` — + "Batch size is a RISK lever, not a token lever: each agent drafts one fn in its own context, so N agents cost ~N× one + agent whether concurrent or serial (~110k tok/drafted fn, measured s14+s15). Concurrency buys wall-clock only. + Therefore **batches of ~5–6**: ~600k tokens of exposure per batch, a measured yield before committing the next, + and a clean stop at any boundary." +- **What happened / what it cost:** This was written after waves of 24 drafters repeatedly produced high draft counts with + low bank rates (`:1973-1980` 24 drafters → 6 banked; `:2478-2489` a queued wave with **no fuel at all**), i.e. large + batches bought no efficiency and lost a full batch of exposure per bad premise. The rule pairs with the measured unit + cost (~110k tok per drafted function; ~460k per banked function before recovery, ~200k after) so batch size becomes a + budget decision rather than a habit. +- **Not banked — greps:** `grep -rn -i -c 'risk lever' ` → 0; `grep -rn -i -c 'concurrency buys' ` → 0; + `grep -rn -i -c 'wall-clock only' ` → 0; `grep -rn -i 'batch size' ` → 2 (decision-log:2333 / cookbook:16413, + both making the *different* point that batch size was not the yield discriminator). + Nearest neighbours checked and distinct: `docs/how-to-ai-decomp/08-models-and-budgets.md:31-33` ("budget per lane"; + "concurrency is a budget" = provider rate limits) and `:70-71` in chapter 08 ("isolated agents for breadth, never serial + in the main loop" = the orchestrator's quadratic context) — neither states the exposure-sizing rule. +- **Proposed home:** DK (a kernel) or `docs/how-to-ai-decomp/08-models-and-budgets.md` Budgets section +- **Portable because:** it is a property of per-agent isolation and of any pay-per-token fan-out, and it converts "how big a + wave?" from taste into an explicit loss cap. + +### C3 — A repair/`--recover` mode whose cost is (exceptions × population) must be gated on a MEASURED exception count; for a broadly divergent set, drop rather than recover +- **Evidence:** `phase-ends/logs/Phase29.md:1635-1639` — + "`dedup_propagate --recover` on func_80169228 (many byte-divergent stragglers) THRASHED >1hr (re-gates the fleet per + excluded straggler = quadratic). Killed + reverted 326 half-mutated files … **LESSON: --recover is only for a FEW + stragglers; a broadly-divergent family must be dropped or capped, never --recover'd.**" +- **What happened / what it cost:** The recovery flag exists to rescue a handful of exceptions; run against a family whose + members genuinely diverge, it re-validated the whole fleet once per exception, burned over an hour, and left 326 + half-mutated files that had to be reverted to the committed baseline. Nothing was lost only because the banks had been + committed before the propagate. Later sessions consistently *dropped* stragglers instead (`:1710-1712`, five dropped, + "`--recover` is the documented thrash hazard and was NOT used"). +- **Not banked — greps:** `grep -rn -i -c 'thrash' ` → 0; `grep -rn -i -c 'quadratic thrash' ` → 0; + `grep -rn -i -c 'half-mutated' ` → 0; `grep -rn -i -c 're-gates the fleet' ` → 0. + (`--recover` itself appears 11× in the cookbook, always as *how to use it* — e.g. `:2451`, `:2471`, `:2499`, `:4193` — + never with its complexity or its stop condition. `straggler` hits in wave-playbook/how-to are about wave tails.) +- **Proposed home:** accelerator (with a cookbook cross-reference at the §55b propagation law) +- **Portable because:** every mass-application tool acquires an "and fix up the exceptions" mode, and the shape + (per-exception full re-validation) is the default naive implementation on any project. + +### C4 — A defect reasoned into a sibling tool is LATENT until a run shows its signature; do not patch it on theory right after that tool produced a clean run +- **Evidence:** `phase-ends/logs/Phase29.md:1191-1196` — + "`jtbl_family_bank` isolates AROUND THE STUB … the same ordering defect just fixed in `harvest_verify`. It did **not** + fire across 274 swept siblings this session … so it is latent, not universal; fixing it on theory after a 274/274 run + would be the very error this phase keeps catching. **Fix it when a sweep fails with the signature (isolation fires → + gate DIFF), not before.**" +- **What happened / what it cost:** After root-causing a real ordering bug (the transform must follow the splice) the same + shape was visible by inspection in a second tool that had just completed 274/274 siblings byte-identical. The judgement + recorded is to *name the signature and wait* rather than mutate a currently-correct mass tool — the phase had repeatedly + paid for edits made on theory (e.g. the fleet-wide `--revert` that broke 138/140, `:97-105`). +- **Not banked — greps:** `grep -rn -i -c 'latent, not universal' ` → 0; `grep -rn -i -c 'on theory' ` → 0; + `grep -rn -i -c 'fix it when.*fires' ` → 0; `grep -rn -i 'did not fire' ` → 2 (cookbook:7615 and :21268, both + about a compiler mechanism not firing, not about deferring a tool fix). +- **Proposed home:** DK (a kernel) — it is the deliberate counterweight to "fix the class, not the instance" +- **Portable because:** every project reaches the moment where one fixed tool implies a bug in its sibling, and the choice + between speculative patch and signature-triggered patch is the same on any codebase. + +### C5 — Sweep for orphaned worker processes at every session boundary; a dead-pipe compiler or a self-matching wait loop holds a core forever and nothing reports it +- **Evidence:** `phase-ends/logs/Phase29.md:1029-1030` — "killed an orphaned `cc1` from the Jul-21 session that had been + burning a full core for **13h23m** (pid 104350, dead pipe)"; and `:1289-1292` — "I left two `while pgrep -f ; do + sleep; done` waiters spinning (one for 4 h) — they self-match their own `bash -c` command line and can never exit." +- **What happened / what it cost:** Two independent leak shapes in one phase — a compiler child left alive by a killed + parent's dead pipe, and shell wait-loops that can never terminate because their pattern matches themselves — each + silently consuming a core across sessions on a machine whose parallelism is the project's throughput. Neither is + reported by any gate, health check or tool: the tree is clean, the fleet is green, and the box is just slower. +- **Not banked — greps:** `grep -rn -i 'orphan' | grep -iE 'process|cc1|worker|pid|core|cpu'` → 0 relevant (all 89 + cookbook hits are orphan *registers/slots/notes*); `grep -rn -i -c 'leftover process' ` → 0; + `grep -rn -i -c 'ps aux' ` → 0; `grep -rn -i -c 'consuming a core' ` → 0. (`R79`/`accelerators.md:197` cover + `pkill -f` killing the *calling shell*, and cookbook:5328 covers a `pkill` pattern killing *sibling runs* — neither is + the leaked-process sweep.) +- **Proposed home:** accelerator (a session-boundary checklist item, next to the tree-clean / fleet-green checks) +- **Portable because:** any project that spawns compilers, searches or workers from an agent harness leaks them the same + way, and the cost scales with how much of the schedule is CPU-bound. + +### C6 — Land a pure rename and a semantic/layout change as separate gated edits, so a gate failure attributes itself +- **Evidence:** `phase-ends/logs/Phase29.md:1959-1962` — "`VECTOR` left UNCHANGED deliberately: ours is 12B vs PsyQ's 16B … + vx/vy/vz offsets already agree so a fix is likely byte-neutral, but it is a LAYOUT change and must not be bundled with a + rename (an R22 failure would then be ambiguous about which caused it). Own commit, later." +- **What happened / what it cost:** In the same task a 205-file pure rename (swapping which name carries the true PsyQ + `MATRIX` layout) went through the full-fleet gate green, precisely because it was *only* a rename; the layout question was + deferred to its own gated change. Recorded as a rule rather than a cost — the phase had, however, just paid for the + opposite shape (`:1680-1685`, a type selector that changed layout for 103 overlays under an edit believed to be a + consistency fix, caught only by the full-fleet gate at 37/140). +- **Not banked — greps:** `grep -rn -i -c 'bundle.*rename' ` → 0; `grep -rn -i -c 'rename.*layout change' ` → 0; + `grep -rn -i -c 'one variable at a time' ` → 0; `grep -rn -i -c 'ambiguous.*R22' ` → 0; `grep -rn -i 'own commit' + ` → 5 (all unrelated: decision-log:3463 the purge commit, cookbook:8680 `tables=` authority, etc.). +- **Proposed home:** G (a rule) or the byte-gate how-to chapter +- **Portable because:** it is a property of any single-bit oracle (byte-identical / not): two changes in one gated edit make + a red result uninterpretable, and byte-neutral refactors are exactly where the temptation to bundle appears. + +## ALREADY-BANKED (one line each) +- A shared fixed scratch path silently clobbers itself under parallelism (14 of 1,752 drafts lost, invisibly, in every parallel wave ever run) — `docs/matching-cookbook.md:4499` and `:9058` +- Wrapper/pipeline exit status lies (`| tail` masks make; `grep -c` returns 1 on no match); read the gate's own output line — `docs/matching-cookbook.md:6620`, §93 `:6956`, §136a `:9010` +- "Twice, identically" is two reads of one contaminated state, not a replication; re-apply a fault from a known-clean tree before recording it as a mechanism — `docs/decision-log.md:1457-1493` +- A failure classifier that scans whole stderr latches onto a benign warning and returns a CONSTANT label, which carries no information — `docs/matching-cookbook.md:4840` +- An undo whose scope is narrower than the stage's write scope destroys work no byte-gate can see; declare and assert each stage's blast radius — `docs/matching-cookbook.md:4864`, `:5075`; `decomp-architect/corpus/decomp-kernels.md:878` +- A per-item gate is structurally blind to shared-state damage (it was RIGHT about its one binary while 137 were broken) — same blast-radius entries, `docs/matching-cookbook.md:4864` +- A truthy `or DEFAULT` fallback made the gate compare one binary against another's hash for a month; two callers of one oracle that disagree ARE the bug report — `docs/decision-log.md:1499-1530` +- Rank remaining work by LIVE (still-open) siblings, not by total reach; `nins×reach` over-counts — `docs/decision-log.md:1636-1672` +- Reconcile toward the byte-truth, not toward the incumbent declaration (the header was the simplified side) — `docs/matching-cookbook.md:4920`; and a shared-header edit is fleet-blind under a per-binary gate — `docs/decision-log.md` §63 UPDATE entries +- A drafting agent must copy its winner back to a canonical path with an existence check (24 winners nearly lost to scratch filenames; recovered from transcripts) — `docs/how-to-ai-decomp/05-cards-lanes-waves.md`, `docs/wave-playbook.md`, `docs/accelerators.md` (`agent_drafts_restore`) +- A pre-filter is evidence ONLY about what it filtered — pre-filter on a unit that FAILED — `docs/decision-log.md:1868` +- R14 applies to your own 3-line scripts: two produced false evidence before any project tool did — `docs/decision-log.md:1867` +- Different types sharing one identifier must be UNIQUIFIED, never merged to a canonical layout (merging broke 103 overlays); types-first does not help matching, a draft-time naming convention (address-suffix inventions; forbid bare generic names) would have saved a session; seed the documented SDK types on day one — `docs/decision-log.md:1875-1903`, kernels `:798-815` +- The carve must FOLLOW the splice — and so must the isolation — `docs/matching-cookbook.md:4699`, `:4836`, `:4869` +- The widest write in a pipeline is the one most likely to be undeclared, and a containment guard measured from `git status` is blind once anything commits — cookbook §66a +- A metric scraped from another tool's prose goes NULL silently on a label change (50 commits recorded `fleet None%`) — cookbook §66b +- A selection tool whose candidate set derives from the wrong directory manufactures both false work and false exhaustion — `docs/matching-cookbook.md:5263` (§66c) +- `pkill -f ` matches every concurrent run and silently kills siblings — `docs/matching-cookbook.md:5328`; the self-kill variant `docs/accelerators.md:197`, `phase-ends/DIGEST.md:295` (R79) +- A carried "do not re-derive" claim can be inverted and its preserved probes vacuous; a probe whose output does not contain the phenomenon proves nothing — `docs/decision-log.md:1369-1403` +- Blockers STACK and the compiler reveals only the FIRST; route by the static oracle's complete list — `docs/matching-cookbook.md:5058`, `:8843` +- Persist a search's winner durably BEFORE the step that consumes it; a tool that has been failing for a long time accretes latent bugs on its success path, so budget for the first success to fail — `docs/matching-cookbook.md:4539-4542` +- Two implementations of one capability, the outer one ineffective AND destructive → delete, don't patch — `docs/matching-cookbook.md:4837` +- Staging ≠ banking; a classifier's `n_templatable`/PURE verdict is a PREDICTION the gate is free to refuse — `docs/matching-cookbook.md:4364-4368`; "is a PREDICTION" also in `docs/accelerators.md`, kernels +- Ladder/scratch dirs that ACCUMULATE across runs silently widen a stage's input set (it banked a function it was never asked to try) — `docs/matching-cookbook.md` ("accumulates across runs", "banked a function it was never asked") +- A revert that restores tracked config but not derived artifacts leaves a tree GIT-CLEAN BUT UNBUILDABLE; a reverted config needs a re-extract — `phase-ends/DIGEST.md:191`, `decomp-architect/templates/registry-E.decomp.md:72`, `docs/accelerators.md:445` +- Regenerate digests on a clean tree — `docs/matching-cookbook.md:5248` +- Cracking GENERATES sweep fuel (sweeps only pay riding a fresh ×1 crack) — and it is family-specific, not a blanket ×N — `docs/decision-log.md:1543-1561` +- Census the corpus SHAPE (duplication, families, reach × size, the unique tail) before choosing a strategy; the ×N economics end and a ×1 tail remains — kernels DK-8 `decomp-architect/corpus/decomp-kernels.md:108-118` +- Attemptable work is bounded by decompiler-cache coverage; high-reach fuel is created by a per-target import (greedy cover), not found — `docs/decision-log.md:1639-1641`, `:1731-1740` +- Never foreground-build while a background clean-fleet verification runs (phantom pass); isolate builds — `docs/how-to-ai-decomp/02-byte-gate.md:43-45`, cookbook concurrent-splice entries `:27784`, `:28729` +- An A/B whose treatment arm receives already-treated input establishes nothing (the ablation is confounded) — `docs/matching-cookbook.md:12938`, `:14277`, `:15921` +- A wrapper that swallows its child's diagnostics makes the failure undiagnosable (swallowed-error class) — `docs/retrospective.md:47`, `docs/how-to-ai-decomp/02-byte-gate.md:30`, `docs/decision-log.md:1141` +- A relocation-masked proxy oracle's validity is CLASS-DEPENDENT (trust it for self-decl fixes, distrust it for callee ones) — cookbook §65c +- A shared-state wall can have a per-unit local escape (de-macroize the instantiation in this overlay's own TU) — cookbook §65b, and the §20 DEF-conflict refutation +- The wave bottleneck is INTEGRATION, not idioms (~92% of drafts byte-correct, ~27% banked) — kernels DK-7 `decomp-architect/corpus/decomp-kernels.md:96-106` +- A tool that cannot handle a class must REFUSE loudly and say that a 0% from that path is a TOOL ARTIFACT, not a wall — R43 / `phase-ends/DIGEST.md` +- Read the FINAL SHA, not the verified count: `None` = no image; a real hash ≠ locked = the edit moved bytes — cookbook §65f +- A "byte-neutral" transform measured neutral 14 of 15 times is not byte-neutral — cookbook §65f +- Commit the cheap verified banks BEFORE the expensive propagate, or a propagate failure takes the banks with it — cookbook §55b (and R42 in `phase-ends/DIGEST.md`) +- Trust the SOURCE, not the report (`grep INCLUDE_ASM`, never the tool's "banked: N") — cookbook §55b trap 4, R66/`phase-ends/DIGEST.md` diff --git a/.run/P33.5/log-mining/Phase29-2of4.md b/.run/P33.5/log-mining/Phase29-2of4.md new file mode 100644 index 0000000000..bb9fb67428 --- /dev/null +++ b/.run/P33.5/log-mining/Phase29-2of4.md @@ -0,0 +1,194 @@ +# Log mining — Phase29-2of4 +Files/ranges: `phase-ends/logs/Phase29.md:2581-5160` (read 2551–5195 incl. ±30 context; reported range only) · Lines read: 2580 of 2580 +Candidates considered: 43 · NEW: 7 · ALREADY-BANKED: 36 + +Slice content: SESSION-17 → SESSION-21 of Phase 29 (the giant/permuter loop, the §73–§91 declaration-plumbing +harvest, the T0 frontier survey, the family-exemplar mass wave). The cookbook covers this stretch densely +(§66d–§91 were distilled in-session under R30), so nearly every *matching* lesson here is already banked; what +survives the greps is planning-, ledger- and prompt-shaped. + +## NEW + +### C1 — A verification flag that short-circuits the tool's WRITE path leaves the stale artefact in place and still exits 0 +- **Evidence:** `phase-ends/logs/Phase29.md:2762-2765` — + > "**⚠️ The sharp edge that hid it:** `worklist.py --assert-partition` **exits at the assertion and never + > rewrites the doc** — so "regenerating" with that flag leaves the stale file in place and still exits 0. + > (Its own assertion printed "160 live stubs, 160 rows → PARTITION OK" while the doc it left behind said 223.)" +- **What happened / what it cost:** The decision spine (`docs/worklist.md` + `.run/fuel_manifest.json`) was **9 days + stale**, claiming 223 live stubs / 870,668 ins where the truth was 160 / 583,077, and it ranked **three + already-banked functions in its top 7**. The session had "regenerated" it — with the assert flag on, which + never reached the write. The two contradictory numbers were on screen simultaneously and the exit code was 0. + Once actually regenerated, the top of the spine changed the whole plan (integration, not the giants). +- **Not banked — greps:** `grep -n -i 'assert-partition' ` → 0; + `grep -n -i 'exits at the assertion' …` → 0; `grep -n -i 'never rewrites the doc' …` → 0; + `grep -n -i 'still exits 0' …` → 0. (`grep -n -i 'plan.*stale' …` → 1, `docs/matching-cookbook.md:2509`, a + volatile-frame codegen idiom — unrelated.) The nearest banked kin, DK-19 / R32, is about a tool asserting its + *denominator*; nothing records a verification flag suppressing the tool's primary side effect. +- **Proposed home:** DK (a kernel), sibling to DK-19 — *"a --check/--assert mode must be additive to the write, + never a substitute for it; a regeneration that printed a verdict and wrote nothing is indistinguishable from a + regeneration that ran."* +- **Portable because:** every project grows `--dry-run`/`--check`/`--assert` variants of its generators, and the + plan is read off the generated artefact — the failure is in the flag design, not in MIPS or gcc. + +### C2 — A diagnosis earns belief when it predicts its own RESIDUAL membership, not when it explains the failures already seen +- **Evidence:** `phase-ends/logs/Phase29.md:3690-3695` — + > "The class-A census (`func_80012ABC`: **73 `s32` vs 7 `s16`**) had predicted the class-B fix would leave exactly + > the class-A overlays behind, and it left **3** — the same 3 the original sweep's blocker breakdown named… + > **A measurement that predicts its own residual to the overlay is the strongest evidence this session produced + > that the classes are real and not a story fitted to the failures.**" +- **What happened / what it cost:** The same session had twice built plausible blocker taxonomies that were stories + fitted to the failures (`3603-3607`: "I ALMOST GENERALIZED FROM ONE SAMPLE"; `3614-3625`: the full-sweep census + *reversed* the ranking, the win ranked first was worth 3 overlays and the one ranked third was worth 132). The + cheap discriminator that settled it was forward: state, before the fix, exactly which members the fix will *not* + clear — then count them. +- **Not banked — greps:** `grep -n -i 'predicts its own residual' ` → 0; + `grep -n -i 'story fitted' …` → 0; `grep -n -i 'fitted to the failure' …` → 0; + `grep -n -i 'predicted the residual' …` → 1 (`docs/matching-cookbook.md:24094`, one instance of a codegen tell, + not the epistemic rule); `grep -n -i 'strongest evidence' …` → 2 (both byte-evidence for specific codegen laws). +- **Proposed home:** DK (a kernel) — pairs with DK-61 ("a known-true case before reading any instrument's output"): + *before* acting on a classification, have it name the members it predicts will still refuse, and check that set. +- **Portable because:** it is a falsification protocol for any bucketing of failures — blocker classes, error + taxonomies, model-failure clusters — and it costs one grep more than the taxonomy already cost. + +### C3 — Make the DRAFTER run the pre-gate guard and return the NAMED banking prerequisite; a batch of diagnosed candidates is worth far more than a batch of opaque MATCHes +- **Evidence:** `phase-ends/logs/Phase29.md:5093-5096` — + > "**⚠️ These are CANDIDATES, not banks** … an earlier 11-core wave had 9/9 match_one MATCHes all gate-fail on + > integration. What makes this wave different is that the agents were *told* to run `symcheck` and report the + > blocker — so instead of 8 opaque MATCHes we have 8 diagnosed ones." +- **What happened / what it cost:** The wave prompt (`4907-4915`) made a `symcheck` run mandatory on any claimed + MATCH and asked for the blocker by name. The result table at `5098-5107` gives, per exemplar, the exact banking + prerequisite — and reading down it revealed that **six of eight were blocked on one mechanical lever** standing + in front of ~153,596 templatable instructions (`5109-5112`). Without the reported blockers that batch would have + read as eight unrelated integration failures. +- **Not banked — greps:** `grep -n -i 'report the blocker' ` → 0; `grep -n -i 'diagnosed ones' …` → 0; + `grep -n -i 'opaque MATCH' …` → 0; `grep -n -i 'mandatory .*symcheck' …` → 0; + `grep -n -iE 'agent (must|runs|shall)|pre-gate' docs/wave-playbook.md docs/how-to-ai-decomp/0{4,5}-*.md` → 1 + (`docs/how-to-ai-decomp/05-cards-lanes-waves.md:127` — the *reconcile lane*, a separate post-draft agent per gate + group, not a requirement on the drafter's own return). `§67a` (`docs/matching-cookbook.md:5439`) banks the guard + as a tool; nothing banks it as a drafter-return contract. +- **Proposed home:** wave playbook / how-to chapter 05 (card & return schema), plus a G-rule: *a drafter's return + is `MATCH + guard verdict + named banking prerequisite`, or it is not a return.* +- **Portable because:** it is a prompt/return-schema rule for any parallel drafting harness — the marginal cost is + one deterministic tool run per agent, and it converts a batch's failures into a single sortable column. + +### C4 — When two blockers are orthogonal, a classifier's if-chain ORDER silently becomes the label — cross-tabulate, never bucket +- **Evidence:** `phase-ends/logs/Phase29.md:4958-4960` — + > "The T2 table above ordered its if-chain with `jr` FIRST, so any family carrying a mid-jr was bucketed as + > "jr → §81 carve chain" **regardless of whether its exemplar needed a crack at all**. That conflated two + > orthogonal axes and under-reported the zero-crack pool by 15 families." +- **What happened / what it cost:** The primary axis is *does the exemplar need a CRACK?*; the sweep path (plain vs + carve chain) is orthogonal to it. Bucketing on first-match collapsed them and hid **15 families / 123,482 + templatable ins** of no-drafting work — the pool read as 45 fams / 85,017 ins instead of 60 / 208,499 + (`4965-4982`). The log itself names the recurrence: this is the same shape as the previous session's "FREE" + label — *"a classification presented as a route"* (`4985-4987`). +- **Not banked — greps:** `grep -n -i 'classification.*route' ` → 0; `grep -n -i 'orthogonal axes' …` → 0; + `grep -n -i 'if-chain' …` → 3 (all `§222`, the switch-vs-if-chain *codegen* idiom); + `grep -n -i 'cross-tabulat|cross tab' …` → 1 (`docs/matching-cookbook.md:30650`, a jump-table reading technique). + `§136h` banks *"do not price a pool without probing a member"*; nothing banks the bucketing defect that produced + the pool's shape in the first place. +- **Proposed home:** DK (a kernel), next to DK-19 — *"a routing table's rows are its if-chain's order; if two + conditions can hold at once, emit a cross-tab and let the reader see both."* +- **Portable because:** every triage script anywhere is an if-chain over overlapping predicates, and the cost is + paid in work that never gets planned rather than in a visible error. + +### C5 — Sequence a phase so the cheapest thing that can invalidate everything below it runs FIRST +- **Evidence:** `phase-ends/logs/Phase29.md:4322-4327` — + > "Not caution — **T0 is the cheapest thing that can invalidate everything below it.** If the families template, + > T2.2 becomes "crack N exemplars and stamp" and most of that bucket evaporates. If they do not, we grind + > *knowingly* instead of hopefully. R35 exists because this project has scoped whole phases against broken + > readings twice; P26 spent its longest phase on a thesis its own tools had already refuted." +- **What happened / what it cost:** The open question (do the remaining functions cluster into templatable families?) + had **flipped three times, every time on TOOLING, never on the compiler** (`4296-4305`) — P25 yes, P26 byte-refuted + at ≈0%, P28 found the refuting probe had a missing carve and the same family banked 89%, P29 found the residual + was an `-O0` compile-flag artefact. It cost **zero agent tokens** to re-measure, and it decided the method for the + largest remaining bucket. Running it first is what made the rest of the plan honest. +- **Not banked — greps:** `grep -n -i 'cheapest thing that can invalidate' ` → 0; + `grep -n -i 'invalidate everything below' …` → 0; `grep -n -i 'sequence.*cheapest|order the work so' …` → 0; + `grep -n -i 'order for discovery' …` → 0. Adjacent but different: `R35` (`phase-ends/DIGEST.md:212`, fix the + instrument before trusting it) and `§3-The` (`docs/cookbook-index.md:1571`, triage cheapest-first *within* a + function's levers) — neither is a phase-ordering rule. +- **Proposed home:** DK (a kernel) or a G-rule for phase planning — *"rank a phase's tasks by (cost) ÷ (how much of + the plan below them the result could delete); the cheapest high-invalidation measurement is task 1."* +- **Portable because:** it is pure plan structure — it applies to any long project whose method choice depends on an + unmeasured property of the remaining work. + +### C6 — A reach-weighted gain figure (size × copies) is not a size; every number must say which of the two it is +- **Evidence:** `phase-ends/logs/Phase29.md:4775-4777` — + > "**⚠️ SIZING TRAP (Drew caught me on this):** `nins × members` is **reach-weighted gain-ins**, NOT a function + > size. `func_80144090` is **154 ins × 136 copies**, not a 20,944-ins monster. **Always label which one you are + > quoting.**" +- **What happened / what it cost:** The endgame map's whole bucket table was priced in gain-ins, and a 154-instruction + function sat in it reading like a 20,944-instruction monster — which mis-prices the *effort* by two orders of + magnitude while the *value* is right. The same conflation runs through the pool tables in this slice (`4370-4376`, + `4991-5006`), where "208,499 ins" is repeatedly cautioned against being read as available work. +- **Not banked — greps:** `grep -n -i 'gain-ins' ` → 0; `grep -n -i 'label which one you are quoting' …` → 0; + `grep -n -i 'not a function size' …` → 0; `grep -n -i 'reach-weighted' …` → 1 + (`docs/matching-cookbook.md:4613`, uses the term in passing, carries no warning). Adjacent: `R41` / + `DK-60` (`decomp-architect/corpus/decomp-kernels.md:752`) require a *denominator* on every rate — this is a + different defect, a **product read as a magnitude**, and neither R41 nor DK-60 names it. +- **Proposed home:** accelerator, or an extension line on DK-60 — *"a weighted total names its weight; size × copies + is never quoted without the multiplier visible."* +- **Portable because:** any decomp/port with duplicated code (overlays, statically-linked libraries, templated + families) prices work as size × instances, and effort tracks the size while value tracks the product. + +### C7 — A repair ladder must probe whether each stage is NEEDED before applying it, or it silently escalates a binary-local bank into a fleet-shared one +- **Evidence:** `phase-ends/logs/Phase29.md:3377-3388` — + > "**⚠️ HYPOTHESIS TO TEST AFTER R22 — the recovery tool may have taken a FLEET-TIER edit it did not need.** + > … So the relaxation looks unnecessary, and it converted a T1 binary-local bank into a **T2 fleet-shared** one + > (blast radius 138 overlays, R22 mandatory) for nothing. … **`recover_integration` should probe whether a stage + > is NEEDED before applying it** — an unrequested tier escalation is exactly the class R32–R35 exist to catch." +- **What happened / what it cost:** Banking one 123-instruction function rewrote `src/shared/engine_core.h` (2 lines) + plus two overlay files, relaxing a prototype whose conflict belonged to the *old, discarded* drafts — the banked + definition already agreed with it. The price of an unneeded T2 write is a mandatory full-fleet R22 cycle and a + 138-overlay blast radius. The same session shows the other half of the bill: the ladder's stages left byte-neutral + edits behind on 0-bank runs (`4473-4481`, `4635-4639`) because "byte-safe" was standing in for "wanted". +- **Not banked — greps:** `grep -n -i 'probe whether a stage is needed' ` → 0; + `grep -n -i 'edit it did not need' …` → 0; `grep -n -i "unnecessary.*fleet.*edit" …` → 0; + `grep -n -i 'tier escalation' …` → 1 (`docs/decision-log.md:281`, a *model* escalation ladder). The closest banked + item is `§89a` (`docs/matching-cookbook.md:6766`, "MEASURE the write set; do not assert its tier") — which detects + the escalation *after* the tool has taken it; nothing requires the tool to establish a stage's necessity first. +- **Proposed home:** DK (a kernel) or a G-rule, as the precondition half of §89a — *"a repair stage runs only after + its blocker is observed; a ladder that applies stages unconditionally buys the widest blast radius it can reach."* +- **Portable because:** any automated fixer with a ladder of increasingly-invasive transforms (codemods, lint + autofix, migration tools) has the same shape: the cheapest stage is skipped, the widest is applied, and the + verification cost scales with the widest write actually taken. + +## ALREADY-BANKED (one line each) +- Assert on the expected SUCCESS STRING, never `$?` after a pipe (`head`'s rc read as `make`'s → a false BYTE-IDENTICAL report) — lives at `docs/matching-cookbook.md:6619` (§85) +- A comparison tool must share its reference oracle's index space exactly (`objdump -dr` elides identical runs; use `-drz`) — `docs/matching-cookbook.md:6806` (§90a) +- Little-endian hex text in a `.s` word field must be byte-swapped before comparison — `docs/matching-cookbook.md:6806ff` (§90a) +- A helper that cannot answer returning an empty set (a library call gated on a CLI-only global) — `decomp-architect/corpus/decomp-kernels.md:896`, `docs/how-to-ai-decomp/02-byte-gate.md:30` (R32/R43/G28) +- A warning that fires on ambiguity is noise, not a signal (~90 spurious fires across 38 tables) — `docs/matching-cookbook.md:6869` +- A do-not-re-buy entry is scoped to its BASE; re-run the negative list after the base moves — `docs/matching-cookbook.md:6304`, `:6316` (§80) +- Crack the smaller family member first; it is a lever library for the larger one — `docs/matching-cookbook.md:6252` +- Budget two passes at behemoth size; a 99% round-1 is on-plan, not a stall — `docs/matching-cookbook.md:6355` +- Templatability is PER-FAMILY: a sample straddling families reports their average and hides bimodality; probe one member then sweep or skip — `docs/matching-cookbook.md:6623`, `:6658` (§86) +- "Byte-neutral" is not "wanted": on a 0-bank group restore the snapshot unconditionally — `docs/matching-cookbook.md:6823` (§90b) +- One member's error names one blocker, not the blocker set — collect the classifier's line across the whole sweep — `docs/matching-cookbook.md:5979` (§75a) +- A diagnostic that asserts the WRONG cause redirects every later session (a mislabelled skip cost ~4 phases of a known mechanical win) — `docs/matching-cookbook.md:5579` (§68) +- Draft QA happens after the wave returns, never during; a stored MATCH is a claim with a timestamp — `docs/matching-cookbook.md:6849`, `:6691` (§90d/§87) +- An agent's CONCLUSION and its EVIDENCE fail independently — re-derive the premise, design the fix yourself — `docs/matching-cookbook.md:6851` (§90e) +- Prefer a bound the program itself declares (the owning function's `sltiu`) over a heuristic boundary — `docs/matching-cookbook.md:6858` (§90e) +- An unverified inherited premise is a hypothesis to test, not a foundation — flagging it costs nothing — `docs/matching-cookbook.md:6470-6472` +- Target byte-VARIANT families to move RE-completeness; high-reach h_exact families move only the display metric — `docs/matching-cookbook.md:6903` +- Count banks from the SOURCE (stub count), never from a tool's own tally — `phase-ends/DIGEST.md:226` (R42) +- 44% of the "near-miss" backlog were partial drafts misfiled with a closeness score; re-gate a sample before planning against any stored-draft pool — `docs/matching-cookbook.md:6695`, `:6697` (§87) +- 91% of the open backlog carried no class label, so the grinder searched undirected — `decomp-architect/corpus/decomp-kernels.md:219`, `docs/how-to-ai-decomp/03-bootstrap-order.md:87` +- `residual_class`'s "structural ⇒ permuter CPU is waste" is measurably false for schedule permutations; instruction-count equality is the tell — `docs/matching-cookbook.md:5525`, `:5549-5551` (§66d-5) +- Read the ILS per-cycle series, never the final best: a repeated score is converged, a still-falling one is budget-limited — `docs/matching-cookbook.md:5344` (§66d-3) +- Profile diversity is a ~0-token lottery ticket, not expected yield (3-for-3 on one giant, 0-for-1 on the next) — `docs/matching-cookbook.md:5525ff` (§66d-4) +- Check every search pass's diff for operand-order regressions — a search cannot see that one of its own edits is locally wrong — `docs/matching-cookbook.md:5501` +- `pgrep -f` self-matches the waiter's own command line; use a captured PID or a bracketed pattern — `docs/wave-playbook.md:721`, `docs/accelerators.md:196`, `phase-ends/DIGEST.md:295` (R79) +- A half-done declaration axis is a guaranteed break, not a smaller win; `0 sites remaining` is the completion assertion — `docs/matching-cookbook.md:6611` (§85) +- A guard defending a crash that was already fixed refuses real work (the §42e pin guard refused 63% of a pool, 0 crashes when overridden) — `docs/matching-cookbook.md:6623-6628` (§86) +- When asking "was this attempted?", glob every draft directory — a hand-listed pair lied in both directions — `docs/cookbook-index.md:1349` (§66c) +- Absolute include paths break a stranger's build and are invisible to every byte gate — `docs/how-to-ai-decomp/11-publishing.md:71`, `docs/matching-cookbook.md:6172` (§77) +- A headline number's denominator can carry an already-done population (main's 59,765 = game code + linked library) — `decomp-architect/corpus/decomp-kernels.md:392` (DK-30) +- A ranked "free" pool's top entries are precisely the known refusers; a byte-derived claim is a prediction, not a bank — `docs/matching-cookbook.md:9274` (§136h), `decomp-architect/corpus/decomp-kernels.md:282` (DK-20) +- Embed each callee's canonical signature in the drafter prompt (33% → 56% → 90% match_one, and the gap is declaration plumbing) — `docs/matching-cookbook.md:1470`, `:4406` (§58) +- Every extraction tool carries a narrow hard-coded preamble set; diff the exemplar's preamble against what the tool emitted, and carry the MINIMAL transitive closure — `docs/matching-cookbook.md:6118` (§77) +- Run the symbol-set guard before paying for a gate; `match_one`/`rtu_match` compile but never link, so an invented extern reads as MATCH — `docs/matching-cookbook.md:5439` (§67a), §87 +- Bare `harvest_verify` is the last rung, not the ladder — `docs/matching-cookbook.md:6529` (§84) +- MEASURE the write set; do not assert its tier — `docs/matching-cookbook.md:6766` (§89a) +- A ledger keyed inconsistently on address vs name yields two "best" records for one function — `docs/matching-cookbook.md:10301` (§3-B) +- Make the tool state its own denominator / a generated artefact must derive every scope figure it prints — `decomp-architect/corpus/decomp-kernels.md:263` (DK-19), `:392` (DK-30) diff --git a/.run/P33.5/log-mining/Phase29-3of4.md b/.run/P33.5/log-mining/Phase29-3of4.md new file mode 100644 index 0000000000..fce21d713d --- /dev/null +++ b/.run/P33.5/log-mining/Phase29-3of4.md @@ -0,0 +1,111 @@ +# Log mining — Phase29-3of4 +Files/ranges: `phase-ends/logs/Phase29.md`:5161-7740 · Lines read: 2580 of 2580 (plus 5131-5160 and 7741-7770 read for context only) +Candidates considered: 36 · NEW: 6 · ALREADY-BANKED: 30 + +Slice content: SESSION-21 T4b→T12, SESSION-22 T13→T30, SESSION-23 T31→T50, SESSION-24 T51→T54 +(the family-sweep / decl-axis / codegen-map-audit arc, 82% → 85.7% instr-weighted). + +Grep file set used for every "already banked?" test (abbreviated `$F` below): +``` +docs/matching-cookbook.md docs/cookbook-index.md docs/decision-log.md docs/accelerators.md +docs/retrospective.md docs/wave-playbook.md docs/how-to-ai-decomp/*.md +decomp-architect/corpus/decomp-kernels.md decomp-architect/templates/registry-E.decomp.md +phase-ends/DIGEST.md +``` + +--- + +## NEW + +### C1 — GNU C sources carry form-feed page separators, and `str.splitlines()` splits on them while `grep`/`sed` do not — so any Python line-number checker over compiler source silently drifts, and blames the agents +- **Evidence:** `phase-ends/logs/Phase29.md:6915-6920` — + > "GNU C sources use **form-feed (`\f`) page separators** — `loop.c` has 47, `cse.c` 36, `reload1.c` 27, `local-alloc.c` 21. **Python's `str.splitlines()` splits on `\f`; `grep`/`sed`/editors do not.** So every line number my checker computed after the first `\f` was shifted (by 20 in local-alloc.c, up to 47 in loop.c) — and 47 is **larger than the checker's own ±40 search window**, which is exactly how a real quote gets reported as FABRICATED." +- **What happened / what it cost:** The checker that validated two whole codegen-map audits (`tools/verify_map_findings.py`) reported "153 NEAR (agents' line arithmetic off by +2..+19)" and "12 FABRICATED" against the audit agents. After the one-line fix (`split("\n")`) the same data reads **180 OK / 0 NEAR / 0 FAB** and **299 OK / 0 NEAR / 0 FAB** — the agents' citations had been exact all along (`6923-6930`). The cost compounded: the false claim was **written into a later agent prompt**, telling agents to be careful about an error that was the instrument's (`6928-6930`), and the follow-up "diagnosis" of the 12 was itself wrong because it grepped the wrong JSON field (`6932-6937`). +- **Not banked — greps:** `grep -n -i -c 'form.feed' $F` → 0; `grep -n -i -c 'splitlines' $F` → 0; `grep -n -i -c 'page separator' $F` → 0; `grep -n -i -c '\\f' $F` → 0. (The adjacent lesson "a wall verdict is only as current as the instrument" is banked at `docs/matching-cookbook.md:10058`, but nothing records this defect or its class.) +- **Proposed home:** accelerator (an instrument-integrity entry) + a one-line note wherever compiler-source citation tooling is described (`docs/how-to-ai-decomp/07-compiler-source.md`). +- **Portable because:** every decomp project that reads a real compiler's source with a Python tool computes line numbers over text that contains `\f`; the failure is silent, it points the blame at the agents, and it survives into the next prompt. + +### C2 — Never restructure a proven tool with blind string replaces at the end of a long session; specify it as the next session's first task instead +- **Evidence:** `phase-ends/logs/Phase29.md:6134-6138` — + > "**⚠️ AND I BROKE `family_sweep` TWICE TRYING TO WIRE THE PARALLEL DEFAULT** (missed import, then a closure-scope error) — on the tool that banked 543 members today. **Reverted, not committed.** Restructuring a proven tool with blind string replaces at the end of a long session is how a working thing gets broken." +- **What happened / what it cost:** Two separate breakages of the single tool that had banked 543 of the day's members; both reverted, the work lost. The same session then twice more deferred structural work for the same stated reason — the `--allow-pins` default flip (`6178-6180`: "deliberately NOT done here, at the end of a long session, because that is exactly how `family_sweep` got broken twice today") and the `--span-tables` archaeology (`6352-6353`) — and both landed cleanly when done fresh. It recurs at `7713-7714`: a deleted `last_err` initializer during a refactor, caught only by re-reading the diff before running. +- **Not banked — greps:** `grep -n -i -c 'end of a long session' $F` → 1 (`docs/decision-log.md:1259`, and there it is an *excuse* for a stale measurement, not a rule about tool surgery); `grep -n -i -c 'structural.*refactor.*session' $F` → 0; `grep -n -i -c 'tired context' $F` → 0; `grep -n -i -c 'late in.*session' $F` → 1 (`docs/accelerators.md:423`, about building a tool late in the *project*, not in a session). +- **Proposed home:** G (a rule) or a kernel — a session-hygiene law beside "commit banked work immediately". +- **Portable because:** every agent-driven project accumulates one or two tools the whole pipeline depends on, and the temptation to "just wire it up" arrives exactly when context is exhausted and the diff is no longer being read. + +### C3 — An oracle that can always be RUN is not always APPLICABLE: state the applicability precondition beside the recipe, or a coarse run returns a large number that reads as a verdict +- **Evidence:** `phase-ends/logs/Phase29.md:7359-7364` — + > "`regalloc.md` §H presents the swap oracle as the way to 'discriminate RC-6 (allocation) from S3 (scheduling) in ONE gdb run'. It only works when **the contested registers are held by PSEUDOS**. Check the `.greg` RTL first: if the diff's registers appear as `(reg/v:SI N )` with N < 68, they are hard already and the oracle cannot move them — a coarse swap will return a large, meaningless number (345 here) that looks like a verdict and is not one." +- **What happened / what it cost:** The `reg_renumber`-swap gdb oracle was built, negative-controlled (a no-op `31↔31` swap reproduced the baseline 13 mismatches exactly, `7328-7329`) and run on the project's single largest prize (51,198 templated instructions). Both contested swaps returned 345 and 97 — numbers that look like allocation verdicts. They were artifacts: `reg_renumber` maps only pseudos (≥ `FIRST_PSEUDO_REGISTER` = 68) and the contested value was already in a hard register at `.greg` time (`7348-7351`). The oracle's framing of the residual was refuted, not confirmed, and the map's §H had to gain a precondition it never stated. +- **Not banked — greps:** `grep -n -i -c 'unstated precondition' $F` → 0; `grep -n -i -c 'looks like a verdict' $F` → 0; `grep -n -i -c 'swap oracle' $F` → 1 (`docs/matching-cookbook.md:3326` — cites the oracle's *result* on another function, states no precondition); `grep -n -i -c 'applicability' $F` → 0. (The negative-control half IS banked, 41 hits for "negative control"; the applicability half is not.) +- **Proposed home:** DK (a kernel) — an oracle-design law next to "negative-control the instrument first". +- **Portable because:** any instrument that reports a scalar will report one when it is structurally inapplicable; the number is then indistinguishable from a measurement, and it is the shape that gets a function written off. + +### C4 — Before a fleet-wide mechanical edit, census the whole population for the exact preconditions the edit assumes; perfect uniformity is the licence to apply it, and non-uniformity is the design input +- **Evidence:** `phase-ends/logs/Phase29.md:7520-7531` — + > "### MEASURED BEFORE BUILDING (R35) — Ran the blocker census over all 132 still-stubbed siblings before writing a line. It is **perfectly uniform**, which is the strongest possible signal that one mechanical edit fixes all of them: … | with the file-scope blocker | **132 of 132**, 0 without | … | file-scope decl statements per (TU, sym) | **exactly 1** — never ambiguous | | file-scope references BELOW the decl | **0** — so the deletion is always safe |" +- **What happened / what it cost:** Every column of that census is a precondition the tool would otherwise have had to guess or guard: "exactly 1 decl" removes an ambiguity branch, "0 references below" is what makes the deletion safe. The tool then applied to 132 files with a diff **uniform to the line (+11/−3)** and banked 132/132 on the follow-up sweep (`7562`, `7599`). The counter-case is in the same slice: `func_80135260`'s family was assumed to share `func_80177DA8`'s blocker without a census and the inference was **WRONG** — "same SC07-only signature, different cause" (`7421-7423`). +- **Not banked — greps:** `grep -n -i -c 'census.*before building' $F` → 0; `grep -n -i -c 'perfectly uniform' $F` → 0; `grep -n -i -c 'one mechanical edit' $F` → 0; `grep -n -i -c 'population.*before' $F` → 4 (`registry-E.decomp.md:172` is regression control — run a new refusal over the population that already SUCCEEDED; `03-bootstrap-order.md:121` is lane staffing; neither is a pre-edit precondition census). `docs/matching-cookbook.md:7341` (§103) banks this lever's *two-step byte verification* and its refusal conditions but not the census. R37/G23 ("probe before costing") grounds an *estimate* on one instance — the opposite direction from censusing all of them. +- **Proposed home:** G (a rule) — beside the two-step verify in §103, or a kernel on fleet-wide edits. +- **Portable because:** any decomp reaches a point where one declaration-shaped edit must be applied across hundreds of files; the census is cheap, it converts every guard the tool would need into a proven fact, and its absence is what turns a mechanical edit into a half-axis. + +### C5 — A status line a script prints unconditionally is not a measurement; derive every conclusion the script emits from the command's own output +- **Evidence:** `phase-ends/logs/Phase29.md:7496` — + > "**Three unconditional `echo` conclusions** (`[shared clean]`, `[none = ...]`) that asserted things the command output contradicted." +- **What happened / what it cost:** Listed by the session itself under "MY ERRORS THIS SESSION (recorded, not buried)", alongside the form-feed bug and two truncated-output reads, under one stated common thread: "**inference from partial output instead of measuring.** Every one was caught by measuring; none by re-reading" (`7499-7500`). No instance is costed individually, so this is the weakest of the six — but the shape is exact and unbanked: a hand-written driver script that prints `[shared clean]` after a command, rather than from it, manufactures a clean verdict on a dirty tree. +- **Not banked — greps:** `grep -n -i -c 'unconditional echo' $F` → 0; `grep -n -i -c 'echo \[' $F` → 0; `grep -n -i -c 'asserted.*contradicted' $F` → 0; `grep -n -i -c 'always prints' $F` → 1 (`docs/wave-playbook.md:80` — a NOTE line that deliberately always prints, the opposite point). The nearest banked rules are "count banks from the SOURCE" (`phase-ends/DIGEST.md:226`) and R66 "write 'banked' only from the tool's printed success line" (`decomp-kernels.md:878`) — both about trusting a *tool's* output, not about a driver script fabricating one. +- **Proposed home:** accelerator (an instrument-integrity entry), or fold as a clause into the existing R66 line. +- **Portable because:** ad-hoc driver scripts are written in every session of an agent-run project, and an unconditional conclusion line is the cheapest possible way to make a red run read green. + +### C6 — Measure what fraction of a cycle a parallelism knob can actually touch before adopting it: `make -j16` bought 12%, because the build was 5 s of a 16 s per-item cycle and the real cost was a four-stage retry ladder +- **Evidence:** `phase-ends/logs/Phase29.md:6115-6124` — + > "**The `-j` theory was WRONG, and measuring said so** (baseline ~18 s/sibling): | `make extract` + `make build`, cold | **~5 s of the 16 s** | | one sibling end-to-end, serial | 16 s | | one sibling end-to-end, `-j16` | **14 s (12%)** | … Make is not the bottleneck. The per-sibling loop tries **up to FOUR stages** (raw → scoped → recovered → reconciled) and **each runs its own `make build`** … `-j16` kept (free, safe, committed) — but it is a 12% win, not 8×." +- **What happened / what it cost:** The session opened this task believing `-j` was the missing 8-16× lever (Drew asked why it wasn't being used). Measuring located the cost in the per-item *ladder*, not the compiler, and located the real 8-16× lever in cross-sibling parallelism — which was then found to be blocked by a specific hazard (`revert()` restores `config/` from git and `config/overlays.mk` is SHARED, so a concurrent revert clobbers peers' carve entries, `6126-6132`). Three sweeps totalling 543 members had already run serially in that session (`6109-6113`). +- **Not banked — greps:** `grep -n -i -c 'not the bottleneck' $F` → 0; `grep -n -i -c 'where the time' $F` → 0; `grep -n -i -c 'each runs its own' $F` → 0; `grep -n -i -c '8-16×' $F` → 1 (`docs/matching-cookbook.md:7287`, §101's stale-default table — it banks *that the parallel farm existed and was never wired*, not the measurement that the obvious knob was worth 12% and why). `docs/accelerators.md:112` (A8) banks processes-vs-threads, longest-first and per-item search, but not "profile the cycle before turning the knob". +- **Proposed home:** accelerator — an addendum to A8. +- **Portable because:** every fleet-scale decomp hits the same instinct (add `-j`), and the same true answer (the per-item retry ladder, not the compiler, dominates); the measurement takes minutes and redirects the whole optimisation. + +--- + +## ALREADY-BANKED (one line each) + +- Re-derive an agent's premise, not just its fix — its conclusion and its evidence fail independently (T4b) — lives at `docs/matching-cookbook.md:6800` (§90e). +- The program declares its own table length (`sltiu N`): prefer an exact self-declared bound to a heuristic, and a warning that fires on ambiguity is noise — `docs/matching-cookbook.md:6800` (§90e). +- The single-table-predecessor inference (a single-table carve's span start IS its table start) — `docs/matching-cookbook.md:8616` (§8e-2, pre-existing). +- Templatability is a per-FAMILY property, bimodal, never a blended rate over a pool — `docs/how-to-ai-decomp/09-economics.md:120` and cookbook §86 (`docs/matching-cookbook.md:6623`). +- Pick targets by the metric you mean to move: byte-VARIANT families move RE-completeness, h_exact families move only the display number — `docs/matching-cookbook.md:7715` (§111). +- Neither "always ladder" nor "never ladder": a verdict from a ladder is a verdict from the ladder; re-run a stubborn reject through the bare gate — `docs/matching-cookbook.md:34698`. +- The bare gate banks better and cleans up worse (no snapshot/restore) — always diff the tree after one — `docs/matching-cookbook.md:8125` (§122 law 2). +- Isolate the pipeline stage; `set -o pipefail` attributes the failure to the LAST stage in the pipe — `docs/matching-cookbook.md:6956` (§93). +- Splice once and dump EVERY cc1 error rather than peeling one conflict per gate cycle — `docs/matching-cookbook.md:7049`. +- gcc-2.7.2 prints hard errors with no `error:` prefix, so a naive grep finds nothing — `docs/matching-cookbook.md:11816` and `:30932`. +- A refused pre-step whose return code is ignored manufactures codegen-flavoured verdicts, and one refusal counted N times inflates the wall count — `docs/matching-cookbook.md:7100` (§97). +- A tool that swallows the underlying error returns a bare fail that reads as a wall — `docs/retrospective.md:47` and `docs/matching-cookbook.md:9730`. +- Diff a rewriting tool's OUTPUT against its INPUT before trusting it — the gate would only have said PLUMBING — `docs/matching-cookbook.md:7234`. +- A completion assertion must be a DELTA; an absolute "at least one exists" check passes vacuously — `docs/matching-cookbook.md:7388` (§103). +- An assertion must be exact about its DOMAIN, not just its condition (it cried wolf on its own by-design skip) — `docs/matching-cookbook.md:7184`. +- PLAN → VALIDATE → WRITE: a refusal path that aborts mid-write creates the half-axis it exists to prevent — `docs/matching-cookbook.md:7187`. +- A PLUMBING verdict is about the DECLARATIONS, never the BODY; a PLUMBING pool is an upper bound on recoverable work, not a count of it — `docs/matching-cookbook.md:7306` (§102). +- A negative that MOVES the failure class is a result, not a null — `docs/matching-cookbook.md:6040`. +- The STALE DEFAULT class: a guard whose cause was removed is a silent skip wearing a safety label (137 skipped, 133 bank) — `docs/matching-cookbook.md:7279` (§101). +- "The fix exists, it just isn't reachable by default" — a lever wired into ONE gate path is a lever most families cannot reach — `docs/matching-cookbook.md:7547` (§107) and §101's table row at `:7287`. +- Before editing anything shared, ask what the smallest scope is that still travels with the body (K&R def / block-scope typedef beat a 524-site conform) — `docs/matching-cookbook.md:7250` (§100) and `:7201` (§99), law at `:7276`. +- A TU that DEFINES a function owns its own declarations; a fleet-wide axis is meaningful only for CONSUMING TUs — `docs/matching-cookbook.md:7147` (§98). +- Persist the MEASUREMENT, derive the POLICY; and prove the re-derivation faithful (1610/1610, table unchanged) BEFORE editing the table — `docs/matching-cookbook.md:7504` (§106). +- A classifier's route that contradicts the knowledge base wastes search CPU (ADDRESSING → permuter, three targets under-delivered) — `docs/matching-cookbook.md:7530` (§106). +- A revert must run on EVERY exit — success, gate-fail, refusal AND the throw; untracked residue survives `git checkout --` — `docs/matching-cookbook.md:7465` (§105). +- Match on masked text, emit by span from the original; a warning that fires 137/137 and is right 0 times masks real causes — `docs/matching-cookbook.md:7429` (§104). +- The codegen map cited the WRONG compiler (2.8.1 vs 2.7.2), drift is not uniform, and each refutation was challenged by an independent agent (119/40/7 of 21) — `docs/how-to-ai-decomp/07-compiler-source.md:44-48`. +- An incremental build pass does not refute a clean-tree failure — `docs/matching-cookbook.md:7197`. +- Truncated output is not exhaustive output (`head -8` of 31 hid the MATCH); and re-measure every stored draft, never a sample — `docs/matching-cookbook.md:10051` and `:10058` (§146). +- A shared scratch directory is a shared blast radius — one subdirectory per agent — `docs/accelerators.md:732`, `docs/matching-cookbook.md:37026`. +- A scoped `git add` is an unverified assertion about a change set's boundary — use `git add -A src/ config/` for a carve bank — `docs/matching-cookbook.md:4523`. +- Two residuals that move in opposite directions under every lever are two symptoms of one starved resource — solve them together — `docs/matching-cookbook.md:34012`. +- A function's past-attempt journal travels with it and prices the next attempt — banked as `past-attempt` history in `docs/wave-playbook.md`, `docs/how-to-ai-decomp/05-cards-lanes-waves.md`, `decomp-architect/corpus/decomp-kernels.md`. +- Never write an unverified diagnosis into an agent's brief as fact — it inherits your wrong search space — `docs/matching-cookbook.md:6979` (§93). +- A draft header's own assertion is a claim, not a fact (R14: verify summaries against the bytes) — `phase-ends/DIGEST.md:174`. +- A verdict says what it was proved against (the honest-coverage-gap habit) — `phase-ends/DIGEST.md` and `docs/how-to-ai-decomp/07-compiler-source.md` (R69/R65). +- Count banks from the SOURCE, never from the report (a diagnostic that prints only failures hides its own successes) — `phase-ends/DIGEST.md:226`, `decomp-architect/templates/registry-E.decomp.md:240`. +- Regenerate the family map after every bank before sweeping — memory `crack-wave-sweep-map-regen`, cited as "the documented regen step" in the log itself at `5240-5244`. +- The -O0 cluster is OVERLAY-LOCAL, so discount an `_o0` exemplar's headline family reach — `docs/matching-cookbook.md:628`, `:1551`, `docs/decision-log.md:1251-1305`. +- A counterfactual byte-gated on a reproduced blocker is the evidence a stage works — `docs/matching-cookbook.md:7405` (§103, AUTOMATED). diff --git a/.run/P33.5/log-mining/Phase29-4of4.md b/.run/P33.5/log-mining/Phase29-4of4.md new file mode 100644 index 0000000000..3b3e276ce2 --- /dev/null +++ b/.run/P33.5/log-mining/Phase29-4of4.md @@ -0,0 +1,88 @@ +# Log mining — Phase29-4of4 +Files/ranges: phase-ends/logs/Phase29.md:7741-10316 (read from 7711 for context; reported range 7741-10316) · Lines read: 2576 of 2576 +Candidates considered: 36 · NEW: 6 · ALREADY-BANKED: 30 + +Slice content: SESSION-24 T54–T78 and SESSION-25 T79–T99 (the family-sweep engine's whole +diagnose-and-unblock arc, four superseded checkpoints, and the Phase-29 close-out burn-down). + +## NEW + +### C1 — A clean-looking verdict that appears immediately AFTER your own repair transform is a SUSPECT, not a result: re-measure the artefact the transform produced before routing the residual +- **Evidence:** `phase-ends/logs/Phase29.md:8207-8215` — + > "`match_one` says **`SIZE-MISMATCH`: draft 58 ins vs target 78** … **`--fix-def-sig` demoted the return `s32` → `void`, and gcc deleted the computation feeding it as dead.** … **A "clean DIFF" that appears right after a transform is a suspect, not a result.**" +- **What happened / what it cost:** T60's `--fix-def-sig` fix moved `func_8016163C` from `param_1 undeclared` to a clean compile with differing bytes, and T60 reported it as "genuine codegen" — the routing verdict that sends a function to the permuter. T61/T62 found the transform had demoted the definition's return type to `void`, so gcc deleted a fifth of the function as dead code; the draft was byte-correct and the *tool* had manufactured the DIFF. The §85 guard added for this only asked whether *callers* consumed the return, never whether the body had `return ;`. The wrong verdict also survived into the knowledge base: `docs/matching-cookbook.md:7657` (§109's table) still records `func_8016163C` as "plumbing fully cleared; genuine codegen". +- **Not banked — greps:** `grep -n -i 'suspect, not a result' ` → 0; `grep -n -i 'after a transform' …` → 0; `grep -n -i 'tool-manufactured' …` → 0; `grep -n -i 'demoting the return' …` → 0; `grep -n -i '8016163C' …` → 1 (`docs/matching-cookbook.md:7657` — and that hit records the *superseded, wrong* conclusion, not the lesson). R40 ("exonerate the instrument") is banked but covers a failing instrument, not a repair pass that silently rewrites the subject into a plausible-looking failure. +- **Proposed home:** DK (a kernel) + a correction to cookbook §109's table +- **Portable because:** every decomp harness grows repair passes that rewrite a draft before the oracle sees it; any of them can convert a correct body into a credible codegen residual, and a length/size verdict is the cheapest tell that it did. + +### C2 — Classify a harness fix as a LOGIC defect or a PATH-REACHABILITY gap and price it accordingly: only the logic defect generalises +- **Evidence:** `phase-ends/logs/Phase29.md:9957-9965`, `9986-9994` — + > "| **§117** symbol-kind | **1,209** | wrong LOGIC — applied fleet-wide | … **Only the logic defect generalised.** The three path-reachability gaps were each worth ~one family. Useful prior for pricing the next fix *before* building it." +- **What happened / what it cost:** Four engine fixes in two sessions looked identical while being built — "an engine defect that unblocks a whole class". Measured after: the one wrong-logic defect (a positional map spelling the sibling's symbol from the exemplar's kind) paid 1,209 members and took 138 families from zero to complete; the three "the lever exists but this path cannot reach it" fixes paid 158, 137 and 136 — each ~one family. The log states the cost of not having the prior explicitly ("Both looked identical before measuring", 9511-9512) and the fix ("Measure the blast radius; do not infer it from the fix's depth"). +- **Not banked — greps:** `grep -n -i 'blast-radius prior' …` → 0; `grep -n -i 'price the next fix' …` → 0; `grep -n -i 'logic defect' …` → 1 and `grep -n -i 'path-reachability' …` → 1 — both the same line, `docs/matching-cookbook.md:8093`, where §120 labels *itself* "targeted … a path-reachability gap, not a logic defect". The label exists per-entry; the PRIOR (the two classes, their measured magnitudes, and using them to price a fix before building it) is nowhere. `grep -n -i 'blast' decomp-architect/corpus/decomp-kernels.md` → 1 (a provenance line only); 0 in `registry-E.decomp.md` and `DIGEST.md`. +- **Proposed home:** DK (a kernel), paired with the existing "verify blast radius, not just the defect" +- **Portable because:** the decision "do I build this fix?" recurs on every harness, and the two classes are distinguishable *before* building — does the code compute something wrong everywhere it runs, or is it correct and simply not on this path? + +### C3 — A wrong prescription left in the knowledge base is worse than no entry: when evidence refutes an entry you wrote, correct THAT entry in place, in the same session, carrying the refutation +- **Evidence:** `phase-ends/logs/Phase29.md:9029-9031` — + > "Cookbook **§116 corrected in place** — it now carries the refutation and the corollary, because a wrong prescription left in the cookbook is worse than no entry: the next session would have spent the same hour." +- **What happened / what it cost:** T79 wrote cookbook §116 with a fix ("move the stub line; byte-neutral by construction"), T80 built the tool, applied it to 133 overlays, and the build refuted it in 56 seconds. The same failure recurred at T92, which recorded a wrong recipe ("strip-if-ambient") **and a phantom second blocker** that were both artefacts of its own bad fix; T93 found both wrong (`9886-9892`) and the wrong recipe had by then been carried as the lead item in two handoff checkpoints (`9832-9844`). Both were corrected in place with named commits (`9a1507462`), and the log states the counterfactual cost: the next session repeats the hour. +- **Not banked — greps:** `grep -n -i 'wrong prescription' …` → 0; `grep -n -i 'worse than no entry' …` → 0; `grep -n -i 'carries the refutation' …` → 0; `grep -n -i 'corrected in place' …` → 3 (all unrelated: a P33.5 memory sweep, a cookbook preamble, one banked-C header comment). `docs/how-to-ai-decomp/06-knowledge-base.md` covers capturing entries while fresh (R30) and mentions "wall proofs and their later refutations" as content, but carries no rule about correcting a refuted entry. +- **Proposed home:** G (a rule) on the knowledge-base chapter, next to R30 +- **Portable because:** any compounding knowledge base is read by future sessions as prescription; an uncorrected wrong entry is a durable negative-value asset, and appending a new entry beside it does not stop the next reader following the old one. + +### C4 — A name grep is not a "defined / banked here" oracle: a DECLARATION carrying the name reads as a definition; use the structured stub oracle, and believe the pipeline's map over your own check +- **Evidence:** `phase-ends/logs/Phase29.md:8080-8087` — + > "My T58 batch-selection test picked the first TU *containing the name* — a declaration — and, seeing no stub in that file, called it banked. **The family map was right all along** … The oracle to use is `corpus.stubs(ov)`, never a name-grep. Their claimed weight (60,720 + 48,576 bytes) was never real fuel." +- **What happened / what it cost:** T58 reported "7 remaining families all have banked exemplars"; two of them (276 + 138 members) had no matched exemplar anywhere. The sweep tool had already excluded them and printed the true count ("6 matched-exemplar families") — a line the log records reading past twice, together with the sibling error at T57 where `--band`'s `substantial` default silently dropped 3 of 5 chosen targets while the tool printed "2 matched-exemplar families" (`7914-7917`). 109,296 bytes of "weight" were ranked as fuel that never existed. +- **Not banked — greps:** `grep -n -i 'name-grep' …` → 0; `grep -n -i 'a declaration, not a definition' …` → 0; `grep -n -i 'the map was right' …` → 0; `grep -n -i "read as .already matched" …` → 1 (`docs/matching-cookbook.md:7904`, §115's aside about `stub_map` and a *curated name form* — a different mechanism); `grep -n -i 'selection line' …` → 0. R41 (assert the denominator) and R43 (a tool must refuse unsupported input) are banked but both address the tool lying or narrowing; here the tool was correct and printed so. +- **Proposed home:** DK (a kernel) or an accelerator +- **Portable because:** every decomp project has both declarations and definitions of the same symbol in the same tree, and the cheap "is it done here?" check is always a grep; the structured oracle exists in every such pipeline and is always the right one. + +### C5 — A yield estimator that counts "unclaimed at the moment it runs" over-projects: it ranks correctly and overstates absolutely; never plan off its absolute numbers +- **Evidence:** `phase-ends/logs/Phase29.md:9350-9354` — + > "I priced these two families at 130 + 129 = **259** new distinct; the measured gain is **125**. The estimator counts a family's h_exact classes that have no matched instance *at the time it runs*, so classes another sweep claims in between are double-counted. **Do not plan off those projections at face value** — they rank families correctly (relative order held) but overstate absolute yield." +- **What happened / what it cost:** The same estimator produced the headline that drove three sessions of target selection (2,962 distinct across 36 families), and a separate correction the same session (`8968-8977`) cut it again by ~40% for a different reason. The measured over-projection factor is ~2×. The rank order held every time, which is exactly why the defect was invisible: the tool kept picking the right next family while pricing it twice too high. +- **Not banked — greps:** `grep -n -i 'estimator' …` → 0; `grep -n -i 'double-count' …` → 1 (`docs/decision-log.md:669`, a stub over-count in a classifier); `grep -n -i 'over-project|over-estimat|ranks correctly' …` → 1 (`docs/decision-log.md:1447`, over-projection from a *staged count*, §57a — a different mechanism); `grep -n -i 'projected yield' …` → 0. +- **Proposed home:** DK (a kernel) or an accelerator on planning/economics +- **Portable because:** every campaign builds a "what is this population worth" estimator, and the natural implementation snapshots the unclaimed set — which decays under the campaign's own concurrent progress. The safe contract (rank with it, never budget with it) is compiler- and console-independent. + +### C6 — Declare a mechanical lever SPENT only on a positive, three-part measurement — every built lever applied and returning zero, the residue split by structure, and the decay curve priced against what the same session banked +- **Evidence:** `phase-ends/logs/Phase29.md:10300-10303` — + > "The final four days decayed **+2.7 → +2.2 → +0.6 → +0.3pp/day with every lever this phase built applied** — and T98 characterised the residue as **80 families / 960 members / 173 distinct** (29 all-STRUCT refused by design + 51 gate-failing at ~3 distinct per independent diagnosis), against 2,713 members banked in the final session. **The mechanical family engine is spent — measured, not felt.**" +- **What happened / what it cost:** SESSION-24 wrote "DO NOT close P29 on ROI" three times over three checkpoints without a criterion for when it *would* be right (`8070`, `8339`, `8554`), so the question re-opened every session. T98/T99 supplied the criterion by construction: probe the top residue family (it was refused for a structural reason, the first still-zero family this phase whose blocker was not the tooling), split the residue cleanly (29 refused-by-design vs 51 needing independent diagnoses), and read the per-day delta from the committed metric history. This project had produced at least four *phantom* exhaustion proofs before (`docs/decision-log.md:1139`), which is precisely why the positive form matters. +- **Not banked — greps:** `grep -n -i 'measured, not felt' …` → 0; `grep -n -i 'engine is spent' …` → 0; `grep -n -i 'when to stop' …` → 1 (`docs/matching-cookbook.md:5475`, about stopping one permuter run); `grep -n -i 'gate-2|gate 2 ' …` → 0. The banked material is all the *negative* form — many entries recording exhaustion claims that were false (retrospective:32, failure-museum row 15, decision-log:804/1139) — and one open instruction with no criterion (`docs/decision-log.md:1680`, "the burn-down floor is still undetermined (needs 3 session-close deltas)"). +- **Proposed home:** DK (a kernel) on economics / phase closure +- **Portable because:** every decomp has a mechanical/templating phase that eventually stops paying, and the failure mode in both directions (grinding a spent lever; abandoning a live one on a phantom proof) is universal. The three parts are all instrument-independent. + +## ALREADY-BANKED (one line each) +- A family-wide `0/N` is a statement about the HARNESS, not the code — lives at `docs/matching-cookbook.md:7627-7630` (§108) +- Diagnose a `0/N` by splicing ONE member and `make -j1` the single object, filtering `warning:` (the `-j16` interleave hides the real cc1 line) — `docs/matching-cookbook.md:7591` +- One serial compile beats an N-way fleet bisect for finding which change broke the gate — `docs/matching-cookbook.md:7591` + `docs/matching-cookbook.md:17052` +- A lever wired into ONE gate path is a lever most families cannot reach — `docs/matching-cookbook.md:7547` (§107) +- Conform a definition to the canonical TYPES but keep the BODY's parameter names; parse with `cdecl`, not a regex — `docs/matching-cookbook.md:7647` (§109) +- A verdict that CHANGES is the signal to re-route; verify a fix by the verdict moving, never by assertion — `docs/matching-cookbook.md:7668` (§109) +- The §85 return-axis precondition: only demote a return type when no caller consumes it — `docs/matching-cookbook.md:7661-7666` +- A unit must define exactly ONE function; "ends in `;`" does not identify the defining line — `docs/matching-cookbook.md:7674` (§110) +- The distinct-code metric is predictable pre-sweep: a byte-identical family is ONE piece of distinct code, a byte-variant family ~N — `docs/matching-cookbook.md:7715` (§111) +- A macro-scoped declaration collides only where the macro is INSTANTIATED — measure the intersection, not the population — `docs/matching-cookbook.md:7764` (§112) +- An ARITY blocker exists only if the macro CALLS the function; an address-taken use imposes none — `docs/matching-cookbook.md:7810` (§113) +- `conflicting types for X` — read X; when X is not the function being templated, the def-signature and header levers are both wrong — `docs/matching-cookbook.md:7877-7880` (§114) +- A partial fix to a name-form assumption produces the exact symptom of no fix, so a correct hypothesis reads as refuted — grep every place the assumption is encoded — `docs/matching-cookbook.md:7906-7909` (§115) +- "Byte-neutral by construction" is a claim about the LINKER; other build stages partition differently — build it before calling it neutral — `docs/matching-cookbook.md:7956-7959` (§116) +- Optimization level is a property of the FILE, not the function; read a family `0/N` against the member's stub home — `docs/matching-cookbook.md:7911` (§116) +- When a MASKED oracle says MATCH and the whole-binary gate says DIFF, suspect a SYMBOL before codegen — `docs/matching-cookbook.md:7993` (§117) +- An immediate-ambiguity check must exclude compiler-synthesised uses from its denominator — `docs/matching-cookbook.md:7997` (§118) +- Two levers on the same axis in opposite directions: "with the flags" and "without" cover HALF the matrix — test the off-diagonal — `docs/matching-cookbook.md:8031` (§119) +- A repair flag that assumes the header is authoritative is a REPAIR, not a default (it breaks drafts whose def is byte-true) — `docs/decision-log.md:1358` +- Before concluding a lever does not work, prove it RAN — diff the staged artefact for the change it should make — `docs/matching-cookbook.md:8088-8090` (§120) +- A guess that turns one error into a DIFFERENT error proves the diagnosis right and the guess wrong — `docs/matching-cookbook.md:8103-8106` (§121) +- Re-probe the exclude/failure list after every tool fix, as part of the fix; a shared-path fix re-opens every population previously booked as failed — `docs/decision-log.md:2785`, `decomp-architect/templates/registry-E.decomp.md:245` +- Probe before COSTING: ground every estimate and attribution on ONE instance; derive counts from `corpus.stubs`, never from what you were looking at — `decomp-architect/templates/registry-E.decomp.md:160` (G23), `phase-ends/DIGEST.md:216` (R37) +- A new precondition/filter must be run against known-good AND known-bad controls before adoption — `decomp-architect/corpus/decomp-kernels.md:762` (DK-61), `docs/retrospective.md:125` +- Family outcomes are all-or-nothing: probe ONE member, then sweep or skip the whole family; a sample straddling families reports their average and hides the bimodality — `docs/matching-cookbook.md:6634` (§3-THE) +- Revert an edit that bought zero banks rather than leave unverified shared-state risk standing — `docs/matching-cookbook.md:32998` (§401 corollary) +- A `CC1 FAIL` verdict whose stdout omits the error costs repeated probes; persist stderr and filter the warnings — `docs/decision-log.md:2000`, `docs/matching-cookbook.md:14045` (§165-13) +- The burn-down/velocity series is DERIVED from the committed metric digest's git history, not built — `phase-ends/DIGEST.md:209` (R33; the log cites R33 itself at `10260`) +- The unswept remainder is enriched in walls — a rate measured on the already-unblocked population does not transfer — `docs/matching-cookbook.md:31109-31113` (§332 / R45 draw policy) +- The economics change when each `0/N` has its own distinct cause (a batch becomes a diagnosis queue) — the actionable core is banked as the per-family probe-then-sweep procedure, `docs/matching-cookbook.md:6650-6659` (§3-THE) diff --git a/.run/P33.5/log-mining/Phase30-1of2.md b/.run/P33.5/log-mining/Phase30-1of2.md new file mode 100644 index 0000000000..10a4e46794 --- /dev/null +++ b/.run/P33.5/log-mining/Phase30-1of2.md @@ -0,0 +1,138 @@ +# Log mining — Phase30-1of2 +Files/ranges: `phase-ends/logs/Phase30.md`:1–2684 (read to 2721 for context; only 1–2684 reported) · Lines read: 2684 of 2684 (assigned range), 2721 of 5368 (file) +Candidates considered: 50 · NEW: 13 · ALREADY-BANKED: 37 + +> Grep corpus for every "already banked?" test below (abbreviated `$FILES`): +> `docs/matching-cookbook.md docs/cookbook-index.md docs/decision-log.md docs/accelerators.md docs/retrospective.md docs/wave-playbook.md docs/how-to-ai-decomp/*.md decomp-architect/corpus/decomp-kernels.md decomp-architect/templates/registry-E.decomp.md phase-ends/DIGEST.md` + +## NEW + +### C1 — An instrument's refusal is a FINDING, not an obstacle: overriding it means explaining why the instrument is wrong, never finding another route +- **Evidence:** `phase-ends/logs/Phase30.md:1369-1371` — "`tools/validate_targets.py` … It fails closed; **an instrument's refusal is a finding, not an obstacle** — S46 routed around it and lost 87 of 119 agents." (also `:1486` "**`wave_snapshot` REFUSED that list** (24 of 57 found) **and I routed around it.** The only instrument warning this session that was correct and overridden"; `:1489-1490`) +- **What happened / what it cost:** `wave_snapshot` refused a 400+-target list, finding only 24 of 57. The refusal was overridden by building the list another way; ~29 of 47 targets were invalid (crossed function↔binary pairings, mid-body addresses, already-banked), **87 of 119 agents wasted, ~9.7M tokens**. It was recorded as the only correct instrument warning of the session and the only one overridden. +- **Not banked — greps:** `grep -n -i "refusal is a finding" $FILES` → 0; `grep -n -i -e "route around" -e "override the refusal" $FILES` → 0; `grep -n -i "routing around" $FILES` → 1 (cookbook:31802, about rerouting around gcc's local-alloc — unrelated); `grep -n -i "an instrument.s refusal" $FILES` → 0 +- **Proposed home:** DK (a kernel) — sibling to "a verdict names its instrument" in `how-to-ai-decomp/04-oracles-and-instruments.md:79` +- **Portable because:** every decomp harness grows fail-closed validators; the temptation to route around one is universal, and the cost is measured here in wasted agent budget. + +### C2 — Prose in an agent prompt is not enforcement: snapshot `git status` around every agent and name the offender +- **Evidence:** `phase-ends/logs/Phase30.md:449-452` — "**3. An agent wrote a TRACKED file** (`src/shared/engine_types.h`, added a typedef) despite the prompt forbidding it twice. … **Prose is not enforcement**: the wave harness should snapshot `git status` before/after each agent and name the offender (TODO)." +- **What happened / what it cost:** during wave 7a an agent edited a tracked, fleet-shared header even though the prompt forbade tree writes twice. `gate_lane`'s entry guard refused to gate on a tree it did not own — caught before any commit, but it cost one gate cycle, and nothing in the harness could say WHICH agent did it. +- **Not banked — greps:** `grep -n -i "prose is not enforcement" $FILES` → 0; `grep -n -i -e "snapshot git status" -e "git status before" $FILES` → 0; `grep -n -i "wrote a tracked" $FILES` → 0. (`decomp-kernels.md:459` DK-35 and `wave-playbook.md:786-793` bank the per-agent WORK DIRECTORY contract and transcript recovery — they do not bank the mechanical check that the contract was obeyed, nor the tracked-tree class.) +- **Proposed home:** G (a rule) or an extension to DK-35 +- **Portable because:** any multi-agent harness that tells agents "do not write outside your directory" needs a cheap mechanical detector; a VCS status diff is one, and it names the culprit. + +### C3 — Read the FIRST ten results of a long run before trusting the other 190; and negative-control any NEW REFUSAL against everything that already succeeded +- **Evidence:** `phase-ends/logs/Phase30.md:250-259` — "**Three runs launched at scale, three stopped early — and every stop was right.** … **Read the first ten results of any long run before trusting the other 190** — and when a check is added to refuse work, NEGATIVE-CONTROL it against everything that already succeeded. That control found two bugs in my own pre-checks (C89 `f()` is UNSPECIFIED parameters, not zero; a definition read as a call to itself), either of which would have silently discarded good drafts." +- **What happened / what it cost:** three scale runs each carried a defect (a selector drafting members a rename cannot reach; a seed name recovered by grepping the first `func_XXXX(` — which in a de-macroized block is a *callee declaration*; a 1,891-line macro decl layer pasted into every destination). The verdict files the tools already wrote named each defect inside the first handful of groups. +- **Not banked — greps:** `grep -n -i "first ten results" $FILES` → 0; `grep -n -i "negative-control.*already succeeded" $FILES` → 0; `grep -n -i -e "refuse work" -e "new refusal" $FILES` → 0 / 8 (all naming specific new refusals, none stating the control rule). `registry-E.decomp.md:106` requires a *planted-fixture* control for the gate; `cookbook:18069` requires a control to use names the system accepts — neither is "control a REFUSAL against the already-succeeding population". +- **Proposed home:** DK (a kernel) +- **Portable because:** a filter that refuses work is the one class of change whose failure mode is silent, and the already-passing corpus is a free, exhaustive control set. + +### C4 — When a tool is repaired, every verdict it produced becomes a hypothesis again — but re-gate only the drafts the repair's BLAST RADIUS plausibly touched, not the whole ledger +- **Evidence:** `phase-ends/logs/Phase30.md:2481-2483` — "**when a tool is repaired, every verdict it produced becomes a HYPOTHESIS again.** The backlog's `closeness` values and residual classes were produced by tooling that has changed materially this session … **Re-measure before respecting any of them.**" and `:2569-2572` — "⚠️ **Do NOT blanket-re-gate the stored backlog — MEASURED this session**: fresh wave-6 drafts **4/6**, unbiased stored sample **1/12**, the two REVERTED overlays **3/17**. … **re-gate the drafts a repair plausibly touched, targeted by its blast radius.**" +- **What happened / what it cost:** four wave-6 drafts recorded as "blocked on a class needing a crack" banked **UNCHANGED** — they had already been freed by the same session's tool repairs; the drafts were correct and the instruments were failing them (the 5th "wall" of the phase to resolve to our own tooling). Two "permanent giant walls" carried since Phase 24/26 (one costing ~477k Fable tokens) then fell to drafts already on disk, worth 50,094 instructions. The measured re-gate rates set the scope: a 1,155-wide blind re-sweep is not justified by ~8%. +- **Not banked — greps:** `grep -n -i -e "after a tool repair" -e "after any tool repair" $FILES` → 0; `grep -n -i -e "targeted by its blast radius" -e "re-gate the drafts" -e "whole ledger" $FILES` → 0; `grep -n -i "becomes a hypothesis" $FILES` → 0; `grep -n -i -E "verdict.*hypothesis" $FILES` → 3 (all "a wall verdict without a pass and a dump line is a hypothesis" — an evidence requirement, not staleness-on-repair); `grep -n -i "stored verdict" $FILES` → 2 (cookbook §137a = verdict-vs-draft *mtime* staleness, a different mechanism) +- **Proposed home:** DK (a kernel) — pairs with the existing "re-probe exclude lists after tool fixes" memory, which this generalises and *bounds* +- **Portable because:** every decomp accumulates a ledger of negative verdicts produced by instruments that keep changing; the law says invalidate them, and the measured rates say how far to spend on re-validation. + +### C5 — A ledger row with no draft artifact is a rumour, not a result +- **Evidence:** `phase-ends/logs/Phase30.md:2352-2354` — "**Treat the 14 as UNVERIFIED** … General rule worth adopting: *a backlog row with no draft artifact is a rumour, not a result* — `backlog.py log` should require a draft path or mark the row unverifiable." +- **What happened / what it cost:** `func_8017C6F4` carried a `closeness=14` backlog row dated 2026-07-01 with `draft: None, klass: None, nins: None, reach: None`; no artifact existed and no draft of it survived on disk. Because `load_best` takes the LOWEST closeness, that unverifiable row **out-ranked a hand-won, committed 63/947 draft in every future target selection** — the same defect class as the Phase-28 `func_80178004 close=0` myth whose truth was 91. +- **Not banked — greps:** `grep -n -i -e "rumour" -e "no draft artifact" -e "unverifiable row" -e "artifact-less" $FILES` → 0 / 0 / 0 / 0; `grep -n -i "require a draft path" $FILES` → 0. (The *sibling* half — the ledger keying on address alone, and closeness being incomparable across sizes — IS banked at `docs/matching-cookbook.md:10301-10316` §3-B; the artifact requirement is not.) +- **Proposed home:** G (a rule) or a schema constraint in the registry template +- **Portable because:** any long project accumulates a target/backlog ledger; a row that outranks real work while carrying no reproducible evidence is a systematic mis-router, and the fix is a schema requirement, not diligence. + +### C6 — Check group identity from the signature files BEFORE probing a family: if the members are structurally identical, a 0% result is a compile-error CERTAINTY, not evidence about codegen +- **Evidence:** `phase-ends/logs/Phase30.md:168-171` — "🚨 **STANDING PRE-PROBE RULE:** check h_norm identity across a family's members from the sig files BEFORE probing. If members are h_norm-identical, a 0% is a **compile-error certainty**, not evidence about codegen." +- **What happened / what it cost:** this was written into the phase's standing rules after repeated sweeps returned total zeros on families whose members were identical by construction — outcomes that read as codegen walls and were plumbing every time (the same line adds "read `.run/hseq_failed.*.classified.txt` before theorising about any sweep failure — 23,211 of them exist and both of S38's zeros were already in there"). +- **Not banked — greps:** `grep -n -i "compile-error certainty" $FILES` → 0; `grep -n -i "check h_norm identity" $FILES` → 0; `grep -n -i "not evidence about codegen" $FILES` → 0; `grep -n -i "0% from the wrong tool" $FILES` → 2 (cookbook:8292, 9872 — §53's general "a 0% from the WRONG TOOL is not evidence", which is about tool ROUTING, not about member identity turning the same zero into a *certainty*) +- **Proposed home:** DK (a kernel) or cookbook +- **Portable because:** any project that groups functions by a structural hash can turn its own grouping invariant into a free, deterministic pre-verdict — "identical inputs cannot produce divergent codegen" holds on any compiler. + +### C7 — Never test a helper by importing its module: a tool with no `__main__` guard runs its whole pipeline on import +- **Evidence:** `phase-ends/logs/Phase30.md:2163-2168` — "`tools/harvest_verify.py` has **no `if __name__ == '__main__'` guard**: `import harvest_verify` runs a full build, splices drafts, and overwrites `.run/harvest_*.txt`. I tripped it unit-testing `classify_fail` … **To test a helper in it, `exec` that function's source — never import the module.**" (repeated at `:1179-1181`, where it ran a full gate on `resident` as a side effect) +- **What happened / what it cost:** tripped twice in the phase, on the project's most load-bearing gate tool. No damage either time (tree clean, binary byte-identical) but each cost a full unplanned build; the file carries a line-14 warning that was not enough, and the fix chosen was to flag it rather than risk a 500-line refactor of the gate. +- **Not banked — greps:** `grep -n -i "__main__" $FILES` → 0; `grep -n -i -e "import the module" -e "runs a build" $FILES` → 0 / 0; `grep -n -i "never import" $FILES` → 1 (`registry-E.decomp.md:65`, about never importing an assembler default — unrelated) +- **Proposed home:** G (a rule) or the tooling section of the how-to +- **Portable because:** decomp tooling is script-shaped and side-effecting by nature; the guard rule and the `exec`-the-function testing workaround transfer to any language with import-time execution. + +### C8 — When A/B-ing a harness knob (model, effort, reasoning level), ship a POSITIVE CONTROL that the knob actually moved +- **Evidence:** `phase-ends/logs/Phase30.md:1541-1547` — "a 3-way on ONE target list … with a **positive control that the effort knob actually moved** (compare per-agent token burn between arms). Without that control, a null result is indistinguishable from the override silently no-opping. ⚠️ Drew believes Workflow agents INHERIT session effort and that only `model` is settable; my tool docs + `docs/effort-map.md` say `opts.effort` overrides. **Unmeasured — treat as open, and design the test so the disagreement resolves itself.**" +- **What happened / what it cost:** the project could not settle whether the effort override on sub-agents was live at all, so every model/effort comparison it ran was uninterpretable in the null direction; the experiment was designed but the control was named as the precondition and the disagreement stayed open. +- **Not banked — greps:** `grep -n -i "positive control" $FILES` → 2 (decision-log:3191 and cookbook:34759 — both specific byte-level controls, neither about a harness setting); `grep -n -i -e "effort override" -e "the setting was applied" $FILES` → 0; `grep -n -i "silently no-op" $FILES` → 2 (decision-log:620, cookbook:2595 — a *tool* no-opping on input it cannot parse, not verifying a harness knob took effect) +- **Proposed home:** DK (a kernel) or the models-and-budgets chapter +- **Portable because:** every agentic project A/Bs model and reasoning settings; a silently-ignored override makes the cheaper arm look equal, which is exactly the wrong conclusion to bank. + +### C9 — Keep the wave harness IN THE REPO with its contracts; a harness rebuilt from memory each run silently goes stale +- **Evidence:** `phase-ends/logs/Phase30.md:688-693` — "**THE HARNESS IS NOW IN THE REPO** … It had lived only in the workflow scratch dir, so each wave rebuilt it from memory — which is how its cookbook citation list went stale at §162 while §163 (5) and §164 (82) were banked in between. **A wave 5 launched from the old script would have re-derived laws already on disk.** The README records the contracts that were paid for in failures: per-agent output dirs, sha1-last verification, prior-notes seeding, size routing, and 'name every banked block in the citation list'." +- **What happened / what it cost:** four waves ran off a script that existed only in scratch; its embedded knowledge-base citation list froze while 87 new idioms were banked, so agents were being briefed against a stale knowledge base — the direct negation of the project's flywheel thesis. +- **Not banked — greps:** `grep -n -i "wave harness" $FILES` → 0; `grep -n -i "rebuilt from memory" $FILES` → 0; `grep -n -i "citation list" $FILES` → 0; `grep -n -i "scratch dir" $FILES` → 15 (all about agent scratch/deliverable directories, none about the harness itself living there) +- **Proposed home:** accelerator or the cards-lanes-waves chapter +- **Portable because:** the fan-out script is the thing that carries the knowledge base into every agent; if it is not versioned alongside the knowledge base it decays exactly as fast as the knowledge base grows. + +### C10 — Do not cross-price two economies: a conversion rate measured on the RESIDUE QUEUE does not price a FRESH wave +- **Evidence:** `phase-ends/logs/Phase30.md:899-901` — "Corollary it also corrected: **do not cross-price the two economies** — 5:1 is a property of the RESIDUE QUEUE, while a FRESH wave converted 81% (and stored MATCH drafts re-gate at ~0%, A10)." (restated at `:1355-1357`) +- **What happened / what it cost:** the 5:1 PLUMBING:DIFF ratio measured on the sweep residue was the basis for the whole "tooling beats volume" doctrine. It was a property of a queue that had already been filtered by everything easy; at session close the residue had crossed over to DIFF 170 of 575 (30%), the declaration-axis vein "was a one-time seam", and the strategy had to flip to "volume with multipliers, not more plumbing". Earlier in the same phase a "~32% recovery" figure projected from wave drafts onto extend drafts produced **0/129**. +- **Not banked — greps:** `grep -n -i "cross-price" $FILES` → 0; `grep -n -i "two economies" $FILES` → 0; `grep -n -i "residue queue" $FILES` → 0 +- **Proposed home:** DK (a kernel) or the economics chapter — a companion to R41 ("quote the denominator") +- **Portable because:** every decomp ends up with at least two populations (fresh targets vs a picked-over backlog) whose conversion rates differ by an order of magnitude; carrying one rate into the other's budget is the standard way to mis-plan a phase. + +### C11 — Derive the target pool from the BUILD'S OWN INVARIANT, not from a reach/similarity metric — reach ranked the most-DONE work first +- **Evidence:** `phase-ends/logs/Phase30.md:1810-1814` — "**Reach-141 identifies the most-DONE work, not the most valuable** — those are the shared engine functions banked over 29 phases, present in each overlay as `DEFINE_func_*` macros (~1,614 per big-3 binary). My '20× leverage' argument was backwards. **Derive targets from the build invariant (R33): `INCLUDE_ASM` in the committed source.**" +- **What happened / what it cost:** the frontier had been priced off a reach metric that counts every *instance* of an already-banked shared function, so the highest-reach rows were the most completed work in the project. Re-deriving from the build's own stub markers gave the real fuel: 1,799 draftable functions, 1,168 already seeded. This was the correction that followed a 50-agent / 2.5M-token wave that banked ZERO. +- **Not banked — greps:** `grep -n -i "build invariant" $FILES` → 0; `grep -n -i "most-DONE" $FILES` → 0; `grep -n -i "derive targets from" $FILES` → 0; `grep -n -i "INCLUDE_ASM in the committed" $FILES` → 0. (`cookbook:8733` banks *which key to rank by* within the family map — a different question from *where the pool comes from*.) +- **Proposed home:** DK (a kernel) — a sharper form of "derive from invariants, don't re-parse" +- **Portable because:** every decomp has a build-visible stub marker that is definitionally the open set; any derived metric can drift from it, and a metric that counts instances of finished work will invert the ranking. + +### C12 — A metric that re-parses source is blind to a body banked through an `#include`; trust the metrics derived from the stub oracle +- **Evidence:** `phase-ends/logs/Phase30.md:2589-2593` — "A body banked by `#include`-ing a shared header is **invisible to fn-count's NUMERATOR** (the definition is not in the `.c`) while its stub leaves the denominator — the whale bank moved fn-count `341186/353718 → 341186/353717`. The **weighted** metrics counted it correctly (+770) because they derive from `corpus.stubs`, not re-parsed C. **Trust the weighted pair.**" +- **What happened / what it cost:** a real +770-instruction bank read as a *fall* in the headline function-count metric, because that metric re-parses the `.c` looking for definitions. The same root cause (a classifier that re-parses C and inherits every parsing blindness) had already produced the phase's stale-digest scare. +- **Not banked — greps:** `grep -n -i "invisible to fn-count" $FILES` → 0; `grep -n -i "trust the weighted" $FILES` → 0; `grep -n -i "include-ing" $FILES` → 0. (Adjacent but distinct: `accelerators.md:664` "every headline % ships with its remainder", about a *denominator* carrying linked-SDK code.) +- **Proposed home:** accelerator or the metrics section of the how-to +- **Portable because:** any project that shares a matched body via a header or macro instead of a per-TU definition will confuse a source-parsing progress metric; the fix (derive from the build's stub set) is compiler- and console-independent. + +### C13 — Keep a glossary line for any term two documents use in OPPOSITE senses +- **Evidence:** `phase-ends/logs/Phase30.md:1343-1345` — "4. **'ZERO-CRACK' MEANS OPPOSITE THINGS** in roadmap §3 T3 ('61 zero-crack = propagation-only') and in the current map/this file (zero-crack = needs its FIRST crack). **30× mis-scope risk.** One glossary line in `family-hseq.md` fixes it." +- **What happened / what it cost:** the same term named 61 families of free propagation work in the roadmap and 1,817 families of unstarted cracking work in the frontier map — a ~30× difference in scope, sitting in two documents a planner reads together. The same audit found the targeting oracle's own scope stamp wrong ("212 OVERLAYS only" for a map containing resident members and 70 modules). +- **Not banked — greps:** `grep -n -i "glossary" $FILES` → 1 (cookbook:137, pointing at the *community's* decomp wiki glossary for compiler terms — not an internal-term glossary); `grep -n -i -e "means opposite" -e "two meanings" -e "ambiguous term" $FILES` → 0 / 0 / 0 +- **Proposed home:** G (a rule) — every derived-metric doc carries a glossary line for its own terms; pairs with the existing "coverage/scope stamp" laws +- **Portable because:** a long project mints its own vocabulary faster than it documents it, and the terms that drift are the ones two independent pipelines both use. + +## ALREADY-BANKED (one line each) +- Never wrap a self-timing tool in a tighter outer timeout; kill process GROUPS, not pids — lives at `docs/matching-cookbook.md:16339` +- Spread a gate slate across destination TUs to avoid per-TU declaration collisions (64% → 92%) — lives at `docs/matching-cookbook.md:16350` (§169) and `:16390` +- A confident WRONG label costs more than a missing one; prefer "CAUSE NOT DETERMINED" + the next check — lives at `docs/accelerators.md:269`, `docs/cookbook-index.md:647` (§423) +- Idiom mining saturates; measure COVERED+UNSOUND and stop when it rises (57→64→76%) — lives at `docs/matching-cookbook.md:15012-15026` (§167) +- Per-agent output directories, sha1-last verification, and transcript replay as the draft backup — lives at `decomp-kernels.md:459` (DK-35), `docs/wave-playbook.md:786-793`, `docs/accelerators.md:734` +- Rank family work by templatable weight (`members × nins`), never by the h_exact-reach worklist — lives at `docs/matching-cookbook.md:8733` +- A bare `except` around a fail-CLOSED oracle reinstates the guess it refuses to make — lives at `docs/matching-cookbook.md:9732`, `docs/decision-log.md:2079` +- A tool built and never wired to its caller is a stale default wearing a safety label — lives at `docs/matching-cookbook.md:7287` (§101, the STALE DEFAULT class) +- Derive a failure class from whether the compile produced an object; do not pattern-match diagnostic prose (the old compiler emits no `error:` token, and `make`'s summary prints last) — lives at `docs/matching-cookbook.md:9201` and `:10319-10325` +- Duplicate/identical-layout anonymous-struct typedefs are distinct, conflicting C types — lives at `docs/matching-cookbook.md:30845` (§321), `:30863` ("the conflict is type IDENTITY") +- The draft-vs-draft data-symbol type clash class (two staged drafts spelling one symbol differently) — lives at `docs/matching-cookbook.md:17931` +- Never hand-type a target or wave args; validate the list — lives at `docs/wave-playbook.md:316` +- Provenance → archive → link → compiler: a band of "walls" may be the vendor's library, not game code — lives at `docs/retrospective.md:113`, `docs/accelerators.md:681`, `docs/how-to-ai-decomp/07-compiler-source.md:94`, `12-failure-museum.md:21` +- A confident NEGATIVE verdict is a claim: date it, name its evidence, re-measure before it parks work — lives at `docs/matching-cookbook.md:10173-10174`, `docs/decision-log.md:2236` +- A control that cannot fail is not a control (a no-op can pass for the wrong reason; a control must use inputs the system accepts) — lives at `docs/matching-cookbook.md:842-845` and `:18069` +- A silently-broken tool invalidates every measurement taken through it (a hand-search floor measured while the permuter was dead) — lives at `docs/matching-cookbook.md:10261`, `:10298` +- Same address + same name ≠ same body; a ledger keyed on address alone mis-ranks, and `closeness` is not comparable across sizes — lives at `docs/matching-cookbook.md:10301-10316` (§3-B) +- Test a ×N claim by renaming the draft and running the per-function oracle against the sibling's asm — lives at `docs/matching-cookbook.md` §148-E (referenced `:10463`, `:11732`) +- The byte difference can live in the CALLERS, which the per-function oracle never compiles — lives at `docs/matching-cookbook.md:10028-10029` +- The byte gate is a null oracle for DOCUMENTS; guard derived metrics with their own audit, comparing integers not percentages — lives at `docs/matching-cookbook.md:9734-9740` +- A growing denominator lowers the headline; that is the honest direction, and percentages compare only within one denominator — lives at `docs/decision-log.md:1135`, `:2200`, `docs/accelerators.md:664` +- Partition the whole medium into buckets that sum, with residue as a DEFECT: a zero-residue partition is a completeness proof, a longer list is only a longer list — lives at `docs/decision-log.md:2173-2180` +- Do the live-capture tour and the onboarding it unblocks in one run, or you pay for the tour twice — lives at `docs/decision-log.md:2197` +- The gate number measures INTEGRATION, not matching; run the recovery ladder before recording a wave's yield — lives at `decomp-kernels.md:530`, `docs/how-to-ai-decomp/05-cards-lanes-waves.md:143`, `registry-E.decomp.md:316` +- SCAN, don't sample: `head -8` of 31 stored drafts missed the match in the 9th — lives at `docs/matching-cookbook.md:10051` +- An exact skeleton hash fragments same-source families into false singletons; similarity clustering recovers them, cousins are SEEDED CRACKS not remaps, rank by unit weight, discount short-function similarity — lives at `docs/matching-cookbook.md:16294-16312` (§168) +- Agent MATCH claims are honest at the cheap tier and should be relied on as a FILTER, never as the arbiter — lives at `docs/matching-cookbook.md:10740` +- Prior-notes / past-attempt seeding is the strongest cheap lever; a NEAR is a resumable state — lives at `docs/matching-cookbook.md:16303`, `docs/wave-playbook.md` (journal_notes wiring) +- A pipeline's exit status is the LAST command's; `set -o pipefail` and never pipe away the evidence — lives at `docs/matching-cookbook.md:6956` (§93), `:9667-9669` +- The splat asm subdir NAMES the destination TU; splicing into the wrong TU is a NO-OP that reads as a codegen failure — lives at `docs/matching-cookbook.md:14958` (§166a); prior notes' TU paths must be re-resolved (`:14701`) +- A usage-limit outage costs only the unfinished work: `resumeFromRunId` replays cached agents and re-runs only the dead ones — lives at `docs/matching-cookbook.md:20755`, `:17175` +- Derive a bank count from the SOURCE / the stub oracle; never accumulate per-batch reports — lives at `phase-ends/DIGEST.md:226`, `docs/matching-cookbook.md:34178`, `docs/decision-log.md:2753` +- Read the agent's own integration notes before diagnosing a gate failure — it has already seen the TU — lives at `docs/matching-cookbook.md:11044` +- Every line-shape decision must route through a comment/string mask; agent prose in a draft is the carrier — lives at `docs/matching-cookbook.md:8745` (§134), `:9774` (§141, the class CLOSED) +- Read the tool's own classified failure payload before theorising about a sweep failure — lives at `docs/matching-cookbook.md:9874` (the standing pre-probe rule) +- C89 `f()` declares UNSPECIFIED parameters, not zero — lives at `docs/matching-cookbook.md:16519` +- Prove a metric CANNOT have fallen (identical sigs + unchanged tools + zero re-added stubs) before hypothesising about a metric move — lives at `docs/matching-cookbook.md:9719-9720`, `docs/decision-log.md:2057`, `:2090` diff --git a/.run/P33.5/log-mining/Phase30-2of2.md b/.run/P33.5/log-mining/Phase30-2of2.md new file mode 100644 index 0000000000..85524e3f99 --- /dev/null +++ b/.run/P33.5/log-mining/Phase30-2of2.md @@ -0,0 +1,90 @@ +# Log mining — Phase30-2of2 +Files/ranges: `phase-ends/logs/Phase30.md:2685-5368` (context read from 2655) · Lines read: 2684 of 2684 +Candidates considered: 44 · NEW: 6 · ALREADY-BANKED: 38 + +Corpus grepped for every candidate (abbreviated `$F` below): +`docs/matching-cookbook.md docs/cookbook-index.md docs/decision-log.md docs/accelerators.md docs/retrospective.md docs/wave-playbook.md docs/how-to-ai-decomp/*.md decomp-architect/corpus/decomp-kernels.md decomp-architect/templates/registry-E.decomp.md phase-ends/DIGEST.md` + +## NEW + +### C1 — Drafting agents must never be able to write the build tree; every agent artifact lands in a scratch directory, so a killed or racing campaign costs build cycles and zero work +- **Evidence:** `phase-ends/logs/Phase30.md:4017-4018` — "Corollary proven twice: **agents must only ever write `.run/`** — that is why both incidents cost build cycles and zero work." Also `:5149-5151` ("all 58 drafts survive untouched in `.run/` because agents never write the tree — the discipline that made this cheap") and `:5215-5222`. +- **What happened / what it cost:** Two tree-corruption incidents landed in one phase — a hardcoded 3600 s propagation timeout killed the driver mid-fleet-write (313 files modified, `config/dedup.us.yaml` never updated, `engine_core.h` half-edited, `check-all` 124/140) and a poll-instead-of-mutex race let a parallel gate, a live propagation loop and a `make clean` overlap (`check-all` 77/140, then 63/140). Both recoveries were a single `git checkout -- src/ config/` back to the last verified commit. Nothing that had been *paid for* was lost, because 58 agent drafts and every wave manifest lived outside the tracked tree. +- **Not banked — greps:** `grep -n -i -E "agents.*tree|never write the tree" $F` → 0; `grep -n -i -E "write only to|writes only to|agents never write|\.run/ only" $F` → 0; `grep -n -i -E "cost build cycles|zero work" $F` → 0; `grep -n -i -E "write scope|write-scope" $F` → 2 (both cookbook §61-family, about a *stage's undo scope being narrower than its write scope*, not about confining agents). The nearest banked neighbours are an ops note (`docs/matching-cookbook.md:16768`, "the drafter workflow makes NO tree writes") and DK-34 (`decomp-architect/corpus/decomp-kernels.md:446`, agents write deliverables *early*) — neither states the confinement as an architectural precondition, and neither carries the incident cost that justifies it. +- **Proposed home:** DK (a kernel), sibling to DK-34; also a G-rule for the day-one lane architecture. +- **Portable because:** it is a property of the *harness*, not the console — any project running concurrent model lanes against a build tree gets the same two incidents, and confinement is what makes recovery a one-line checkout instead of lost work. + +### C2 — A stored verdict can be stale because the INSTRUMENT changed, not the draft: re-gate the drafts a tool repair plausibly touched, scoped by the repair's blast radius — never the whole ledger +- **Evidence:** `phase-ends/logs/Phase30.md:4660` — "**⚠️ THE FINDING: all four banked UNCHANGED — no new work on the drafts.** … The drafts were correct; the instruments were failing them. **⇒ RE-GATE STORED DRAFTS AFTER ANY TOOL REPAIR before treating a stored verdict as a fact about the code**" and, in the same entry, the measurement that bounds it: "fresh wave-6 drafts **4/6** · unbiased stored sample **1/12** · the two REVERTED overlays **3/17** … The real rule: **re-gate the drafts a repair plausibly touched, targeted by its blast radius — not the whole ledger.**" +- **What happened / what it cost:** Six wave-6 drafts had been recorded as blocked on "a class needing a fix". A later session's tool repairs (an alias-deletion blindness in two splitters, plus a corpus-reload defect) had already freed them; four then banked byte-identical **with no edit at all** (+1,905 instructions). The prior session had written the block down as a property of the code. This was counted in the log as the fifth "wall" of that phase to resolve to the project's own tooling. +- **Not banked — greps:** `grep -n -i -E "after a tool (fix|repair)|tool repair|repaired.*re-run.*drafts" $F` → 0; `grep -n -i -E "banked UNCHANGED|no new work on the draft|the drafts.*the instruments|instrument.*failing them" $F` → 0; `grep -n -i -E "re-gate.*blast radius" $F` → 0; `grep -n -i -E "stored verdict" $F` → 2 — both are cookbook §137a, which enumerates exactly two staleness axes (the draft got worse; the draft got *better*) and does not contain the instrument axis. The banked memory `reprobe-exclude-lists-after-tool-fixes` and failure-museum row 20 cover the *exclude list*, i.e. targets never attempted — not stored drafts already judged and filed. +- **Proposed home:** DK (a kernel) — pairs with the banked "an exclude list records what the tooling could not do". +- **Portable because:** any project that stores model output plus a verdict accumulates verdicts whose truth depends on tool versions; the scoping half (blast radius, not the ledger) is what stops the rule from costing a 1,155-build sweep. + +### C3 — Count agent completions from the run journal's result records, never from artifact existence: an agent writes its deliverable early and then iterates, so the file proves nothing +- **Evidence:** `phase-ends/logs/Phase30.md:4001-4003` — "*(Counting note, R14: my first count said "2 remaining" because I measured draft-FILE existence. An agent writes its draft early and then iterates, so a file proves nothing about completion — the journal's `result` records are the truth. Drew's "12" was right and my measure was wrong.)*" Context: `:3990-4000`, 70 agents launched, 58 returned, 12 killed mid-run, "10 of the 12 have a **PARTIAL draft** on disk from the killed run — a starting point, NOT a verified result". +- **What happened / what it cost:** After a wave was cut short, the orchestrator counted surviving work by listing draft files and reported 2 outstanding agents; the true figure was 12. The human caught it. Ten of the twelve had partial drafts on disk that read as finished work, and every one of them had to be re-verified with the match tool before any claim. +- **Not banked — greps:** `grep -n -i -E "file proves nothing|draft-FILE|file existence" $F` → 0; `grep -n -i -E "counted files|file count|progress by artifact|an artifact that exists" $F` → 0; `grep -n -i -E "journal.*result records" $F` → 1 (`docs/accelerators.md:502`, about an *unread* journal corpus, a different lesson). DK-34 (`decomp-architect/corpus/decomp-kernels.md:446`) banks the agent-side practice ("the draft file first, the verdict last") — this is its unbanked orchestrator-side consequence: because drafts are written early, artifact counts are not completion counts. +- **Proposed home:** DK — an addendum line on DK-34. +- **Portable because:** the write-early convention is itself recommended day-one practice, so every project that adopts it inherits this counting trap; and a partial artifact that reads as a finished one is how a wave silently loses work. + +### C4 — A derived claim outranks a heuristic verdict; when two heuristics disagree, take the UNION and queue the disagreements — under-reporting hides work, over-reporting only costs review +- **Evidence:** `phase-ends/logs/Phase30.md:4657-4666` — "**A claim outranks a heuristic.** I let the statistical verdict override a SHA match … Onboarded bucket was understated by **14.5 MB**. … **Whole-payload averaging DILUTES code** … **80 disagreements, all this shape.** Resolution: a claim wins outright; otherwise take the **UNION** of the two oracles — over-reporting code puts a payload in a review queue, under-reporting it hides code, which is the exact failure mode that produced the three surprises. Unclaimed code went **1.70 MB → 3.56 MB** once fixed." +- **What happened / what it cost:** The whole-medium audit's first layer classified payloads by a statistical code test and let that verdict beat a payload's own build-hash claim, so binaries the project had already built were filed as data (14.5 MB understated). The second oracle disagreed on 80 payloads, all with the same cause (averaging a code head against a large data tail). Fixing the precedence and switching to a union roughly doubled the measured unclaimed-code figure — the very quantity the audit existed to find. +- **Not banked — greps:** `grep -n -i -E "a claim outranks a heuristic|over-report|averaging|fail.*safe.*direction" $F` → 0; `grep -n -i -E "union of the two oracles" $F` → 0; `grep -n -i -E "review queue" $F` → 1 — `docs/decision-log.md:2188`, the *pre-implementation design intent* ("disagreements become the review queue"), which records neither the precedence law nor the measured cost of getting it backwards. R34 (a second, disagreeing oracle) is banked everywhere; how to *resolve* a disagreement is not. +- **Proposed home:** DK — an extension of the R34 kernel (oracle precedence + error-direction asymmetry). +- **Portable because:** every completeness audit ends with two oracles disagreeing, and the choice of which one wins, and which way to err, is where the audit's answer actually comes from. + +### C5 — A failure that will not reproduce earns a negative-control-proven DETECTOR, not a speculative fix; and every abort path must PROVE its revert by diffing the worktree against a baseline captured at start +- **Evidence:** `phase-ends/logs/Phase30.md:4064` — "**I did not 'fix the bug'; I made the next occurrence name itself.**" And `:4077-4081` — "**An incomplete 'REVERTED'** … struct_check restored only `touched`, leaking every Part-B reconcile kept on disk. New `_abort()` undoes `touched` **and** every kept reconcile, then **diffs the worktree against a baseline captured at start and says what survived**. A dirty tree nobody knows about is the expensive failure: it makes every later byte-gate report `near`, so its verdicts are void and get misread as draft failures (that cost two batches …)." +- **What happened / what it cost:** A 92-minute propagation run died with a terse "not instantiated — REVERTED" that named no mechanism, and the failure did not reproduce at HEAD. Rather than guess, the session shipped three things: an explicit gap report naming what each unplaced site actually is, a complete abort, and a baseline worktree diff — then proved all of it with a negative control that recreated the exact condition (fail-closed, exit 1, "tree restored to baseline; no residue"). The incomplete revert it replaced had already cost two batches of misread gate verdicts. +- **Not banked — greps:** `grep -n -i -E "does not reproduce|name itself|next occurrence|speculative fix" $F` → 4, none about a non-reproducing failure (two are codegen shapes, one is a coverage assertion, one a residual tell); `grep -n -i -E "incomplete revert|partial revert" $F` → 0; `grep -n -i -E "diffs the worktree against a baseline|baseline captured at start|says what survived|restored to baseline" $F` → 0. The *consequence* is banked ("a gate that starts on a dirty tree cannot tell your edits from its own", `docs/how-to-ai-decomp/02-byte-gate.md:66`, `docs/wave-playbook.md:21`); the *cause* — an abort that asserts a revert instead of proving one — is not. +- **Proposed home:** DK (a kernel) — the abort/revert half; the detector half is a G-rule. +- **Portable because:** every campaign tool that mutates a tree has an abort path, and an unproven revert is the standard way a project acquires a dirty tree nobody knows about; the "instrument, don't guess" half applies to any non-reproducing failure in any harness. + +### C6 — Never let model-authored prose reach the shell inside double quotes: a backticked command in a `git commit -m` message EXECUTED +- **Evidence:** `phase-ends/logs/Phase30.md:2893-2894` — "**A backticked `` `make extract` `` in a `-m` commit message EXECUTED** — corrupted the message and ran a real extract. Use quoted heredocs. (No damage; re-committed.)" Recorded again at `:2707` and `:3051-3053` ("what I used everywhere else and lapsed on once"). +- **What happened / what it cost:** A commit message written in the project's own prose style — tool names in backticks — was passed as a double-quoted `-m` argument. The shell ran the command substitution: the message was corrupted and a real `make extract` executed against the tree. Cost was tokens and a re-commit, not correctness, but the same lapse against a destructive command would not have been. +- **Not banked — greps:** `grep -n -i -E "command substitution|shell expansion|double-quoted|git commit -m|quoted heredoc|heredoc" $F` → 0; `grep -n -i -E "backtick" $F` → 2 (both `docs/how-to-ai-decomp/12-failure-museum.md:45` / `docs/decision-log.md:3501`, about backtick-bracketing as an *AI tell in outward prose* — a style rule, not a shell-safety rule); `grep -n -i -E "-m .*execut|commit message.*execut" $F` → 0. +- **Proposed home:** G (a rule) — one line in the day-one governance/ops rules. +- **Portable because:** every agent-run project has a model writing commit messages and PR bodies full of backticked identifiers, and every one of them passes those strings to a shell. + +## ALREADY-BANKED (one line each) +- The same line-shape parsing defect recurred in six tools; route every such decision through one masking oracle instead of a seventh patch — lives at `docs/matching-cookbook.md:9777` (§141) and `docs/cookbook-index.md:706`. +- A committed-stale generated digest manufactured a phantom regression; a metric is not a measurement until it reproduces from the committed tree (`make audit-digest`) — `docs/matching-cookbook.md:9694` (§140), `docs/decision-log.md:2039`. +- A prep step that silently returns its input on failure is indistinguishable from a search that found nothing; assert the OUTPUT — `docs/matching-cookbook.md:10285` (§149-A). +- A hand-search floor measured while a tool was silently broken is not a floor — `docs/matching-cookbook.md:10298`. +- Same address + same name ≠ same body; a ledger keyed on address alone merges two functions, and absolute closeness is not comparable across sizes — `docs/matching-cookbook.md:10300-10312` (§149-B). +- A null field defaulting to a plausible value made a real result invisible (`binary: null`; absent scored as done) — `docs/matching-cookbook.md:10310`. +- The build tool's own summary line is not a diagnosis; take the FIRST real diagnostic, and a label identical for every input carries no information — `docs/matching-cookbook.md:10317-10325` (§149-C). +- A carried "cheap fuel" item with a price and no probe behind it is a guess wearing a number; probe one member before scheduling — `docs/matching-cookbook.md:10327-10340` (§149-D). +- Bimodal bank rates (all-or-nothing per group) are a tooling signature, not a codegen one — `docs/matching-cookbook.md:8763`. +- An exactly-zero result is a decoder/structural gap until proven otherwise — `decomp-architect/corpus/decomp-kernels.md:254` (DK-18), `docs/accelerators.md:82`. +- `.DELETE_ON_ERROR` — a truncated object outlived its own compile error and the next build linked it — `docs/matching-cookbook.md:8637`. +- A pipelined fan-out (refill each slot on completion) beats a batched one; 16 targets in 82 min at 3.8× vs a barrier — `docs/wave-playbook.md:448`, `docs/how-to-ai-decomp/05-cards-lanes-waves.md:106-109`. +- Rank by pool realisation in instructions, not bank rate in heads — `decomp-architect/corpus/decomp-kernels.md:542`, `docs/how-to-ai-decomp/09-economics.md:83`. +- Rank leverage by LIVE (unmatched) reach, not total sharers — `docs/decision-log.md:1666`. +- Guard the CAMPAIGN, not the process; a process poll is a sampling test on a gappy signal — `docs/matching-cookbook.md:8866`, `:16344`. +- `pgrep -f` self-matches, so the waiter never exits — `docs/accelerators.md:196-201`, `docs/wave-playbook.md:721`. +- A killed process performs no undo — `docs/matching-cookbook.md:9056`. +- A pipeline's exit status is the last command's; use `pipefail`/`PIPESTATUS` — `docs/matching-cookbook.md:9667`, `docs/cookbook-index.md:1538` (§93). +- A masked comparison tool cannot validate an edit to the fields it masks (a wrong symbol map still reports MATCH) — `docs/matching-cookbook.md:10478-10479`. +- Count banks from the SOURCE (the stub's presence), never from a gate report — `phase-ends/DIGEST.md:226` (R42), `docs/matching-cookbook.md:34178`. +- The byte gate is a perfect correctness oracle and a null coverage oracle (green at 0% decompiled) — `docs/how-to-ai-decomp/02-byte-gate.md:30`, `docs/decision-log.md:703`. +- Every target passes a validity gate before a wave; a coverage assertion that refuses must be asked why, not routed around — `decomp-architect/corpus/decomp-kernels.md:366-373`, `docs/accelerators.md:155-173`. +- Paste wave arguments from a derived manifest, never retype them — `docs/matching-cookbook.md:8861`. +- When N drafts share one TU, forbid concurrent agent builds — `docs/matching-cookbook.md:8864`. +- A fixed temp path is a correctness bug the day two workers run (fake isolation); processes not threads for CPU-bound sweeps — `docs/accelerators.md:137-138`, `decomp-architect/corpus/decomp-kernels.md:435`. +- Batch every failure the sweep already computed: convergence ~138 rounds → 1–3 with no extra builds — `docs/accelerators.md:123`. +- Read the recorded failure verdicts before designing an experiment — `phase-ends/DIGEST.md:218` (R38), `docs/matching-cookbook.md:9874`. +- A model-tier boundary set by extrapolation was wrong by a factor of two; a throttled run is an instrument failure, not a model verdict — `docs/how-to-ai-decomp/08-models-and-budgets.md:9-11`, `:43`. +- One declaration conflict hides the next, so a plumbing verdict says nothing about the body — `docs/matching-cookbook.md:9475`, `:9358`. +- A tool repair's blast radius is verified by diffing its output against the pre-fix tool over the whole population — `docs/matching-cookbook.md:8767`. +- A gate result measured against stale asm (or during a rebuild, or on a dirty tree) is not a measurement — `docs/matching-cookbook.md:9676-9678`, `docs/how-to-ai-decomp/02-byte-gate.md:66`. +- A known gap that announces itself is not the same defect as a silent one — `docs/matching-cookbook.md:8790`. +- A carried list (exclude list / wall list) records what the tooling could not do; the draw refuses a stale one — `docs/how-to-ai-decomp/12-failure-museum.md:28`, `docs/wave-playbook.md:37-41`. +- A whole-medium residue-0 partition with `claimed-by` derived from the build contracts, so unclaimed code is enumerated rather than stumbled into — `docs/how-to-ai-decomp/12-failure-museum.md:20`, `docs/decision-log.md:2210`. +- A module with top-level side effects cannot be shared or tested by import; extract the predicate — `docs/matching-cookbook.md:30935` (§323b). +- Stale is worse than absent: refresh the replayable checkpoint block — `docs/how-to-ai-decomp/01-governance.md:81`, `docs/wave-playbook.md:715`. +- Enumerate banked twins before drawing a wave (some of the band is deterministic tool work, not agent work) — `docs/how-to-ai-decomp/10-integration-and-propagation.md:42`. +- An agent's standalone MATCH is not a bank (measured 88–92% match_one → 60–71% banked) — `phase-ends/DIGEST.md` R-rules / memory `standalone-match-is-not-a-bank`; cookbook `docs/matching-cookbook.md:1661`. diff --git a/.run/P33.5/log-mining/Phase31-1of3.md b/.run/P33.5/log-mining/Phase31-1of3.md new file mode 100644 index 0000000000..0465ad2f20 --- /dev/null +++ b/.run/P33.5/log-mining/Phase31-1of3.md @@ -0,0 +1,135 @@ +# Log mining — Phase31-1of3 +Files/ranges: `phase-ends/logs/Phase31.md`:1-2902 · Lines read: 2902 of 2902 (read through 2955 for context per the brief) +Candidates considered: 52 · NEW: 12 · ALREADY-BANKED: 40 + +Slice covers S52–S63 (2026-08-14 → 2026-08-27): the atlas/lane build-out (T0–T9), the overnight +wave campaign (waves A–ZZ, ox/DeepSeek era), the S60/S61 harness-defect sessions, the red-binary +surgeries, and the T4/T5 Claude-ladder waves. + +## NEW + +### C1 — More output-token budget is NOT more quality; measure it as a paired A/B and treat truncation as recoverable, not as a defect signal +- **Evidence:** `phase-ends/logs/Phase31.md:2292-2302` — "**BANKABLE: 8k 10/10 · 16k 10/10 · 24k 8/10 · 32k 8/10** … Truncated turns 6/1/0/0 — truncation recovers across the turn loop and does not cost banks … MORE OUTPUT BUDGET IS NOT MORE QUALITY on this model/population — budget bought nothing the turn loop didn't already provide." +- **What happened / what it cost:** The project spent most of a session (S60) arguing 8k-vs-16k output budget from historical waves and could not settle it, because token budget and population exhaustion had moved together (`:357-362`). Raising the budget had earlier caused a *real* regression that was actually a harness interaction — at ~30 tok/s a 16k generation runs ~530 s while the straggler grace was 120 s, so agents were guillotined mid-thought with no draft at all (`:350-356`). The settling experiment (S61) was a paired four-arm run over the same 10 functions with identical warm-start inputs, byte-judged; its first table was itself wrong (an R53 false green) and the corrected table was 8k 6 · 16k 7 · 24k 5 · 32k 3 (`:2304-2317`). Also measured twice: truncated-turn rate is *inversely* correlated with bank rate (`:308-311`, `:2294-2295`). +- **Not banked — greps:** `grep -n -i 'output token' ` → 2 (both about token accounting by phase); `grep -n -i 'MAXTOK' ` → 8 (all provenance mentions of the wave tag `ab8/ab16/…`, no result); `grep -n -i 'truncat' ` → 14 (compiler truncation, corpus-prep truncation, int truncation — none about agent output budget); `grep -n -i 'bigger budget|more budget|budget is not more quality|max_tokens|completion budget|inversely correlated|truncated turn'` → 0. `docs/how-to-ai-decomp/08-models-and-budgets.md` "Budgets" covers per-lane caps (R46), concurrency and steady-state, not the output-length dial. DK-43 covers the model ladder, not the budget dial. +- **Proposed home:** DK (a kernel), beside DK-43 — "budget per lane" gains a second half: *the output-length dial is not a quality dial; settle it with one paired within-wave A/B on one population, and treat a truncated turn as recovered work.* +- **Portable because:** every agent harness exposes a max-output setting, every project is tempted to raise it when drafts look cut off, and the experiment design (hold the population constant, split within a single wave) is the only one that isolates it. + +### C2 — In a pipelined drafter/gater, "still open" is not "not yet attempted": consecutive waves re-drafted the wave still in flight +- **Evidence:** `phase-ends/logs/Phase31.md:258-260` — "**Every second wave re-drafted the wave still in flight** — `--retry-unbanked` returns still-open cards and the pre-draw runs WHILE a wave drafts, so ck->cl were 239/239 identical, co->cp 238/238. Yield alternated 47.6% / 3.8% / 35.3% / 3.6%. Fixed by deriving finished-ness from BOTH gate logs." +- **What happened / what it cost:** Once drafting and gating ran as independent lanes, the draw for wave N+1 fired before wave N's cards had reached a gate, so the "still open" filter returned wave N's slate verbatim — three consecutive pairs at 239/239, 238/238, 222/222 (`:431-434`). Half of every pair of waves was pure waste, and the alternating yield read as population volatility rather than as duplication until the identity of the card sets was checked. +- **Not banked — greps:** `grep -n -i 'retry-unbanked'` → 0; `grep -n -i 'duplicate wave|identical cards|re-draw'` → 6 (all "recover before you re-draw", DK-41/G50 — a different rule about recovering a gate's failures before drawing again); `grep -n -i 'not yet gated|already drafted but not yet gated|still in flight'` → 1 (about an in-flight wall proof). +- **Proposed home:** DK (a kernel) or the wave playbook's draw section — *a draw filter keyed on "banked" double-counts everything in flight; derive finished-ness from the gate's own log, and assert the new slate is disjoint from the live one.* +- **Portable because:** any project that overlaps generation with verification has the same race, and the failure is silent — both waves look healthy and the duplication only shows in a set comparison. + +### C3 — Snapshot every target's disassembly BEFORE gating: a successful bank prunes it, and every downstream step needs both sides +- **Evidence:** `phase-ends/logs/Phase31.md:1612` — "Snapshot every target's `.s` BEFORE gating — banking prunes it and recovery/harvest need both sides." (also `:1689-1690`, `:785` where a snapshot dir is what saved a dead harvest) +- **What happened / what it cost:** The build system stops emitting a function's assembly once it matches, so the moment a draft banks, the artifact every later step compares against is gone — the idiom harvest, the near-miss autopsy and the adversarial verifiers all need the target's original bytes. Verifiers in one harvest had to rebuild their targets from the ROM because the `.s` had been pruned on bank (`:1765-1767`). The habit became a standing step (`.run/wave*_asm_snapshot/` with a MANIFEST.sha1). +- **Not banked — greps:** `grep -n -i 'snapshot'` over the corpus → 2 (a dated-snapshot documentation rule; one accelerator about a coverage assertion refusing a list); `grep -n -i "target's .s BEFORE|before gating"` → 4 (all about *checking* a slate before gating, not preserving evidence); `grep -n -i 'prunes|the target disappears|verifiers need both sides'` → 0. +- **Proposed home:** DK (a kernel) — *the byte gate destroys its own evidence on success; capture the "before" artifact before the step that consumes it.* +- **Portable because:** any decomp whose extractor stops emitting matched functions has this, and so does any pipeline whose success step deletes its input — the loss is invisible until a later pass needs the pair. + +### C4 — Draft first, then carve: a build-unit split is only safe when the new unit is immediately populated with proven bodies +- **Evidence:** `phase-ends/logs/Phase31.md:922-923` — "**Draft-first ordering is what makes carves safe: populated-at-carve is 137/137; stubs-in-new-object was 1/4 and 2/2.**" (repeated as a standing note at `:991-992`) +- **What happened / what it cost:** Splitting a jump-table or `-O0` region into its own object changes layout, padding and symbol placement. Done against a still-stubbed region there is nothing to verify the new layout with, and the ledger shows it converting at 1/4 and 2/2. Done with drafts already in hand — so the first build of the new unit is also the byte proof — it went 137/137. The ordering, not the carve tooling, was the variable. +- **Not banked — greps:** `grep -n -i 'populated-at-carve|stubs-in-new-object|draft-first ordering'` → 0; `grep -n -i '137/137'` → 6 (all §94/§100 type-carry family sweeps, unrelated); `grep -n -i 'carve'` over `decomp-kernels.md` → 8 (carve chain, carve-refused verdicts, gitignored registry — the ordering rule is absent). +- **Proposed home:** DK (a kernel) or accelerator — *sequence layout surgery after the drafts that will fill it; an empty new unit has no oracle.* +- **Portable because:** it is a general refactor law under a byte gate — any change to how code is grouped into objects must land together with content that proves the grouping. + +### C5 — A cracked idiom transfers WITHIN its family and not across it: price a lane by families, not by class size +- **Evidence:** `phase-ends/logs/Phase31.md:1300-1302` — "**§206 transfers WITHIN a family, NOT across.** exemplar 40 turns/5 oracle -> within-family **11 turns/3 oracle MATCH** -> cross-family **56 turns/0 compiles, FAILED**. So jtbl costs ~40 turns of learning **per family** (191 families), not per class. The jtbl quest is a project, not a lane — defer it." +- **What happened / what it cost:** A free-tier model cracked a jump-table exemplar and wrote a publishable idiom, which made the whole 245-member class look like a lane about to open. The transfer test — same idiom, a *different* family — burned 56 turns and produced zero compiling drafts. Multiplying the exemplar's 40-turn learning cost by 191 families rather than by 1 class turned a "lane" into a project and it was correctly deferred. The class was still open two sessions later. +- **Not banked — greps:** `grep -n -i 'transfers within|within a family, not across|cross-family'` → 0; `grep -n -i 'per family'` → 6 (§3-THE "all-or-nothing per family", `family_sweep` mechanics — about member banking, not about the learning cost of an idiom); `grep -n -i 'learning cost|learned on one|40 turns'` → 0. +- **Proposed home:** DK (a kernel), next to the model-ladder/economics kernels — *before scaling a lane off one cracked exemplar, run the same recipe on a member of a different family and count the turns; the ratio is the lane's real price.* +- **Portable because:** it is the general question of whether a learned technique amortises, and the cheap two-target test that answers it works for any idiom, any compiler, any harness. + +### C6 — Sibling count and never-drafted count are different denominators; conflating them overstated the free-remap leverage ~3× and hid that the remaining mass was singletons +- **Evidence:** `phase-ends/logs/Phase31.md:298-304` — "My '~3,900 siblings behind ~334 skeletons' conflated two different populations — the never-drafted stub count with the sibling count — and overstated remap leverage by ~3x. Most remaining work is singletons that each need their own crack." +- **What happened / what it cost:** The campaign's strategy rested on "the untouched functions sit BEHIND those skeletons and bank by mechanical remap once an exemplar cracks — they are not waiting for a draft" (`:465-467`). An independent read-only audit at session end measured the atlas honestly: 1,292 of 1,772 groups were singletons carrying **57% of the open instruction mass**, and only 480 groups were multi-member. The wrong denominator had been steering wave selection all session, and the corrected picture is what forced the pivot to the deterministic integration lane. +- **Not banked — greps:** `grep -n -i 'remap leverage|siblings behind|carrying 57|instruction mass'` → 0; `grep -n -i 'conflated'` → 5 (unblocked-vs-unmatched, hash confusion, a stale ×1 handoff — different conflations); `grep -n -i 'two populations'` → 1 (the wave exclude list's two unbankable populations); `grep -n -i 'singleton'` → 8 (accelerator #17 on the similarity join's *tiers*, and a wave-playbook note — none about the leverage arithmetic). +- **Proposed home:** DK (a kernel) or accelerator — *state the group-size distribution and the singleton share of instruction mass before pricing any "one crack banks N" strategy; the two counts look alike and differ by 3×.* +- **Portable because:** every decomp with duplicated code plans around family leverage, and the same two counts (instances vs. never-attempted) exist in every such project. + +### C7 — "Never write the tree" in a drafting prompt is a request, not an enforcement +- **Evidence:** `phase-ends/logs/Phase31.md:2879-2882` — "**A wave agent WROTE to `src/` and then `git checkout`-reverted it** … It was harmless ONLY because R42 meant every bank was already committed. The prompt's 'never modify src/' is a request, not an enforcement — candidate: *a drafting agent must run where it cannot write the tree, or the harness must detect and refuse the write*." +- **What happened / what it cost:** It happened at least twice in the slice. The earlier instance (`:1912`) had a wave agent write its body straight into `src/boot.c`; a dirty tree ABORTS the gate for the whole wave, so one agent's stray write can void every other agent's work in that wave. The response both times was to harden the prompt's wording — which is exactly the fix that cannot hold, since a later wave did it again despite the hardened prompt. The second incident was harmless only by accident of the commit-immediately rule. +- **Not banked — greps:** `grep -n -i 'never write into the tree|never modify src|request, not an enforcement|cannot write the tree'` → 0; `grep -n -i 'dirty tree ABORTS'` → 0; `grep -n -i 'sandbox'` → 6 (all about re-probing a CC1-FAIL wall in a sandbox TU, and worktree isolation making a worker see *less* — none about denying a drafter write access). +- **Proposed home:** G (a rule) — drafting agents run read-only or in an isolated copy; the harness detects and refuses a write to the shared tree rather than asking the model not to. +- **Portable because:** it is a harness-architecture rule for any multi-agent setup where a shared build tree is the verification surface, and prompts are never enforcement anywhere. + +### C8 — A free or preview model tier can be withdrawn mid-campaign without notice; a fleet-wide 404 is an epoch event, not N model failures +- **Evidence:** `phase-ends/logs/Phase31.md:2319-2321` — "**THE OX WINDOW CLOSED — 2026-08-26 07:55** (probed: HTTP 404 on stealth/ox-alpha; stealth/* gone from the model list). The free-drafting era ended mid-m0b (its 29 agents all 404'd at turn 0 — a harness-epoch event, not 29 model failures, R40)." +- **What happened / what it cost:** The campaign's whole drafting economy had been built on a free stealth model over several sessions. When it vanished the in-flight wave's 29 agents lost the race "by a minute" (`:2337`) and drafting anything new immediately required a paid decision that had to go to the human. What saved the position was that the session had *raced the door deliberately* — every never-attempted function was drafted before the window closed (196/196 overlays, 196/225 main). The R40 attribution point is banked; the planning point (a free tier is a closing window, so spend it on coverage of never-attempted work rather than on refinement) is not. +- **Not banked — greps:** `grep -n -i 'window closed'` → 2 (one about a shipping decision, one a provenance line "the night the ox window closed" with no lesson); `grep -n -i 'free model|free-drafting|delisted|vanish|provider withdr'` → 0 relevant; `grep -n -i 'stealth'` → 1 (a provenance credit only). +- **Proposed home:** DK (a kernel) beside the economics kernels, or accelerator — *treat a free/preview tier as a window with an unknown close date: spend it on never-attempted coverage, keep a paid fallback configured, and read a fleet-wide provider error as an epoch event before reading it as failure.* +- **Portable because:** free and preview model tiers are a normal part of any current agent budget, and they are withdrawn without notice everywhere. + +### C9 — A metered API key's own cap is a separate limit from the account's credit +- **Evidence:** `phase-ends/logs/Phase31.md:2380-2383` — "**DeepSeek push paused by a KEY CAP, not the budget (10:4x, R40-corrected):** the OpenRouter key carries a $60 LIFETIME limit; usage hit $60.21 mid-ds2 and every request 403s ('Key limit exceeded (total limit)') while the ACCOUNT still holds ~$10.6 credit." +- **What happened / what it cost:** A paid drafting run died mid-wave and "two wrong theories were burned en route" — resumed-turn budgets and a generic harness fault — before the 403 body was read (`:2386`). The account balance, the obvious thing to check, was healthy the whole time. +- **Not banked — greps:** `grep -n -i 'lifetime limit|key limit|lifetime cap|account still'` → 0; `grep -n -i '403'` → 5 (all cookbook section numbers §403, unrelated). +- **Proposed home:** accelerator (ops) — record every spend limit that exists at each layer (key, account, org, per-model) at setup time, and read the provider's error *body* before theorising. +- **Portable because:** multi-layer spend caps are standard on every commercial model gateway, and the symptom (a hard stop with credit remaining) reads as a harness bug. + +### C10 — Incremental gates never exercise the extraction/regeneration step, so regeneration rot is undated and invisible for weeks +- **Evidence:** `phase-ends/logs/Phase31.md:2376-2378` — "The general lesson repeats §61c with a new edge: INCREMENTALLY-GREEN HIDES EXTRACT ROT — a periodic `make extract` sweep per binary (not just check-all builds) would have dated defect (a) precisely." +- **What happened / what it cost:** One binary was found carrying two stacked defects; the older one was "an extract/ld_interleave inconsistency of UNKNOWN, older date (never surfaced because nothing ran a full extract for this binary between its introduction and this morning's probes — gates build incrementally)" (`:2370-2372`). Five surgical repair attempts were spent partly because the defect could not be dated. The same shape recurred at the next full sweep: `make clean → extract-all → check-all` "failed SIX binaries that no incremental gate had flagged" including one `[EXTRACT FAIL]` that only appeared under the parallel extract (`:2545-2551`). A related ops hazard from the same task: "never read `asm/` while a sweep's extract-all runs" (`:2601`). +- **Not banked — greps:** `grep -n -i 'extract rot|incrementally-green|hides extract'` → 0; `grep -n -i 'incremental build'` → 6 (all R22/§130 — *verify a match from a clean rebuild*, which is the build half); `grep -n -i 'full extract|regeneration'` → 4 (none about periodically re-running the extraction step as its own check). +- **Proposed home:** accelerator, as a second row on the two-independent-paths table (`docs/accelerators.md:338`) — *"is the input tree still derivable?" path A: the incremental gate (never runs it) · path B: a periodic per-unit re-extract.* +- **Portable because:** every decomp has a generation step (splat/extract/disassembly) upstream of the build that the verification loop never re-runs; the same is true of any codegen'd input in any project. + +### C11 — A probe that splices the tree in order to measure it must restore it, or the progress oracle counts the splices as banks +- **Evidence:** `phase-ends/logs/Phase31.md:2598-2601` — "Instrument lessons of the task: a stale classification file re-labels every fn 'CARVE-REFUSED' until re-gated (verify from the gate, not the ledger); an autopsy script without its scratch dir leaves TUs spliced and `corpus` then reads them as banked (R40 twice tonight); never read `asm/` while a sweep's extract-all runs." +- **What happened / what it cost:** The project's matched-ness oracle derives from the source tree, so a diagnostic tool that substitutes a draft into a TU and then dies makes the progress counter report banks that never happened — twice in one task, each time diagnosed as a subject failure first (R40). This is distinct from the banked "shared scratch directory is a shared blast radius": here the tool corrupts *the metric*, not another agent's files. +- **Not banked — greps:** `grep -n -i 'leaves TUs spliced|reads them as banked|left the TU'` → 0; `grep -n -i 'scratch dir'` → 5 (the agent-verdict-vs-scratch-dir oracle pair; the shared-scratch blast radius — neither is the self-poisoning-metric case); `grep -n -i "an instrument's own write path"` → R57 in `phase-ends/DIGEST.md:242`, which covers a *repair* tool corrupting what it measures — the read-only-probe variant that fakes progress is not stated. +- **Proposed home:** accelerator, or a sharpening note under R57 — *a probe that mutates the measured tree restores it in a trap/finally, and the progress counter is re-read after any aborted probe.* +- **Portable because:** every project whose completion metric is derived from the working tree can have that metric written by a crashed tool. + +### C12 — Keep prompt/law text in data, not inside the launcher's source template +- **Evidence:** `phase-ends/logs/Phase31.md:1702-1703` — "A wave script's LAWS block is a JS template literal: **backticks inside the law text terminate it**. Wave U's first launch died on exactly that; write law text with single quotes." +- **What happened / what it cost:** The wave launcher embedded the drafting laws as a JS template literal, so a backtick in a law (i.e. any inline code span, which technical law text is full of) terminated the string and killed the launch. It was hit at least twice and is carried as a standing watch-for (`:1489`, `:1606`). +- **Not banked — greps:** `grep -n -i 'template literal'` → 0; `grep -n -i 'backtick'` → 2 (both about outward-facing prose style — "no backtick-bracketing" in upstream PR text — unrelated); `grep -n -i 'LAWS block|law text'` over `docs/wave-playbook.md decomp-architect/corpus/decomp-kernels.md docs/how-to-ai-decomp/05-cards-lanes-waves.md` → 0. +- **Proposed home:** accelerator (minor), or the wave playbook's launch section — the knowledge fed to agents lives in a data file the launcher reads, never inline in the launcher's own syntax. +- **Portable because:** every agent harness assembles prompts in some host language, and technical prompt text always contains that language's string delimiters. + +## ALREADY-BANKED (one line each) +- A byte-diff against a build that failed to include your draft always reads MATCH — lives at `docs/matching-cookbook.md:21583` +- The null-draft control: a defect that reproduces with zero drafts is not caused by drafts — `docs/matching-cookbook.md:16836` +- A detector is advisory, the whole-binary gate is the arbiter; never withhold a standalone-MATCH draft on a flag alone — `docs/matching-cookbook.md:16815` +- Bisect with a null control first; a true binary search beats the built-in linear bisect (7 steps/176 s vs 3 hours) — `docs/matching-cookbook.md:17966` +- The card is the cheapest place in the pipeline to put a fact; every field costs zero tokens per wave forever — `docs/matching-cookbook.md:20594` +- Before adding an agent, an attempt or a prompt paragraph, ask what the tree already knows — `docs/matching-cookbook.md:20593` +- A project idiom the tools cannot parse is an idiom that silently costs work — `docs/matching-cookbook.md:21184` +- A script that parses argv at import cannot be shared; extract the predicate — `docs/matching-cookbook.md:30935` (§323b) +- A 0% from a broken tool and a 0% from a working one are the same number and opposite facts — `docs/matching-cookbook.md:4125` +- Every cost/rate/yield/effort number ships with its denominator (covers "a sampling filter ships its denominator") — `phase-ends/DIGEST.md:224` (R41) +- The gate spends ~3 whole-binary builds per failing draft; gate cost is measured in groups/wall-clock — `docs/decision-log.md:2504`, `docs/how-to-ai-decomp/09-economics.md:88` +- Gate cost scales with (binary, TU) groups, not drafts; concentration is a draw-time choice — `docs/how-to-ai-decomp/09-economics.md:88`, `decomp-architect/corpus/decomp-kernels.md:541` (DK-42) +- The harvest's biggest number is "already covered" — that is a RETRIEVAL problem, not a knowledge gap — `docs/matching-cookbook.md:18475` +- Reconcile before the first gate; a parked draft gets harder to bank, not easier — `docs/matching-cookbook.md:17087` (§3-C2), `decomp-architect/corpus/decomp-kernels.md:526` (DK-41) +- One adversarial verifier per candidate, defaulting to REJECT; a law from one wave is a first draft — `docs/matching-cookbook.md:18470` +- Stopping a wave mid-flight costs the in-flight tail (and how much is recoverable) — `docs/matching-cookbook.md:17160` (§176j) +- Closeness must be counted, not read off the first differing index — `docs/matching-cookbook.md:17198` +- A clean pre-gate check is a licence to build, not a prediction of success — `docs/matching-cookbook.md:17155` +- A mask coarser than the linker's is not weak evidence, it is no evidence — `docs/matching-cookbook.md:18206` +- A guard downstream of the failure is not a guard, and a guard that is not running is not a guard — `phase-ends/DIGEST.md:239` (R54) +- The straggler tail: a batch cannot gate until its slowest agent lands; stream instead — `docs/how-to-ai-decomp/05-cards-lanes-waves.md:107`, `docs/wave-playbook.md:441` +- The draw sets the request rate; card supply, not model capacity, is the binding constraint; measure the steady state not the launch — `docs/how-to-ai-decomp/08-models-and-budgets.md:31-36` +- An option nothing ever passes is a dead feature (`match_one --o0` existed for a year) — `docs/matching-cookbook.md:24851` +- A tool's arm that has never fired over N opportunities (`decl_prior` %hi/%lo, 0 of 1,210) — `docs/matching-cookbook.md:21879` (§204-E) +- The permuter's search space excludes external-callee/arity residuals, and it rejects register pins — `docs/how-to-ai-decomp/07-compiler-source.md:48` +- Bank rate by function size (the sub-50 economics and the cheap-tier price per function) — `docs/how-to-ai-decomp/09-economics.md:22`, `decomp-architect/corpus/decomp-kernels.md:555` (DK-43) +- A running lane script does not read your edit; know which of code/args/defaults takes effect when — `docs/accelerators.md:217`, `decomp-architect/corpus/decomp-kernels.md:310` +- `pgrep -f` self-matches; the bracket trick protects only the pattern, not the command line — `docs/matching-cookbook.md:17990` (§180d), `docs/accelerators.md:196` +- A backlog draft path is not a stable original; snapshot the text you mean to re-gate — `docs/matching-cookbook.md:21569` +- Batched drafts must agree with EACH OTHER, not just with the file — `docs/matching-cookbook.md:16848` +- A TU retype is byte-neutral only if every existing use site still compiles; edit the side that is cheap to verify — `docs/matching-cookbook.md:18112` (§185) +- The A10 stored-verdict law: a blind stored-draft re-gate converts at ~0–8%, so fresh-fix lanes outrank resurrection lanes — `docs/matching-cookbook.md:16948` +- A caller-saved register pin can be a correctness bug, not just a scheduling choice (it silently deletes a store) — `docs/matching-cookbook.md:16776` (§175) +- A lane banking far below the rest is a harness fault until proven otherwise; every status check covers every lane — `docs/how-to-ai-decomp/04-oracles-and-instruments.md:86`, `decomp-architect/templates/registry-E.decomp.md:222` +- A blanket refusal on a chronically-true condition silently disables the check (the fleet sweep that never ran; the draw that always fell back) — `docs/how-to-ai-decomp/12-failure-museum.md:34` (row 26) +- A workflow killed by a usage limit resumes with `resumeFromRunId`; never rebuild a wave by hand — `docs/matching-cookbook.md:20755` +- Matching is solved; integration/plumbing is the dominant spend (≈92% of drafts byte-correct, ≈27% banked) — `docs/how-to-ai-decomp/09-economics.md:103`, `docs/retrospective.md:82` +- Harness defects wearing model-failure costumes were about half a late session — `docs/how-to-ai-decomp/04-oracles-and-instruments.md:111`, `docs/accelerators.md:351` +- The S60/S61 rule block (R44–R60): card levers must resolve in the knowledge base, draw-time bankability, budget per lane, consume every verdict layer, never key by bare fn name, a soft error in a success envelope, periodic whole-fleet verification, derived-property-as-config staleness, blanket committers, verify a build from its exit code, unattended lanes leave evidence, gate verdicts need a green baseline, an instrument's write path, session-close "clean" quotes the fleet's green count, blanket-committing another lane's mid-gate tree, carve-state files — `phase-ends/DIGEST.md:228-246` +- A mis-scoped sweep's small yield is not evidence (3 vs 50 banked, same tree, same day) — memory `crack-wave-sweep-map-regen`; `docs/accelerators.md` "silently narrowed tool scope" family diff --git a/.run/P33.5/log-mining/Phase31-2of3.md b/.run/P33.5/log-mining/Phase31-2of3.md new file mode 100644 index 0000000000..6cb6274d60 --- /dev/null +++ b/.run/P33.5/log-mining/Phase31-2of3.md @@ -0,0 +1,166 @@ +# Log mining — Phase31-2of3 +Files/ranges: `phase-ends/logs/Phase31.md:2903-5804` (context read 2873-5834) · Lines read: 2902 of 2902 (assigned range; 2962 including the ±30 context) +Candidates considered: 52 · NEW: 11 · ALREADY-BANKED: 41 + +**Grep harness.** `$D` below = the ten distilled paths named in the BRIEF: +`docs/matching-cookbook.md docs/cookbook-index.md docs/decision-log.md docs/accelerators.md docs/retrospective.md docs/wave-playbook.md docs/how-to-ai-decomp/*.md decomp-architect/corpus/decomp-kernels.md decomp-architect/templates/registry-E.decomp.md phase-ends/DIGEST.md`. +Harness verified against a known-true case first (`grep -c -i 'byte gate' $D` → 41 across 12 files; +`grep -n -i 'exonerate the instrument' $D` → 5) before any zero was believed. + +--- + +## NEW + +### C1 — The health suite must assert that a tool DID ITS WORK, not only that the data is intact: zero inputs, an impossible wall-clock and a missing persistent effect are each a DEFECT, not a result +- **Evidence:** `phase-ends/logs/Phase31.md:5122-5128` — + > `* assert_inputs — zero readable inputs is a DEFECT, not a zero-yield result. **"0 of 0" is a fact` + > ` about the HARNESS; "0 of 57" is a fact about the SUBJECT**, and reporting the first as the second` + > ` cost a 35-binary batch. * assert_floor — work claiming a compile/gate cannot beat physics.` +- **What happened / what it cost:** `parallel_gate` resolved a relative `--drafts` path inside its worktree, found 0 drafts, banked 0 and exited **rc=0** — 35 binaries / 57 drafts "banked 0" in **1–2 seconds each** while the same drafts gated in-tree banked 15/16 (`:4901-4909`). The only tell was the runtime. `make tools-health` ran five audits and **every one asserted DATA integrity; none asserted that a tool did the work it claims** (`:5106-5108`). The answer was one module — `assert_inputs` / `assert_floor` / `assert_effect` — negative-controlled 11/11 in both directions and wired into the gate tools and the health target (`:5122-5136`). +- **Not banked — greps:** `grep -n -i 'work_evidence' $D` → 0; `grep -n -i 'assert_floor' $D` → 0; `grep -n -i 'cannot beat physics' $D` → 0; `grep -n -i 'fact about the harness' $D` → 0; `grep -n -i 'too fast' $D` → 0; `grep -n -i 'zero readable inputs' $D` → 0. (`grep -n -i '0 of 0' $D` → 1, cookbook:29606, a different subject — an agent wave filing 0 verdicts.) Accelerator #15's differential-oracle harness runs the same QUESTION down two paths; it has no "did the tool run at all" pair, and its table (`docs/accelerators.md:330-337`) does not contain one. +- **Proposed home:** DK (a kernel, section 2 "Instruments"), with a G-rule pointer from G19/G28 +- **Portable because:** every decomp harness has tools whose "0 banked / exit 0" is indistinguishable from real work; inputs, a runtime floor and a persistent trace are assertable on any platform, from day one. + +### C2 — A health check that cannot finish is not a check: keep the health target sampled and fast, and put the exhaustive form behind its own name +- **Evidence:** `phase-ends/logs/Phase31.md:5175-5181` — + > `**S70 — `make tools-health` WAS UNRUNNABLE AND IS NOW 333s GREEN.** … `audit-cdecl` re-parsed` + > `**every declaration in all 4,168 TUs** … ~787s of pure-Python collection before the first cc1` + > `call … `--limit` already existed, its own help calls it "a fast smoke run", and nothing used it.` +- **What happened / what it cost:** a full-corpus regression test was living inside a health target, so `tools-health` **had never once completed** in the project's life; sampling it by default took it to 61 s and the whole target to 333 s green. The companion defect: the sample takes the FIRST 60 translation units, so the smoke run always tests the same files (`:5239-5240`), and parallelising the target exposed a latent fixed-temp-filename race that was safe only while serial (`:5183-5187`). +- **Not banked — greps:** `grep -n -i 'tools-health' $D` → 5 (all are "X now runs in tools-health", none about its runnability); `grep -n -i 'never completed' $D` → 0; `grep -n -i 'health target' $D` → 0; `grep -n -i 'smoke run' $D` → 0. +- **Proposed home:** DK (a kernel) or accelerator; the G-rule neighbourhood is G30 ("a guard not running is not a guard") +- **Portable because:** every decomp accumulates audits; the first one that costs ten minutes silently converts the whole suite into something nobody runs, and the fix (sample by default, exhaustive form as a separate target, randomise the sample) is platform-free. + +### C3 — "Independent" names the INSTRUMENT, not the input: two refusals of two separately-written drafts from one tool is one test repeated +- **Evidence:** `phase-ends/logs/Phase31.md:3589-3590` — + > `**"Independent" means a DIFFERENT INSTRUMENT, not a different input.** Two runs of one tool on two` + > `drafts is one test repeated. R40, sharpened.` +- **What happened / what it cost:** `ov_SC06_022:func_80181664` was ledgered a WALL after "two independent gate refusals" — **both came from `parallel_gate`, the one gate that structurally cannot host a jump-table carve** (its worktree `asm/` is a symlink). Running the class through the SERIAL gate banked **16 of 32**, including functions `parallel_gate` had refused (`:3582-3588`, `:3671-3676`). +- **Not banked — greps:** `grep -n -i 'different instrument' $D` → 2 (both unrelated: a compiler-dump battery, and a cookbook lever "same instrument as §153"); `grep -n -i 'not a different input' $D` → 0; `grep -n -i 'two runs of one tool' $D` → 0; `grep -n -i 'two independent gate refusals' $D` → 0. Adjacent but not the same: G21 (a second *disagreeing* oracle) and G33 (a verdict names its instrument) both leave "two refusals from one gate" reading as corroboration; `docs/decision-log.md:1196` asks whether corroboration came "through the same tool you just fixed" — a tool that CHANGED, not a tool that was always the wrong instrument for the class. +- **Proposed home:** G (sharpen G21/G33 with the independence test) and DK-17 +- **Portable because:** it is the acceptance criterion for every "wall" verdict on any project — count instruments, not attempts. + +### C4 — A drafting agent must run where it CANNOT write the source tree; "never modify src/" in a prompt is a request, not an enforcement +- **Evidence:** `phase-ends/logs/Phase31.md:2879-2882` — + > `* **A wave agent WROTE to `src/` and then `git checkout`-reverted it** (SYS_OBJ_2264, self-reported).` + > `It was harmless ONLY because R42 meant every bank was already committed. The prompt's "never modify` + > `src/" is a request, not an enforcement` +- **What happened / what it cost:** a drafting agent edited a live source file and then reverted it with a blanket `git checkout` — the exact reflex that destroyed 61 banks in an earlier session. It cost nothing only because every bank was already committed; the agent self-reported it, so the harness had no independent way to know. Raised as a rule candidate twice (`:2879-2883`, `:2986-2988`) and never ratified. +- **Not banked — greps:** `grep -n -i 'must not be able to write' $D` → 0; `grep -n -i 'cannot write the tree' $D` → 0; `grep -n -i "agent .{0,20}wrote to .{0,10}src" $D` → 0; `grep -n -i 'read-only' $D` → 6 (the purge probe R57, read-only survey agents, "never an agent that…" in chapter 08 — none grants or denies a drafting agent write access to the tree). DK-35 bounds an agent's *cleanup* to its own directory; it does not deny write access to the source tree. +- **Proposed home:** G (a conduct/harness rule beside G47/G48) and DK-35 +- **Portable because:** any harness that fans out drafting agents over a repository can give them a sandbox, a worktree or a read-only mount; the alternative is trusting a paragraph of prose with the only copy of the banked work. + +### C5 — A proven transform that is not a rung of the ladder the agents' drafts actually pass through does not exist for those drafts +- **Evidence:** `phase-ends/logs/Phase31.md:3810-3814` — + > ``scope_data_externs.fix()` (§8d) has been byte-proven since Phase 26 and is used by `family_sweep` /` + > ``bank_exemplar` / `jtbl_family_bank` — but **nothing in `gate_stage`'s ladder ever called it**, so a` + > `draft written by a wave agent had never seen it. … went from **21 residuals to ZERO**.` +- **What happened / what it cost:** a fix proven five phases earlier and wired into three mechanical tools was absent from the one path every agent-written draft travels. Wiring it as an `_xform` rung took the `conflicting types for D_*` class from 21 residuals to zero, and recovered automatically the very function a previous session's audit had named as its byte-proven instance. +- **Not banked — greps:** `grep -n -i 'ladder ever called' $D` → 0; `grep -n -i 'not wired into' $D` → 0; `grep -n -i 'every proven transform' $D` → 0; `grep -n -i 'scope_demote' $D` → 4 (all usage instructions: which rung to run for which symptom, never the audit lesson); `grep -n -i 'residuals to zero' $D` → 0. Adjacent: DK-24 / G40 require every *verdict layer* to be consumed — this is about a *repair* that no lane applied. +- **Proposed home:** DK (extend DK-24 from verdict layers to repair rungs) or G40 +- **Portable because:** every decomp accumulates deterministic repairs faster than it wires them; the audit ("is each proven transform reachable from the pipeline the agents' output enters?") is a one-time enumeration on any project. + +### C6 — The knowledge-harvest selector must be able to see FAILED attempts; a filter that can only read banked work learns from the easy half +- **Evidence:** `phase-ends/logs/Phase31.md:3236-3240` — + > `**The distill novelty selector was INVERTED** … `'no cookbook lever'` matched "no cookbook lever` + > `*needed*" (a TRIVIAL note) and was the only pick of 24 … **and it structurally could not see` + > `UNBANKED functions at all**, which is where the hardest functions write their richest notes` +- **What happened / what it cost:** the flywheel's own harvesting tool selected exactly the wrong note out of 24 and was blind by construction to every failed attempt — the population that produces the new laws. Fixed with `--with-unbanked`, carrying `banked=False` through to the verifier so what comes from a gate-refused draft is banked marked UNPROVEN rather than dropped (`:3236-3240`, `:3495-3496`). +- **Not banked — greps:** `grep -n -i 'only ever learns' $D` → 0; `grep -n -i 'unbanked' $D` → 6 (all are counts of unbanked drafts in specific waves, not the selector lesson); `grep -n -i 'novelty' $D` → 6 (harvest rounds, false-novelty claims, the "retrieval not content" reading — none names the selector's blindness). Note the tension worth recording with it: **G45** says "harvest only from byte-proven results", which is right for what may ENTER the base and wrong as a filter on what may be READ. +- **Proposed home:** G (a qualifier on G45) + DK-32 +- **Portable because:** every project that has agents write notes will build a selector over them; if it keys on success it will never see the wall classes, which is the only place new laws come from. + +### C7 — A refusal names the branch the caller entered, not the subject — make the applier consult the classifier it already has +- **Evidence:** `phase-ends/logs/Phase31.md:3346-3348` — + > `**None was a real wall. Each was a tool describing its own confusion in the language of a limit** —` + > `§401's law … generalised: *a refusal names the branch you entered, not the function you asked about.*` +- **What happened / what it cost:** the §322b carve route had **three stacked blockers**, and three full gate passes booked the same functions CARVE-REFUSED. The terminal refusal ("the carve model covers jump tables only, not an island of mixed included data") read as a permanent toolchain wall and was a ROUTING error: `island_probe` **already classified that same function `'tail'`** and its own detail named the right lane, but the refusal path never consulted it (`:5324-5338`). Consulting the probe first banked, immediately, the function three gate passes had refused. The same shape appeared inside the tools: "the tool's own error message pointed at the layer that was working" (`:5594-5595`). +- **Not banked — greps:** `grep -n -i 'names the branch' $D` → 0; `grep -n -i 'the language of a limit' $D` → 0; `grep -n -i 'routing error' $D` → 0; `grep -n -i "wearing a .{0,20}wall" $D` → 0. Adjacent and inverse: DK-26 makes the cheap probe call the real planner; this is the applier failing to call the probe. +- **Proposed home:** DK (pair it with DK-26) and the failure museum +- **Portable because:** every decomp pipeline grows multi-branch appliers with a cheap classifier beside them; the class of "wall" that is really an unrouted branch is universal, and the audit is to make every refusal name the branch and cite the classifier's verdict. + +### C8 — Every status claim in an agent's context must be expiry-checked against live state, or agents will report it back to you as an observation +- **Evidence:** `phase-ends/logs/Phase31.md:5394-5396` — + > `**stale BASELINE-RED verdict** — OPEN. THREE agents reported a red baseline on binaries I verified` + > `BYTE-IDENTICAL; the third revealed the source: *"the pack's last gate verdict was BASELINE-RED"* —` + > `they read it off the card. Cost me two phantom-regression chases.` +- **What happened / what it cost:** a historical gate verdict printed on a wave card came back as three independent agent reports of a live regression, and cost two phantom-regression investigations. The shipped repair (`:5511-5524`) reads the same live red union the gate itself consults and reports an expired claim as EXPIRED — measured over the real ledgers: **173 expired claims retired, 981 pairs gained a best-measured residual they were previously denied**. +- **Not banked — greps:** `grep -n -i 'BASELINE-RED' $D` → 4 (all about `gate_stage` refusing red-listed binaries' drafts — the producer side, never the card republishing a stale verdict); `grep -n -i 'read it off the card' $D` → 0; `grep -n -i "stale .{0,15}card" $D` → 0. +- **Proposed home:** G (extend G44 — a card names only what the knowledge base contains — to "and every status claim on it is expiry-checked") + DK-39 +- **Portable because:** any agent given a status line will treat it as an observation and report it; the fix is that a card carries no present-tense claim it cannot re-derive at build time. + +### C9 — A validity stamp must be honoured by every downstream consumer; and a wrong write-side label cannot be repaired by a correct read key +- **Evidence:** `phase-ends/logs/Phase31.md:5437-5442` — + > ``gate_feedback` gated on `shape=='MATCH'` when reloc_identity's binding condition is **`aligned`** …` + > `Below that bar reloc_identity itself downgrades status to `MISMATCH?` and stamps the row **ADVISORY**` + > `— and gate_feedback republished it as a binding per-index instruction.` +- **What happened / what it cost:** **15 of 130** wave cards carried the block; **15 of 15 were not index-aligned**, so each "the target references 0x…" line was arithmetic against a different stream — **55 of 66 printed lines (83%) named a value that is not an address**, and 4 of 4 checkable cases named symbols the target never relocates. The second half stayed open upstream: the row's binary was stamped from a bare-name dict, so 2,317 of 42,655 drafts (5.4%) carried the wrong binary — **"a correct read key cannot repair a wrong write-side stamp"**, which is why the shipped fix validates against the target's own bytes rather than the label (`:5447-5451`, `:5481-5489`). +- **Not banked — greps:** `grep -n -i 'advisory' $D` → 3 (cookbook §176e's own producer-side "mark advisory when the shapes differ", and a rewrite rule — none about a consumer stripping the stamp); `grep -n -i 'republish' $D` → 0; `grep -n -i 'write-side' $D` → 0; `grep -n -i 'aligned=False' $D` → 0. G41 banks the bare-name key; it does not bank the consumer-side stamp obligation. +- **Proposed home:** DK (instruments) or a G-rule beside G41 +- **Portable because:** any pipeline that computes a hint with a precondition will eventually have a consumer that reads the value and drops the precondition — and the consumer here is the model's own context. + +### C10 — Yield is clustered by binary, not spread over the fleet: draw per-binary once two independent lanes concentrate in the same place +- **Evidence:** `phase-ends/logs/Phase31.md:5040-5043` — + > `**20 of 52 twin remaps banked (38%)** … and **all 20 in `ov_SC06_011`**, the same binary that carried` + > `15 of the 21 standalone banks. Two lanes, same concentration … the fleet's remaining work is NOT` + > `uniformly distributed. **Draw future waves per-binary, not fleet-wide.**` +- **What happened / what it cost:** **35 of the session's 44 banks came from one binary**, reached by two mechanically independent lanes (a standalone-match sweep and a twin-remap sweep). The forecast built on the first cluster probed had been ~45–55 banks from 86 candidates and delivered 22, because that cluster was the least representative (`:4880-4882`). The next-session instruction became "find the next cluster by twin density + open-stub count per binary rather than drawing fleet-wide" (`:5235-5236`). +- **Not banked — greps:** `grep -n -i 'per-binary, not fleet-wide' $D` → 0; `grep -n -i 'not uniformly distributed' $D` → 0; `grep -n -i 'clustered, not uniform' $D` → 0; `grep -n -i "draw .{0,12}per.binary" $D` → 0; `grep -n -i 'concentrat' $D` → 6 (DK-42 / chapter 09 concentrate targets per binary to cut **rebuild wall-clock** — a cost argument, not a yield-distribution argument). +- **Proposed home:** DK-42 (add the yield half) or the wave playbook's draw step +- **Portable because:** on any fleet of overlays/objects sharing an engine, the residue concentrates where one module's tail is nearly finished; a per-binary draw finds it and a fleet-wide draw averages it away. + +### C11 — (minor) A build-system conditional that expands at PARSE time makes its own negative control vacuous +- **Evidence:** `phase-ends/logs/Phase31.md:5155-5158` — + > `*Make gotcha worth keeping: my first patch used `ifeq ($(filter $*,...))`, which make evaluates at` + > `PARSE time when `$*` is empty — it would have silently always taken the maspsx branch and the` + > `"byte-inert" result would have been vacuous. `$(if ...)` expands per-target*` +- **What happened / what it cost:** the per-object assembler switch that unlocked six "compiler wall" functions was verified BYTE-INERT against the whole fleet — a verification that would have proved nothing, because the conditional selecting the new branch could never be true. Caught before shipping; the rule's value is that the *control* was the thing at risk, not the change. +- **Not banked — greps:** `grep -n -i 'parse time' $D` → 0; `grep -n -i 'ifeq' $D` → 1 (a cookbook line describing an existing `ifeq ($(BINARY),main)` block, not the hazard); `grep -n -i 'vacuous' $D` → 4 (P31 S58's vacuous *checks* and P29's vacuous *probes* — neither is the make-expansion-time hazard). +- **Proposed home:** accelerator / cookbook (build-system section) +- **Portable because:** every matching decomp drives a per-TU flag matrix through make, and a byte-inert claim measured under a branch that never fires is the cheapest possible false green. + +--- + +## ALREADY-BANKED (one line each) + +- A self-report is not a bank; read the judge's artifact, never the agent's note — `decomp-architect/corpus/decomp-kernels.md:868` (DK-68); `docs/how-to-ai-decomp/12-failure-museum.md` #33 +- A standalone closeness-0 proves the BODY, never that the TU accepts the SIGNATURE (§376) — `docs/accelerators.md:362` (#16), DK-20, G10 +- A sampling filter ships with its denominator; a headline % ships its remainder — `decomp-architect/templates/registry-E.decomp.md:183` (G27), DK-30 +- A stage that REWRITES the artifact it measures must gate the ORIGINAL first (raw-first ladder, §313) — `docs/matching-cookbook.md:30199` +- A killed/aborted gate's dirty tree is adjudicated by the byte gate, never blind-reverted — `decomp-architect/templates/registry-E.decomp.md:238` (G37) +- The half-fix that manufactures a false wall ("K&R converts only 4/26" — those 22 were the half-fix) — `docs/matching-cookbook.md:30966` +- A fleet tool that composes a path from a binary NAME encodes the layout; pass the fact you have, never reconstruct it (§363) — `docs/matching-cookbook.md:31707-31722` +- A cheap probe that does not model the applier's real step (the carve) is optimistic; the gate is the arbiter — `decomp-architect/corpus/decomp-kernels.md:353` (DK-26); `docs/accelerators.md:287` (#14) +- A verdict produced inside a worktree describes the worktree; a gitignored input turns a whole class into "refused" — `decomp-architect/corpus/decomp-kernels.md:291` (DK-21); `docs/accelerators.md:462` (#19) +- A worker's per-function verdict rows die with its worktree unless copied out (produced-but-not-consumed) — `decomp-architect/corpus/decomp-kernels.md:377` (DK-28) +- The five things a fresh worktree lacks; negative-control an UNMODIFIED binary first — `docs/wave-playbook.md:620-628` +- A carve writes three things and the merge must carry all three; splice one binary's block of the shared make fragment — `docs/wave-playbook.md:626-631`; G43 +- Never `xargs -P` a tool that mutates shared state; one writer and one committer per shared file — `decomp-architect/corpus/decomp-kernels.md:316` (DK-23); `docs/accelerators.md:246` +- Never pattern-match the process table for a string your own command line contains; monitor by artifact — `docs/accelerators.md:196` (harness wound 2); DK-22 +- Launch detached and wait in a separate invocation — a shared process group hands the job to your tool timeout — `docs/wave-playbook.md:738-739` +- Read tool output unfiltered (`tail`), never through a keyword grep that can swallow a traceback — `docs/accelerators.md:203-217`; DK-22 +- Pass the payload the tool computed; never hand-type a path/arg into a workflow launch — `docs/wave-playbook.md` (`wave_args` step); G49 +- `-j` on every build; parallelism across binaries is a different knob — `decomp-architect/corpus/decomp-kernels.md:433` (DK-33) +- Distinguish "judged and failed" from "not judged" (the gater ledgering a draft at STAGE time, skipped forever after) — `decomp-architect/templates/registry-E.decomp.md:226` (G35); `phase-ends/DIGEST.md:247` (R61) +- A rate-limited `NO-DRAFT` is a harness verdict, not a verdict on the target — `docs/how-to-ai-decomp/12-failure-museum.md` #36; `docs/accelerators.md:322` +- The agents did not have the laws file (`SYS.md` one directory below where they were told to read) — `docs/how-to-ai-decomp/12-failure-museum.md` #21; DK-39 +- "No banked twin" on a card that is false of the world; rescan twins after every bank — museum #22; G44, G46 +- An exclude/wall list records what the TOOLING could not do and goes stale the day the tooling improves — museum #20; G38 +- The byte gate is structurally blind to LINKED library subsegs — a draft there gates GREEN while wrong — `docs/wave-playbook.md:87`; `docs/accelerators.md:310` +- A carve-config bank is RED until you re-extract, and that looks exactly like a false bank (§384) — `docs/cookbook-index.md:832` +- `match_one`/`rtu_match` compare `.text` ONLY, so a jump table is invisible to them (§405-A) — `docs/matching-cookbook.md:33099`, `:35010` +- A scan over `asm/` is a scan over UNMATCHED code only; "no banked function does X" is an empty-world answer — `docs/matching-cookbook.md:33206-33207` +- A declaration fix that only turns CC1-FAIL into DIFF has bought nothing; revert it — `docs/matching-cookbook.md:32995` +- The build is the batch verdict: one bad draft in a unit's slate fails its siblings — `decomp-architect/corpus/decomp-kernels.md:465` (DK-36) +- Difficulty is the residual class, not instruction count — route the model tier off history (§413) — `docs/cookbook-index.md:1787` +- Your own matched corpus is a labelled ground-truth corpus — mine the draft→final diff for the fix that actually worked, instead of authoring rules from prose (with the 1,062-sections-vs-13-coded-rules asymmetry and the fire-rate-by-band measurement) — `docs/decision-log.md:2624-2665` +- A shape census is symmetric and prices both polarities; re-derive the selector from the mine-vs-target residual (§406/§408) — `docs/decision-log.md:2666-2700` +- Read the recorded verdicts/journals before designing a new probe — `decomp-architect/templates/registry-E.decomp.md:166` (G24) +- A ledger's ordering/tie-break IS part of the instrument (recency sort discarding an earlier real measurement) — `decomp-architect/corpus/decomp-kernels.md:839` (DK-66) +- Pair every finding with an adversarial skeptic told to refute it — `docs/matching-cookbook.md:3708`, `:11051`; DK-43 +- A byte gate is a null oracle for "is this C?" (a draft that restores its own `INCLUDE_ASM` builds identical and counts as a bank) — `decomp-architect/corpus/decomp-kernels.md:384` (DK-29) +- Extrapolating a class's yield from n=1 / the least representative member — `docs/decision-log.md:1162`; `docs/matching-cookbook.md:9316` +- A restore that prints success while its consumer sees nothing; a confident zero from a fresh scanner — `docs/accelerators.md:297-337` (#15) +- Verify a fix FIRES before reporting it landed (a fix parsing a line the tool never prints) — `docs/how-to-ai-decomp/12-failure-museum.md` #33 (R66/R68); DK-68 +- Enumerate every lane that may write the tree before a fleet verification; liveness is recorded, never inferred — `decomp-architect/templates/registry-E.decomp.md:199` (G30, R54/R55) +- Harvest before the next wave, then toolify the mechanical idioms — `decomp-architect/templates/registry-E.decomp.md:285` (G45) diff --git a/.run/P33.5/log-mining/Phase31-3of3.md b/.run/P33.5/log-mining/Phase31-3of3.md new file mode 100644 index 0000000000..ed79152ef5 --- /dev/null +++ b/.run/P33.5/log-mining/Phase31-3of3.md @@ -0,0 +1,109 @@ +# Log mining — Phase31-3of3 +Files/ranges: `phase-ends/logs/Phase31.md`:5805-8705 (read from 5775 for context) · Lines read: 2931 of 2931 (assigned range 2901 of 2901) +Candidates considered: 30 · NEW: 7 · ALREADY-BANKED: 23 + +> `$F` below = `docs/matching-cookbook.md docs/cookbook-index.md docs/decision-log.md docs/accelerators.md +> docs/retrospective.md docs/wave-playbook.md docs/how-to-ai-decomp/*.md decomp-architect/corpus/decomp-kernels.md +> decomp-architect/templates/registry-E.decomp.md phase-ends/DIGEST.md`. Every grep was +> `grep -n -i -c -- '' $F` (hit counts per file), followed by reading the surrounding lines of any hit. + +## NEW + +### C1 — Never wrap a project tool in a `timeout` shorter than its own internal budget; you pre-empt its documented recovery handler and lose its buffered output +- **Evidence:** `phase-ends/logs/Phase31.md:7024-7027` — + "**Never wrap a project tool in a shorter `timeout` than its own budget.** My `timeout 2400` beat + `gate_stage`'s 3600 s budget, SIGTERM'd the tree mid-propagation, and Python lost its buffered + stdout — three gates with NO verdict and three half-applied, never-byte-gated propagations." +- **What happened / what it cost:** The session wrapped each gate in `timeout 2400` while `gate_stage`'s own budget was 3600 s. When `dedup_propagate` ran long the wrapper killed the tree; three gates produced no verdict at all and three propagations were left half-applied and never byte-gated (`:6891-6898`). `gate_stage` has a handler for exactly that state ("*the fleet is HALF-PROPAGATED and the tree is DIRTY. Revert, then re-gate with `--no-propagate`*") and the shorter external deadline pre-empted it. Recovery cost the evening; the documented recovery then ran clean, 5/5 in minutes each. +- **Not banked — greps:** `grep -n -i -c -- 'shorter timeout' $F` → 0; `grep -n -i -c -- 'timeout 2400' $F` → 0; `grep -n -i -c -- 'buffered stdout' $F` → 0; `grep -n -i -c -- 'own budget' $F` → 1 (cookbook:33494, about agents running to their own token budget); `grep -n -i -c -- 'SIGTERM' $F` → 1 (cookbook:2362, run `dedup_propagate` in the background — the opposite direction, no rule about nesting deadlines). +- **Proposed home:** DK (a kernel) — belongs beside the "a slow gate is a bug" / harness-hygiene kernels. +- **Portable because:** every decomp harness wraps long tools (gates, propagation, extraction) in supervisors, CI steps and agent timeouts; a nested deadline that fires before the tool's own defeats every graceful-recovery path the tool was given. + +### C2 — A programmatic edit to a long-lived knowledge document silently truncates or duplicates it; verify the sections, never the commit +- **Evidence:** `phase-ends/logs/Phase31.md:6529-6533` — + "**§429 had been SILENTLY DELETED from the cookbook.** My §428a rewrite wrote `t[:start] + new` + instead of `t[:start] + new + t[end:]`, truncating everything below it … **When editing a doc by + index slicing, re-read the tail.**" +- **What happened / what it cost:** One index-slice rewrite dropped the whole tail of the 3.5 MB cookbook; §429 was gone for the rest of the session and had to be restored from an old commit. The same class recurred twice more: `:7307-7308` "**§462, §463 and §464 silently vanished from the cookbook** after their commits. Restored; … **Verify each section, not the commit.**", and `:8351-8352` "the S78 and S79 blocks share every section heading, and `str.index` on a heading **duplicated a region twice this session**". Sections written that day were the same ones cracking functions that afternoon, so the loss was live knowledge, not archive. +- **Not banked — greps:** `grep -n -i -c -- 'index slicing' $F` → 0; `grep -n -i -c -- 're-read the tail' $F` → 0; `grep -n -i -c -- 'verify each section' $F` → 0; `grep -n -i -c -- 'duplicated a region' $F` → 0; `grep -n -i -c -- 'silently deleted' $F` → 1 and `'silently vanished'` → 1 (cookbook:27967 a DCE'd inline-asm draft; cookbook:35044 nop-padded tails vanishing from a disassembly — neither is about editing the knowledge base). +- **Proposed home:** G (a rule) for the knowledge-base chapter — the flywheel's own write path needs a read-back assertion. +- **Portable because:** any project whose knowledge base is one enormous append-only markdown file will edit it by offset with a script or an agent; the failure is silent, the loss is the most recent (most valuable) content, and the commit looks perfect. + +### C3 — The live hand-off block must be strictly APPENDED at the end of its file: file order is the only recency signal a fresh session has +- **Evidence:** `phase-ends/logs/Phase31.md:6126-6128` — + "> **The \"last block is the live one\" rule was BROKEN when this session started** — S70 FINAL-5 sat + > below S71 CLOSE in file order while being a day older. This block is appended at the END, which + > restores the rule. Keep appending." +- **What happened / what it cost:** Checkpoint blocks were written into the phase log in a place other than the end, so a newer block sat *above* an older one. A fresh session that follows the standing "read the last 🛑 block" instruction would have inherited a day-old state as current. The fix was mechanical (always append) plus, later in the file, an editing rule for the same file: `:8459-8460` "EDIT THE LIVE BLOCK THROUGH A SLICE FROM ITS OWN HEADER (`s.rindex('## 🛑 SESSION CHECKPOINT — S79')`), the blocks share headings". +- **Not banked — greps:** `grep -n -i -c -- 'last block is the live one' $F` → 0; `grep -n -i -c -- 'SUPERSEDES every earlier block' $F` → 0; `grep -n -i -c -- 'supersedes' $F` → 9 (cookbook/playbook/decision-log, all about one cookbook section superseding another); `grep -n -i -c -- 'checkpoint block' $F` → 1 (registry-E:374 G59 — the checkpoint is written to be replayed: content, not position); `grep -n -i -c -- 'last block' $F` → governance:77 states the discipline but not the ordering invariant or its failure. +- **Proposed home:** G (a rule) — a one-line amendment to G59. +- **Portable because:** every long agent project hands off through an append-only log; the "read the last block" convention is worthless the moment anything is inserted, and the failure is invisible to the writer and fatal to the reader. + +### C4 — A long stateless batch must persist each confirmed result the moment it is confirmed +- **Evidence:** `phase-ends/logs/Phase31.md:7121-7124` — + "**`gate_main` writes banks only at the END.** `try_batch` is stateless — every attempt is + `git checkout` main's TUs → `make extract` → substitute → build — so a 34-minute bisection holds + its result in memory and a kill loses all of it. Writing each confirmed match immediately is the + single highest-value gate improvement available." +- **What happened / what it cost:** main's gate re-derives its whole world per attempt and only commits at the end, so a bisection that takes over half an hour is one interruption away from total loss — and interruptions were routine that week (context exhaustion, an over-short `timeout`, a harness kill). The session named it the highest-value gate improvement on the board, and the durability journal shipped a session later (`:7288` "`gate_main` (durability journal + verbatim refusal)"). +- **Not banked — greps:** `grep -n -i -c -- 'writes banks only at the end' $F` → 0; `grep -n -i -c -- 'holds its result in memory' $F` → 0; `grep -n -i -c -- 'durability journal' $F` → 0; `grep -n -i -c -- 'a kill loses' $F` → 0; `grep -n -i -c -- 'try_batch' $F` → 1 (accelerators:238, the `try_batch([])` NULL-INPUT control — a different lesson). +- **Proposed home:** DK (a kernel) — a design property of any expensive gate. +- **Portable because:** every byte-gate loop on a large binary is minutes-to-hours per batch and runs inside agent sessions that die; a result that exists only in a process's memory is a result you will pay for twice. + +### C5 — "Idempotent by SKIPPING" seals an artifact against later evidence; make regenerated artifacts idempotent by REPLACEMENT +- **Evidence:** `phase-ends/logs/Phase31.md:6171-6174` — + "**`journal_notes.py`** — was idempotent by SKIPPING, so a pack with one old note could never + receive a newer one, and `claude_wave_packs` calls it at build time, so **every pack with any + history was sealed against later evidence**. Now idempotent by replacement." +- **What happened / what it cost:** The project's whole flywheel thesis is that each session's new evidence reaches the next wave's packs. A skip-if-present idempotence guard meant that any target that had ever received a note was frozen at its FIRST note — silently, at pack-build time, for every historical target. The fix paid within the hour: "the §428 escalation note reached `func_8001B0D4`'s pack only because of this" (`:6174-6175`). +- **Not banked — greps:** `grep -n -i -c -- 'idempotent by skipping' $F` → 0; `grep -n -i -c -- 'idempotent by replacement' $F` → 0; `grep -n -i -c -- 'sealed against' $F` → 0; `grep -n -i -c -- 'idempoten' $F` → 15 (playbook:308 "back-fill … (idempotent)", cookbook:33397 "Idempotent (never appends twice)", etc. — all assert idempotence as a virtue, none records that the skipping form freezes stale content). +- **Proposed home:** G (a rule) or a kernel line beside the pack/journal machinery. +- **Portable because:** every AI-decomp harness regenerates derived context (packs, briefs, notes, indexes) and reaches for an idempotence guard; the skipping form is the obvious one and it silently converts a compounding knowledge loop into a write-once one. + +### C6 — A verification target too slow to complete is not a check: sample it by default and keep the exhaustive form as a separate target +- **Evidence:** `phase-ends/logs/Phase31.md:6027-6029` — + "**`make tools-health`** — was UNRUNNABLE (>15 min, never once completed). `audit-cdecl` was a + full-corpus regression test in a health target (~787s of pure-Python collection before the first + cc1 call). Now sampled (`CDECL_AUDIT_TUS ?= 60`, 61s); `audit-cdecl-full` keeps the exhaustive + form. **333s green.**" +- **What happened / what it cost:** The project's own tool-health gate had never once run to completion, so every tool defect it would have caught shipped unchecked — in a phase whose defining finding was that essentially every "codegen wall" was an instrument defect. The remedy was not a faster machine but a split: sample in the health target, keep the exhaustive run behind its own name. +- **Not banked — greps:** `grep -n -i -c -- 'unrunnable' $F` → 0; `grep -n -i -c -- 'never once completed' $F` → 0; `grep -n -i -c -- 'a health target' $F` → 0; `grep -n -i -c -- 'too slow' $F` → 0; `grep -n -i -c -- 'audit-cdecl' $F` → 2 (decision-log:1574, cookbook:3870 — both about what the check *does*, not about it being unrunnable). The nearest banked relative is `decomp-kernels.md:439` "a slow gate is a bug", which is about gate throughput, not about a verification target that has literally never finished and therefore verifies nothing. +- **Proposed home:** accelerator (or a clause appended to the "a slow gate is a bug" kernel). +- **Portable because:** every decomp grows a `tools-health`/CI aggregate; the temptation to put the exhaustive corpus regression inside it is universal, and the failure mode is a green-looking check nobody has ever seen finish. + +### C7 — Never adopt a subagent's worktree wholesale: it is a snapshot of an older tree and may predate a bank; re-gate its artifacts against HEAD +- **Evidence:** `phase-ends/logs/Phase31.md:6808-6809` — + "**Never adopt an agent's worktree wholesale.** One predated a bank of `func_8017F9C0`; copying its + TU would have destroyed it. Re-gate against HEAD with the fixed tools instead." +- **What happened / what it cost:** Drafting agents ran in git worktrees provisioned at their launch time. Merging a finished agent's TU back wholesale would have reverted a function banked in the main tree after that worktree was cut — a silent destruction of byte-proven work that no gate would have flagged (the merged TU builds fine, it just loses a match). The safe move is to take only the agent's draft artifact and re-gate it against current HEAD. +- **Not banked — greps:** `grep -n -i -c -- "agent's worktree" $F` → 0; `grep -n -i -c -- 'predated' $F` → 0; `grep -n -i -c -- 're-gate against HEAD' $F` → 0; `grep -n -i -c -- 'adopt.*worktree' $F` → 1 (accelerators:715-718 — `parallel_gate` banking *into* its worktree and adopting nothing, the opposite direction); `grep -n -i -c -- 'stale worktree' $F` → 3 (DIGEST:291 / registry-E:117, R77: worktrees pinning old commits through a `gc` — a git-hygiene lesson, not a merge-safety one). +- **Portable because:** worktree-per-agent is the standard way to parallelise gating, banking continues in the main tree while agents run, and "just take the agent's tree" is the obvious integration shortcut. + +## ALREADY-BANKED (one line each) +- A byte gate is a null oracle for "is this C?" — a verbatim `__asm__` body matches by construction and a scoper ranks it highest-value — lives at `decomp-architect/corpus/decomp-kernels.md:384` (DK-29), `docs/how-to-ai-decomp/12-failure-museum.md:24`, `docs/accelerators.md:702`. +- A documented lever wired into no code path runs for zero targets (`neighbor_ref`, the misplaced laws file) — lives at `docs/wave-playbook.md:231`, `decomp-architect/corpus/decomp-kernels.md:504` (DK-39), `docs/how-to-ai-decomp/12-failure-museum.md:29`. +- A law from a NEAR is a hypothesis; a law from a MATCH is evidence — lives at `docs/matching-cookbook.md:34101`. +- A yield table is evidence; the story about WHY needs its own negative control (§479 rewritten three times in a day) — lives at `docs/decision-log.md:3063`, `docs/matching-cookbook.md:35860`. +- Partition along the structure the original PRODUCER used (symbols), not the one your measurement grouped by (bases) — turned a certified wall into six pieces — lives at `docs/accelerators.md:673-677`. +- Provenance → archive → link → compiler, in that order: a "compiler wall" in bytes no archive you hold can place is a provenance question first — lives at `docs/matching-cookbook.md:36175`, `docs/retrospective.md:31`. +- Reconcile a headline % against its independently-derived remainder; a snapshot denominator drifts (main 59.8% → 91.8%) — lives at `docs/decision-log.md:3109-3113`. +- §474 as the template for a wall claim (name the pass, cite file:line, measure each escape); a NEAR citing a pass + file:line is a wall-proof candidate, not a redraft — lives at `docs/matching-cookbook.md:35577`, `docs/decision-log.md:3028`. +- Triage is free: read the draft's own header before aiming anything at it; the mismatch count predicts nothing — lives at `docs/matching-cookbook.md:35864`. +- When you fix a blindness, enumerate the consumers (the 4th consumer of the verbatim blindness; two worktree provisioners with no shared list) — lives at `docs/matching-cookbook.md` "enumerate the consumers" / "fourth consumer", `docs/decision-log.md` "found twice". +- Assert the scan measured something — a bare `except: continue` gave a confident FALSE verdict over 2,603/2,603 — lives at `docs/matching-cookbook.md:8351` (§126a) and `:34714`. +- A failing gate must preserve its artifact BEFORE any control or cleanup rebuilds over it (the R40 control in the wrong order parked 11 functions) — lives at `docs/matching-cookbook.md` ("save its artifact" / "before any control"), `docs/decision-log.md` (red image). +- A uniform failure shape across independent drafts (same delta, same first mover) is a LAYOUT signature, not N codegen walls — lives at `docs/matching-cookbook.md:33892`. +- Positive-control a fix against a deliberately broken input, not only negative-control it on the unchanged population — lives at `docs/matching-cookbook.md:34758-34760`, `docs/decision-log.md:3191`. +- A second, independent oracle is the only thing that can see a class its sibling is structurally blind to (and it must derive from different evidence) — lives at `docs/decision-log.md:826-830, 873-876`, `docs/matching-cookbook.md:3751`. +- Every bank cascades: it gives the TU a real definition that contradicts the stale `extern` later drafts carry — re-run the sync, do not conclude the draft went bad — lives at `docs/wave-playbook.md` / `docs/matching-cookbook.md` ("every bank gives its TU"). +- A tool reading text it should not (comment prose as code, a mid-line block comment, dead `#ifdef` branches as live) was the phase's dominant defect class — lives at `docs/matching-cookbook.md:34446-34453` (§437) and `:34681-34685`. +- Per-TU optimisation-level / build-config gaps block byte-correct bodies (a `-O0` island; `$(filter)`'s exact stem match) — lives at `docs/matching-cookbook.md` ("build-config gap", "compiler-inexpressible") and the -O0 material in `docs/accelerators.md` / `docs/how-to-ai-decomp/03-bootstrap-order.md`. +- A gate's verdict parser is a witness, not a judge (a pre-existing warning read as the error) — lives at `docs/matching-cookbook.md:36363`. +- `config/overlays.mk` + the splat yamls are shared carve STATE; a rejected gate damaged them while `git status src/` said nothing was wrong — lives at `phase-ends/DIGEST.md:245` (R60), `decomp-architect/templates/registry-E.decomp.md:273` (G43), `docs/matching-cookbook.md:34729` (§444). +- Never kill a running workflow to relaunch it differently — add the new tier alongside the dying agents — lives at `docs/wave-playbook.md:554`. +- Run the reject to ground before respelling the body; a reject whose cause you have not named is not evidence about the C — lives at `docs/matching-cookbook.md:34411-34414`. +- Usage limits are a wall class: three frontier agents died on an account-wide limit; check the quota before routing a wave, and treat the death as an instrument failure — lives at `docs/how-to-ai-decomp/08-models-and-budgets.md:42-44`, `docs/wave-playbook.md:507`. +- A claim about another project (the sotn duplicate-function precedent) resting on your own paraphrase must be verified at the source before it becomes doctrine — lives at `docs/decision-log.md:2934-2936`. +- An empty draw pool read as coverage (`draw_waves --main` never iterated main at all) — lives at `docs/decision-log.md:2972`, `:2769-2771`. +- Probe one instance before pricing (the assumed 1–2 min gate cycle measured 16 s) — lives at `phase-ends/DIGEST.md:216` (R37). +- A masked score is not a closeness until the diff is read; agent-tool drafters outlive their session and their verdict is the transcript's last JSON — lives at `phase-ends/DIGEST.md:251` (R63) / `decomp-architect/templates/registry-E.decomp.md:231` (G36), `docs/accelerators.md:720`, `docs/matching-cookbook.md` ("last JSON"). diff --git a/.run/P33.5/log-mining/Phase33.md b/.run/P33.5/log-mining/Phase33.md new file mode 100644 index 0000000000..23788867b0 --- /dev/null +++ b/.run/P33.5/log-mining/Phase33.md @@ -0,0 +1,280 @@ +# Log mining — Phase33 +Files/ranges: `phase-ends/logs/Phase33.md:1-1372` (whole file: the phase header, the 41-task list, the Log +2026-09-06 S86 → 2026-09-07 S89, the 🛑 checkpoint, and the VERBATIM approved plan at :894-1372) · Lines read: 1372 of 1372 +Candidates considered: 37 · NEW: 16 · ALREADY-BANKED: 21 + +> **Grep set** (used for every candidate; `$DOCS` below): +> ``` +> DOCS="docs/matching-cookbook.md docs/cookbook-index.md docs/decision-log.md docs/accelerators.md \ +> docs/retrospective.md docs/wave-playbook.md docs/how-to-ai-decomp/*.md \ +> decomp-architect/corpus/decomp-kernels.md decomp-architect/templates/registry-E.decomp.md phase-ends/DIGEST.md" +> ``` +> Hit counts below are totals over that whole set (`grep -n -i -c … | awk` sum), and every non-zero hit was read in context +> before the verdict. + +## NEW + +### C1 — A Ghidra script directory compiles as ONE bundle: a single non-compiling script disables every script in it +- **Evidence:** `phase-ends/logs/Phase33.md:186-190` — "the runs after the import failed with `Failed to get OSGi bundle + containing script: …/ApplySymbols.java` (same for ExportAnnotations) — Ghidra compiles the script DIRECTORY as one bundle, + so ONE file that does not compile breaks every script in it" +- **What happened / what it cost:** B5 (the Ghidra text-export + rebuild-proof task) was blocked at 87% context and had to be + checkpointed as a WIP commit. `ExportAnnotations.java` had already exported 129 programs successfully; adding a *new, + unrelated* file (`ImportAnnotations.java`, with a bad `LocalVariableImpl` ctor) silently broke `ApplySymbols` and + `ExportAnnotations` too. The failure names the *working* script, not the broken one, so the error message points away + from the cause. +- **Not banked — greps:** `grep -n -i 'OSGi' $DOCS` → 0; `grep -n -i 'script directory' $DOCS` → 0; + `grep -n -i 'bundle containing script' $DOCS` → 0; `grep -n -i 'ExportAnnotations\|ImportAnnotations' $DOCS` → 0; + `grep -n -i 'ghidra script' $DOCS` → 1 (a tool-count row in `09-economics.md`, not this). +- **Proposed home:** accelerator (a Ghidra-scripting section) + a failure-museum row — the error message misattributes. +- **Portable because:** every decomp that scripts Ghidra keeps its scripts in one directory; the same one-bad-file-breaks-all + bundling applies to any plugin host that compiles a directory as a unit. + +### C2 — Ghidra refuses a project path containing a component that starts with `.` — a scratch project cannot live under `.run/` +- **Evidence:** `phase-ends/logs/Phase33.md:185-186` — "its import step works: the scratch project must live under + `build/ghidra_rebuild/proj` — Ghidra refuses a path component starting with '.'" +- **What happened / what it cost:** the whole project convention is that scratch lives under `.run/` (a memory-enforced rule + here). The rebuild-proof harness had to be relocated to `build/` instead, discovered by a failing run rather than by + reading a doc. +- **Not banked — greps:** `grep -n -i "path component starting with\|starting with '\.'" $DOCS` → 0; + `grep -n -i 'refuses a path\|dot-director' $DOCS` → 0; `grep -n -i 'ExportAnnotations\|ImportAnnotations' $DOCS` → 0. +- **Proposed home:** accelerator (same Ghidra-scripting section as C1). +- **Portable because:** any project that adopts a dot-prefixed scratch directory will collide with the tools that refuse one; + the general form is "pick the scratch directory name only after checking what your RE tool will accept". + +### C3 — `git check-ignore` is SILENT for tracked paths: an ignore-coverage audit run before the untracking passes vacuously +- **Evidence:** `phase-ends/logs/Phase33.md:236-237` — "ignore coverage proven with `git check-ignore --no-index` on every + path (the plain form is BLIND to tracked files — it reported nothing)" +- **What happened / what it cost:** in the preparatory purge commit (B9/C3) the audit that proves "every purged path is now + ignored" is run *while the paths are still tracked*. The plain form printed nothing, which reads as "no path is ignored" + and is indistinguishable from a run that found nothing to say. `--no-index` was required for the assertion to mean + anything. A vacuous pass here means a later blanket `git add -A` re-adds ROM bytes to a public history. +- **Not banked — greps:** `grep -n -i 'reports nothing for tracked' $DOCS` → 0; `grep -n -i 'ignore rule.*tracked' $DOCS` → 0; + `grep -n -i 'check-ignore' $DOCS` → 1 — `docs/decision-log.md:3473` records the *incantation and its timing* + ("proven with `git check-ignore --no-index` on every path BEFORE the paths became untracked") but not the trap: nothing + says the plain form is silent for tracked paths, i.e. that the naive audit **passes without checking anything**. +- **Proposed home:** accelerator, or a line appended to the existing publishing chapter's ignore step; it is a silent-false-pass + instance of the project's dominant defect class. +- **Portable because:** every project that moves files out of git before a rewrite runs exactly this audit at exactly this + moment, on any host. + +### C4 — A coverage instrument that infers its denominator from OPEN work INVERTS at 100% — carry the scanned denominator in the artifact, and test the instrument at both endpoints +- **Evidence:** `phase-ends/logs/Phase33.md:126-129` — "Regenerating `.run/family_hseq.json` made the `audit-binaries` warning + WORSE (6 → 217 'missing'): at 100% the map's `families` list is empty and CHECK 4 inferred coverage from family members, so + a complete map read as empty; the P32-close warning had been a stale pre-onboarding file." +- **What happened / what it cost:** the check had been *silently green because the map was stale*; regenerating it turned one + warning into 217. The fix was to make the artifact carry its own denominator (`"binaries": 217`, `"open_instances": 0`) and + make the consumer read that, plus a "predates the coverage field" warning for old-format maps. The instrument was correct at + every intermediate percentage and wrong at both endpoints. +- **Not banked — greps:** `grep -n -i 'infer.*denominator\|denominator.*infer' $DOCS` → 0; + `grep -n -i 'complete map read as empty\|read as empty' $DOCS` → 0; + `grep -n -i 'carries its own coverage\|carry the denominator' $DOCS` → 0; `grep -n -i 'family_hseq' $DOCS` → 13 (campaign + selection and an R32 gap, never this inversion). R32 ("assert your COVERAGE") and R41 ("every number ships with its + denominator") are the family this belongs to; neither states the endpoint inversion or the "denominator lives in the + artifact" fix. +- **Proposed home:** DK (a kernel under the oracles/instruments family) + a line in `04-oracles-and-instruments.md`. +- **Portable because:** every decomp builds progress/coverage instruments and runs them for months at 1–99%; the day the + project finishes is the day the untested endpoint fires, and the same is true of a fresh project at 0%. + +### C5 — An annotator that writes into the text it reads must never treat its own output as evidence +- **Evidence:** `phase-ends/logs/Phase33.md:499-503` — "Three instrument defects, found by its own controls/verify before any + tag was written … (3) after writing, a neighbouring cite's tag read as a `2.8.1` cue — the instrument reading its own + output — caught by `--verify`." +- **What happened / what it cost:** `gccmap_cites.py` derives each source citation's provenance tag from cues near the cite. + Once it had written `[2.8.1 pm]` next to one cite, that written tag became a "cue" for the *next* cite, so a second run + would have drifted the tags. Caught only because the tool shipped a `--verify` that re-derives every written tag from + scratch and a `--controls` mode over known-true cases; the first two defects (span pairing inside a ±160-char window, + fenced code blocks inverting the pairing) came from the same controls. +- **Not banked — greps:** `grep -n -i 'reading its own output' $DOCS` → 0; + `grep -n -i 'its own output as evidence\|own output as input' $DOCS` → 0; + `grep -n -i 'idempotent.*annotat\|annotator' $DOCS` → 0. +- **Proposed home:** DK (a kernel) — the in-place-annotation corollary to R57 ("an instrument's own write path is part of the + instrument"): an in-place annotator needs a re-derivation check, not just idempotence. +- **Portable because:** in-place annotation of one's own documents (provenance tags, section ids, cross-refs, cookbook + indices) is a standard decomp housekeeping tool. + +### C6 — Regex-extracted evidence needs a structural marker, or prose becomes data +- **Evidence:** `phase-ends/logs/Phase33.md:503-504` — "Bare ALL-CAPS prose words (`NOT`, `AND`, `DEST`) had also passed as + evidence: identifiers now need an underscore, as every real gcc macro/function cited has." +- **What happened / what it cost:** the citation tagger's identifier heuristic accepted ordinary emphasised English words as + gcc identifiers, so prose voted on provenance. The fix was a structural predicate derived from the corpus itself (every + real cited gcc macro/function contains `_`), not a longer stopword list. +- **Not banked — greps:** `grep -n -i 'ALL-CAPS' $DOCS` → 0 (run as part of the C5 batch); + `grep -n -i 'idempotent.*annotat\|annotator' $DOCS` → 0; `grep -n -i 'reading its own output' $DOCS` → 0. +- **Proposed home:** accelerator (a line under the instrument-controls material), or folded into C5's kernel. +- **Portable because:** every decomp mines its own prose (cookbooks, logs, decision records) with regexes; the general rule is + "derive the acceptance predicate from a property the true population provably has". + +### C7 — `objdump -dr` interleaves relocation records only for OBJECT files; a linked ELF lists them separately and shifts the address column +- **Evidence:** `phase-ends/logs/Phase33.md:526-529` — "`objdump -dr` interleaves relocation records only for OBJECT files — + a linked ELF lists them separately (`-r`, section-relative offsets) … and a linked listing puts the address at column 0 (an + object listing indents it) — the instruction regex is `^\s*`." +- **What happened / what it cost:** two gotchas in building the xsig test fixtures, each of which silently produces an + *empty or wrong* parse rather than an error: relocation-aware signing over a linked ELF sees no relocations at all, and an + instruction regex tuned on object listings matches nothing on a linked listing. +- **Not banked — greps:** `grep -n -i 'linked ELF' $DOCS` → 0; `grep -n -i 'relocation records' $DOCS` → 0; + `grep -n -i 'objdump -dr' $DOCS` → 10 (all about normalised instruction diffing and `masked_diff` on objects, never the + object-vs-linked difference). +- **Proposed home:** cookbook (the tooling/objdump area) or accelerator. +- **Portable because:** any decomp writing reloc-masked scoring, cross-project signatures, or a differ has to parse both + object and linked listings, on any binutils target. + +### C8 — A derive-then-apply pipeline over a live repository needs a freshness guard, and a stated sequencing law +- **Evidence:** `phase-ends/logs/Phase33.md:285-288` — "One more R43 guard added to `run_filter.py`: it refuses a dictionary + whose main count/HEAD differ from the clone's (a stale dictionary would drop rows from the public map) … **Sequencing law + for C4:** the dictionary + ids are rebuilt from the FINAL tree right before the clone (4 s + 3 min) — any commit after that + invalidates them (the guard enforces it)." +- **What happened / what it cost:** the hash dictionary and the blob-id strip list are derived from the repository, then applied + to a clone of it; every commit made between derivation and application silently invalidates them, and the failure mode + (rows missing from the *public* commit map) is invisible at run time. Trial #2 had matched "by construction", i.e. by luck of + ordering — the guard was added so the property is enforced rather than observed. +- **Not banked — greps:** `grep -n -i 'stale dictionary\|rebuilt from the FINAL tree' $DOCS` → 0; + `grep -n -i 'invalidates' $DOCS` → 8 (all compiler-pass semantics in the cookbook); + `grep -n -i 'created EMPTY\|never a fork' $DOCS` → 1 (`11-publishing.md:29`, a different step of the same procedure). +- **Proposed home:** DK (a kernel) or an added step in the publishing chapter's rewrite recipe; it is an R43 instance with a + named sequencing law. +- **Portable because:** the derive-a-map-then-apply-it-to-a-copy shape recurs far beyond history rewrites (symbol maps, splat + configs, dedup registries applied to a worktree copy). + +### C9 — A content-hash "no ROM bytes" audit collides on zero-length files +- **Evidence:** `phase-ends/logs/Phase33.md:205-207` — "controls: the current tree FAILS with exactly the purge set — 255 + offender rows — a clean subset OK, a renamed EXE copy caught by content; empty-file SHA1 collision with the zero-length + SC04/SC05 `FILE_029/1.6` payloads found and exempted" +- **What happened / what it cost:** `audit_public.py` flags any tracked file whose SHA1 appears in the extracted-ROM manifest. + The disc contains zero-length payloads, so the empty-file SHA1 is in the manifest — and every empty tracked file in the + repository then reads as ROM-derived. Found by running the control, not by reasoning. +- **Not banked — greps:** `grep -n -i 'zero-length\|empty-file\|SHA1 collision' $DOCS` → 0; + `grep -n -i 'empty blob' $DOCS` → 5 (all the *git* empty-blob strip-list defect, a different mechanism in a different tool). +- **Proposed home:** accelerator, or a line in the publishing chapter's no-ROM-audit description. +- **Portable because:** every project that gates publication on "no tracked file's hash appears in the ROM manifest" inherits + this collision, since discs and archives routinely contain zero-length entries. + +### C10 — `git push --mirror` does not push `refs/stash` +- **Evidence:** `phase-ends/logs/Phase33.md:295-296` — "`git ls-remote archive` == local refs except **`refs/stash`, which + `--mirror` does not push** (the bundle holds it; Drew may push it as a branch)." +- **What happened / what it cost:** the archive repository was the one snapshot of the pre-rewrite history and was verified + ref-by-ref against the local repo. The stash — three phase-26 WIP entries — was not in it; only the separately made + `--all --reflog` bundle held it. Had the bundle not existed, an "identical mirror" check would have passed while losing work. +- **Not banked — greps:** `grep -n -i 'refs/stash' $DOCS` → 0; `grep -n -i 'stash' $DOCS` → 5 (all MIPS register-stashing in + the cookbook); `grep -n -i 'created EMPTY\|never a fork' $DOCS` → 1 (the neighbouring archive step, silent on this). +- **Proposed home:** accelerator, or a line in the publishing chapter's archive step ("mirror **and** bundle; verify both"). +- **Portable because:** it is a property of git, and every project archiving a history before a rewrite does exactly this push. + +### C11 — Route a host purge request through the flow that actually exists; the obvious form is a trap +- **Evidence:** `phase-ends/logs/Phase33.md:635-641` — "via the Support portal's **Virtual Agent 'Clear cached views'** flow — + the route that actually works … The static 'Repositories' form's 'Deletes' sub-option is a trap: it is the + delete-the-whole-repository flow — never submit it." +- **What happened / what it cost:** the public flip is gated on the host garbage-collecting force-pushed-away objects, which + only Support can do. The form a reasonable person picks (Repositories → Deletes) would have destroyed the repository. The + working path is a specific chatbot flow with a specific sequence of answers and a ~500-character reason field, and it is now + written down in the runbook with the answers. +- **Not banked — greps:** `grep -n -i 'Virtual Agent\|cached views' $DOCS` → 0; + `grep -n -i 'delete.*whole repository\|delete the repository' $DOCS` → 0; + `grep -n -i 'Support ticket' $DOCS` → 6 (the ticket is cited as a *cost* in the retrospective and economics chapters, and + the publishing chapter says "a Support ticket and a daily probe" — none records the route or the trap). +- **Proposed home:** accelerator + the publishing chapter (with an explicit "this is host-specific and dated" caveat). +- **Portable because:** the *shape* transfers to any host — the destructive option and the wanted option live under the same + menu word; write the working route down the day you find it, because you find it once and need it under time pressure. + +### C12 — A host feature can be gated on the very flip it was meant to precede: read the settings page, don't infer +- **Evidence:** `phase-ends/logs/Phase33.md:692-694` — "~~The wiki can be pushed NOW~~ — **WRONG (R14, corrected minutes + later):** GitHub's settings page reads 'Upgrade or make this repository public to enable Wikis'; on the free plan wikis + exist only on public repos, so the wiki waits for the flip." +- **What happened / what it cost:** F3 authored 25 wiki pages and a sync script on the assumption they could be published + during the purge wait; the claim was made and retracted within minutes, and the sync script's message plus the owner's + checklist item had to be corrected back. Cheap only because it was checked. +- **Not banked — greps:** `grep -n -i 'wikis\? only on public\|enable Wikis\|free plan' $DOCS` → 0; + `grep -n -i 'Support ticket' $DOCS` → 6 (the wait, not the feature gating); `grep -n -i 'fresh clone\|fresh-clone' $DOCS` + → 13 (unrelated). +- **Proposed home:** accelerator (one line under the publishing sequencing), or a failure-museum row. +- **Portable because:** the general law — verify a host/service capability against its own settings page before sequencing + work behind it — applies to badges, pages, discussions, artifact hosting, and any tier-gated feature. + +### C13 — A public scratch service's compiler image is NOT your pinned toolchain; rebuild it locally and prove byte-identity before asking for a preset +- **Evidence:** `phase-ends/logs/Phase33.md:645-649` — "the `gcc2.7.2-psx` image is old-gcc **0.13** + maspsx **`86ccd7d8`** + (not our 0.17 + `874855c5`; SETUP's 'same pinned commit we use' was stale; three rows corrected, R14) … presets: **no create + button in the UI** … `name`/`platform` immutable, no owner delete (405) → prove before requesting." +- **What happened / what it cost:** the phase plan itself (`:1232-1236`) said presets "are created in-browser by any logged-in + user (`POST /api/preset/`)" — wrong; the frontend has no create call and maintainers create them from an issue template. And + the project's own SETUP had asserted the service used our pinned commits, which was false in both components. The response + was `tools/decompme_replica.sh`: rebuild *their* image locally (tarball sha256-pinned, maspsx at their commit, their `as` + wrapper verbatim), run their two backend commands, and compare words — BYTE-IDENTICAL on all 26 words of one function, with + our pipeline as the control and one-component-at-a-time attribution. +- **Not banked — greps:** `grep -n -i 'no create button\|created in-browser\|preset-request\|maintainers create' $DOCS` → 0; + `grep -n -i 'decomp\.me' $DOCS` → 7 (the preset named as a deliverable, the service's own "do not hook up an LLM" policy, + a layout probe) ; `grep -n -i 'old-gcc 0\.13\|86ccd7d8' $DOCS` → 1 — `phase-ends/DIGEST.md:141` is the Phase-33 synopsis + and records the measured FACT plus the local proof ("decomp.me = old-gcc 0.13 + maspsx 86ccd7d8, rebuilt locally, 26/26 + words"). **NEW is the day-one law, not the fact:** nothing records that presets cannot be created from the UI, that + `name`/`platform` are immutable with no owner delete, or that the project's own SETUP had asserted the service used our + pins and was wrong — i.e. that the correct move is to rebuild THEIR image and prove byte-identity BEFORE requesting. +- **Proposed home:** DK or accelerator + the publishing chapter — "publishing a preset is a *proof*, not a form". +- **Portable because:** every decomp eventually wants a scratch-service preset, and the service's image drifts from the + project's pins independently; an unprovable, immutable, undeletable preset is a permanent public error. + +### C14 — A miner over your own records finds only what its pattern anticipates: measure the widened pattern's yield +- **Evidence:** `phase-ends/logs/Phase33.md:429-431` — "`tools/mine_hindsight.py` (… the decision-log's `Hindsight` bullets + AND `### Hindsight` sections — 19 over 79 entries after widening the pattern, the first cut found 11 —, the two 'What we + believed' sections, 237 deviation rows over 32 PhaseEnds)" +- **What happened / what it cost:** the retrospective's entire input is what this miner returns. The first pattern found 11 + hindsight entries; widening it to the second syntactic form found 19 over the same 79 entries — 42% of the corpus was + invisible to the first cut, and nothing in the output would have said so. +- **Not banked — greps:** `grep -n -i 'widening the pattern\|first cut found' $DOCS` → 0; + `grep -n -i 'recall' $DOCS` → 5 (a symbol-join recall floor in the cookbook, and card authoring in the decision log — a + related idea for a different instrument); `grep -n -i 'mine_hindsight' $DOCS` → 2 (`docs/retrospective.md:8` names the tool + as the retrospective's input and `phase-ends/DIGEST.md:143` lists it as a Phase-33 deliverable; neither records the + widened-pattern yield or that the first cut saw 11 of 19). +- **Proposed home:** DK (a kernel) or accelerator — the "assert your denominator" law (R32/R41) applied to *text mining of your + own records*, which is how the retrospective, the cookbook index and this very pass are produced. +- **Portable because:** the record formats drift over a long project (this one had two hindsight syntaxes and two progress + formats); any future project mining its own logs inherits the same silent under-recall. + +### C15 — A harness's low-memory guard silently kills a long BACKGROUND job — and workers from closed phases survive for days +- **Evidence:** `phase-ends/logs/Phase33.md:508-511` — "the first tools-health run was KILLED by the harness's low-memory guard + during the report step (a transient spike; 29 GB available afterwards) — and the process table held **8 orphaned + `tools/permuter/run_masked.py` workers from a closed phase, 49 h old (parent PID 18)**, stopped by PID … the foreground + re-run passed." +- **What happened / what it cost:** the project's full health check reads as a failure when it is actually a harness kill, and + the memory pressure that triggers the kill was manufactured by the project's own abandoned workers from a *previous phase*. + Two independent instrument errors compounding: run long checks in the foreground, and audit the process table across phase + boundaries. +- **Not banked — greps:** `grep -n -i 'low-memory guard\|memory guard' $DOCS` → 0; + `grep -n -i 'background.*killed\|killed.*background' $DOCS` → 0; `grep -n -i 'orphan.*worker\|orphaned' $DOCS` → 14 (an + orphaned dedup reconcile, orphaned per-binary reports — never a live process). +- **Proposed home:** accelerator (an ops/lane-hygiene line) + a failure-museum row (an R40 instance: exonerate the harness). +- **Portable because:** any agent harness that supervises long jobs has a resource guard, and any project that runs unattended + worker fleets leaks them across phases. + +### C16 — A `cd` in one agent shell call persists into the next +- **Evidence:** `phase-ends/logs/Phase33.md:660-661` — "Gotcha, recorded: a `cd` in one Bash call persists into the next — the + first download landed inside `tools/maspsx/` (moved out; submodule clean)." +- **What happened / what it cost:** a downloaded tarball was written into a git *submodule*, which dirties a pinned dependency + rather than the project's own tree — the kind of pollution a `git status` in the superproject reports only as "modified + content". Cleaned by hand; the correction cost a step, not a session. +- **Not banked — greps:** `grep -n -i 'cd in one\|cd persists\|persists into the next' $DOCS` → 0; + `grep -n -i 'working directory persists' $DOCS` → 0. +- **Proposed home:** accelerator (one line) — minor, but it belongs with the other harness-shell gotchas (R79's `pkill`). +- **Portable because:** it is a property of agent harnesses that keep a persistent shell; the durable form is "absolute paths + in every tool invocation, and never rely on the working directory across calls". + +## ALREADY-BANKED (one line each) +- A linked worktree's HEAD is a ref — audit `git worktree list` before any gc/purge (12 stale worktrees pinned 3,729 commits) — lives at `decomp-architect/corpus/decomp-kernels.md:692` (DK-54), `phase-ends/DIGEST.md:291` (R77), `docs/how-to-ai-decomp/11-publishing.md:116`, `docs/how-to-ai-decomp/12-failure-museum.md:37`, `decomp-architect/templates/registry-E.decomp.md:115` (G15) +- A probe or guard must never write into the repository it guards (the purge probe re-imported 5.97 GiB) — lives at `phase-ends/DIGEST.md:298` (R81), `decomp-architect/templates/registry-E.decomp.md:119` (G16), `docs/how-to-ai-decomp/04-oracles-and-instruments.md:74` +- A rewritten history is not private until the objects are gone from the HOST; the Activity view publishes every pre-force-push tip — lives at `phase-ends/DIGEST.md:300` (R82), `decomp-architect/corpus/decomp-kernels.md:708` (DK-56), `decomp-architect/templates/registry-E.decomp.md:125` +- Outward text is rewritten the way a developer writes, never a model draft with the tells removed; read the target's AI policy first — lives at `phase-ends/DIGEST.md:303` (R83), `docs/decision-log.md:3500`, `docs/how-to-ai-decomp/12-failure-museum.md:45`, `decomp-architect/corpus/decomp-kernels.md:727`, `decomp-architect/templates/registry-E.decomp.md:421` +- `--strip-blobs-with-ids` with the shared EMPTY blob's id silently undid every "file emptied" change in history — lives at `docs/retrospective.md:71`, `docs/how-to-ai-decomp/11-publishing.md:32`, `docs/how-to-ai-decomp/12-failure-museum.md:36`, `decomp-architect/corpus/decomp-kernels.md:916` +- A byte-identical commit keeps its hash across a rewrite and trips old-hash assertions — lives at `docs/retrospective.md:35`, `docs/how-to-ai-decomp/12-failure-museum.md:36` +- Rehearse every irreversible repository operation on a scratch copy and prove it pair by pair with positive assertions — lives at `docs/how-to-ai-decomp/11-publishing.md:114` +- A metric's denominator comes from the artifact you control (the build), not the analysis tool — Ghidra left 3,628 words owned by no function — lives at `docs/decision-log.md:3450-3461` +- When two instruments disagree by a systematic offset, the one corrected instance is usually a class — lives at `docs/decision-log.md:3459-3461` +- No ROM-derived bytes in ANY published artifact — test fixtures, JSON, badges, reports included — lives at `docs/how-to-ai-decomp/11-publishing.md:112`, `phase-ends/DIGEST.md:284` +- Published numbers are generated, never typed (R51 applied to documents; the README's numbers were Phase-19 stale) — lives at `docs/how-to-ai-decomp/11-publishing.md:113` +- A link checker that widens its document set must classify a missing *promised* page as PENDING, never BROKEN — lives at `docs/how-to-ai-decomp/11-publishing.md:116`, `:100-102` +- Build the test fixture from your own C at two link addresses so a ROM-facing tool ships publishable tests — lives at `phase-ends/DIGEST.md:284`, `docs/how-to-ai-decomp/11-publishing.md:106` +- The archive repository is created EMPTY — never a fork or import, which share the host's object store — lives at `docs/how-to-ai-decomp/11-publishing.md:29` +- Other clones reset to the new history and never `git pull` (an 8,000-commit merge of two lineages) — lives at `docs/how-to-ai-decomp/11-publishing.md:46` +- Every third-party licence read from its own source of truth; two upstreams publish none and the table says so; never the phrase "clean-room" — lives at `docs/how-to-ai-decomp/11-publishing.md:90-94` +- An upstream interactive tool's author does not want an unattended tool's objective function — ask what the maintainer's workflow needs before offering — lives at `docs/decision-log.md:3504`, `:3496` +- `pkill -f` never with a literal the calling shell's own command line contains (exit 144) — lives at `phase-ends/DIGEST.md:295` (R79), `docs/accelerators.md:197`, `docs/wave-playbook.md:730` +- To prove an RE database regenerable from text, subtract a baseline exported from the freshly rebuilt program (analysis-origin rows) — lives at `docs/how-to-ai-decomp/11-publishing.md:83-85` +- A fresh-clone reproducibility proof only proves independence from what is actually absent — B3's clone still contained the vendor SDK, so the no-SDK path needed its own control — lives at `docs/decision-log.md:3475`, and the corrected claim at `docs/how-to-ai-decomp/11-publishing.md:69` +- A verification harness that writes its evidence logs into the tree it checks fails its own cleanliness step (`00_tree.log`; the tracked A5 logs) — covered by the same class at `phase-ends/DIGEST.md:298` (R81) and `:242` (R57 — an instrument's own write path is part of the instrument) diff --git a/.run/P33.5/log-mining/Phase7.md b/.run/P33.5/log-mining/Phase7.md new file mode 100644 index 0000000000..429da2b589 --- /dev/null +++ b/.run/P33.5/log-mining/Phase7.md @@ -0,0 +1,70 @@ +# Log mining — Phase7 +Files/ranges: phase-ends/logs/Phase7.md:1-201 · Lines read: 201 of 201 +Candidates considered: 28 · NEW: 5 · ALREADY-BANKED: 23 + +Grep file set used for every "already banked?" test (abbreviated `$F` below): +`docs/matching-cookbook.md docs/cookbook-index.md docs/decision-log.md docs/accelerators.md docs/retrospective.md +docs/wave-playbook.md docs/how-to-ai-decomp/*.md decomp-architect/corpus/decomp-kernels.md +decomp-architect/templates/registry-E.decomp.md phase-ends/DIGEST.md` + +## NEW + +### C1 — A decompiler's `unaff_` output means it decompiled ONE entry path of a multi-entry function and handed you a fragment; that is an instrument limit, never evidence the function is hard +- **Evidence:** `phase-ends/logs/Phase7.md:109` — "**SaveLoadRoutine DEFERRED to Q#5** (Drew) — it is a 1139-ins multi-entry save/memcard blob (single `jr $ra`, 3 saveHeaderTemplate entry points, Ghidra returns only an unaff_-register fragment)" +- **What happened / what it cost:** Phase 7 parked the single largest loader function on the shape of the Ghidra output — one `jr $ra` with three entry points, so the decompiler returned only an `unaff_`-register fragment. The project later measured the cost of that parking: `docs/decision-log.md:2883,2903` records `SaveLoadRoutine` (1,165 ins) carried as "an unbreakable wall" for a whole phase — "9.2% of everything left in the project ... while its body was byte-identical the entire time". The `unaff_` tell that started the parking is recorded nowhere; only the much later instrument-verdict post-mortem is. +- **Not banked — greps:** `grep -n -i -c 'unaff_' $F` → 0 (18 hits for the substring `unaff` in the cookbook are all the English word "unaffected"); `grep -n -i -c 'ghidra.*fragment' $F` → 0; `grep -n -i -c 'multi-entry' $F` → 0; `grep -n -i -c 'multiple entry point' $F` → 1 (`docs/matching-cookbook.md:34078`, about gcc refusing a loop-invariant hoist — unrelated) +- **Proposed home:** DK (a kernel), extending DK-19 / the failure museum: a decompiler artefact is not a difficulty signal. Also a one-line cookbook tell. +- **Portable because:** every decompiler (Ghidra, IDA, m2c) emits an "undefined/unaffected register" placeholder when a routine has entry points its CFG did not reach; on any console, any compiler, that placeholder is the tool's boundary, not the function's. + +### C2 — A milestone counted in "matched functions" must exclude the splitter's auto-generated empty bodies; define the bar on substantive matches at the moment you set it +- **Evidence:** `phase-ends/logs/Phase7.md:12` — "**≥25 = real substantive matches** (the 42 splat-auto empties do NOT count). 14 real now → need ≥11 more." +- **What happened / what it cost:** splat auto-generates `void f(void){}` for empty functions, and every naive source-side progress scanner counts them as decompiled C. At the top of Phase 7 the raw number was 56 while the real number was 14 — a 4× inflation on the phase's own exit criterion. The locked decision had to be written by hand into the phase's "Locked decisions" block, and the authoritative baseline was then restated in three separate places (`:20`, `:25`, `:32`) because the numerator kept being contested. +- **Not banked — greps:** `grep -n -i -c 'splat-auto' $F` → 1 (`docs/matching-cookbook.md:710`, about auto-empties *regenerating identically* across a resegment, not about counting); `grep -n -E -i -c 'empt(y|ies).*not count' $F` → 0; `grep -n -i -c 'do NOT count' $F` → 2 (`matching-cookbook.md:12998` fall-through predecessors, `:16699` atlas labels — both unrelated). DK-19 and DK-30 (`decomp-kernels.md:263,392`) govern the *denominator*; this is a contaminated numerator. +- **Proposed home:** DK (a kernel) alongside DK-19/DK-30, plus a line in `docs/how-to-ai-decomp/03-bootstrap-order.md` §"the honest census". +- **Portable because:** every splitter in this family (splat, and its equivalents) synthesises trivial bodies for empty/no-op functions; any project that reports "% decompiled" from the source tree inherits the same free-match inflation on day one. + +### C3 — Duplicate/family census run before the vendor library is linked out is contaminated: the duplicate groups are library fragments and epilogues, not game functions +- **Evidence:** `phase-ends/logs/Phase7.md:26` — "Dedup-collapse skipped (the sig dup groups are mostly PsyQ library fragments/epilogues, not real funcs)." and `:25` — "dup leverage 8 h_exact + 44 h_norm redundant" +- **What happened / what it cost:** Phase 7 built the duplicate report (`tools/dup_report.py`), measured a tiny leverage (8 exact + 44 normalized), and then discarded the whole dedup-collapse plan after inspecting what the groups actually were — SDK library fragments and shared epilogues. The census was built, run, and thrown away because it was measured on an image that still contained ~350 un-linked vendor-library functions. +- **Not banked — greps:** `grep -n -i -c 'dedup leverage' $F` → 0; `grep -n -i -c 'library fragment' $F` → 0; `grep -n -i -c 'dedup.*librar' $F` → 1 (`docs/how-to-ai-decomp/09-economics.md:13`, the final results table: "2,220 dedup groups; 1,256 vendor-library functions linked" — a result line, not the ordering lesson); `grep -n -i -c 'duplicate.*epilogue' $F` → 3 (all `matching-cookbook.md` codegen entries about gcc duplicating an epilogue — unrelated) +- **Proposed home:** `docs/how-to-ai-decomp/03-bootstrap-order.md` §"Phase 1 — the honest census" (which lists byte-identical duplication and structural families as the first measurements and does not warn about this contaminant), and a kernel. +- **Portable because:** any target that statically links a vendor SDK (PsyQ, SDK libs on N64/Saturn/Dreamcast, a vendored libc) carries thousands of library bodies whose leaves and epilogues dominate any signature-duplication census taken before those objects are identified. + +### C4 — A byte-identical baseline is only proven when the whole gate is green from clean across several INDEPENDENT sessions; one green run hides a nondeterministic extract +- **Evidence:** `phase-ends/logs/Phase7.md:75` — "NOTE: Gen1 exit needs ≥3 SESSIONS of green `make check` — **satisfied** (A, B, C, D, E, F, +G)" and `:18` — "fixes a LATENT NON-REPRODUCIBILITY: spimdisasm 1.41.0 auto-detection of that 8-byte inter-fn blob is unstable across clean extracts" +- **What happened / what it cost:** the phase made "≥3 sessions of green `make check`" a milestone gate and kept a per-session green log (`:101-109`, 7 entries, each recording the full `clean && extract && build && check` and the exact hash `143dbb89…`). The bar earned itself immediately: re-running the extract from clean in a fresh session exposed an 8-byte inter-function blob whose auto-detection was unstable, i.e. a build that had been "green" was not reproducible. A single green run in one session would not have found it. +- **Not banked — greps:** `grep -n -i -c 'zero-regression' $F` → 0; `grep -n -i -c 'sessions of green' $F` → 0; `grep -n -i -c 'green across' $F` → 0; `grep -n -i -c 'independent session' $F` → 0; `grep -n -i -c 'three sessions' $F` → 2 (`accelerators.md:683` wall-sweeping verdicts, `decision-log.md:3324` "three sessions running the finding is the same" — neither is the gate-repetition bar). `how-to-ai-decomp/03-bootstrap-order.md:19` and `02-byte-gate.md:15` require "deterministic extraction" and a manifest but state no repetition-across-sessions criterion. +- **Proposed home:** `docs/how-to-ai-decomp/02-byte-gate.md` (the baseline's acceptance criterion) + a kernel. +- **Portable because:** every decomp pipeline has a nondeterministic step somewhere (auto-detected boundaries, dict ordering, timestamps); the cheapest detector is repetition from clean in a fresh process/session, and it costs nothing to make it the milestone's definition. + +### C5 — When a foundation task hits a structural wall mid-phase, reorder: bank the tractable wins first and re-scope the wall as its own focused sub-project +- **Evidence:** `phase-ends/logs/Phase7.md:3-4` — "**REORDERED 2026-06-14** (Drew): bank non-switch wins first; the rodata-island foundation + LZSS match are DEFERRED to a focused sub-project after the easier tasks." and `:14` — "rodata foundation hit a structural wall … do reports + harvest + non-switch loader FIRST, then a focused LZSS/rodata sub-project." +- **What happened / what it cost:** the phase opened on the rodata-island foundation, hit a wall, and was reordered rather than ground on. The deferred work came back as "Task 2′ — Focused LZSS + rodata-island sub-project" (`:33`) and closed completely (LZSS byte-for-byte, the rodata carve, cookbook §10) *after* the report machinery, the harvest and the -O0 split mechanism existed. In between, the tractable path banked 14 → 43 real matches, and the wall itself was re-root-caused (`:113` "RESOLVED end-to-end") with the tooling the intervening tasks had built. +- **Not banked — greps:** `grep -n -i -c 'easy wins first' $F` → 0; `grep -n -i -c 'focused sub-project' $F` → 0; `grep -n -i -c 'defer.*sub-project' $F` → 0; `grep -n -i -c 'blocked foundation' $F` → 0; `grep -n -i -c 'order of attack' $F` → 0. The nearest banked statement is the opposite axis — `docs/how-to-ai-decomp/03-bootstrap-order.md:93` "Order by leverage, not difficulty" (how to schedule cracking), not what to do when the scheduled foundation task is blocked. +- **Proposed home:** kernel (sequencing) / `03-bootstrap-order.md`. +- **Portable because:** it is a scheduling law about a blocked prerequisite, independent of console or compiler; the project's own banked law ("every structural wall resolved to our own tooling", `docs/decision-log.md:1495`) is exactly why deferring a wall until the tooling has grown is the right move rather than a retreat. + +## ALREADY-BANKED (one line each) +- Link the vendor SDK's real library objects byte-exact instead of hand-decompiling them (~350 free functions, and it fixes the per-object alignment) — lives at `docs/matching-cookbook.md:630-660` (§9.x) +- Relocation-masked search locates a library object's `.text` in the image; a raw byte search false-negatives on relocated functions — `docs/matching-cookbook.md:640` and `:645` +- Recover an object's unresolved externals straight out of the EXE's already-resolved relocations (R_MIPS_26; HI16+LO16) — `docs/matching-cookbook.md:651-654` +- A uniform +4 section offset is an alignment fault (obj-parser emits align 2**3, the original is 4-aligned); fix with `objcopy --set-section-alignment` — `docs/matching-cookbook.md:655-657` +- Place library data with linker `NOLOAD` sections so no data carving is needed — `docs/matching-cookbook.md` (§9.3, 19 hits) and `docs/accelerators.md` +- The second library integration exposes per-library state baked into the first: namespace the NOLOAD section names and pass the sibling `*_externals.ld` to the trial link or you get a spurious "UNRESOLVED" alarm — `docs/matching-cookbook.md:740-745` +- A resegment shifts the disassembler's auto-detected function/data boundaries in *unchanged* regions; fix by declaring the real symbols and carving the data-in-text table — `docs/matching-cookbook.md:745-751` +- Disambiguate near-identical SDK objects (PRESET2/PRESET3, OBJT2/OBJT3) and `.text`-identical aliases by the `.data`/`.rdata` byte test; hardcode the winners in a committed regen script — `docs/matching-cookbook.md:753-757` +- Short objects need an explicit placement window or they are "ambiguous" and drop out of the placement map — `docs/matching-cookbook.md:772-777` +- Count OBJECTS, not stubs: the splitter over-segments library code into ~5× more `INCLUDE_ASM` stubs than real functions — `docs/matching-cookbook.md:779` +- The finer resegmentation must be split-deterministic across two clean extracts — `docs/matching-cookbook.md:780-782` +- The build must stay byte-identical WITH or WITHOUT the gitignored SDK objects (fresh-clone stub fallback) — `docs/matching-cookbook.md:783` +- Extra `.align 3` jump-table padding means the original had separate translation units; the disassembler even prints file-split suggestions — `docs/matching-cookbook.md:357`, and set TU boundaries at the jump-table spans at segmentation time — `docs/accelerators.md:530` (#20) +- Per-module optimization-level divergence is real; detect an `-O0` module by its frame-pointer prologue (`21F0A003`) — `docs/how-to-ai-decomp/03-bootstrap-order.md:24`, `docs/cookbook-index.md:24` +- A floor-free `.text`-only metric measures progress when the permuter's score has a rodata/jump-table floor — `docs/matching-cookbook.md:865` (also `:261`, `:1400`) +- Web-research the real compiler source for compiler-quirk residuals, an escalation tier above the permuter (Phase 7's R17, from the LZSS cross-jump barrier) — `phase-ends/DIGEST.md` R17 and `docs/matching-cookbook.md` §3a/§5a +- Regenerating a curated `.c` from a fresh extract drops its comments; surgically insert the generated lines instead — `docs/matching-cookbook.md:349` +- Logically-correct-but-unmatched C lives under a `NON_MATCHING` guard with the assembly stub still in the default build — `docs/how-to-ai-decomp/02-byte-gate.md:19`, `decomp-architect/templates/registry-E.decomp.md:38` +- The progress scanner miscounted (`extern …(…);` forward declarations swallowed the next function and double-counted) — the instrument-integrity law is banked at `docs/retrospective.md:47`, `docs/how-to-ai-decomp/02-byte-gate.md:30` (R32/R34) and `decomp-architect/corpus/decomp-kernels.md:263` (DK-19) +- Carry each function's past-attempt residual notes with the function so a later pass resumes from them — `docs/how-to-ai-decomp/05-cards-lanes-waves.md`, `06-knowledge-base.md`, `decomp-kernels.md` (journal notes) +- Harvest the trivial accessor leaves (getters/setters of globals) for the first cheap real matches — banked *and superseded*: "Order by leverage, not difficulty. 'Smallest, simplest leaf first' is the lowest-leverage order" — `docs/how-to-ai-decomp/03-bootstrap-order.md:93` +- Preserve the phase worklog out of the load order instead of deleting it, so in-flight work can resume (R19) — `phase-ends/DIGEST.md:184` (and `:32`) +- A "structural wall" resolves to your own tooling (the rodata-island foundation resolved end-to-end in session C) — `docs/decision-log.md:1495`, `docs/matching-cookbook.md:8611` diff --git a/.run/P33.5/log-mining/Phase8-13.md b/.run/P33.5/log-mining/Phase8-13.md new file mode 100644 index 0000000000..351fc73474 --- /dev/null +++ b/.run/P33.5/log-mining/Phase8-13.md @@ -0,0 +1,158 @@ +# Log mining — Phase8-13 +Files/ranges: `phase-ends/logs/Phase8.md`:1-53 · `Phase9.md`:1-55 · `Phase10.md`:1-102 · `Phase11.md`:1-96 · `Phase12.md`:1-65 · `Phase13.md`:1-29 · Lines read: 400 of 400 +Candidates considered: 38 · NEW: 4 · ALREADY-BANKED: 34 + +> Grep corpus used for every "already banked?" test (`$C`), verbatim: +> `docs/matching-cookbook.md docs/cookbook-index.md docs/decision-log.md docs/accelerators.md docs/retrospective.md docs/wave-playbook.md docs/how-to-ai-decomp/{00..12}-*.md decomp-architect/corpus/decomp-kernels.md decomp-architect/templates/registry-E.decomp.md phase-ends/DIGEST.md` +> +> **Headline:** Phases 8-13 are the most thoroughly distilled span I could have drawn — cookbook §9.x/§9.7/§11/§12/§13 +> were written *from* these logs, so ~90% of every candidate is already banked, often verbatim. The residue is four +> items, and one of them (C1) is not merely unbanked: **the distilled record still asserts the claim these logs refuted.** + +## NEW + +### C1 — A toolchain/SDK-version detector's verdict is a HYPOTHESIS until a placement COUNT backs it; when a version stamp and a byte-probe disagree, the probe wins — and the refuted stamp must be un-banked everywhere it was written +- **Evidence:** `phase-ends/logs/Phase10.md:44-54` — + > "DetectPsyQ on the imported resident program reports **PsyQ Version = 470** (the EXE is 4.0.0) … the one PsyQ-signature hit in-range is `DsMix`" + > "**Implication:** the resident's PsyQ library linking (Phase 11/12) uses **4.7**, not the EXE's 4.0 — so the 4.7 `.LIB`s are a needed asset" + + and the refutation, `phase-ends/logs/Phase12.md:41-46` — + > "**Result: NIL library footprint.** … 4.7 libsnd **1/226** (a 4-ins coincidence `ut_rev_2.o`), libspu **0/134**, libgte **0/509**" + > "R24's 'resident is 4.7 → link its 4.7 libs' is **moot** — the DetectPsyQ 4.7 signal was one coincidental DsMix-region pattern, not a linked footprint." + > "the resident uses the Makefile's default pinned triple (no override) … **Two byte-exact matches** confirm it" +- **What happened / what it cost:** A signature scanner reported a second SDK version for the second binary on the strength of + **one** in-range hit. That single number was promoted to a project rule (R24), carried forward in the PhaseEnd's "Notes for + Future Phases" (`Phase10.md:52-54`), acted on by sourcing/converting/checksumming a whole second SDK + (`Phase11.md:10` — "PsyQ 4.7 = sha-record only … `tools/psyq/conv47/` + `psyq-4.7-converted.zip`"), and used to shape an + entire Block A of the Phase-12 plan around linking those libs. Phase 12's T1 then measured the footprint with a + denominator — 1/226, 0/134, 0/509, 0/61 — and T2's two byte-exact matches proved the binary's triple was **identical to + the one already in use**. Every 4.7 task was removed mid-phase (`Phase12.md:19,59`). +- **Not banked — greps:** + `grep -n -i -E 'DetectPsyQ' $C` → **0**; + `grep -n -i -E 'nil footprint|NIL library' $C` → **0**; + `grep -n -i -E 'signature hit|single hit|version detect|footprint survey' $C` → **0**; + `grep -n -i -E 'resident is 4\.7|resident.*PsyQ 4\.7|refuted|moot' $C` → 12 hits, **all about other refutations** (the + producer-census, the pin-crash wall, the §42e wall) — none about the SDK version; + `grep -n -i 'dsmix' $C` → 1 hit (`docs/decision-log.md:2260`, an unrelated phantom-function bug). + The failure museum's 37 exhibits (`docs/how-to-ai-decomp/12-failure-museum.md:9-45` — one row per exhibit) and the retrospective's belief + table (`docs/retrospective.md:20-36`) both cover Phase 12 — but only for the script-VM belief (exhibit #3), never this one. +- **⚠ LIVE STALE ASSERTION (worth more than the lesson):** `docs/matching-cookbook.md:978-981`, the §11 heading + **"Per-binary toolchain provenance (R24)"**, still reads *"the EXE is PsyQ 4.0, the **resident is 4.7** + (`tools/psyq/conv47/`…). Never assume one binary's SDK applies to another — the 4.0 libs won't byte-match the + resident's 4.7 objects."* Every clause after the first is false on the bytes: the resident links **no** stock objects of + **either** version, and its compiler triple is the EXE's. `phase-ends/DIGEST.md:76` repeats it ("it detects PsyQ 4.7"). + A future project inheriting the cookbook inherits the error. (I am read-only; flagging, not editing.) +- **Proposed home:** DK (a kernel, paired with DK-8's "second oracle on anything that steers strategy") **+** a failure-museum + exhibit **+** a correction to cookbook §11 / DIGEST R24. The kernel's prescription: *a version/SDK stamp is a lead, not a + finding — before it steers a plan, run the placement survey and quote its denominator, and run the detector against a + binary whose version you already know (Phase 12 did exactly this: "4.0 libsnd vs the EXE snd region = **35/163** placed — + the tool works + is version-sensitive", `Phase12.md:42`).* It also **sharpens the banked bootstrap order** + (`docs/how-to-ai-decomp/03-bootstrap-order.md:22-25`, "Library version stamps, **then** idiom-revealing probe + functions"): the two steps can disagree, and the ranking is not stated there — the probe function's bytes win. +- **Portable because:** every console decomp starts by fingerprinting an SDK from signatures, on every binary it finds; a + one-hit positive with no denominator is the default output of every signature matcher ever written. + +### C2 — Histogram a secondary binary's `jal` targets BY ADDRESS RANGE before assuming it carries its own copy of anything +- **Evidence:** `phase-ends/logs/Phase12.md:45` — + > "Call-target scan: **61 distinct EXE-range `jal` vs 37 internal** — the resident calls the EXE's SDK/engine. + > **Resident = game code, no linkable library. Verified by both oracles.**" + + corroborated at `Phase12.md:43`: "the PsyQ SDK lives in the **EXE** (959 LINKED); the **resident is entirely custom engine + code** that calls the EXE's resident SDK + engine fns via fixed addresses (no RAM-wasting SDK duplication in an + always-loaded blob)." +- **What happened / what it cost:** The whole 4.7 detour (C1) is answered in one cheap scan that needs no SDK, no signature + database and no decompiler: bucket every `jal` target by which binary's address range it lands in. A blob whose calls + leave its own range is *linked against* the primary image, so it cannot contain the library code you are about to go + source. On a memory-constrained console this is the expected architecture, not the exception — an always-resident blob + that duplicated the SDK would waste the RAM the design exists to save. +- **Not banked — greps:** + `grep -n -i -E 'call-target|call target' $C` → 3 hits, **all** about a *wrong call target being invisible to a + relocation-masked diff* (`docs/matching-cookbook.md:5120`, `docs/decision-log.md:1988`) or a DESTPTR cross-check + (`:36769`) — a different subject; + `grep -n -i -E 'jal.*range|cross-range|calls into the (EXE|main)' $C` → 1 hit + (`docs/matching-cookbook.md:11125`, a liveness-across-`jal` question — unrelated); + `grep -n -i -E 'no.*duplicat.*SDK|SDK duplication|shares the (EXE|main)\x27s' $C` → **0**. + Cookbook §11 records the *consequence* ("overlays *call*, don't embed, the resident", `:968`) but never the + **instrument** that measures it, and never as a step to run on a newly-discovered binary. +- **Proposed home:** accelerator + a line in the bootstrap-order chapter's Phase-1 census (alongside "duplication, + families, reach × size, the unique tail"): *for every non-primary binary, the call-target range histogram, before any + library survey.* +- **Portable because:** any multi-binary game (overlays, DLLs, a kernel + modules, a resident + transients) answers + "does this binary contain library code or borrow it?" from its own call targets, with a disassembler and a histogram. + +### C3 — A raw blob's load address is a hypothesis; the free confirmation is arithmetic against the NEXT known segment's base +- **Evidence:** `phase-ends/logs/Phase10.md:13` — + > "Load **vram 0x800CEDF8** (Phase-3 T6b proven); `VRAM_BASE`=0x800CEDF8 (fileoff 0→vram). Computed **end vram + > 0x80128154** (4 B under overlay slot 0x80128158 — boundary corroboration)." +- **What happened / what it cost:** Before carving a single subsegment of a headerless 365,404-byte blob, the phase + checked `base + size` against the *already-known* base of the segment that loads next. It landed 4 bytes under it. That + is an independent confirmation of the load address — obtained from arithmetic, at zero cost, ahead of the byte-match + iteration that would otherwise have been the first thing to discover a wrong base (and would have presented as a + mystery diff, not as "your base is wrong"). The blob had **no header segment** (`Phase10.md:14`), so nothing in the file + itself carried the address. +- **Not banked — greps:** + `grep -n -i -E '(verify|confirm|corroborat|prove).{0,40}(load address|vram base|base address)' $C` → **0**; + `grep -n -i -E '(load address|vram base|base address).{0,40}(corroborat|cross-check|neighbou?r)' $C` → **0**; + `grep -n -i -E 'corroborat' $C` → 78 hits, **all** cookbook card cross-references ("corroborated by wave dl…") — a + different sense of the word; + `grep -n -i -E 'end vram|code end|end address' $C` → 3, all about function epilogues / loop back-edges. + The flat-blob recipe is banked (`docs/matching-cookbook.md:1091-1096`, SETUP §6.7) but it starts *from* a known base. +- **Proposed home:** accelerator, or a bullet in the flat-blob recipe (cookbook §11 "notes for reuse"). +- **Portable because:** every console has a documented memory map with adjacent, known segment bases, and every + headerless payload's load address starts as an inference from a loader trace. + +### C4 — ⚠ LOW VALUE, recommend dropping: during a parameterization refactor whose oracle is a byte-locked binary, generalize only what a second target actually needs +- **Evidence:** `phase-ends/logs/Phase9.md:26` — + > "**Refinement vs plan:** SDK-region vars (LIB*_ELF…) left un-namespaced — they're already main-only by the ifeq gate; + > namespacing deferred to when a 2nd binary needs SDK regions (avoids speculative churn, consistent with de-risked scope)." +- **What happened / what it cost:** The Phase-9 plan called for namespacing every Makefile variable; the executor + namespaced only what the second binary would actually resolve, on the grounds that each unnecessary edit is diff + surface against a byte-locked oracle carrying no proof with it. It held — Phase 10's resident and Phase 13's four + overlays never needed those vars. +- **Not banked — greps:** + `grep -n -i -E 'speculative' $C` → **0**; + `grep -n -i -E 'second instance|rule of three|generali[sz]e (only )?(at|on) the second' $C` → 9 hits, all cookbook cards + citing "a second instance of §187/§343" (evidence corroboration, not scope discipline); + `grep -n -i -E 'de-risked scope|speculative churn' $C` → **0**. +- **Proposed home:** none — **recommend NOT banking.** It is generic YAGNI wearing a decomp costume; its one + decomp-specific edge (an edit with no negative control attached is risk without proof) is already covered by the banked + negative-control law at `docs/matching-cookbook.md:842-845`. Listed only so the parent can see it was considered and + judged, rather than missed. It also sits in tension with DK-5 (`decomp-architect/corpus/decomp-kernels.md:68`, "Build + propagation the moment a second binary exists") and would need that boundary drawn before it could be stated safely. +- **Portable because:** (weakly) any refactor under a byte-locked oracle; not decomp-specific. + +## ALREADY-BANKED (one line each) +- An object's `.text` size is the PADDED/aligned size, not its instruction count — a boundary set from the nominal size shifts every byte after it (libc2 SETJMP.o: 30 ins = 0x78, `.text` = 0x80, +8 global shift) — lives at `docs/matching-cookbook.md:796` +- Two libraries whose objects interleave must be linked as ONE combined region; disambiguate an aliased address by byte-matching the *linked* `.text` — `docs/matching-cookbook.md:801-805` +- An object whose `.bss` commons the original linker scattered can't be reproduced by one NOLOAD base → exclude it BY ADDRESS and bank the other 60 (the GS_001 class) — `docs/matching-cookbook.md:806-811` +- Never byte-check an incremental build — `psyq_integrate` rewrites the `.ld` in place and a rebuild can transiently mis-resolve a sibling library (a harvest falsely diffed in libmcrd) — `docs/matching-cookbook.md:812-815`; R22; `docs/how-to-ai-decomp/12-failure-museum.md:18` (exhibit 10) +- Byte-correct-but-not-hand-written code needs its own progress category (LINKED), never counted as a stub — the metric jumped 20.31% → 50.24% with no bytes changed — `docs/matching-cookbook.md:816-818` +- Derive the progress metric from the build's own single source of truth (the Makefile's `psyq_integrate` calls), and land the parser change in the SAME commit as the call-site change — `docs/matching-cookbook.md:816-817`, `:846-852` +- Transitional-default technique: give each tool a param DEFAULTING to the kept global, update callers one green commit at a time, then a final commit removes the globals → required params; refactor LEAF-FIRST so a missed call site fails loud — `docs/matching-cookbook.md:834-841` +- A pure no-op refactor proves nothing (the param may be accepted-but-ignored) — ALSO pass a deliberately WRONG value and require the build to DIVERGE — `docs/matching-cookbook.md:842-845` +- Required params with no defaults, so a second binary can never silently inherit the first's values — `docs/matching-cookbook.md:826-833` +- The dual gate: byte-identical WITH the vendor objects AND WITHOUT them (the fresh-clone stub fallback), or the public tree rots — `docs/matching-cookbook.md:727-729` (§9.3); `docs/how-to-ai-decomp/11-publishing.md:69-70`; the cost of only ever checking it by hand is at `docs/decision-log.md:3090-3095` +- A resegment/trim must preserve the existing matched C and `#ifdef NON_MATCHING` blocks — splat will NOT overwrite an existing `.c`, so a stale one mis-places everything — `docs/matching-cookbook.md:704-710` +- A flat blob's leading data word before code fights `section_order` → emit it as a no-dot `rodata` subseg; `build_path` must match the Makefile's object rules — `docs/matching-cookbook.md:1091-1096`, `:36731` +- Ghidra's raw-binary auto-analysis finds only the `jal`-reachable subset — seed it with splat's validated boundaries (`DefineFunctions.java`) — `docs/matching-cookbook.md:1113-1116` +- Call-graph BFS FAILS on overlays that dispatch through function-pointer tables; use a linear partition bounded by a validity-based `detect_code_end` — `docs/matching-cookbook.md:970-976` +- Function-boundary rule for a linear sweep: the first `jr $ra`(+delay) at/after all forward branch targets; non-contiguous bodies are the inherent residual — `docs/matching-cookbook.md:970-972`, `:975-976` +- Game-code dedup is SOURCE-LEVEL, not an object swap — the linker cannot excise bytes interior to an object, so author the body once as a macro and instantiate it per site — `docs/matching-cookbook.md:945-950`; `docs/how-to-ai-decomp/10-integration-and-propagation.md:38-39` +- Dedup value lives among peers that share a LOAD ADDRESS (134 overlays at one vram), not across binaries with different roles; schedule highest reach × size first — `docs/matching-cookbook.md:929-931`, `:968`; `docs/how-to-ai-decomp/03-bootstrap-order.md:94` +- `h_norm` should be self-consistent, NOT a byte-exact replica of another tool's normalization — the byte gate is the acceptance, so chasing the third-party tool's own inconsistencies is low-value — `docs/matching-cookbook.md:965-968` +- Collapsing duplicates whose members are already individually matched has no recovery value — leave the dedup backlog alone; propagate deliberately — `docs/how-to-ai-decomp/10-integration-and-propagation.md:64` +- Never `git add -A` in a tree with a live reverse-engineering database (or a dirty gate) — commit the named files; DB rename churn is noise — `phase-ends/DIGEST.md:193` (R23); `docs/matching-cookbook.md:6827` +- A draft's inlined scalar `typedef` is a C89 REDEFINITION error — a compile fail, not a byte miss; strip them and keep the types in `common.h` — `docs/matching-cookbook.md:1042`, `:1344`, `:2759-2760` +- A standalone match is not a bank: the single-TU build rejects bodies on conflicting shared-symbol extern types; recover by unifying widths/signatures, don't redraft — `docs/matching-cookbook.md:1660`, `:2496`; `docs/wave-playbook.md:356-395` +- Parallel-draft + an incorruptible deterministic byte-gate scales blind drafting safely (resident 1.4% → 71.7% in one session); agent over-claims cost nothing — `docs/matching-cookbook.md:983-1010` (§12) +- The non-4-aligned-overlay gotcha (`bin` subseg carve + `.incbin` asset rule + `objcopy --set-section-alignment`), ≈75% of the fleet — `docs/matching-cookbook.md:1098-1106` +- Two binaries with the SAME sha1 are one target: `src/ov_B/ov_B.c` can `#include "../ov_A/ov_A.c"` and inherit every match from one source — `docs/matching-cookbook.md:1134-1136` +- Detect `-O0` by the frame-pointer prologue signature `21F0A003` (`addu $fp,$sp,$zero`); opt level is a property of the FILE — `docs/matching-cookbook.md:274-276`, `:1521`; `docs/cookbook-index.md:24` +- Pin the compiler/assembler/flags by evidence from the binary; never inherit a sibling's triple; expect per-module variation — `docs/how-to-ai-decomp/03-bootstrap-order.md:22-25` *(C1 above sharpens the ordering of its two clauses)* +- "The engine holds a bytecode script VM" was false — it is compiled-MIPS dispatch; a written determination replaced a phase of work — `docs/retrospective.md:22`; `docs/how-to-ai-decomp/12-failure-museum.md:11` (exhibit 3) +- A tool that writes a placeholder and exits 0 on missing input is the R43 failure class — *(origin visible in this slice: `Phase10.md:81-82` installs exactly that behaviour deliberately, for the not-yet-existing second-binary sig file)* — `docs/how-to-ai-decomp/12-failure-museum.md:16` (exhibit 8); `docs/decision-log.md:769` +- A report tool silently crediting binary A's linked libraries to binary B (`progress.py` before `linked_subsegs()` was scoped) is the assert-the-denominator class — `phase-ends/DIGEST.md:224` (R41); `docs/how-to-ai-decomp/12-failure-museum.md:33` (exhibit 25) +- Census the corpus shape with instruments verified against a disagreeing oracle, with a known-true case checked first, before choosing a strategy — `decomp-architect/corpus/decomp-kernels.md:108-118` (DK-8); `docs/how-to-ai-decomp/03-bootstrap-order.md:27-40` +- A deferral/exclude list records what the TOOLING could not do on the day it was written, not a property of the functions — *(this slice labels its deferrals "low-value" at `Phase8.md:28,44`, which is precisely the property-claim G38 forbids)* — `decomp-architect/templates/registry-E.decomp.md:243-247` (G38); `docs/how-to-ai-decomp/12-failure-museum.md:28` (exhibit 20) +- Renaming is byte-neutral — a type or symbol name emits no code — so naming/RE work never risks the byte gate — `docs/matching-cookbook.md:2568`, `:5012`, `:29834` +- A verdict is the table with every row accounted for ("zero UNEXPLAINED"), not a headline percentage — `docs/accelerators.md:805`; `docs/retrospective.md:55` diff --git a/.run/P33.5/log-mining/bank.py b/.run/P33.5/log-mining/bank.py new file mode 100644 index 0000000000..dfb7b25c31 --- /dev/null +++ b/.run/P33.5/log-mining/bank.py @@ -0,0 +1,319 @@ +#!/usr/bin/env python3 +"""Bank the log-mining harvest into the kit's kernels (DK-69 … DK-80) and emit the harvest table for the phase log. + +Reads HARVEST.md (slice → candidate → cited log line), applies the ASSIGN map below (every candidate exactly once, or DROP with +a reason), writes the kernel section text with a generated `provenance:` line per kernel, inserts it into decomp-kernels.md +before the museum, updates the Coverage lines, and writes HARVEST_TABLE.md (candidate · verdict · where). +""" +import pathlib, re, sys + +D = pathlib.Path(__file__).resolve().parent +REPO = D.parents[2] +K = REPO / "decomp-architect" / "corpus" / "decomp-kernels.md" + +# ---- the assignment: kernel -> [(slice, C)] ; DROP -> [(slice, C, reason)] +ASSIGN = { + "DK-69": [("Phase26","C1"),("Phase26","C2"),("Phase26","C3"),("Phase29-2of4","C1"),("Phase29-3of4","C5"),("Phase33","C4"),("Phase30-1of2","C12"),("Phase29-4of4","C4"),("Phase33","C5"),("Phase33","C6")], + "DK-70": [("Phase30-2of2","C2"),("Phase30-1of2","C4"),("Phase30-1of2","C5"),("Phase30-2of2","C3"),("Phase29-4of4","C1"),("Phase23-27","C6"),("Phase31-2of3","C1"),("Phase31-2of3","C2"),("Phase31-3of3","C6"),("Phase31-1of3","C10"),("Phase31-2of3","C9"),("Phase31-2of3","C8")], + "DK-71": [("Phase29-3of4","C3"),("Phase31-2of3","C3"),("Phase29-2of4","C2"),("Phase29-1of4","C4"),("Phase30-2of2","C5"),("Phase30-2of2","C4"),("Phase30-1of2","C1"),("Phase30-1of2","C3"),("Phase31-2of3","C7"),("Phase29-4of4","C2"),("Phase28-32","C3"),("Phase8-13","C1")], + "DK-72": [("Phase29-1of4","C1"),("Phase29-2of4","C6"),("Phase31-1of3","C6"),("Phase29-4of4","C5"),("Phase30-1of2","C10"),("Phase29-2of4","C4"),("Phase29-3of4","C6"),("Phase7","C2"),("Phase7","C3"),("Phase28-32","C2"),("Phase30-1of2","C13"),("Phase21","C3"),("Phase21","C2")], + "DK-73": [("Phase17-18","C1"),("Phase25","C1"),("Phase30-1of2","C11"),("Phase31-2of3","C10"),("Phase31-1of3","C5"),("Phase19-20-22","C1"),("Phase17-18","C6"),("Phase15-16","C5"),("Phase15-16","C1"),("Phase29-4of4","C6"),("Phase29-2of4","C5"),("Phase7","C5"),("Phase24","C2"),("Phase19-20-22","C2")], + "DK-74": [("Phase23-27","C2"),("Phase23-27","C3"),("Phase23-27","C4"),("Phase31-1of3","C1"),("Phase30-1of2","C8"),("Phase23-27","C1"),("Phase31-2of3","C6"),("Phase24","C1"),("Phase24","C3"),("Phase17-18","C4"),("Phase29-2of4","C3"),("Phase15-16","C3"),("Phase31-1of3","C12")], + "DK-75": [("Phase23-27","C5"),("Phase15-16","C6"),("Phase31-3of3","C4"),("Phase31-3of3","C1"),("Phase29-1of4","C5"),("Phase33","C15"),("Phase21","C1"),("Phase29-1of4","C2"),("Phase29-1of4","C3"),("Phase31-1of3","C8"),("Phase31-1of3","C9"),("Phase31-1of3","C2"),("Phase25","C2")], + "DK-76": [("Phase30-2of2","C1"),("Phase31-1of3","C7"),("Phase31-2of3","C4"),("Phase30-1of2","C2"),("Phase31-3of3","C7"),("Phase28-32","C6"),("Phase31-1of3","C11"),("Phase31-1of3","C3"),("Phase30-2of2","C6"),("Phase30-1of2","C9"),("Phase33","C16"),("Phase30-1of2","C7")], + "DK-77": [("Phase24","C4"),("Phase29-2of4","C7"),("Phase29-1of4","C6"),("Phase31-1of3","C4"),("Phase31-2of3","C5"),("Phase15-16","C4"),("Phase15-16","C8"),("Phase31-3of3","C5"),("Phase28-32","C4"),("Phase29-3of4","C4"),("Phase28-32","C5"),("Phase7","C4"),("Phase31-2of3","C11"),("Phase28-32","C1")], + "DK-78": [("Phase15-16","C2"),("Phase15-16","C7"),("Phase7","C1"),("Phase30-1of2","C6"),("Phase29-3of4","C1"),("Phase17-18","C3"),("Phase17-18","C2"),("Phase8-13","C2"),("Phase8-13","C3"),("Phase33","C7")], + "DK-79": [("Phase17-18","C5"),("Phase29-4of4","C3"),("Phase23-27","C7"),("Phase31-3of3","C2"),("Phase31-3of3","C3"),("Phase29-3of4","C2"),("Phase33","C14"),("Phase23-27","C8"),("Phase28-32","C7"),("Phase33","C8")], + "DK-80": [("Phase33","C3"),("Phase33","C9"),("Phase33","C10"),("Phase33","C11"),("Phase33","C12"),("Phase33","C13"),("Phase33","C1"),("Phase33","C2"),("Phase25","C3")], +} +DROP = [("Phase8-13", "C4", "the miner's own verdict: low value — a refactoring generality with no measured cost")] + +TITLES = { + "DK-69": "An instrument's blind spots: the tests that pass by construction", + "DK-70": "A verdict has three staleness axes, and the health suite asserts the work was done", + "DK-71": "What earns belief: applicability, independence, prediction, and the refusal as a finding", + "DK-72": "Denominators, units and labels — the number must say what it counts", + "DK-73": "Leverage is not tractability; scope a campaign by verdicts, invariants and kill criteria", + "DK-74": "Models and prompts: targeted context, named degenerate outputs, two-sided caps, tiers by what they can learn", + "DK-75": "The unattended run: a crash is a pause, a stop is a file, a limit is an epoch", + "DK-76": "Agents and the tree: write-isolation is architecture, never a sentence in a prompt", + "DK-77": "Edits that keep their proofs: repair the caller, land changes separately, draft before you carve", + "DK-78": "The search harness and the compiler as evidence: same context, own corpus, a sibling's silence proves nothing", + "DK-79": "Maintaining the knowledge base and the record: contradictions are work items, edits are verified by section", + "DK-80": "Hosts and services: the small facts that each cost an hour", +} + +BODIES = { +"DK-69": """- **Kernel:** a check that cannot fail is not a check, and several common shapes cannot. A round-trip selftest of a + partition or rewrite tool is a serialisation check, not a coverage check — it passes by construction when a missed item is + absorbed into its neighbour's span, so every such tool needs an independent detector of items it failed to anchor. A set + that gates work must be reconstructible from committed artifacts; a roster kept in ignored scratch is an unversioned oracle, + silently wrong for anything it was not named after and blind after a fresh clone. A guard allowed to sit red and uncalled + does not exist: its value is zero until it is green on the head commit and invoked by the standing report, and a docstring + claiming it is wired is not wiring. A verification flag that short-circuits the tool's write path leaves the stale artifact + in place and still exits zero. A status line a script prints unconditionally is not a measurement — derive every conclusion + the script emits from the command's own output. A coverage instrument that infers its denominator from *open* work inverts + at 100% (a complete map read as everything missing): carry the scanned denominator in the artifact and test the instrument + at both endpoints. A metric that re-parses source is blind to a body banked through an include; trust the metric derived + from the stub oracle. A name grep is not a "defined here" oracle — a declaration carrying the name reads as a definition. + An annotator that writes into the text it reads must never treat its own output as evidence, and regex-extracted evidence + needs a structural marker or prose becomes data. +- **When it applies:** every selftest, health line, coverage figure and "is it banked" query — at the moment it is written. +- **Cost:** a parser defect that survived eight sessions under a green selftest; a rename-drift failure undetected across two + phases beside a red detector; a decision document nine days stale under a green flag.""", +"DK-70": """- **Kernel:** a stored verdict can be stale because the DRAFT changed, because the BASELINE changed — or because the + INSTRUMENT changed. When a tool is repaired, every verdict it produced becomes a hypothesis again; re-gate the drafts the + repair's blast radius plausibly touched, scoped by that radius and never the whole ledger. A ledger row with no draft + artifact is a rumour, not a result. Count agent completions from the run journal's result records, never from artifact + existence — an agent writes its deliverable early and then iterates, so the file proves nothing. A clean-looking verdict + that appears immediately after your own repair transform is a suspect, not a result: re-measure the artifact the transform + produced before routing the residual. The aggregate check target is proven fail-closed before checks are added to it, and + every audit oracle has a dependent that would notice its absence. The health suite asserts that a tool DID its work, not + only that the data is intact: zero inputs, an impossible wall-clock and a missing persistent effect are each a defect. A + health check that cannot finish is not a check — keep the health target sampled and fast, and put the exhaustive form behind + its own name. Incremental gates never exercise the extraction step, so regeneration rot is undated and invisible for weeks — + sweep it on a schedule. A validity stamp must be honoured by every downstream consumer, and a wrong write-side label cannot + be repaired by a correct read key. Every status claim in an agent's context is expiry-checked against live state, or agents + report it back as an observation. +- **When it applies:** after any tool repair; in every health target; in every ledger read by a fresh session. +- **Cost:** four byte-correct drafts banked unchanged a month late; a health target that had never completed; three agents + reporting a red baseline that was a stale note in their pack.""", +"DK-71": """- **Kernel:** an oracle that can always be RUN is not always APPLICABLE — state the applicability precondition beside the + recipe, or a coarse run returns a large number that reads as a verdict. "Independent" names the instrument, not the input: + two refusals of two separately-written drafts from one tool is one test repeated. A diagnosis earns belief when it predicts + its own residual membership, not when it explains the failures already seen. A defect reasoned into a sibling tool is latent + until a run shows its signature; do not patch on theory right after that tool produced a clean run. A failure that will not + reproduce earns a negative-control-proven detector, not a speculative fix — and every abort path proves its revert by diffing + the worktree against a baseline captured at the start of the run. A derived claim outranks a heuristic verdict; when two + heuristics disagree, take the union and queue the disagreements — under-reporting hides work, over-reporting only costs + review. An instrument's refusal is a finding, not an obstacle: overriding it means explaining why the instrument is wrong, + never finding another route. Read the first ten results of a long run before trusting the other hundreds, and + negative-control any new refusal against everything that already succeeded. A refusal names the branch the caller entered, + not the subject — make the applier consult the classifier it already has. Classify a harness fix as a logic defect or a + path-reachability gap and price it accordingly; only the logic defect generalises. Re-verify a task's premise in the code at + execution time — roadmap lines, audit findings and even an audit's own correction footer go stale, and the document that + named a defect is usually the first to. A toolchain-version detector's verdict is a hypothesis until a placement count backs + it; when a version stamp and a byte probe disagree, the probe wins and the refuted stamp is un-banked. +- **When it applies:** every diagnosis, every disagreement between two instruments, every premise inherited from a document. +- **Cost:** a "genuine codegen" verdict that was a size mismatch; a classifier verdict that overrode a hash match and + understated a whole bucket; a version stamp banked for a phase against the bytes.""", +"DK-72": """- **Kernel:** a number that does not say what it counts will be read as the wrong thing. A stop/continue instrument + aggregates at exactly the unit the decision is made in; one that averages a finer unit manufactures a false "we are at the + floor". A reach-weighted gain (size × copies) is not a size — every figure says which of the two it is. Sibling count and + never-drafted count are different denominators; conflating them overstates free leverage and hides that the remaining mass + is singletons. A yield estimator that counts "unclaimed at the moment it runs" ranks correctly and over-projects absolutely; + never plan off its absolute numbers. Do not cross-price two economies: a conversion rate measured on the residue queue does + not price a fresh wave. When two blockers are orthogonal, a classifier's if-chain order silently becomes the label — + cross-tabulate, never bucket. Measure what fraction of a cycle a parallelism knob can actually touch before adopting it. A + milestone counted in matched functions excludes the splitter's auto-generated empty bodies, defined at the moment the bar + is set. A duplicate census run before the vendor library is linked out is contaminated — the groups are library fragments + and epilogues. The file a function lives in is not evidence of its class; read the recorded attribute, never the hosting + split. Keep a glossary line for any term two documents use in opposite senses. Say which currency a wave buys — percentage + or idioms — before launching it, and judge it in that currency; and a class-distribution assessor only sees the population + already attempted, so "analyse all remaining work" is a cheap triage pass, not a static analysis. +- **When it applies:** every plan figure, every ledger column, every verdict of "at the floor". +- **Cost:** a phase nearly closed on an artifact of the wrong unit; a leverage estimate three times too high; a thirty-fold + mis-scope risk from one term meaning two things.""", +"DK-73": """- **Kernel:** the most-duplicated functions are systematically the hardest — leverage and tractability are + anti-correlated — so a leverage-first queue front-loads hand-tier work, and its early bank rate is not a harness fault. + Carry a measured closeness read per target and never let reach × size stand in for "crackable"; the reach ranking finds the + most-DONE work first, so derive the target pool from the build's own invariant. Yield clusters by binary, not across the + fleet — draw per binary once two independent lanes concentrate in the same place. A cracked idiom transfers within its + family and not across it: price a lane by families, not by class size. An open-ended grind phase's milestone is invariants + held plus a clean checkpoint, never a percentage; a research phase is scoped by a per-class verdict (a validated lever or a + falsifiable wall verdict per class), not by a percentage either. Write numeric kill criteria into the plan before the data + exists, and let them fire. Measure a pipeline's yield on the residual, not on solved functions: a known-answer ladder (revert + a match to a stub, make the pipeline re-derive it) sets the ceiling, and the gap to the unmatched tail is the real number. + Declare a mechanical lever spent only on a positive, three-part measurement — every built lever applied and returning zero, + the residue split by structure, and the decay curve priced against what remains. Sequence a phase so the cheapest thing that + can invalidate everything below it runs first; when a foundation task hits a structural wall mid-phase, bank the tractable + wins and re-scope the wall as its own sub-project. Inside one leverage class, schedule by measured remaining effort and pull + the payoff-dominating outlier out for an immediate cheap triage. Choose the exemplar for cracking a codegen class by the + size of its residual: the one-instruction mismatches are the cleanest real-function isolates. +- **When it applies:** the draw, the phase plan, the campaign's close. +- **Cost:** a phase priced by percentage that could only be closed by verdicts; a mega-leverage "freebie" that was a + stack-switcher; a lane priced by class that cost forty turns of learning per family.""", +"DK-74": """- **Kernel:** give a drafting model a targeted slice of the knowledge base, never the whole — full context measurably made + a model worse. Name the degenerate output in the prompt: an empty body compiles, so "translate every instruction, never an + empty body" is a required instruction. The output-token cap is a two-sided knob and both failure modes read as "the model is + bad"; more budget is not more quality — measure it as a paired A/B and treat truncation as recoverable, not as a defect + signal. When A/B-ing any harness knob, ship a positive control that the knob actually moved. Mine new idioms from fresh + cracks, never from the failed backlog — the failure pile re-teaches what you already know — while the harvest SELECTOR must + still be able to see failed attempts, or it learns from the easy half. Escalations to the expensive tier run strictly serial + with idiom-banking between them; only the tier that cannot learn is run in parallel. The wall-breaker tier is a match tier, + not a plumbing tier: work a deterministic arbiter can judge does not need the expensive model. Accelerate the stage that is + the bottleneck — a byte-exact search loop costs a compile plus a whole-binary gate per candidate, so hardware brute force + buys nothing. Make the drafter run the pre-gate guard and return the NAMED banking prerequisite; a batch of diagnosed + candidates is worth more than a batch of opaque matches. Before concluding a pipeline is weak, histogram the compiler's + error text: most failures were one missing declaration, fixed once. Keep prompt and law text in data, never inside the + launcher's source template. +- **When it applies:** the pack builder, the routing table, every model experiment. +- **Cost:** a local model made worse by more context; eight diagnosed matches revealing one lever the opaque batch had + hidden; a wave killed at launch by a quote character inside a template.""", +"DK-75": """- **Kernel:** build the unattended campaign so a crash is a pause — probe the dependency at the top of each cycle, commit + per cycle, persist the tried-set and each confirmed result the moment it is confirmed; a long stateless batch that writes + only at the end loses everything to a kill. Design it for a human with no agent session: a STOP file honoured at a safe + boundary, a supervisor that tells a clean exit from a crash, a status one-liner, and crash-resume proven by a deliberate kill + before the first real run. Never wrap a project tool in a timeout shorter than its own budget — you pre-empt its recovery + handler and lose its buffered output. Sweep for orphaned worker processes at every session boundary: a dead-pipe compiler + holds a core forever and nothing reports it, and a harness's low-memory guard silently kills long background jobs. A run + that looks throttled is usually blocked on an interactive approval prompt — check the pending prompt before diagnosing the + provider. Batch size is a risk lever, not a token lever: isolated agents cost about N times one agent whether concurrent or + serial, so size a batch by the unverified spend you are willing to lose before the next measurement. A repair mode whose + cost is exceptions × population is gated on a measured exception count; for a broadly divergent set, drop rather than + recover. A free or preview model tier can be withdrawn mid-campaign without notice — a fleet-wide 404 is an epoch event, not + N model failures — and a metered key's own cap is a separate limit from the account's credit. In a pipelined + drafter/gater, "still open" is not "not yet attempted": consecutive waves re-drafted the wave still in flight. A fan-out + script generated by an orchestrator runs sandboxed without the repository: it is self-contained, so target selection + belongs to the generator, not the workers. +- **When it applies:** every lane that runs while nobody watches. +- **Cost:** a thirteen-hour orphaned compiler; three gates with no verdict and half-applied propagations from one timeout; + a six-hour "throttle" that was a permission prompt.""", +"DK-76": """- **Kernel:** drafting agents must never be ABLE to write the build tree; every agent artifact lands in a scratch directory, + so a killed or racing campaign costs build cycles and zero paid work. "Never modify the source tree" in a prompt is a + request, not an enforcement — snapshot the tree status around every agent and name the offender. Never adopt a subagent's + worktree wholesale: it is a snapshot of an older tree and may predate a bank; re-gate its artifacts against the head. The + generated disassembly tree is shared mutable state — a fleet verify/clean chain and the per-function instruments cannot run + at the same time. A probe that splices the tree in order to measure it must restore it, or the progress oracle counts the + splices as banks. Snapshot every target's disassembly before gating: a successful bank prunes it, and harvest and recovery + need both sides. Never let model-authored prose reach the shell inside double quotes — a backticked command in a commit + message executed; use a quoted heredoc. Keep the wave harness in the repository with its contracts; a harness rebuilt from + memory each run silently goes stale. A `cd` in one agent shell call persists into the next. Never test a helper by importing + its module: a tool with no main guard runs its whole pipeline on import. +- **When it applies:** the day the first agent is launched, as architecture; then every wave. +- **Cost:** two tree corruptions recovered with one checkout only because the drafts lived outside the tree; a real extract + run by a commit message; a download landing inside a submodule.""", +"DK-77": """- **Kernel:** when a byte-proven body will not integrate, the repair moves the CALLER's declaration to the definition's + signature — never the definition to the caller's — and only where the change is width-compatible. A repair ladder probes + whether each stage is needed before applying it, or it silently escalates a binary-local bank into a fleet-shared edit. Land + a pure rename and a semantic or layout change as separate gated edits, so a gate failure attributes itself. Draft first, + then carve: a build-unit split is safe only when the new unit is immediately populated with proven bodies. A proven + transform that is not a rung of the ladder the drafts actually pass through does not exist for those drafts. Re-run the + deterministic declaration canonicaliser over old quarantined drafts after every large bank — recovery odds rise as the + banked corpus grows and the pile costs nothing to keep. An idempotency guard keyed on presence freezes every record created + before the system matured; key it on completeness, and make regenerated artifacts idempotent by replacement, never by + skipping. A hard-coded assumption fixed in one tool survives in its siblings: grep the tree for the literal and fix every + twin in the same change. Before a fleet-wide mechanical edit, census the whole population for the exact preconditions the + edit assumes — uniformity is the licence, non-uniformity the design input. A batch gate that bisects on failure re-runs the + singleton against an unchanged baseline: special-case one, or pay a duplicate build on the hot path. A byte-identical + baseline is proven only when the whole gate is green from clean across several independent sessions. A build-system + conditional that expands at parse time makes its own negative control vacuous. Never round-trip a curated configuration + through a serializer: every oracle you own measures bytes, so a formatting-destructive write is invisible to all of them. +- **When it applies:** every integration repair, every carve, every fleet-wide edit. +- **Cost:** a registry's forty-seven comment lines destroyed under green gates; a fleet edit taken by a bank that needed a + local one; a duplicate build on every single-draft gate.""", +"DK-78": """- **Kernel:** a search harness must compile in the SAME declaration context as the real build; an isolated context does not + merely fail to verify, it makes the search converge on the wrong answer. The search unit is the C expression — it cannot + freeze the instructions already right, because register allocation couples them. A decompiler's "unaffected register" + output means it decompiled one entry path of a multi-entry function and handed you a fragment: an instrument limit, never + evidence the function is hard. Check group identity from the signature files before probing a family — if the members are + structurally identical, a zero result is a compile-error certainty, not evidence about codegen. Vendor compiler sources + carry form-feed page separators, which a scripting language's line splitter honours and grep does not, so a line-number + checker over the source drifts and blames the wrong line. Your own corpus of byte matches is an experiment already run on + the toolchain: settle "is my rebuilt compiler faithful?" from it before installing the original vendor tools. Another + project's unmatched stubs are a record of what they did not crack, never proof that a class is uncrackable — + cross-project corroboration multiplies confidence in a wrong verdict as readily as a right one. Histogram a secondary + binary's call targets by address range before assuming it carries its own copy of anything; a raw blob's load address is a + hypothesis whose free confirmation is arithmetic against the next known segment's base. A relocation-interleaved + disassembly is produced only for object files; a linked image lists relocations separately with a shifted address column. +- **When it applies:** the permuter's base file, the first probe of any family, every cross-project citation. +- **Cost:** an overnight run that "closed" forty per cent of its near-misses and gated zero; a class written off on a + neighbour's silence.""", +"DK-79": """- **Kernel:** a contradiction between two entries of your own knowledge base is a work item, not noise — replay the levers + already written down, under the correct oracle, before commissioning new research. A wrong prescription left in the base is + worse than no entry: when evidence refutes an entry, correct that entry in place, in the same session, carrying the + refutation. A document that cites a repository path is an untested claim about the repository — lint it. A programmatic + edit to a long-lived knowledge document silently truncates or duplicates it; verify the sections, never the commit. The + live hand-off block is strictly appended at the end of its file — file order is the only recency signal a fresh session + has. Never restructure a proven tool with blind string replaces at the end of a long session; specify it as the next + session's first task. A miner over your own records finds only what its pattern anticipates — measure the widened pattern's + yield. Assert that the work ledger partitions the live work, and treat a row the invariant refutes as a lie a fresh session + will act on. Write the phase synthesis in a fresh session that re-reads the committed state cold; the cold read is what + catches stale artifacts. A derive-then-apply pipeline over a live repository needs a freshness guard and a stated sequencing + law. +- **When it applies:** every harvest, every close, every programmatic edit of a document that outlives the session. +- **Cost:** a cookbook section silently deleted for a session; a day-older hand-off read as the live one; a class re-researched + because two entries disagreed and nobody replayed either.""", +"DK-80": """- **Kernel:** small facts about hosts and services, each learned at the cost of an hour. `git check-ignore` is silent for + tracked paths, so an ignore-coverage audit run before the untracking passes vacuously — use its no-index form. A + content-hash "no forbidden bytes" audit collides on zero-length files. A mirror push does not push the stash ref. Route a + host purge request through the flow that actually exists; the obvious form is a trap. A host feature can be gated on the + very flip it was meant to precede — read the settings page, do not infer. A public scratch service's compiler image is not + your pinned toolchain; rebuild it locally and prove byte-identity before asking for a preset. A disassembler's script + directory compiles as one bundle, so a single non-compiling script disables every script in it and the error names a + working one. The same tool refuses a project path containing a component that starts with a dot, so a scratch project + cannot live under a dot-directory. Run reference-compiler dump passes from a scratch working directory, or the dumps land + in the repository root and later read as committed artifacts. +- **When it applies:** the first time each host or service is touched. +- **Cost:** an hour each, and one flip-gate misread.""", +} + + +def parse_harvest(): + text = (D / "HARVEST.md").read_text(encoding="utf-8") + out = {} + slice_ = None + cur = None + for ln in text.splitlines(): + m = re.match(r"^=+ (\S+):", ln) + if m: + slice_ = m.group(1); continue + m = re.match(r"^ (C\d+) — (.*)", ln) + if m: + cur = (slice_, m.group(1)); out[cur] = {"title": m.group(2), "log": None}; continue + m = re.match(r"^ LOG (phase-ends/logs/\S+?):(\d+):", ln) + if m and cur: + out[cur]["log"] = f"{m.group(1)}:{m.group(2)}" + return out + + +def main(): + cands = parse_harvest() + assigned = {} + for dk, lst in ASSIGN.items(): + for s, c in lst: + key = (s, c) + assert key in cands, f"unknown candidate {key}" + assert key not in assigned, f"{key} assigned twice ({assigned[key]} and {dk})" + assigned[key] = dk + for s, c, why in DROP: + assert (s, c) in cands and (s, c) not in assigned + assigned[(s, c)] = f"DROP: {why}" + missing = [k for k in cands if k not in assigned] + assert not missing, f"unassigned candidates: {missing}" + print(f"bank: {len(cands)} candidates, {sum(1 for v in assigned.values() if v.startswith('DK'))} banked into {len(ASSIGN)} kernels, {len(DROP)} dropped") + # section text + sec = [] + for dk in ASSIGN: + cites = "; ".join(f"{cands[(s, c)]['log']} ({s} {c})" for s, c in ASSIGN[dk]) + sec.append(f"### {dk} — {TITLES[dk]}\n{BODIES[dk]}\nprovenance: BFM phase-log mining pass (P33.5 task 14.5, S92; the accelerators entry \"P33.5 S92\"): {cites}\n") + section = "\n".join(sec) + k = K.read_text(encoding="utf-8") + old_head = "## 8. Added by the coverage pass — the record's residue\n\n*The kit's coverage check derives every rule and every hindsight entry of the source project and refuses one that is neither\ncited by a provenance line nor dispositioned. These are the lessons that check found in the record and nowhere in the kit.*\n" + assert k.count(old_head) == 1 + new_head = ("## 8. Added by the coverage and log-mining passes — the record's residue\n\n" + "*The kit's coverage check derives every rule and every hindsight entry of the source project and refuses one that is neither\n" + "cited by a provenance line nor dispositioned; the three kernels after it are what that check found in the record and nowhere in\n" + "the kit. The twelve after those come from a pass that read every phase worklog of the source project once more (some thirty\n" + "thousand lines, one read-only agent per slice) for lessons banked in none of its distilled records — 777 candidates, 634 already\n" + "banked, 143 new, clustered here by theme. Each provenance line names the worklog lines the lessons came from.*\n") + k = k.replace(old_head, new_head) + anchor = "\n---\n\n## 9. The failure museum, condensed — what looked right at the time\n" + assert k.count(anchor) == 1 + k = k.replace(anchor, "\n" + section + anchor) + old_cov = "- **Added by the coverage pass:** DK-66 … DK-68 — 3 kernels.\n- **In all:** DK-1 … DK-68 — 68 kernels" + assert k.count(old_cov) == 1 + k = k.replace(old_cov, "- **Added by the coverage pass:** DK-66 … DK-68 — 3 kernels.\n- **Added by the log-mining pass:** DK-69 … DK-80 — 12 kernels (143 worklog lessons, clustered).\n- **In all:** DK-1 … DK-80 — 80 kernels") + K.write_text(k, encoding="utf-8") + n = len(re.findall(r"^### DK-", k, re.M)) + print("kernels now:", n) + # harvest table + rows = ["| Slice | C | Lesson | Log line | Verdict |", "|---|---|---|---|---|"] + for (s, c), v in sorted(cands.items(), key=lambda kv: (kv[0][0], int(kv[0][1][1:]))): + rows.append(f"| {s} | {c} | {v['title'][:110]} | `{v['log']}` | {assigned[(s, c)]} |") + (D / "HARVEST_TABLE.md").write_text("# Harvest table — every NEW candidate of the log-mining pass, its verdict and its home\n\n" + "\n".join(rows) + "\n", encoding="utf-8") + print("table rows:", len(rows) - 2) + + +if __name__ == "__main__": + main() diff --git a/.run/P33.5/log-mining/harvest.py b/.run/P33.5/log-mining/harvest.py new file mode 100644 index 0000000000..312a18fdcd --- /dev/null +++ b/.run/P33.5/log-mining/harvest.py @@ -0,0 +1,40 @@ +#!/usr/bin/env python3 +"""Harvest the log-mining slices: list every NEW candidate with its cited log line VERBATIM (R14 — verify the claim against the +bytes, not the agent's summary), and print the denominators. Usage: harvest.py [--full] [slice ...]""" +import pathlib, re, sys + +D = pathlib.Path(__file__).resolve().parent +REPO = D.parents[2] +full = "--full" in sys.argv +names = [a for a in sys.argv[1:] if not a.startswith("--")] +files = sorted(p for p in D.glob("*.md") if p.name not in ("BRIEF.md", "HARVEST.md") and (not names or p.stem in names)) +tot_c = tot_n = tot_b = 0 +lines_read = 0 +for f in files: + text = f.read_text(encoding="utf-8") + hdr = re.search(r"Candidates considered: (\d+) · NEW: (\d+) · ALREADY-BANKED: (\d+)", text) + lr = re.search(r"Lines read: (\d+) of (\d+)", text) + c, n, b = (int(x) for x in hdr.groups()) if hdr else (0, 0, 0) + tot_c += c; tot_n += n; tot_b += b + lines_read += int(lr.group(1)) if lr else 0 + new = text.split("## NEW", 1)[1].split("## ALREADY-BANKED", 1)[0] if "## NEW" in text else "" + cands = re.split(r"^### ", new, flags=re.M)[1:] + print(f"\n=================== {f.stem}: considered {c} · NEW {n} (found {len(cands)}) · banked {b} · lines read {lr.group(1) if lr else '?'} of {lr.group(2) if lr else '?'}") + for cand in cands: + title = cand.splitlines()[0] + ev = re.search(r"\*\*Evidence:\*\*\s*`?(phase-ends/logs/Phase[\d.]+\.md):(\d+)(?:[–-](\d+))?`?", cand) + home = re.search(r"\*\*Proposed home:\*\*\s*(.+)", cand) + print(f" {title[:200]}") + if ev: + path, a, bb = ev.group(1), int(ev.group(2)), int(ev.group(3) or ev.group(2)) + src = (REPO / path).read_text(encoding="utf-8", errors="replace").splitlines() + if a - 1 < len(src): + quote = " ".join(src[a - 1:min(bb, a + 2)]) + print(f" LOG {path}:{a}: {quote[:260]}") + else: + print(f" LOG {path}:{a}: *** LINE OUT OF RANGE ({len(src)} lines) ***") + else: + print(" LOG: *** no parsable evidence line ***") + if full and home: + print(f" home: {home.group(1)[:200]}") +print(f"\nTOTAL over {len(files)} slices: considered {tot_c} · NEW {tot_n} · already-banked {tot_b} · lines read {lines_read}") diff --git a/decomp-architect/README.md b/decomp-architect/README.md index 9ee7b66921..e618150987 100644 --- a/decomp-architect/README.md +++ b/decomp-architect/README.md @@ -71,7 +71,7 @@ session and say `Begin Phase 1`. | **The firewall pack** (`templates/gitignore.decomp`, the audit template and its config, a planted fixture, the CI workflow) | no game-derived bytes in git from commit one; an audit that derives its forbidden set, asserts its coverage and fails on the fixture before it is trusted; CI on every push | | **The layout and its READMEs** (`docs/`, `.run/`, the ops-setup overlay) | where each kind of knowledge goes; scratch under the repository with dated tracked exceptions; the ops reference a fresh machine rebuilds from | | **The overlays** (`templates/pa-overlays.md`, `templates/CLAUDE.decomp-overlay.md`) | marked-section appends to `CLAUDE.md` (the decomp fail-safes and session-start extras), the effort map (Max on the plan, the PhaseEnd, the compiler pin, the segmentation decision, any wall verdict), the cookbook (the idiom entry shape, the triage table), ops-setup; the digest, the replayable checkpoint block and the PhaseEnd narrative axis | -| **The corpus** (`corpus/decomp-kernels.md`) | DK-1 … DK-68 — what the source project learned late, each with when it applies and what it cost, and the failure museum | +| **The corpus** (`corpus/decomp-kernels.md`) | DK-1 … DK-80 — what the source project learned late, each with when it applies and what it cost, and the failure museum | | **The three dictionaries** (`corpus/tools/`, `corpus/cookbook/`, `corpus/record/`) | the source project's tools, verbatim, by ladder phase, behind an index keyed by the NEED each answers (what it does, what proved it, what to adapt); its knowledge base — the cookbook, its symptom index and the codegen map — verbatim, behind a front page that says what transfers to another compiler; and its record — the how-to, the decision log, the accelerators, the retrospective, the story, the playbook, the effort doctrine, the readability charter and every phase-end — verbatim, behind a front page that says what each is, so that every rule's and kernel's `provenance:` line can be followed to its evidence. All three are regenerated from the source project and asserted equal to it; a new project installs the index and the front pages and keeps the folder as its reference shelf | | **The memory seed** (`memory-seed/`) | working agreements and harness facts learned on the source project, appended to the memory ProjectArchitect configured | | **The skeletons** | LICENSE, `src/NOTICE.md`, README, CONTRIBUTING (with the AI-conduct section), `.clang-format` and a `make format` snippet | diff --git a/decomp-architect/corpus/decomp-kernels.md b/decomp-architect/corpus/decomp-kernels.md index 9fe4658d9b..0a24fb733d 100644 --- a/decomp-architect/corpus/decomp-kernels.md +++ b/decomp-architect/corpus/decomp-kernels.md @@ -831,10 +831,13 @@ provenance: BFM decision-log "P33.5 S91-b" (the hindsight on types, 2026-09-07); --- -## 8. Added by the coverage pass — the record's residue +## 8. Added by the coverage and log-mining passes — the record's residue *The kit's coverage check derives every rule and every hindsight entry of the source project and refuses one that is neither -cited by a provenance line nor dispositioned. These are the lessons that check found in the record and nowhere in the kit.* +cited by a provenance line nor dispositioned; the three kernels after it are what that check found in the record and nowhere in +the kit. The twelve after those come from a pass that read every phase worklog of the source project once more (some thirty +thousand lines, one read-only agent per slice) for lessons banked in none of its distilled records — 777 candidates, 634 already +banked, 143 new, clustered here by theme. Each provenance line names the worklog lines the lessons came from.* ### DK-66 — A ledger's tie-break, a checker's widening and a blanket commit are part of the instrument - **Kernel:** three bookkeeping choices around an oracle that changed its verdicts without anyone reading them as part of it. @@ -877,6 +880,244 @@ provenance: BFM accelerators P33.5 S91 (1); the kit's dry-run run 1 (Step 3.7's - **Cost:** false banks, false walls, a plan built on a count nobody had checked. provenance: BFM R14 (verify recon/sub-agent summary counts against the bytes — a summarised signal is a claim), R66 (write "banked" only from the tool's printed line), the S82 coordinator that read prose results, the memory "verify blast radius, not just the defect" +### DK-69 — An instrument's blind spots: the tests that pass by construction +- **Kernel:** a check that cannot fail is not a check, and several common shapes cannot. A round-trip selftest of a + partition or rewrite tool is a serialisation check, not a coverage check — it passes by construction when a missed item is + absorbed into its neighbour's span, so every such tool needs an independent detector of items it failed to anchor. A set + that gates work must be reconstructible from committed artifacts; a roster kept in ignored scratch is an unversioned oracle, + silently wrong for anything it was not named after and blind after a fresh clone. A guard allowed to sit red and uncalled + does not exist: its value is zero until it is green on the head commit and invoked by the standing report, and a docstring + claiming it is wired is not wiring. A verification flag that short-circuits the tool's write path leaves the stale artifact + in place and still exits zero. A status line a script prints unconditionally is not a measurement — derive every conclusion + the script emits from the command's own output. A coverage instrument that infers its denominator from *open* work inverts + at 100% (a complete map read as everything missing): carry the scanned denominator in the artifact and test the instrument + at both endpoints. A metric that re-parses source is blind to a body banked through an include; trust the metric derived + from the stub oracle. A name grep is not a "defined here" oracle — a declaration carrying the name reads as a definition. + An annotator that writes into the text it reads must never treat its own output as evidence, and regex-extracted evidence + needs a structural marker or prose becomes data. +- **When it applies:** every selftest, health line, coverage figure and "is it banked" query — at the moment it is written. +- **Cost:** a parser defect that survived eight sessions under a green selftest; a rename-drift failure undetected across two + phases beside a red detector; a decision document nine days stale under a green flag. +provenance: BFM phase-log mining pass (P33.5 task 14.5, S92; the accelerators entry "P33.5 S92"): phase-ends/logs/Phase26.md:764 (Phase26 C1); phase-ends/logs/Phase26.md:742 (Phase26 C2); phase-ends/logs/Phase26.md:101 (Phase26 C3); phase-ends/logs/Phase29.md:2762 (Phase29-2of4 C1); phase-ends/logs/Phase29.md:7496 (Phase29-3of4 C5); phase-ends/logs/Phase33.md:126 (Phase33 C4); phase-ends/logs/Phase30.md:2589 (Phase30-1of2 C12); phase-ends/logs/Phase29.md:8080 (Phase29-4of4 C4); phase-ends/logs/Phase33.md:499 (Phase33 C5); phase-ends/logs/Phase33.md:503 (Phase33 C6) + +### DK-70 — A verdict has three staleness axes, and the health suite asserts the work was done +- **Kernel:** a stored verdict can be stale because the DRAFT changed, because the BASELINE changed — or because the + INSTRUMENT changed. When a tool is repaired, every verdict it produced becomes a hypothesis again; re-gate the drafts the + repair's blast radius plausibly touched, scoped by that radius and never the whole ledger. A ledger row with no draft + artifact is a rumour, not a result. Count agent completions from the run journal's result records, never from artifact + existence — an agent writes its deliverable early and then iterates, so the file proves nothing. A clean-looking verdict + that appears immediately after your own repair transform is a suspect, not a result: re-measure the artifact the transform + produced before routing the residual. The aggregate check target is proven fail-closed before checks are added to it, and + every audit oracle has a dependent that would notice its absence. The health suite asserts that a tool DID its work, not + only that the data is intact: zero inputs, an impossible wall-clock and a missing persistent effect are each a defect. A + health check that cannot finish is not a check — keep the health target sampled and fast, and put the exhaustive form behind + its own name. Incremental gates never exercise the extraction step, so regeneration rot is undated and invisible for weeks — + sweep it on a schedule. A validity stamp must be honoured by every downstream consumer, and a wrong write-side label cannot + be repaired by a correct read key. Every status claim in an agent's context is expiry-checked against live state, or agents + report it back as an observation. +- **When it applies:** after any tool repair; in every health target; in every ledger read by a fresh session. +- **Cost:** four byte-correct drafts banked unchanged a month late; a health target that had never completed; three agents + reporting a red baseline that was a stale note in their pack. +provenance: BFM phase-log mining pass (P33.5 task 14.5, S92; the accelerators entry "P33.5 S92"): phase-ends/logs/Phase30.md:4660 (Phase30-2of2 C2); phase-ends/logs/Phase30.md:2481 (Phase30-1of2 C4); phase-ends/logs/Phase30.md:2352 (Phase30-1of2 C5); phase-ends/logs/Phase30.md:4001 (Phase30-2of2 C3); phase-ends/logs/Phase29.md:8207 (Phase29-4of4 C1); phase-ends/logs/Phase27.md:17 (Phase23-27 C6); phase-ends/logs/Phase31.md:5122 (Phase31-2of3 C1); phase-ends/logs/Phase31.md:5175 (Phase31-2of3 C2); phase-ends/logs/Phase31.md:6027 (Phase31-3of3 C6); phase-ends/logs/Phase31.md:2376 (Phase31-1of3 C10); phase-ends/logs/Phase31.md:5437 (Phase31-2of3 C9); phase-ends/logs/Phase31.md:5394 (Phase31-2of3 C8) + +### DK-71 — What earns belief: applicability, independence, prediction, and the refusal as a finding +- **Kernel:** an oracle that can always be RUN is not always APPLICABLE — state the applicability precondition beside the + recipe, or a coarse run returns a large number that reads as a verdict. "Independent" names the instrument, not the input: + two refusals of two separately-written drafts from one tool is one test repeated. A diagnosis earns belief when it predicts + its own residual membership, not when it explains the failures already seen. A defect reasoned into a sibling tool is latent + until a run shows its signature; do not patch on theory right after that tool produced a clean run. A failure that will not + reproduce earns a negative-control-proven detector, not a speculative fix — and every abort path proves its revert by diffing + the worktree against a baseline captured at the start of the run. A derived claim outranks a heuristic verdict; when two + heuristics disagree, take the union and queue the disagreements — under-reporting hides work, over-reporting only costs + review. An instrument's refusal is a finding, not an obstacle: overriding it means explaining why the instrument is wrong, + never finding another route. Read the first ten results of a long run before trusting the other hundreds, and + negative-control any new refusal against everything that already succeeded. A refusal names the branch the caller entered, + not the subject — make the applier consult the classifier it already has. Classify a harness fix as a logic defect or a + path-reachability gap and price it accordingly; only the logic defect generalises. Re-verify a task's premise in the code at + execution time — roadmap lines, audit findings and even an audit's own correction footer go stale, and the document that + named a defect is usually the first to. A toolchain-version detector's verdict is a hypothesis until a placement count backs + it; when a version stamp and a byte probe disagree, the probe wins and the refuted stamp is un-banked. +- **When it applies:** every diagnosis, every disagreement between two instruments, every premise inherited from a document. +- **Cost:** a "genuine codegen" verdict that was a size mismatch; a classifier verdict that overrode a hash match and + understated a whole bucket; a version stamp banked for a phase against the bytes. +provenance: BFM phase-log mining pass (P33.5 task 14.5, S92; the accelerators entry "P33.5 S92"): phase-ends/logs/Phase29.md:7359 (Phase29-3of4 C3); phase-ends/logs/Phase31.md:3589 (Phase31-2of3 C3); phase-ends/logs/Phase29.md:3690 (Phase29-2of4 C2); phase-ends/logs/Phase29.md:1191 (Phase29-1of4 C4); phase-ends/logs/Phase30.md:4064 (Phase30-2of2 C5); phase-ends/logs/Phase30.md:4657 (Phase30-2of2 C4); phase-ends/logs/Phase30.md:1369 (Phase30-1of2 C1); phase-ends/logs/Phase30.md:250 (Phase30-1of2 C3); phase-ends/logs/Phase31.md:3346 (Phase31-2of3 C7); phase-ends/logs/Phase29.md:9957 (Phase29-4of4 C2); phase-ends/logs/Phase28.md:143 (Phase28-32 C3); phase-ends/logs/Phase10.md:44 (Phase8-13 C1) + +### DK-72 — Denominators, units and labels — the number must say what it counts +- **Kernel:** a number that does not say what it counts will be read as the wrong thing. A stop/continue instrument + aggregates at exactly the unit the decision is made in; one that averages a finer unit manufactures a false "we are at the + floor". A reach-weighted gain (size × copies) is not a size — every figure says which of the two it is. Sibling count and + never-drafted count are different denominators; conflating them overstates free leverage and hides that the remaining mass + is singletons. A yield estimator that counts "unclaimed at the moment it runs" ranks correctly and over-projects absolutely; + never plan off its absolute numbers. Do not cross-price two economies: a conversion rate measured on the residue queue does + not price a fresh wave. When two blockers are orthogonal, a classifier's if-chain order silently becomes the label — + cross-tabulate, never bucket. Measure what fraction of a cycle a parallelism knob can actually touch before adopting it. A + milestone counted in matched functions excludes the splitter's auto-generated empty bodies, defined at the moment the bar + is set. A duplicate census run before the vendor library is linked out is contaminated — the groups are library fragments + and epilogues. The file a function lives in is not evidence of its class; read the recorded attribute, never the hosting + split. Keep a glossary line for any term two documents use in opposite senses. Say which currency a wave buys — percentage + or idioms — before launching it, and judge it in that currency; and a class-distribution assessor only sees the population + already attempted, so "analyse all remaining work" is a cheap triage pass, not a static analysis. +- **When it applies:** every plan figure, every ledger column, every verdict of "at the floor". +- **Cost:** a phase nearly closed on an artifact of the wrong unit; a leverage estimate three times too high; a thirty-fold + mis-scope risk from one term meaning two things. +provenance: BFM phase-log mining pass (P33.5 task 14.5, S92; the accelerators entry "P33.5 S92"): phase-ends/logs/Phase29.md:1283 (Phase29-1of4 C1); phase-ends/logs/Phase29.md:4775 (Phase29-2of4 C6); phase-ends/logs/Phase31.md:298 (Phase31-1of3 C6); phase-ends/logs/Phase29.md:9350 (Phase29-4of4 C5); phase-ends/logs/Phase30.md:899 (Phase30-1of2 C10); phase-ends/logs/Phase29.md:4958 (Phase29-2of4 C4); phase-ends/logs/Phase29.md:6115 (Phase29-3of4 C6); phase-ends/logs/Phase7.md:12 (Phase7 C2); phase-ends/logs/Phase7.md:26 (Phase7 C3); phase-ends/logs/Phase28.md:182 (Phase28-32 C2); phase-ends/logs/Phase30.md:1343 (Phase30-1of2 C13); phase-ends/logs/Phase21.md:475 (Phase21 C3); phase-ends/logs/Phase21.md:63 (Phase21 C2) + +### DK-73 — Leverage is not tractability; scope a campaign by verdicts, invariants and kill criteria +- **Kernel:** the most-duplicated functions are systematically the hardest — leverage and tractability are + anti-correlated — so a leverage-first queue front-loads hand-tier work, and its early bank rate is not a harness fault. + Carry a measured closeness read per target and never let reach × size stand in for "crackable"; the reach ranking finds the + most-DONE work first, so derive the target pool from the build's own invariant. Yield clusters by binary, not across the + fleet — draw per binary once two independent lanes concentrate in the same place. A cracked idiom transfers within its + family and not across it: price a lane by families, not by class size. An open-ended grind phase's milestone is invariants + held plus a clean checkpoint, never a percentage; a research phase is scoped by a per-class verdict (a validated lever or a + falsifiable wall verdict per class), not by a percentage either. Write numeric kill criteria into the plan before the data + exists, and let them fire. Measure a pipeline's yield on the residual, not on solved functions: a known-answer ladder (revert + a match to a stub, make the pipeline re-derive it) sets the ceiling, and the gap to the unmatched tail is the real number. + Declare a mechanical lever spent only on a positive, three-part measurement — every built lever applied and returning zero, + the residue split by structure, and the decay curve priced against what remains. Sequence a phase so the cheapest thing that + can invalidate everything below it runs first; when a foundation task hits a structural wall mid-phase, bank the tractable + wins and re-scope the wall as its own sub-project. Inside one leverage class, schedule by measured remaining effort and pull + the payoff-dominating outlier out for an immediate cheap triage. Choose the exemplar for cracking a codegen class by the + size of its residual: the one-instruction mismatches are the cleanest real-function isolates. +- **When it applies:** the draw, the phase plan, the campaign's close. +- **Cost:** a phase priced by percentage that could only be closed by verdicts; a mega-leverage "freebie" that was a + stack-switcher; a lane priced by class that cost forty turns of learning per family. +provenance: BFM phase-log mining pass (P33.5 task 14.5, S92; the accelerators entry "P33.5 S92"): phase-ends/logs/Phase17.md:247 (Phase17-18 C1); phase-ends/logs/Phase25.md:473 (Phase25 C1); phase-ends/logs/Phase30.md:1810 (Phase30-1of2 C11); phase-ends/logs/Phase31.md:5040 (Phase31-2of3 C10); phase-ends/logs/Phase31.md:1300 (Phase31-1of3 C5); phase-ends/logs/Phase19.md:24 (Phase19-20-22 C1); phase-ends/logs/Phase18.md:93 (Phase17-18 C6); phase-ends/logs/Phase16.md:24 (Phase15-16 C5); phase-ends/logs/Phase16.md:36 (Phase15-16 C1); phase-ends/logs/Phase29.md:10300 (Phase29-4of4 C6); phase-ends/logs/Phase29.md:4322 (Phase29-2of4 C5); phase-ends/logs/Phase7.md:3 (Phase7 C5); phase-ends/logs/Phase24.md:29 (Phase24 C2); phase-ends/logs/Phase20.md:14 (Phase19-20-22 C2) + +### DK-74 — Models and prompts: targeted context, named degenerate outputs, two-sided caps, tiers by what they can learn +- **Kernel:** give a drafting model a targeted slice of the knowledge base, never the whole — full context measurably made + a model worse. Name the degenerate output in the prompt: an empty body compiles, so "translate every instruction, never an + empty body" is a required instruction. The output-token cap is a two-sided knob and both failure modes read as "the model is + bad"; more budget is not more quality — measure it as a paired A/B and treat truncation as recoverable, not as a defect + signal. When A/B-ing any harness knob, ship a positive control that the knob actually moved. Mine new idioms from fresh + cracks, never from the failed backlog — the failure pile re-teaches what you already know — while the harvest SELECTOR must + still be able to see failed attempts, or it learns from the easy half. Escalations to the expensive tier run strictly serial + with idiom-banking between them; only the tier that cannot learn is run in parallel. The wall-breaker tier is a match tier, + not a plumbing tier: work a deterministic arbiter can judge does not need the expensive model. Accelerate the stage that is + the bottleneck — a byte-exact search loop costs a compile plus a whole-binary gate per candidate, so hardware brute force + buys nothing. Make the drafter run the pre-gate guard and return the NAMED banking prerequisite; a batch of diagnosed + candidates is worth more than a batch of opaque matches. Before concluding a pipeline is weak, histogram the compiler's + error text: most failures were one missing declaration, fixed once. Keep prompt and law text in data, never inside the + launcher's source template. +- **When it applies:** the pack builder, the routing table, every model experiment. +- **Cost:** a local model made worse by more context; eight diagnosed matches revealing one lever the opaque batch had + hidden; a wave killed at launch by a quote character inside a template. +provenance: BFM phase-log mining pass (P33.5 task 14.5, S92; the accelerators entry "P33.5 S92"): phase-ends/logs/Phase23.md:24 (Phase23-27 C2); phase-ends/logs/Phase23.md:91 (Phase23-27 C3); phase-ends/logs/Phase23.md:54 (Phase23-27 C4); phase-ends/logs/Phase31.md:2292 (Phase31-1of3 C1); phase-ends/logs/Phase30.md:1541 (Phase30-1of2 C8); phase-ends/logs/Phase23.md:93 (Phase23-27 C1); phase-ends/logs/Phase31.md:3236 (Phase31-2of3 C6); phase-ends/logs/Phase24.md:57 (Phase24 C1); phase-ends/logs/Phase24.md:61 (Phase24 C3); phase-ends/logs/Phase17.md:96 (Phase17-18 C4); phase-ends/logs/Phase29.md:5093 (Phase29-2of4 C3); phase-ends/logs/Phase16.md:40 (Phase15-16 C3); phase-ends/logs/Phase31.md:1702 (Phase31-1of3 C12) + +### DK-75 — The unattended run: a crash is a pause, a stop is a file, a limit is an epoch +- **Kernel:** build the unattended campaign so a crash is a pause — probe the dependency at the top of each cycle, commit + per cycle, persist the tried-set and each confirmed result the moment it is confirmed; a long stateless batch that writes + only at the end loses everything to a kill. Design it for a human with no agent session: a STOP file honoured at a safe + boundary, a supervisor that tells a clean exit from a crash, a status one-liner, and crash-resume proven by a deliberate kill + before the first real run. Never wrap a project tool in a timeout shorter than its own budget — you pre-empt its recovery + handler and lose its buffered output. Sweep for orphaned worker processes at every session boundary: a dead-pipe compiler + holds a core forever and nothing reports it, and a harness's low-memory guard silently kills long background jobs. A run + that looks throttled is usually blocked on an interactive approval prompt — check the pending prompt before diagnosing the + provider. Batch size is a risk lever, not a token lever: isolated agents cost about N times one agent whether concurrent or + serial, so size a batch by the unverified spend you are willing to lose before the next measurement. A repair mode whose + cost is exceptions × population is gated on a measured exception count; for a broadly divergent set, drop rather than + recover. A free or preview model tier can be withdrawn mid-campaign without notice — a fleet-wide 404 is an epoch event, not + N model failures — and a metered key's own cap is a separate limit from the account's credit. In a pipelined + drafter/gater, "still open" is not "not yet attempted": consecutive waves re-drafted the wave still in flight. A fan-out + script generated by an orchestrator runs sandboxed without the repository: it is self-contained, so target selection + belongs to the generator, not the workers. +- **When it applies:** every lane that runs while nobody watches. +- **Cost:** a thirteen-hour orphaned compiler; three gates with no verdict and half-applied propagations from one timeout; + a six-hour "throttle" that was a permission prompt. +provenance: BFM phase-log mining pass (P33.5 task 14.5, S92; the accelerators entry "P33.5 S92"): phase-ends/logs/Phase23.md:92 (Phase23-27 C5); phase-ends/logs/Phase16.md:89 (Phase15-16 C6); phase-ends/logs/Phase31.md:7121 (Phase31-3of3 C4); phase-ends/logs/Phase31.md:7024 (Phase31-3of3 C1); phase-ends/logs/Phase29.md:1029 (Phase29-1of4 C5); phase-ends/logs/Phase33.md:508 (Phase33 C15); phase-ends/logs/Phase21.md:731 (Phase21 C1); phase-ends/logs/Phase29.md:2387 (Phase29-1of4 C2); phase-ends/logs/Phase29.md:1635 (Phase29-1of4 C3); phase-ends/logs/Phase31.md:2319 (Phase31-1of3 C8); phase-ends/logs/Phase31.md:2380 (Phase31-1of3 C9); phase-ends/logs/Phase31.md:258 (Phase31-1of3 C2); phase-ends/logs/Phase25.md:487 (Phase25 C2) + +### DK-76 — Agents and the tree: write-isolation is architecture, never a sentence in a prompt +- **Kernel:** drafting agents must never be ABLE to write the build tree; every agent artifact lands in a scratch directory, + so a killed or racing campaign costs build cycles and zero paid work. "Never modify the source tree" in a prompt is a + request, not an enforcement — snapshot the tree status around every agent and name the offender. Never adopt a subagent's + worktree wholesale: it is a snapshot of an older tree and may predate a bank; re-gate its artifacts against the head. The + generated disassembly tree is shared mutable state — a fleet verify/clean chain and the per-function instruments cannot run + at the same time. A probe that splices the tree in order to measure it must restore it, or the progress oracle counts the + splices as banks. Snapshot every target's disassembly before gating: a successful bank prunes it, and harvest and recovery + need both sides. Never let model-authored prose reach the shell inside double quotes — a backticked command in a commit + message executed; use a quoted heredoc. Keep the wave harness in the repository with its contracts; a harness rebuilt from + memory each run silently goes stale. A `cd` in one agent shell call persists into the next. Never test a helper by importing + its module: a tool with no main guard runs its whole pipeline on import. +- **When it applies:** the day the first agent is launched, as architecture; then every wave. +- **Cost:** two tree corruptions recovered with one checkout only because the drafts lived outside the tree; a real extract + run by a commit message; a download landing inside a submodule. +provenance: BFM phase-log mining pass (P33.5 task 14.5, S92; the accelerators entry "P33.5 S92"): phase-ends/logs/Phase30.md:4017 (Phase30-2of2 C1); phase-ends/logs/Phase31.md:2879 (Phase31-1of3 C7); phase-ends/logs/Phase31.md:2879 (Phase31-2of3 C4); phase-ends/logs/Phase30.md:449 (Phase30-1of2 C2); phase-ends/logs/Phase31.md:6808 (Phase31-3of3 C7); phase-ends/logs/Phase32.md:255 (Phase28-32 C6); phase-ends/logs/Phase31.md:2598 (Phase31-1of3 C11); phase-ends/logs/Phase31.md:1612 (Phase31-1of3 C3); phase-ends/logs/Phase30.md:2893 (Phase30-2of2 C6); phase-ends/logs/Phase30.md:688 (Phase30-1of2 C9); phase-ends/logs/Phase33.md:660 (Phase33 C16); phase-ends/logs/Phase30.md:2163 (Phase30-1of2 C7) + +### DK-77 — Edits that keep their proofs: repair the caller, land changes separately, draft before you carve +- **Kernel:** when a byte-proven body will not integrate, the repair moves the CALLER's declaration to the definition's + signature — never the definition to the caller's — and only where the change is width-compatible. A repair ladder probes + whether each stage is needed before applying it, or it silently escalates a binary-local bank into a fleet-shared edit. Land + a pure rename and a semantic or layout change as separate gated edits, so a gate failure attributes itself. Draft first, + then carve: a build-unit split is safe only when the new unit is immediately populated with proven bodies. A proven + transform that is not a rung of the ladder the drafts actually pass through does not exist for those drafts. Re-run the + deterministic declaration canonicaliser over old quarantined drafts after every large bank — recovery odds rise as the + banked corpus grows and the pile costs nothing to keep. An idempotency guard keyed on presence freezes every record created + before the system matured; key it on completeness, and make regenerated artifacts idempotent by replacement, never by + skipping. A hard-coded assumption fixed in one tool survives in its siblings: grep the tree for the literal and fix every + twin in the same change. Before a fleet-wide mechanical edit, census the whole population for the exact preconditions the + edit assumes — uniformity is the licence, non-uniformity the design input. A batch gate that bisects on failure re-runs the + singleton against an unchanged baseline: special-case one, or pay a duplicate build on the hot path. A byte-identical + baseline is proven only when the whole gate is green from clean across several independent sessions. A build-system + conditional that expands at parse time makes its own negative control vacuous. Never round-trip a curated configuration + through a serializer: every oracle you own measures bytes, so a formatting-destructive write is invisible to all of them. +- **When it applies:** every integration repair, every carve, every fleet-wide edit. +- **Cost:** a registry's forty-seven comment lines destroyed under green gates; a fleet edit taken by a bank that needed a + local one; a duplicate build on every single-draft gate. +provenance: BFM phase-log mining pass (P33.5 task 14.5, S92; the accelerators entry "P33.5 S92"): phase-ends/logs/Phase24.md:119 (Phase24 C4); phase-ends/logs/Phase29.md:3377 (Phase29-2of4 C7); phase-ends/logs/Phase29.md:1959 (Phase29-1of4 C6); phase-ends/logs/Phase31.md:922 (Phase31-1of3 C4); phase-ends/logs/Phase31.md:3810 (Phase31-2of3 C5); phase-ends/logs/Phase15.md:57 (Phase15-16 C4); phase-ends/logs/Phase15.md:51 (Phase15-16 C8); phase-ends/logs/Phase31.md:6171 (Phase31-3of3 C5); phase-ends/logs/Phase28.md:36 (Phase28-32 C4); phase-ends/logs/Phase29.md:7520 (Phase29-3of4 C4); phase-ends/logs/Phase28.md:58 (Phase28-32 C5); phase-ends/logs/Phase7.md:75 (Phase7 C4); phase-ends/logs/Phase31.md:5155 (Phase31-2of3 C11); phase-ends/logs/Phase28.md:181 (Phase28-32 C1) + +### DK-78 — The search harness and the compiler as evidence: same context, own corpus, a sibling's silence proves nothing +- **Kernel:** a search harness must compile in the SAME declaration context as the real build; an isolated context does not + merely fail to verify, it makes the search converge on the wrong answer. The search unit is the C expression — it cannot + freeze the instructions already right, because register allocation couples them. A decompiler's "unaffected register" + output means it decompiled one entry path of a multi-entry function and handed you a fragment: an instrument limit, never + evidence the function is hard. Check group identity from the signature files before probing a family — if the members are + structurally identical, a zero result is a compile-error certainty, not evidence about codegen. Vendor compiler sources + carry form-feed page separators, which a scripting language's line splitter honours and grep does not, so a line-number + checker over the source drifts and blames the wrong line. Your own corpus of byte matches is an experiment already run on + the toolchain: settle "is my rebuilt compiler faithful?" from it before installing the original vendor tools. Another + project's unmatched stubs are a record of what they did not crack, never proof that a class is uncrackable — + cross-project corroboration multiplies confidence in a wrong verdict as readily as a right one. Histogram a secondary + binary's call targets by address range before assuming it carries its own copy of anything; a raw blob's load address is a + hypothesis whose free confirmation is arithmetic against the next known segment's base. A relocation-interleaved + disassembly is produced only for object files; a linked image lists relocations separately with a shifted address column. +- **When it applies:** the permuter's base file, the first probe of any family, every cross-project citation. +- **Cost:** an overnight run that "closed" forty per cent of its near-misses and gated zero; a class written off on a + neighbour's silence. +provenance: BFM phase-log mining pass (P33.5 task 14.5, S92; the accelerators entry "P33.5 S92"): phase-ends/logs/Phase16.md:71 (Phase15-16 C2); phase-ends/logs/Phase16.md:43 (Phase15-16 C7); phase-ends/logs/Phase7.md:109 (Phase7 C1); phase-ends/logs/Phase30.md:168 (Phase30-1of2 C6); phase-ends/logs/Phase29.md:6915 (Phase29-3of4 C1); phase-ends/logs/Phase18.md:64 (Phase17-18 C3); phase-ends/logs/Phase18.md:52 (Phase17-18 C2); phase-ends/logs/Phase12.md:45 (Phase8-13 C2); phase-ends/logs/Phase10.md:13 (Phase8-13 C3); phase-ends/logs/Phase33.md:526 (Phase33 C7) + +### DK-79 — Maintaining the knowledge base and the record: contradictions are work items, edits are verified by section +- **Kernel:** a contradiction between two entries of your own knowledge base is a work item, not noise — replay the levers + already written down, under the correct oracle, before commissioning new research. A wrong prescription left in the base is + worse than no entry: when evidence refutes an entry, correct that entry in place, in the same session, carrying the + refutation. A document that cites a repository path is an untested claim about the repository — lint it. A programmatic + edit to a long-lived knowledge document silently truncates or duplicates it; verify the sections, never the commit. The + live hand-off block is strictly appended at the end of its file — file order is the only recency signal a fresh session + has. Never restructure a proven tool with blind string replaces at the end of a long session; specify it as the next + session's first task. A miner over your own records finds only what its pattern anticipates — measure the widened pattern's + yield. Assert that the work ledger partitions the live work, and treat a row the invariant refutes as a lie a fresh session + will act on. Write the phase synthesis in a fresh session that re-reads the committed state cold; the cold read is what + catches stale artifacts. A derive-then-apply pipeline over a live repository needs a freshness guard and a stated sequencing + law. +- **When it applies:** every harvest, every close, every programmatic edit of a document that outlives the session. +- **Cost:** a cookbook section silently deleted for a session; a day-older hand-off read as the live one; a class re-researched + because two entries disagreed and nobody replayed either. +provenance: BFM phase-log mining pass (P33.5 task 14.5, S92; the accelerators entry "P33.5 S92"): phase-ends/logs/Phase18.md:26 (Phase17-18 C5); phase-ends/logs/Phase29.md:9029 (Phase29-4of4 C3); phase-ends/logs/Phase27.md:71 (Phase23-27 C7); phase-ends/logs/Phase31.md:6529 (Phase31-3of3 C2); phase-ends/logs/Phase31.md:6126 (Phase31-3of3 C3); phase-ends/logs/Phase29.md:6134 (Phase29-3of4 C2); phase-ends/logs/Phase33.md:429 (Phase33 C14); phase-ends/logs/Phase27.md:45 (Phase23-27 C8); phase-ends/logs/Phase28.md:86 (Phase28-32 C7); phase-ends/logs/Phase33.md:285 (Phase33 C8) + +### DK-80 — Hosts and services: the small facts that each cost an hour +- **Kernel:** small facts about hosts and services, each learned at the cost of an hour. `git check-ignore` is silent for + tracked paths, so an ignore-coverage audit run before the untracking passes vacuously — use its no-index form. A + content-hash "no forbidden bytes" audit collides on zero-length files. A mirror push does not push the stash ref. Route a + host purge request through the flow that actually exists; the obvious form is a trap. A host feature can be gated on the + very flip it was meant to precede — read the settings page, do not infer. A public scratch service's compiler image is not + your pinned toolchain; rebuild it locally and prove byte-identity before asking for a preset. A disassembler's script + directory compiles as one bundle, so a single non-compiling script disables every script in it and the error names a + working one. The same tool refuses a project path containing a component that starts with a dot, so a scratch project + cannot live under a dot-directory. Run reference-compiler dump passes from a scratch working directory, or the dumps land + in the repository root and later read as committed artifacts. +- **When it applies:** the first time each host or service is touched. +- **Cost:** an hour each, and one flip-gate misread. +provenance: BFM phase-log mining pass (P33.5 task 14.5, S92; the accelerators entry "P33.5 S92"): phase-ends/logs/Phase33.md:236 (Phase33 C3); phase-ends/logs/Phase33.md:205 (Phase33 C9); phase-ends/logs/Phase33.md:295 (Phase33 C10); phase-ends/logs/Phase33.md:635 (Phase33 C11); phase-ends/logs/Phase33.md:692 (Phase33 C12); phase-ends/logs/Phase33.md:645 (Phase33 C13); phase-ends/logs/Phase33.md:186 (Phase33 C1); phase-ends/logs/Phase33.md:185 (Phase33 C2); phase-ends/logs/Phase25.md:337 (Phase25 C3) + --- ## 9. The failure museum, condensed — what looked right at the time @@ -941,7 +1182,8 @@ reading, the denominator on every number, and a decision log that records each o - **Governance and sessions:** DK-59 … DK-63 — 5 kernels. - **Readability at day one:** DK-64 … DK-65 — 2 kernels. - **Added by the coverage pass:** DK-66 … DK-68 — 3 kernels. -- **In all:** DK-1 … DK-68 — 68 kernels (the installer's check compares `grep -c '^### DK-'` against this figure). +- **Added by the log-mining pass:** DK-69 … DK-80 — 12 kernels (143 worklog lessons, clustered). +- **In all:** DK-1 … DK-80 — 80 kernels (the installer's check compares `grep -c '^### DK-'` against this figure). - **The failure museum:** 37 exhibits, condensed. - Conduct rules are not duplicated here; they are the registry seed's E.7 group. The generic engineering kernels of ProjectArchitect's own corpus apply unchanged and are not repeated. diff --git a/decomp-architect/corpus/record/docs/accelerators.md b/decomp-architect/corpus/record/docs/accelerators.md index bd08002a97..dd20dd6c45 100644 --- a/decomp-architect/corpus/record/docs/accelerators.md +++ b/decomp-architect/corpus/record/docs/accelerators.md @@ -866,3 +866,29 @@ coverage asserted both ways turned that into a need-keyed index in one task, and whose successors had existed for phases. Accelerator: from the first phase that has ten tools, adding a tool means adding its dictionary row (need · phase · what it hard-codes) or the health check fails; the index is generated, and the kit's corpus is generated from the same row. + +## P33.5 S92 (2026-09-07) — the log-mining pass: the distillation had never read the worklogs, and a coverage check found what the kit lacked + +Drew asked whether the kit held "the whole of our experience". Measured from the kit's own provenance lines: the tools and the cookbook +were in verbatim; the rules were a distillation citing 57 of 83; the kernels cited 40 of 53 accelerator entries; **no PhaseEnd and no +phase worklog was cited anywhere** — the kit had been written from the summaries (this ledger, the decision log, the retrospective, the +how-to), never from the 30,510 lines of `phase-ends/logs/`. Three things followed in one task. **(1) A coverage check for a +distillation** (`tools/kit_coverage.py`): derive both source populations (every rule from the digest, every entry of this ledger at +numbered-item granularity) and refuse one that no kit provenance line cites unless an authored map says where it went. Its first run +found 26 uncited rules and 21 uncited entries; three were genuine gaps and became kernels (a ledger's tie-break, a checker's widening and +a blanket commit are part of the instrument; the ignore file's directory-form wall; a summarised signal is a claim, not ground truth). +**Accelerator: a distillation ships with a coverage check against the populations it claims to distil, or it silently drops the lessons +nobody remembered to cite.** **(2) The record as a verbatim dictionary** (`decomp-architect/corpus/record/`): the how-to, this ledger, the +decision log, the retrospective, the story, the playbook, the effort doctrine, the readability charter, the digest and every PhaseEnd, +asserted equal on every health check, so that a rule's or kernel's provenance line leads to its evidence. **(3) The worklogs, read once +more:** 21 read-only agent slices over the 26 logs (the three giants split by line range), each briefed to extract only lessons that a +grep over every distilled record could not find, with the greps recorded — **777 candidates, 634 already banked, 143 new**, every cited +log line verified to exist and to say what the candidate claims, clustered into twelve kernels (DK-69–DK-80: instrument blind spots; +verdict staleness and the health suite; what earns belief; denominators and units; leverage versus tractability; models and prompts; +the unattended run; agents and the tree; edits that keep their proofs; the search harness and the compiler as evidence; maintaining the +knowledge base; hosts and services), each provenance line naming the worklog lines. **Accelerator: the "capture while it hurts" rule +does not capture everything — an end-of-project pass over the raw worklogs, with an "already banked?" grep per candidate, recovers the +lessons that were fixed but never generalised; on this project one in five candidates was such a lesson.** Cost (R41): 21 Opus agents, +~4.3M tokens, ~13 min wall each in two batches; the coordinator's cost was the brief, the harvest script, the clustering and the +provenance generation. + diff --git a/decomp-architect/corpus/record/docs/decision-log.md b/decomp-architect/corpus/record/docs/decision-log.md index 0e7c4a6122..8c883fb822 100644 --- a/decomp-architect/corpus/record/docs/decision-log.md +++ b/decomp-architect/corpus/record/docs/decision-log.md @@ -3617,3 +3617,31 @@ register lever (accelerators (16)), and treat "PROVED" as "proved against this l (`lift_types`, `canon_sig_reconcile`, `sync_tu_decls`, `decl_prior`, `conform_decls`) point at Phase 6, not only Phase 10, and the cookbook front page names the width class as the one place a type moves bytes; (e) the methodology's integration section states the two-sided verdict in one paragraph. The source project's own Gen3 (Phase 35+) is the proof the kit will later cite. + +## P33.5 S92 (2026-09-07) — the kit's distillation gets a coverage check, a verbatim record and a pass over the worklogs (Drew's question: "the whole of our experience?") + +- **Context and belief.** Task 14 closed with the kit's wiki page written and the belief that the kit held the project's experience: the + tools and the cookbook verbatim (task 13.5), the rules G1–G67 and the kernels DK-1–DK-65 distilled from the retrospective, the + accelerators, the failure museum, the how-to and the memories (task 10), the S91-b types hindsight implemented (task 14). +- **What the question exposed.** Drew asked whether the PhaseEnds, the logs, the rules and the memories had all gone in. Measured from the + kit's provenance lines (not from memory): 57 of 83 rules cited, 40 of 53 accelerator entries, 0 PhaseEnds, 0 worklogs; the record + itself (the how-to, the decision log, the accelerators, the retrospective, the playbook, the digest) absent verbatim, so a reader of + the kit could not follow a provenance line to its evidence. The distillation had been honest about its sources and silent about its + coverage — the same defect class as a scanner without a coverage assertion (R32), applied to a document. +- **The pivot (plan amendment 3, task 14.5, Drew: "do all three").** (1) `decomp-architect/corpus/record/` — the record verbatim behind a + front page, regenerated and asserted equal like the other two corpora. (2) `tools/kit_coverage.py` — both populations derived (rules + from the digest, accelerator entries from the ledger at item granularity), each entry cited by a kit provenance line or dispositioned in + `config/kit_coverage_map.tsv` with a checked vocabulary; in tools-health. Its first run: 26 + 21 uncovered → DK-66/67/68 written for the + three genuine gaps, 41 dispositions authored (15 ProjectArchitect's own, 3 environment, the rest folds into named G/DK entries). + (3) The worklog pass: 21 read-only Opus agents over 30,510 lines, each with the "already banked?" grep protocol, deliverables early; + 777 candidates, 634 already banked, 143 new, every cited line verified by script and read; clustered into DK-69–DK-80 with generated + provenance lines; one dropped on the miner's own verdict. The harvest table is `.run/P33.5/log-mining/HARVEST_TABLE.md`. +- **Why it was right, measured.** The "capture while it hurts" rule (R30/R31) captured 634 of 777 lessons the agents found — 82% — and + missed 18%, concentrated in things that were FIXED and never GENERALISED (a selftest blind by construction, a roster in ignored scratch, + a red guard nobody wired, a timeout shorter than a tool's budget). The instrument-shaped lessons dominate the new kernels, which is the + retrospective's own finding restated from the raw record. A summary-only distillation would have shipped without them. +- **Hindsight — the better path.** Build the coverage check the day the distillation is designed (task 10), not after it ships; and + schedule the worklog pass as the LAST harvest of the project by design, since its yield (one new lesson per five candidates) is far + above any mid-campaign harvest's. Both are now the kit's own rules: the coverage check is in the kit's health target, and the record's + front page tells the next project that its logs deserve one final read. + diff --git a/decomp-architect/corpus/tools/P9/kit_coverage.py b/decomp-architect/corpus/tools/P9/kit_coverage.py index 30a83d83ec..c4f8ce0e65 100644 --- a/decomp-architect/corpus/tools/P9/kit_coverage.py +++ b/decomp-architect/corpus/tools/P9/kit_coverage.py @@ -99,7 +99,7 @@ def accel_cited(key, prov_text): elif num: pat = rf"\b{tok}\b(?: [A-Z]\w*)? \({num}\b" else: - pat = rf"(?/` + `corpus/cookbook/` + `corpus/record/` (the how-to, the decision log, the accelerators, the retrospective, the story, the playbook, the effort map, the gen3 docs, the digest, every PhaseEnd — task 14.5) (`--corpus`; superseded tools as pointer files; sha1-equal to their sources). `--check` in `tools-health`; `make kit-corpus` = `--all`. `--consumers FILE` is the referrer census before any `git mv` of a tool. | | | `tools/kit_coverage.py` | **(P33.5 task 14.5)** The kit's DISTILLATION coverage: derives the rule population (every `- **R` of DIGEST §3, asserted contiguous) and the hindsight population (every `## ` heading of `docs/accelerators.md` at numbered-item granularity — 58 at S92) and asserts each is cited by a `provenance:` line of the registry seed / the kernels OR dispositioned in `config/kit_coverage_map.tsv` (`G` / `DK-` / `FOLDED:G` / `ENV` / `PA` / `SEED:` / `KIT:` / `RECORD` / `COOKBOOK` / `NOT-PORTABLE`; unknown ids refused, R43); counts with denominators (R41); rc 1 on any gap. In `tools-health` after `tool_census --check`. Its first run found 26 uncited rules and 21 uncited entries → three new kernels (DK-66–DK-68) and 41 authored dispositions. | -| | `decomp-architect/` (the day-one decomp kit) | **(P33.5 tasks 9–14)** The package a new matching-decomp project installs as Phase 0.5 on ProjectArchitect 2.0: `README.md` (the three steps), `intake.decomp.md` (ProjectArchitect's twelve items pre-answered + the phase ladder + the six readability inversions), `SETUP.md` (the installer, Step 0 contract … Step 10 verify + hard stop; `answers: ` for unattended runs), `decomp-architect.md` (the methodology), `templates/` (the firewall pack — `gitignore.decomp`, `firewall.txt`, `audit_public.template.py`, the planted fixture, `no-rom.template.yml` —, the READMEs, `pa-overlays.md`, `registry-E.decomp.md` G1–G67, the skeletons, `PLACEHOLDERS.md`, `layout-contract.md`), `corpus/decomp-kernels.md` (DK-1 … DK-68), the three dictionaries `corpus/tools//` + `corpus/cookbook/` + `corpus/record/` (generated by `make kit-corpus`), `memory-seed/` (18), `tools/MANIFEST.md` (generated). **How it is checked:** `tools/kit_lint.py` (de-specialisation, placeholders, syntax, the gitignore-template diff) + `tools/tool_census.py --check` (the corpora equal their sources) + `tools/gitignore_template_check.py`, all in `tools-health`; the dry-run harness under `.run/P33.5/kit-dryrun/` (`answers.md`, `expected-manifest.txt`, `judge.py`, the install logs and verdicts of runs 1–4 — a kit change is re-verified by a resume on the last throwaway `repo/`, a fresh full run only when SETUP's steps change; the judge compares the real tree's dirty PATH SETS before/after). Wiki page: `docs/wiki/Start-a-new-decomp-project.md`. Split into its own repository after the flip. | +| | `decomp-architect/` (the day-one decomp kit) | **(P33.5 tasks 9–14)** The package a new matching-decomp project installs as Phase 0.5 on ProjectArchitect 2.0: `README.md` (the three steps), `intake.decomp.md` (ProjectArchitect's twelve items pre-answered + the phase ladder + the six readability inversions), `SETUP.md` (the installer, Step 0 contract … Step 10 verify + hard stop; `answers: ` for unattended runs), `decomp-architect.md` (the methodology), `templates/` (the firewall pack — `gitignore.decomp`, `firewall.txt`, `audit_public.template.py`, the planted fixture, `no-rom.template.yml` —, the READMEs, `pa-overlays.md`, `registry-E.decomp.md` G1–G67, the skeletons, `PLACEHOLDERS.md`, `layout-contract.md`), `corpus/decomp-kernels.md` (DK-1 … DK-80), the three dictionaries `corpus/tools//` + `corpus/cookbook/` + `corpus/record/` (generated by `make kit-corpus`), `memory-seed/` (18), `tools/MANIFEST.md` (generated). **How it is checked:** `tools/kit_lint.py` (de-specialisation, placeholders, syntax, the gitignore-template diff) + `tools/tool_census.py --check` (the corpora equal their sources) + `tools/gitignore_template_check.py`, all in `tools-health`; the dry-run harness under `.run/P33.5/kit-dryrun/` (`answers.md`, `expected-manifest.txt`, `judge.py`, the install logs and verdicts of runs 1–4 — a kit change is re-verified by a resume on the last throwaway `repo/`, a fresh full run only when SETUP's steps change; the judge compares the real tree's dirty PATH SETS before/after). Wiki page: `docs/wiki/Start-a-new-decomp-project.md`. Split into its own repository after the flip. | | **Verification** | `tools/verify_contract.sh` | **(P33 A5/C8) THE recorded contract run**: 00 tree · 01 check-env · 02 family_hseq · 03 `make clean && extract-all && check-all` · 04 sdk-dual (or a recorded SKIP) · 05 tools-health (zero `[warn]`) · 06 audit-frontier · 07 audit-disc · 08 report; one log per step ending `EXIT=`, abort on the first red (R53), every step asserted by its contract line (R49), `SUMMARY.md` generated → `.run/P33/verify/` (tracked evidence, quoted by `docs/verification.md` §2); step 00 ignores its own output dir. ≈14 min on 32 CPUs. | | **Publishing** | `tools/progress.py --json \| --readme [--check]` | **(P33 D1/D3)** The same numbers as DATA: `--json` → `docs/progress.json` (schema 1: the four metrics with numerator/denominator/pct, the counts, 218 per-binary rows incl. instruction totals; no run date) + `docs/badges/{fleet_instr,fleet_fn,distinct,binaries}.json` (shields endpoint format; the README references `fleet_instr` + `binaries` by name); `--readme` rewrites the README's `` block (refuses a README without the markers); `--check` asserts JSON + block + badges are fresh (in `make audit-digest`). Run by `make report BINARY=main`. | | | `tools/wiki_render.py OUT_DIR \| --list \| --selftest` | **(P33 F3)** Render `docs/wiki/*.md` + `docs/how-to-ai-decomp/*.md` into GitHub-wiki page names with every relative link rewritten deterministically (wiki page → its name; a chapter → `How-to-AI-decomp-NN-name`; any other repo path → a `blob/main` / `tree/main` / raw URL; URLs, mailto and anchors untouched; **a dead link is an error**, R43). `--selftest` = the 12-case fixture incl. the dead-link negative control **plus (P33.5 task 7) the reachability assertion: every `docs/wiki/*.md` except `_Sidebar`/`_Footer`/`Home` is linked from `_Sidebar.md`, every chapter from `_Sidebar.md` AND `How-to-AI-decomp.md`** — a published page nobody can navigate to fails here; in `make tools-health`. | diff --git a/docs/accelerators.md b/docs/accelerators.md index bd08002a97..dd20dd6c45 100644 --- a/docs/accelerators.md +++ b/docs/accelerators.md @@ -866,3 +866,29 @@ coverage asserted both ways turned that into a need-keyed index in one task, and whose successors had existed for phases. Accelerator: from the first phase that has ten tools, adding a tool means adding its dictionary row (need · phase · what it hard-codes) or the health check fails; the index is generated, and the kit's corpus is generated from the same row. + +## P33.5 S92 (2026-09-07) — the log-mining pass: the distillation had never read the worklogs, and a coverage check found what the kit lacked + +Drew asked whether the kit held "the whole of our experience". Measured from the kit's own provenance lines: the tools and the cookbook +were in verbatim; the rules were a distillation citing 57 of 83; the kernels cited 40 of 53 accelerator entries; **no PhaseEnd and no +phase worklog was cited anywhere** — the kit had been written from the summaries (this ledger, the decision log, the retrospective, the +how-to), never from the 30,510 lines of `phase-ends/logs/`. Three things followed in one task. **(1) A coverage check for a +distillation** (`tools/kit_coverage.py`): derive both source populations (every rule from the digest, every entry of this ledger at +numbered-item granularity) and refuse one that no kit provenance line cites unless an authored map says where it went. Its first run +found 26 uncited rules and 21 uncited entries; three were genuine gaps and became kernels (a ledger's tie-break, a checker's widening and +a blanket commit are part of the instrument; the ignore file's directory-form wall; a summarised signal is a claim, not ground truth). +**Accelerator: a distillation ships with a coverage check against the populations it claims to distil, or it silently drops the lessons +nobody remembered to cite.** **(2) The record as a verbatim dictionary** (`decomp-architect/corpus/record/`): the how-to, this ledger, the +decision log, the retrospective, the story, the playbook, the effort doctrine, the readability charter, the digest and every PhaseEnd, +asserted equal on every health check, so that a rule's or kernel's provenance line leads to its evidence. **(3) The worklogs, read once +more:** 21 read-only agent slices over the 26 logs (the three giants split by line range), each briefed to extract only lessons that a +grep over every distilled record could not find, with the greps recorded — **777 candidates, 634 already banked, 143 new**, every cited +log line verified to exist and to say what the candidate claims, clustered into twelve kernels (DK-69–DK-80: instrument blind spots; +verdict staleness and the health suite; what earns belief; denominators and units; leverage versus tractability; models and prompts; +the unattended run; agents and the tree; edits that keep their proofs; the search harness and the compiler as evidence; maintaining the +knowledge base; hosts and services), each provenance line naming the worklog lines. **Accelerator: the "capture while it hurts" rule +does not capture everything — an end-of-project pass over the raw worklogs, with an "already banked?" grep per candidate, recovers the +lessons that were fixed but never generalised; on this project one in five candidates was such a lesson.** Cost (R41): 21 Opus agents, +~4.3M tokens, ~13 min wall each in two batches; the coordinator's cost was the brief, the harvest script, the clustering and the +provenance generation. + diff --git a/docs/decision-log.md b/docs/decision-log.md index 0e7c4a6122..8c883fb822 100644 --- a/docs/decision-log.md +++ b/docs/decision-log.md @@ -3617,3 +3617,31 @@ register lever (accelerators (16)), and treat "PROVED" as "proved against this l (`lift_types`, `canon_sig_reconcile`, `sync_tu_decls`, `decl_prior`, `conform_decls`) point at Phase 6, not only Phase 10, and the cookbook front page names the width class as the one place a type moves bytes; (e) the methodology's integration section states the two-sided verdict in one paragraph. The source project's own Gen3 (Phase 35+) is the proof the kit will later cite. + +## P33.5 S92 (2026-09-07) — the kit's distillation gets a coverage check, a verbatim record and a pass over the worklogs (Drew's question: "the whole of our experience?") + +- **Context and belief.** Task 14 closed with the kit's wiki page written and the belief that the kit held the project's experience: the + tools and the cookbook verbatim (task 13.5), the rules G1–G67 and the kernels DK-1–DK-65 distilled from the retrospective, the + accelerators, the failure museum, the how-to and the memories (task 10), the S91-b types hindsight implemented (task 14). +- **What the question exposed.** Drew asked whether the PhaseEnds, the logs, the rules and the memories had all gone in. Measured from the + kit's provenance lines (not from memory): 57 of 83 rules cited, 40 of 53 accelerator entries, 0 PhaseEnds, 0 worklogs; the record + itself (the how-to, the decision log, the accelerators, the retrospective, the playbook, the digest) absent verbatim, so a reader of + the kit could not follow a provenance line to its evidence. The distillation had been honest about its sources and silent about its + coverage — the same defect class as a scanner without a coverage assertion (R32), applied to a document. +- **The pivot (plan amendment 3, task 14.5, Drew: "do all three").** (1) `decomp-architect/corpus/record/` — the record verbatim behind a + front page, regenerated and asserted equal like the other two corpora. (2) `tools/kit_coverage.py` — both populations derived (rules + from the digest, accelerator entries from the ledger at item granularity), each entry cited by a kit provenance line or dispositioned in + `config/kit_coverage_map.tsv` with a checked vocabulary; in tools-health. Its first run: 26 + 21 uncovered → DK-66/67/68 written for the + three genuine gaps, 41 dispositions authored (15 ProjectArchitect's own, 3 environment, the rest folds into named G/DK entries). + (3) The worklog pass: 21 read-only Opus agents over 30,510 lines, each with the "already banked?" grep protocol, deliverables early; + 777 candidates, 634 already banked, 143 new, every cited line verified by script and read; clustered into DK-69–DK-80 with generated + provenance lines; one dropped on the miner's own verdict. The harvest table is `.run/P33.5/log-mining/HARVEST_TABLE.md`. +- **Why it was right, measured.** The "capture while it hurts" rule (R30/R31) captured 634 of 777 lessons the agents found — 82% — and + missed 18%, concentrated in things that were FIXED and never GENERALISED (a selftest blind by construction, a roster in ignored scratch, + a red guard nobody wired, a timeout shorter than a tool's budget). The instrument-shaped lessons dominate the new kernels, which is the + retrospective's own finding restated from the raw record. A summary-only distillation would have shipped without them. +- **Hindsight — the better path.** Build the coverage check the day the distillation is designed (task 10), not after it ships; and + schedule the worklog pass as the LAST harvest of the project by design, since its yield (one new lesson per five candidates) is far + above any mid-campaign harvest's. Both are now the kit's own rules: the coverage check is in the kit's health target, and the record's + front page tells the next project that its logs deserve one final read. + diff --git a/docs/wiki/Start-a-new-decomp-project.md b/docs/wiki/Start-a-new-decomp-project.md index 8e30407dfe..1006504dac 100644 --- a/docs/wiki/Start-a-new-decomp-project.md +++ b/docs/wiki/Start-a-new-decomp-project.md @@ -40,7 +40,7 @@ the installer, it never defaults. Every git call is by explicit path; the instal | The registry seed (`templates/registry-E.decomp.md`) | rules G1–G67 in seven groups — the oracles and the gate, the ROM firewall, the instruments, the campaign, compiler walls, publishing and the record, the use of AI — each with a `provenance:` line naming the failure that earned it | | The firewall pack | a copyable `.gitignore` (the same block as [The ROM firewall](The-ROM-firewall.md), asserted identical in this project's health check), an audit that derives its forbidden set from a config and fails on a planted fixture before it is trusted, and a CI workflow — no game-derived bytes in git from commit one | | The layout, the overlays, the skeletons | the `docs/` and `.run/` conventions from [Docs and scratch conventions](Docs-and-scratch-conventions.md); marked-section appends to `CLAUDE.md`, the effort map, the cookbook, the ops reference; the session-start digest, the replayable checkpoint block and the PhaseEnd narrative axis; LICENSE, NOTICE, README and CONTRIBUTING skeletons, a `.clang-format` and a `make format` snippet | -| The kernels (`corpus/decomp-kernels.md`) | DK-1 … DK-68: what this project learned late, each with when it applies and what it cost, plus the failure museum | +| The kernels (`corpus/decomp-kernels.md`) | DK-1 … DK-80: what this project learned late, each with when it applies and what it cost, plus the failure museum | | The memory seed (`memory-seed/`) | eighteen working agreements and harness facts, de-specialised, appended to the memory ProjectArchitect configured | **It installs no tools.** A byte gate, a splitter config, a permuter harness, a decompiler context, a differential @@ -60,8 +60,9 @@ is **three dictionaries**, all generated from this tree and asserted equal to it every PhaseEnd, verbatim, behind a front page that says what each is and how to read it. The kit's rules and kernels were distilled from these; a `provenance:` line on any of them leads back here. A coverage check asserts that every one of this project's rules and every entry of its accelerators ledger is either cited by such a line or explicitly dispositioned, so - the distillation has no silent gap. The phase worklogs are not included; a dedicated pass read every one of them for lessons - banked nowhere else, and those became kernels. + the distillation has no silent gap. The phase worklogs are not included; a dedicated pass read every one of them (some thirty thousand + lines, one read-only agent per slice) for lessons banked nowhere else — 143 of 777 candidates were — and those became twelve + kernels, each pointing at the worklog lines it came from. ## The phase ladder diff --git a/tools/kit_coverage.py b/tools/kit_coverage.py index 30a83d83ec..c4f8ce0e65 100644 --- a/tools/kit_coverage.py +++ b/tools/kit_coverage.py @@ -99,7 +99,7 @@ def accel_cited(key, prov_text): elif num: pat = rf"\b{tok}\b(?: [A-Z]\w*)? \({num}\b" else: - pat = rf"(?