Files
BFM-decomp/phase-ends/CURRENT_PHASE.md
T
Drew T 0840eda5bb feat(phase-31 T8): LEN+N lane — 587 near-misses routed; 345 wrong-drafts reclassified; detectors live
- match_one --emit-streams (additive; stdout-identity NC'd): word streams for the
  len lane
- family_align.addr_true_rel: reloc-vs-constant range discriminator — FULL
  conservative set kept for pair semantics (NC-1 157/157 regression), address-
  true subset for indel eligibility only (a constant li-cluster must not read as
  reloc-in-indel); synthetic probes green both directions
- tools/len_tells.py: aligned classification + §172b tell tagging (EXTPAIR/
  SELECT/NOP) on target-side indels; detectors imported from atlas_features
  (R33); cookbook text embedded in cards
- tools/lenmiss_route.py: pool-parallel (A8) — 587 audit LEN rows re-verified
  live + routed in 24s: redraft 345 (frac>0.35, APPEND-ONLY backlog
  reclassification — near-miss metrics stop lying) / permuter-length 49 (grinder
  fuel) / cards 192 incl 14 tell-tagged (the audit's own detectors had emitted
  ZERO) / mechanical 0 — an HONEST NULL: stored drafts rarely get constants
  wrong; LEN drift is shape, family_align's value here is classifier/detector
- R32 accounting 587/587
2026-08-14 19:54:57 -06:00

63 lines
17 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CURRENT PHASE — Phase 31: The Frontier Atlas & Wide-Tolerance Campaign
**Started:** 2026-08-14 · **Plan approved:** 2026-08-14 (gate 1; Drew) · **Effort doctrine:** xHigh default / Max deep (T5, T7, synthesis) / Ultracode waves (R26/R27 prompts) / Fable-tier only for new wall classes.
**Approved plan:** `/home/musashi/.claude/plans/fable-5-set-max-goofy-seahorse.md` (the full design; this file is the crash-recovery log).
**Approval also ratified R37 (probe before costing), R38 (read recorded failure verdicts first), R39 (negative-control new refusal-checks) — now binding.**
## The re-charter (one paragraph)
Instead of roadmap-v2 P31's per-function grind, Phase 31 organizes the 12,059 remaining stubs (main 1,041 · resident 14 · ov 9,832 · md 1,179) into **crack groups**: a deterministic per-function feature layer + multi-tier similarity atlas (li-normalized exact tier, seed sweep vs the 2,719 matched skeletons, calibrated warm tier for the 3,238-unit cold tail + main, kNN neighborhood graph), evidence-joined to a lever label per group; plus **widened mechanical lanes** (LEN-tolerant aligned remap with a fully-mechanical LI class, §172b EXTPAIR/SELECT detectors + routing, PLUMBING campaign, permuter cluster warm-start, weak-seed cards); then a **campaign loop to ceiling** — deterministic $0 lanes first, agents only for exemplars, velocity-ledgered, closed on measured decay. Milestone shape = P30 (campaign to ceiling; every remaining stub on a named ledger at close). Main fully included from day one.
## Task checklist
- [x] **T0 — Pivot log + freshness + hygiene** — DONE 2026-08-14. (xHigh)
- [x] **T1 — Integration quick-bank sweep** — DONE 2026-08-14 (pending final R22 log line). **8 banked, 0 agent tokens.** (xHigh)
- [x] **T2 — References** — DONE 2026-08-14. (xHigh)
- [x] **T3 — Main enablement** — DONE 2026-08-14. (xHigh)
- [x] **T4 — atlas_features.py** — DONE 2026-08-14. (xHigh)
- [x] **T5 — atlas.py** — DONE 2026-08-14. **THE ATLAS EXISTS.** (Max)
- [x] **T6 — PLUMBING campaign** — DONE 2026-08-14. **+9 banked (probe 64%); recipe + 3 laws distilled (§173); sweep proved the no-draft majority routes to family lanes.** (xHigh)
- [x] **T7 — family_align.py built + NC'd; the mechanical-cousin premise REFUTED by its probe (0/26)** — engine re-scoped to T8's LEN+N same-function pile; no driver built (correctly). (Max)
- [x] **T8 — LEN+N lane** — DONE 2026-08-14. **Pile routed 587/587; 345 wrong-drafts reclassified; 49 permuter + 192 card fuel staged; mechanical lane = honest null.** (xHigh)
- [ ] **T9 — Warmstart + weak-cards**: warmstart.py + grinder admission widening + --weak-cards; NCs + probes. (xHigh)
- [ ] **T10+ — Campaign loop to ceiling** (repeating sessions; velocity ledger; close on measured decay). (Ultracode waves / xHigh orchestration / Fable new-walls)
- [ ] **Tclose — PhaseEnd** (gate 2). (Max)
## Standing verification (every task)
R22 clean-fleet **213/213** after every banked batch · tools-health green · 0 NON_MATCHING (G4) · dedup-check 0 failed · R32 coverage assertions on every new scanner · R39 negative controls on every refusal check · R37 probes before pricing · R38 ledgers before experiments · commit per task (task + this log in the same commit; Drew pushes).
## Progress log
- 2026-08-14 — Phase planned and approved (3 Explore + 2 Plan agents; full design in the plan file). Task list built (harness tasks #1–#12). T0 started.
- 2026-08-14 — **T0 COMPLETE.** (1) R31 decision-log entry (the re-charter WHY + R37–R39 ratification). (2) `harvest_verify.py` import guard: a bare import now RAISES instead of running a gate (verified both directions; CLI behavior unchanged). (3) **Resident ±1 RESOLVED + FIXED**: `--bootstrap`'s linear partition had fused the +0 data word with `func_800CEDFC` (row `0x800CEDF8` nins=18) and dropped `func_800D33E0` past a glued tail — the true denominator is **145** (progress was right, the sig wrong). `sig-resident` now ELF-seeds (S45 pattern: unique 4-aligned T-symbol addrs inside the `resident_TEXT_START/END` markers → exactly 145; bootstrap fresh-clone fallback). All three oracles now agree (sig 145 · corpus matched 131 · progress byte-ident 131); `audit-corpus` 0 PHANTOM + 0 TRUNCATED. (4) Family maps regenerated at HEAD `commit:2161`: **11,025 open non-main members = 12,059 − main's 1,034 EXACT** (the stale map's 102 phantoms cleared); cousins totals now A-prop 1,125 / seeded 1,293 / cousin-multi 5,373 / cold 3,234 inst; adapt cards 704, aprop cards **204 (full emission)**. (5) **Main fuel gap is DEAD**: 2,001/2,002 main stubs have cached Ghidra-C (only `func_80049600` missing) — the roadmap's "0/2,096" note was stale. (6) `make tools-health` → OK (dedup 2,063/0; C1 254,521/254,521; audit-digest confirms the fleet digest; resident fix moved instr num+denom by the same +9).
- 2026-08-14 — **T1 COMPLETE: 8 banked for 0 agent tokens.** R38-first: partitioned the MATCH-108 pile against current stubs (75 still open) and against the S50 gate history (62 gated-and-failed with verdicts · 13 never-gated). The lanes and their measured outcomes:
- **Never-gated 13** → gate_lane: 0/13, but the verdicts decomposed to 11× `undefined reference to D_*` = the §171 stale-seed-symbol class. **Extended `aprop_symfix` with STALE-DELTA** (n:n uniform-delta rebase; R39 synthetic + snapshot NCs, zero false positives; the delta test even refused a pair my hand-check wrongly accepted) → 4 rebased, **4/4 banked** (`func_8016BCC0`, `func_8017F1C8`, `func_80186BD8`, `func_80186BF8`). Cookbook **§171-D** written in-session.
- **SELF-decl PLUMBING 7** → `recover_integration --stages demacroize --max-tier binary`: **4/7 banked** (`func_80139BE0`, `func_8014ED28`, `func_80161D88`, `func_801659DC`); 3 stay near.
- **Stored-draft re-gates** (no-verdict 7 + close=0 8 + diff 1 + 9 STALE→clean world-motion drafts): **0/23-ish banked** — the ~8% A10 stored-verdict law held again; all re-verdicted fresh.
- **Handed forward with fresh classifications**: CALLEE-decl 15 + CC1-FAIL 14 → T6 (cast-callees/tu-scope stages); UNDEF-DATA/OTHER 9 → the §171b-1 data-definition carry (T7/T8); md CARVE-REFUSED 8 → campaign side-quest ledger. immfix pile: fully consumed (0 open). fix20: 19/20 consumed in S50.
- Tool fixes landed: `gate_lane` propagate-commit tag now derives from GATE_PHASE (was hardcoded phase-30 S49).
- Rate lesson for the velocity ledger: fresh-fix lanes (STALE-DELTA 4/4, demacroize 4/7) vastly outperform blind stored re-gates (0/23) — the campaign loop's L2 ordering is confirmed by measurement.
- 2026-08-14 — **T2 COMPLETE (references).** (1) **PsyQ dev-CD extracted**: walked the on-disk Track-1 image (MODE2/2352) with the frozen `tools/bfm_extract/iso9660.py` (R33 — no new extractor; walker = iter_directory/read_extent with out-of-range extents skipped) → `tools/reference/psyq-sdk/` (gitignored): 2,374 files / 231.6 MB, **400 C sources (373 in PSX/SAMPLE/** across CD/GRAPHICS/SOUND/MODULE/CMPLR/…); only 7 out-of-track `.DA` audio skipped. (2) **Provenance find**: `GNU/SNGNUVER.TXT` = SN Systems' gcc build history (`2.7.2.SN32.3.7.0002`, 14.5.97) with per-build changelog of SN's patches vs vanilla — only `UNROLL.C` (parameterised max unroll insns) is codegen-relevant; recorded in the idiom notes as the first-look suspect if a loop-unroll residual ever defies the vanilla model. (3) **gcc-2.7.2 reference completed**: +6 files from GNU ftp (`calls.c` + `caller-save.c` — both cited by §172's producer model, previously missing — + integrate/optabs/varasm/recog), tarball sha256 `7cd8bce5…` recorded. (4) `docs/psyq-sample-idioms.md` seeded (inventory, provenance, first style conventions, the lane hook); SETUP §5.6 rows added (R21).
- 2026-08-14 — **T3 COMPLETE (main enablement).** (1) `sig_image` seed-ends extension: `--seeds` now accepts `0xADDR NINS` lines (and jsonl `nins`); a seeded nins is authoritative — the slice is exactly `[addr, addr+4·nins)`, bypassing `func_end` (whose heuristic mis-sliced 3/40 main samples). (2) `corpus.s_ins_count()` factored from audit() (R33, one counter) + `corpus.py <bin> --seed-ends` CLI. (3) **`make sig-main`**: 2,002 main stubs signed at splat-true lengths → `.run/sig.main.jsonl` (deliberately splat-SEEDED — the atlas needs the boundaries a match must hit; NOT the independent second oracle, which stays deferred per `docs/second-oracle.md`). (4) **Full word cross-check: 2,002/2,002 slices byte-faithful** (EXE bytes vs `.s` comment-column words; 0 SLICE-SUSPECT; note the `.s` word field is byte-order hex, not the LE-decoded value — first checker draft compared wrong and read 1,999 false suspects). (5) `family_remap` main special-cases: `vram_of('main')` = code-segment `vram − start` from `splat.us.exe.yaml` = 0x8000F800; `img_path('main')` from the same yaml's target_path (R33). `stream_words('main')` == `.s` words on **25/25** samples. (6) Regression: `sig-resident` re-run **byte-identical** after the shared `read_seeds` change.
- 2026-08-14 — **T4 COMPLETE (`tools/atlas_features.py`).** Per-function features memoized per distinct `h_exact` (`.run/feat_memo.json`) and fanned out 1:1 with sigs to `.run/feat.<bin>.jsonl` ×213 (main reads the splat-true `sig.main.jsonl`; registry untouched). **363,525 rows / 92,855 distinct bodies in 21 s** (vs the 3–6 min estimate). Features: nins/band · o0 prologue tell · frame/callee-saved set+order/fp · CFG skeleton (nblk/ncond/nback/ret_n via branch-target scan; jal = call, never an edge — position-independence proof in the docstring) · mid_jr/jalr · stable-call-sequence hash (fixed main+resident ranges) · reloc-kind sequence hash · 16-bucket ophist · **§172b tells as importable shared functions (extpair/dupselect/sign_mix/magic_div — T8's len_tells imports these, R33)** · **li-normalized skeleton `h_seqn`** (non-anchor lui dropped, ori→addiu class). Verifies: R32 join asserted per binary at write; determinism 0 mismatches on 200 re-derived bodies; **mid_jr cross-check vs family_hseq's independent oracle: 6,444/6,444 exemplars agree** (first verifier draft compared ZERO rows — a string-vs-int addr type mismatch, the R32 silent-no-op class caught in my own verifier; now indexed + fails loud if compared==0). o0-tell vs `corpus.is_o0` on 13,019 open fns: 41 both / **94 tell-only = candidate undiscovered -O0 fns in -O2 TUs (§116 — o0_subsplit carve fuel for the campaign)** / 14 src-only outliers.
- 2026-08-14 — **T5 COMPLETE (`tools/atlas.py` + `make atlas`) — THE FRONTIER ATLAS EXISTS.** Survey in 92 s, all assertions green: **5,139 groups cover ALL 12,058 open instances / 613,710 ins** (12,058 = progress's 12,051 stubs + 7 NON_MATCHING exactly — the +1-blob and +8-banked discrepancies were both chased and resolved: data blobs excluded, T1 banks confirmed absent; main open reconciles to classify()'s stubs+NM with a hard assert). Tiers: **calibration measured** (positives = cousin ≥0.85-raw merges, negatives = size-matched cross-unit pairs; the li-normalized metric is so discriminative that recall ~99% holds down to 0.55; rule = smallest t with neg-accept ≤0.2% ∧ recall ≥95% → **THRESH_WARM=0.70** at 99.1%/0.18% — chosen because a false warm-merge wastes an exemplar crack, a miss only routes cheaper) · T1.5 h_seqn 2 merges · **warm tier 1,019 merges** · **seed sweep: 4,702/7,247 open skeletons (65%) carry a ≥0.55 matched seed** from the 2.7k-skeleton pool · kNN graph (top-8 open + matched neighbors) · tiny-direct. **Lever table (the strategic map):** head-crack 1,283g/186.9k ins · UNKNOWN 1,964g/138.6k (honest) · **extend-tell 575g/76.7k** · redraft 46.8k · **jtbl-carve 190g/45.7k** · integration 575 inst/23.4k · seeded-crack 23.4k · len-vein 778 inst/16.8k · swaprepeat 9.2k · plumbing 161 inst/8.1k · o0-lane 6.6k · cc1 6.4k. Top group: 268 drifted single-member skeletons unified (per-location ~23-ins family, 0.72 seed). Evidence joins accounted (audit/backlog/ledgers/cards; residuals EXCLUDED-STALE until regenerated); `--targets` resolves 12/12 `.s`. `make atlas` = the full regen chain. docs/frontier-atlas.md committed.
- 2026-08-14 — **T6 IN PROGRESS (PLUMBING campaign) — probe converged after two R14 corrections.** `plumbing_groups.py` derived the honest pool: **237 still-open** PLUMBING rows (the "1,217" was ledger-vintage inflation): SELF 109 (3 concentrated binaries) · CALLEE 48 · OTHER 48 · DATA 32. Probe (ov_SC03_107 SELF-21 → 14 with drafts): first run **0/14 with a PHANTOM shared error** — root-caused to **cross-group TU-edit poisoning** (a demacroize edit in TU-A persists while TU-B's drafts gate; every whole-binary build compiles ALL TUs) → fixed `recover_integration` with per-group isolation (git-checkout binary TUs between groups, gate_lane's proven pattern; engine_core.h untouched so fleet-tier arity persists) + added the **`macro-externs` draft stage** (§121: one draft's guessed `extern int f()` vs the TU's `DEFINE_`-macro definition — lifted from family_sweep, R33) + **`tu-scope` stage** (§103 STU, the sweep-only lever). Isolated re-run exposed the TRUE class: `undefined reference to D_*` = **the §171 stale-seed-symbol class** (rtu_match MATCHes these drafts — it is blind to relocation names; the R34 two-oracle disagreement exactly as documented) → **symfix-first**: 11/11 rebased (mixed 1:1 + n:n deltas incl. the S49 Δ0x4128) → **9/14 banked (64%)** where the raw path scored 0. The standing recipe: `aprop_symfix --fix → recover_integration --stages macro-externs,demacroize,tu-scope` (per-group isolated). Full sweep over the remaining ~200 rows now running.
- 2026-08-14 — **T6 COMPLETE.** Sweep over the remaining ~200 rows: **the decisive finding is Law 3 (§173)** — the biggest groups (ov_SC02_037 44 rows, most of ov_MAIN_012's 42) have **zero stored drafts**: their PLUMBING verdicts came from transient family-sweep remaps never persisted to the backlog. A recovery lane can only recover what was stored — verdict-only rows are family-lane fuel (already atlas-labeled), not recovery fuel; the true stored-draft class was largely consumed by the probe (**9 banked, 64%**). Singleton tail with drafts: 0/~20 (fresh per-fn verdicts recorded). Cookbook **§173** (symfix-first · per-group isolation · verdicts-without-drafts) + index regenerated (518). **R22 clean fleet: 213/213** with all T6 banks. Phase total so far: **17 banked, 0 agent tokens** (stubs 12,059 → 12,042).
- 2026-08-14 — **T7 COMPLETE (honest refutation).** `tools/family_align.py` built: aligned classifier (SequenceMatcher over FC.tok; LEN-LI/LEN-NOP/LEN-JTBL/LEN-STRUCT/STRUCT-ALIGNED/PURE/IMM verdicts) + li-cluster reconstructor (with the split-cluster absorb for the rs-changed addiu partner) + the aligned imm engine. **NC-1: 157/157 verdict-equivalence** with classify_member on banked pairs — the NC itself caught two real classifier gaps (R-type non-shift sa = STRUCT; registers tested BEFORE the reloc skip — a reloc-slot word with a different register is STRUCT). NC-2 parity 21/21. **The R37 probe then killed the planned driver before it was built: 0/26 LI-ONLY cards classify mechanically** (regfields drift ×19) — cousins are 0.85-similar DIFFERENT functions; §168 law 1 ("a cousin is a seeded crack, never a remap") re-derived by measurement. The engine's true consumer is **T8's LEN+N pile** (draft vs its own target = same function). Parked for T8: the reloc-vs-constant range discriminator for lui-bearing clusters. Decision-log entry written (R31).
- 2026-08-14 — **T8 COMPLETE (LEN+N lane).** Built: `match_one --emit-streams` (additive, stdout-identity NC'd) · `family_align.addr_true_rel` (the reloc-vs-constant range discriminator — full conservative set kept for pair semantics/NC-1 [157/157 regression], address-true subset for indel eligibility; synthetic probes: lui-bearing constant cluster now LEN-LI, address-anchored indel still refuses) · `tools/len_tells.py` (per-draft aligned classification + §172b tell tagging on target-side indels, detectors imported from atlas_features R33, cookbook text embedded in cards) · `tools/lenmiss_route.py` (pool-parallel per A8: **587 audit LEN rows re-verified live and routed in 24 s**). **Routing:** redraft **345** (frac>0.35 — the draft is not the function; reclassified APPEND-ONLY in the backlog, so near-miss metrics stop lying) · permuter-length **49** (|Δ|≤2 clean drift — grinder fuel) · cards **192** incl. **14 EXTPAIR/SELECT/NOP-tagged** (the audit's own detectors had emitted zero — these are new signal) · **mechanical 0 — an honest null**: stored drafts rarely get CONSTANTS wrong (agents copy them from asm); LEN drift is shape, so the LI-cluster swap lane has no fuel in this pile and family_align's value here is the classifier/detector. R32 accounting 587/587.
## Blockers
(none)
## 🛑 SESSION CHECKPOINT
T0+T1 done (commits through the t6-recover gates; T1 close commit pending the R22 run). NEXT = T2 (reference expansion: PsyQ Track-1 SAMPLE/ extraction → tools/reference/psyq-sdk/; gcc-2.7.2 calls.c + caller-save.c fetch; SETUP rows; idiom-notes seed). If resuming fresh: read the approved plan file above; check .run/t1_r22_check.log for the fleet verdict; then start T2.