- tools/atlas.py: cousin units baseline + T1.5 h_seqn merges + CALIBRATED warm tier (measured: li-norm metric holds ~99% recall to 0.55; rule = smallest t with neg-accept<=0.2% AND recall>=95% -> THRESH_WARM=0.70 @ 99.1%/0.18% — false merges waste exemplar cracks, misses only route cheaper) + seed sweep (65% of open skeletons carry a >=0.55 matched seed) + kNN graph + tiny-direct + evidence joins (audit/backlog/ledgers/cards; unparsable=fatal) + lever labels with confidence measured>ledger>tell>default>UNKNOWN - partition ASSERTED: 12,058 = progress stubs 12,051 + NM 7 EXACTLY (chased the +1: data blobs now excluded, reconciled against classify() buckets; T1 banks confirmed absent); every instance in exactly one group; main joins at the atlas layer only (family maps stay non-main — 4 silent-skip hazards) - warm tier merged 1,019; top group unifies 268 drifted per-location skeletons - lever table: head-crack 186.9k ins / UNKNOWN 138.6k (honest) / extend-tell 76.7k / redraft 46.8k / jtbl-carve 45.7k / integration 23.4k / seeded 23.4k / len-vein 16.8k / swaprepeat 9.2k / plumbing 8.1k / o0 6.6k / cc1 6.4k - atlas_features: li_norm_toks exported (shared with atlas, R33; hash-stable); mid_jr verifier fixed (compared ZERO rows — R32 silent no-op; now 6,444/6,444) - make atlas = full regen chain (~10-15 min, zero tokens); --targets emits crack slates (12/12 .s resolved); survey 92 s - SETUP rows (R21); docs/frontier-atlas.md committed
12 KiB
CURRENT PHASE — Phase 31: The Frontier Atlas & Wide-Tolerance Campaign
Started: 2026-08-14 · Plan approved: 2026-08-14 (gate 1; Drew) · Effort doctrine: xHigh default / Max deep (T5, T7, synthesis) / Ultracode waves (R26/R27 prompts) / Fable-tier only for new wall classes.
Approved plan: /home/musashi/.claude/plans/fable-5-set-max-goofy-seahorse.md (the full design; this file is the crash-recovery log).
Approval also ratified R37 (probe before costing), R38 (read recorded failure verdicts first), R39 (negative-control new refusal-checks) — now binding.
The re-charter (one paragraph)
Instead of roadmap-v2 P31's per-function grind, Phase 31 organizes the 12,059 remaining stubs (main 1,041 · resident 14 · ov 9,832 · md 1,179) into crack groups: a deterministic per-function feature layer + multi-tier similarity atlas (li-normalized exact tier, seed sweep vs the 2,719 matched skeletons, calibrated warm tier for the 3,238-unit cold tail + main, kNN neighborhood graph), evidence-joined to a lever label per group; plus widened mechanical lanes (LEN-tolerant aligned remap with a fully-mechanical LI class, §172b EXTPAIR/SELECT detectors + routing, PLUMBING campaign, permuter cluster warm-start, weak-seed cards); then a campaign loop to ceiling — deterministic $0 lanes first, agents only for exemplars, velocity-ledgered, closed on measured decay. Milestone shape = P30 (campaign to ceiling; every remaining stub on a named ledger at close). Main fully included from day one.
Task checklist
- T0 — Pivot log + freshness + hygiene — DONE 2026-08-14. (xHigh)
- T1 — Integration quick-bank sweep — DONE 2026-08-14 (pending final R22 log line). 8 banked, 0 agent tokens. (xHigh)
- T2 — References — DONE 2026-08-14. (xHigh)
- T3 — Main enablement — DONE 2026-08-14. (xHigh)
- T4 — atlas_features.py — DONE 2026-08-14. (xHigh)
- T5 — atlas.py — DONE 2026-08-14. THE ATLAS EXISTS. (Max)
- T6 — PLUMBING campaign: plumbing_groups.py + recover_integration --stages cast-callees,tu-scope; 20-fn probe then sweep. (xHigh)
- T7 — LEN-LI mechanical lane: family_align.py + adapt_autodraft.py; NC-1/2/3; 1-card probe → 27-card pile. (Max)
- T8 — LEN+N lane: len_tells.py + match_one --emit-streams + lenmiss_route.py; per-class probes; redraft reclassification. (xHigh)
- T9 — Warmstart + weak-cards: warmstart.py + grinder admission widening + --weak-cards; NCs + probes. (xHigh)
- T10+ — Campaign loop to ceiling (repeating sessions; velocity ledger; close on measured decay). (Ultracode waves / xHigh orchestration / Fable new-walls)
- Tclose — PhaseEnd (gate 2). (Max)
Standing verification (every task)
R22 clean-fleet 213/213 after every banked batch · tools-health green · 0 NON_MATCHING (G4) · dedup-check 0 failed · R32 coverage assertions on every new scanner · R39 negative controls on every refusal check · R37 probes before pricing · R38 ledgers before experiments · commit per task (task + this log in the same commit; Drew pushes).
Progress log
-
2026-08-14 — Phase planned and approved (3 Explore + 2 Plan agents; full design in the plan file). Task list built (harness tasks #1–#12). T0 started.
-
2026-08-14 — T0 COMPLETE. (1) R31 decision-log entry (the re-charter WHY + R37–R39 ratification). (2)
harvest_verify.pyimport guard: a bare import now RAISES instead of running a gate (verified both directions; CLI behavior unchanged). (3) Resident ±1 RESOLVED + FIXED:--bootstrap's linear partition had fused the +0 data word withfunc_800CEDFC(row0x800CEDF8nins=18) and droppedfunc_800D33E0past a glued tail — the true denominator is 145 (progress was right, the sig wrong).sig-residentnow ELF-seeds (S45 pattern: unique 4-aligned T-symbol addrs inside theresident_TEXT_START/ENDmarkers → exactly 145; bootstrap fresh-clone fallback). All three oracles now agree (sig 145 · corpus matched 131 · progress byte-ident 131);audit-corpus0 PHANTOM + 0 TRUNCATED. (4) Family maps regenerated at HEADcommit:2161: 11,025 open non-main members = 12,059 − main's 1,034 EXACT (the stale map's 102 phantoms cleared); cousins totals now A-prop 1,125 / seeded 1,293 / cousin-multi 5,373 / cold 3,234 inst; adapt cards 704, aprop cards 204 (full emission). (5) Main fuel gap is DEAD: 2,001/2,002 main stubs have cached Ghidra-C (onlyfunc_80049600missing) — the roadmap's "0/2,096" note was stale. (6)make tools-health→ OK (dedup 2,063/0; C1 254,521/254,521; audit-digest confirms the fleet digest; resident fix moved instr num+denom by the same +9). -
2026-08-14 — T1 COMPLETE: 8 banked for 0 agent tokens. R38-first: partitioned the MATCH-108 pile against current stubs (75 still open) and against the S50 gate history (62 gated-and-failed with verdicts · 13 never-gated). The lanes and their measured outcomes:
- Never-gated 13 → gate_lane: 0/13, but the verdicts decomposed to 11×
undefined reference to D_*= the §171 stale-seed-symbol class. Extendedaprop_symfixwith STALE-DELTA (n:n uniform-delta rebase; R39 synthetic + snapshot NCs, zero false positives; the delta test even refused a pair my hand-check wrongly accepted) → 4 rebased, 4/4 banked (func_8016BCC0,func_8017F1C8,func_80186BD8,func_80186BF8). Cookbook §171-D written in-session. - SELF-decl PLUMBING 7 →
recover_integration --stages demacroize --max-tier binary: 4/7 banked (func_80139BE0,func_8014ED28,func_80161D88,func_801659DC); 3 stay near. - Stored-draft re-gates (no-verdict 7 + close=0 8 + diff 1 + 9 STALE→clean world-motion drafts): 0/23-ish banked — the ~8% A10 stored-verdict law held again; all re-verdicted fresh.
- Handed forward with fresh classifications: CALLEE-decl 15 + CC1-FAIL 14 → T6 (cast-callees/tu-scope stages); UNDEF-DATA/OTHER 9 → the §171b-1 data-definition carry (T7/T8); md CARVE-REFUSED 8 → campaign side-quest ledger. immfix pile: fully consumed (0 open). fix20: 19/20 consumed in S50.
- Tool fixes landed:
gate_lanepropagate-commit tag now derives from GATE_PHASE (was hardcoded phase-30 S49). - Rate lesson for the velocity ledger: fresh-fix lanes (STALE-DELTA 4/4, demacroize 4/7) vastly outperform blind stored re-gates (0/23) — the campaign loop's L2 ordering is confirmed by measurement.
- Never-gated 13 → gate_lane: 0/13, but the verdicts decomposed to 11×
-
2026-08-14 — T2 COMPLETE (references). (1) PsyQ dev-CD extracted: walked the on-disk Track-1 image (MODE2/2352) with the frozen
tools/bfm_extract/iso9660.py(R33 — no new extractor; walker = iter_directory/read_extent with out-of-range extents skipped) →tools/reference/psyq-sdk/(gitignored): 2,374 files / 231.6 MB, 400 C sources (373 in PSX/SAMPLE/ across CD/GRAPHICS/SOUND/MODULE/CMPLR/…); only 7 out-of-track.DAaudio skipped. (2) Provenance find:GNU/SNGNUVER.TXT= SN Systems' gcc build history (2.7.2.SN32.3.7.0002, 14.5.97) with per-build changelog of SN's patches vs vanilla — onlyUNROLL.C(parameterised max unroll insns) is codegen-relevant; recorded in the idiom notes as the first-look suspect if a loop-unroll residual ever defies the vanilla model. (3) gcc-2.7.2 reference completed: +6 files from GNU ftp (calls.c+caller-save.c— both cited by §172's producer model, previously missing — + integrate/optabs/varasm/recog), tarball sha2567cd8bce5…recorded. (4)docs/psyq-sample-idioms.mdseeded (inventory, provenance, first style conventions, the lane hook); SETUP §5.6 rows added (R21). -
2026-08-14 — T3 COMPLETE (main enablement). (1)
sig_imageseed-ends extension:--seedsnow accepts0xADDR NINSlines (and jsonlnins); a seeded nins is authoritative — the slice is exactly[addr, addr+4·nins), bypassingfunc_end(whose heuristic mis-sliced 3/40 main samples). (2)corpus.s_ins_count()factored from audit() (R33, one counter) +corpus.py <bin> --seed-endsCLI. (3)make sig-main: 2,002 main stubs signed at splat-true lengths →.run/sig.main.jsonl(deliberately splat-SEEDED — the atlas needs the boundaries a match must hit; NOT the independent second oracle, which stays deferred perdocs/second-oracle.md). (4) Full word cross-check: 2,002/2,002 slices byte-faithful (EXE bytes vs.scomment-column words; 0 SLICE-SUSPECT; note the.sword field is byte-order hex, not the LE-decoded value — first checker draft compared wrong and read 1,999 false suspects). (5)family_remapmain special-cases:vram_of('main')= code-segmentvram − startfromsplat.us.exe.yaml= 0x8000F800;img_path('main')from the same yaml's target_path (R33).stream_words('main')==.swords on 25/25 samples. (6) Regression:sig-residentre-run byte-identical after the sharedread_seedschange. -
2026-08-14 — T4 COMPLETE (
tools/atlas_features.py). Per-function features memoized per distincth_exact(.run/feat_memo.json) and fanned out 1:1 with sigs to.run/feat.<bin>.jsonl×213 (main reads the splat-truesig.main.jsonl; registry untouched). 363,525 rows / 92,855 distinct bodies in 21 s (vs the 3–6 min estimate). Features: nins/band · o0 prologue tell · frame/callee-saved set+order/fp · CFG skeleton (nblk/ncond/nback/ret_n via branch-target scan; jal = call, never an edge — position-independence proof in the docstring) · mid_jr/jalr · stable-call-sequence hash (fixed main+resident ranges) · reloc-kind sequence hash · 16-bucket ophist · §172b tells as importable shared functions (extpair/dupselect/sign_mix/magic_div — T8's len_tells imports these, R33) · li-normalized skeletonh_seqn(non-anchor lui dropped, ori→addiu class). Verifies: R32 join asserted per binary at write; determinism 0 mismatches on 200 re-derived bodies; mid_jr cross-check vs family_hseq's independent oracle: 6,444/6,444 exemplars agree (first verifier draft compared ZERO rows — a string-vs-int addr type mismatch, the R32 silent-no-op class caught in my own verifier; now indexed + fails loud if compared==0). o0-tell vscorpus.is_o0on 13,019 open fns: 41 both / 94 tell-only = candidate undiscovered -O0 fns in -O2 TUs (§116 — o0_subsplit carve fuel for the campaign) / 14 src-only outliers. -
2026-08-14 — T5 COMPLETE (
tools/atlas.py+make atlas) — THE FRONTIER ATLAS EXISTS. Survey in 92 s, all assertions green: 5,139 groups cover ALL 12,058 open instances / 613,710 ins (12,058 = progress's 12,051 stubs + 7 NON_MATCHING exactly — the +1-blob and +8-banked discrepancies were both chased and resolved: data blobs excluded, T1 banks confirmed absent; main open reconciles to classify()'s stubs+NM with a hard assert). Tiers: calibration measured (positives = cousin ≥0.85-raw merges, negatives = size-matched cross-unit pairs; the li-normalized metric is so discriminative that recall ~99% holds down to 0.55; rule = smallest t with neg-accept ≤0.2% ∧ recall ≥95% → THRESH_WARM=0.70 at 99.1%/0.18% — chosen because a false warm-merge wastes an exemplar crack, a miss only routes cheaper) · T1.5 h_seqn 2 merges · warm tier 1,019 merges · seed sweep: 4,702/7,247 open skeletons (65%) carry a ≥0.55 matched seed from the 2.7k-skeleton pool · kNN graph (top-8 open + matched neighbors) · tiny-direct. Lever table (the strategic map): head-crack 1,283g/186.9k ins · UNKNOWN 1,964g/138.6k (honest) · extend-tell 575g/76.7k · redraft 46.8k · jtbl-carve 190g/45.7k · integration 575 inst/23.4k · seeded-crack 23.4k · len-vein 778 inst/16.8k · swaprepeat 9.2k · plumbing 161 inst/8.1k · o0-lane 6.6k · cc1 6.4k. Top group: 268 drifted single-member skeletons unified (per-location ~23-ins family, 0.72 seed). Evidence joins accounted (audit/backlog/ledgers/cards; residuals EXCLUDED-STALE until regenerated);--targetsresolves 12/12.s.make atlas= the full regen chain. docs/frontier-atlas.md committed.
Blockers
(none)
🛑 SESSION CHECKPOINT
T0+T1 done (commits through the t6-recover gates; T1 close commit pending the R22 run). NEXT = T2 (reference expansion: PsyQ Track-1 SAMPLE/ extraction → tools/reference/psyq-sdk/; gcc-2.7.2 calls.c + caller-save.c fetch; SETUP rows; idiom-notes seed). If resuming fresh: read the approved plan file above; check .run/t1_r22_check.log for the fleet verdict; then start T2.