Commit Graph

773 Commits

Author SHA1 Message Date
Drew T 3b31508a24 feat(phase-27): the honest frontier — fix the instruments, audit the disc, dissolve a wall (v1.26.0)
- INSTRUMENTS FIXED: Makefile fail-closed (report's gates were swallowed); one cdecl typedef-strip
  primitive (was 6 regexes); scanners derived not hand-listed (difficulty/exemplar_miner); the
  second boundary oracle extended to resident. Each fix CHANGED an answer the old tool hid.
- DISC AUDITED HONEST: 4 hidden SC07 overlays onboarded (136->140, code at PAC entry 1) + 39
  un-onboarded type-1 code modules found (resident-class, load-address RE pending). The byte-gate
  is blind to un-onboarded code (R34); game-code TRUE 100% now spans 140 + ~39. Instr 68.9->67.0%
  (denominator correction, not regression).
- PIN-CRASH WALL DISSOLVED: the §42e "cc1 SIGABRTs the sibling TU" wall is the extract_unit macro-
  drop (sched.c:2725), fixed (T5 _carry_macros); pinned families stage 133/133 clean -> P31 open.
- HONEST FRONTIER: worklist --assert-partition (R32, caught 5 stale rows); ledger corruption fixed;
  calibration.md (the templatability swing: h_exact cores ~xN, h_seq families ~0% -> B2 refuted).
- FABLE5 SPRINT: 4 cracks + the SIGABRT, 0 banks, but 3 wall reclassifications + the wall dissolved
  + ~9 pin-free levers distilled (cookbook §42e/§44 + regalloc/cse_expr §H + decision-log R31).
- 140/140 byte-identical (R22), 0 NON_MATCHING (G4), audit gates green + fail-closed. No tools
  installed. rules R35 (fix the instrument before trusting its measurement). bumps 1.25.0 -> 1.26.0.
2026-07-15 20:42:32 -06:00
Drew T d1ef983af0 docs(phase-27 T1 close): func_80176734 crack + distill — the CSE address-fold antidote
The last Fable5 pass of the sprint (Drew capped further waves at 86% context). func_80176734 (371 ins,
fresh un-drafted core): NO bank (mine=370 vs 371, 5 permuter-shaped clusters — entry-schedule tie,
caller-saved shuffles, a combine-merge missing insn, qty ties), pin-free, honestly handed off (P9;
match_one confirms the DIFF). Draft -> decomp-permuter warm-start (P29).

Idiom harvest (cookbook cse_expr §H):
- THE CSE ADDRESS-FOLD ANTIDOTE (zero asm): find_best_addr's cost-ungated qty-const fold + from_plus
  re-association eat reg-based global accesses on every cse walk; a balanced if/else DIAMOND makes the
  merge label barrier-preceded -> fresh cse table -> both folds die with no #APP. Replaced two asm dials.
- update_equiv_regs doubles live_length for single-set REG_EQUIV pseudos (local-alloc.c:1064) — a 2nd
  set forfeits the doubling, ~4x the allocno priority; explains a "my dial broke the $s-order" class.
- record_jump_equiv fall-through delete (cse.c:7511) — a recognition tell for genuine dead source logic.

T1 sprint COMPLETE: 4 cracks + the SIGABRT characterization, 0 direct banks, but 3 wall reclassifications
+ 2 cracked roots + the pin-crash wall dissolved + ~9 new pin-free levers. Fable5 DISCOVERS, cheap-Opus
APPLIES — the ROI is idioms, not banks (docs/calibration.md).
2026-07-15 20:34:35 -06:00
Drew T ce88baf365 feat(phase-27 T9): calibration — the templatability swing measured (structural families are NOT cheap)
docs/calibration.md — the byte-gate-grounded rates that size P28/P29 (roadmap §6 held yield
projections until this).

- VELOCITY: instr 68.9->67.0% (a T7 denominator re-baselining DOWN, not a regression) + ~0 matches
  banked (an infrastructure/findings phase). The honest flip-checkpoint read: denominator correction
  + unblocking findings, NOT 0 progress/session — velocity resumes at P29, re-measure there.
- THE TEMPLATABILITY SWING (decisive for P28/P29): h_exact reach-N cores propagate ~xN near-100%
  (§52: 5 cores -> 670 instances) vs h_seq/h_norm structural families ~0% (0x8017BEBC: 106/112 stage
  but 0/8 bank, all genuine DIFF). So remaining yield = per-member cracking + mechanical xN for the
  h_exact cores, NOT "template x120 the 986 families" — B1/B2's cheap-harvest hope is byte-refuted.
  The 223-stub frontier: 101 reach-134 (xN-able if cracked) + 119 reach-1.
- COST/TIER: Fable5 ~230k tok/fn, 0 banks / 5 — ROI is idioms + the pin-crash wall dissolved, not
  banks (the doctrine held). cheap-Opus is the banking tier; permuter tail exhausted; local-v3 $0/<=15.
- HONEST GAP: the headline member-adapt close-rate on register-drift members needs P28's member_adapt
  tool (chicken-and-egg) -> P28 opens by measuring it on a byte-gated sample, per the risk register.
- NEW un-projected fuel: ~20 PINS-class stubs now harvestable (pin-crash dissolved, T5).
2026-07-15 19:40:08 -06:00
Drew T 0f68a83cfa feat(phase-27 T8): worklist --assert-partition (R32) + honest re-scan + ledger corruption fixed
- worklist --assert-partition: the audit's literal R32 prescription (tooling-audit.md:1173) —
  enumerate live stubs from corpus.stubs (the invariant, R33), assert the fuel manifest partitions
  its source overlay's stubs, exactly one row each. Scoped honestly (worklist's universe is ONE
  overlay ~223 stubs, not the fleet's 53k — a fleet partition is a scope change, not a flag). PROVEN:
  it caught 5 stale rows (pin-free cores Phase-26 banked, manifest never re-derived) -> FAIL exit 1.
- honest re-scan: build_fuel_manifest on the fixed tools + 140 binaries. Giants re-verified reach-138
  (was 134 — the SC07 overlays now counted). Partition PASSES 223==223 after refresh.
- ledger corruption fixed: func_80178004's 2 false `close=0 "MATCH"` records (a Phase-26-retracted
  myth — the seed's best was 5 pinned, and a real close=0 whole-binary match BANKS; it is still a
  stub) -> corrected to the honest close=91 regalloc wall. func_8012E364 already honest (close=23 —
  the "stale closeness" flag was itself stale). No real duplicate rows (load_best dedups by addr;
  the uniq hits were func names in where_stuck prose). docs/worklist.md + docs/backlog.md regenerated.
- the 1,670-untriaged near-miss triage SCOPED TO P29 (P5d): Phase-21 automation leftovers whose class
  labels re-derive at harvest, and the pin-crash finding re-buckets the PINS class — an Ultracode
  fan-out buys low-durable labels; the gate's residue map is the partition + the class summary, done.
2026-07-15 19:37:28 -06:00
Drew T 54bdf96218 docs(phase-27 T1): distill the Fable5 wave — the pin-crash wall refuted + 3 RC-6 downgrades
The flywheel step (R16/R30): turn the wave-1 + SIGABRT byte-proven findings into cookbook/codegen-map
knowledge, in the producing session. The distillation REWRITES wall verdicts, so accuracy is load-bearing.

- cookbook §42e-CORRECTION: the "pin-crash wall" (register-pin-heavy families "SIGABRT the sibling TU,
  ov077-TU-context-specific, NOT ×134-recoverable") is REFUTED. The SIGABRT is real (sched.c:2725
  create_reg_dead_note, a sched1 REG_DEAD-note conservation bug) but was TRIGGERED by extract_unit
  dropping file-scope #define macros (the T5 bug) -> implicit-call GTE ops -> caller-saved pins in the
  fatal shape. Only 1 of 4 families genuinely crashed; 3 were exit-33 plumbing folded into one crash
  bucket. Fixed, all 4 stage 133/133 clean. Per-pin predicate recorded. P31's pin route is OPEN.
- cookbook §44-Lever-5: the 3 functions it cited as intrinsic (func_8014D820/8016CBC0/801670E4) are
  each oracle-refuted (2 cracked roots + 1 RC-6-not-S3). Corrected the "NEVER ship pinned, it SIGABRTs"
  claim per §42e-CORRECTION.
- gcc-2.7.2-map/regalloc.md §H: THE reg_renumber-swap oracle (discriminate RC-6 allocation from S3
  scheduling in one gdb run — patch reg_renumber at reload entry, swap the contested regs; byte-exact =
  pure allocation), RC-14 reused-load-temp serialization (the MERGE pole; pin-free, cheap-Opus), RC-15
  the density dial across a floor_log2 boundary (subsumes "coalescing knife-edge"), and the local-vs-
  global allocation tie as a precisely-named honest sub-class. Continues the §F/§G RC-6-downgrade series.
- decision-log.md R31: the 3 Phase-27 strategic findings (disc is bigger: 140 + 39 modules; a wall was
  our tool again; a cheap win is dead) + the through-line — the roadmap's numbers were red-teamed, the
  tools under them were not, until this phase.

func_80176734 (fresh-core wave-2 agent) still running; its findings fold in before the PhaseEnd.
2026-07-15 19:30:12 -06:00
Drew T ee4b3a02e8 feat(phase-27 T5): extract_unit carries file-scope #define macros — honest 0x8017BEBC probe + the pin-crash wall dissolved
family_remap.extract_unit dropped the file-scope function-like #define macros a body references
(the gte_* C inline-asm GTE-op macros live ABOVE the function; the backward preamble walk stopped at
the first #define/continuation line). A staged sibling saw every GTE op as an implicit-declaration
CALL. Two consequences in one bug:
- staging failed: 0x8017BEBC's 112 members all CC1-FAIL'd -> a FAKE 0% probe that reads
  "mechanical harvest dead" when the tool was broken (a 4th phantom exhaustion proof, exactly the
  26-A audit class).
- the SIGABRT: with a caller-saved register PIN present, the phantom call pushes cc1's sched1 into
  create_reg_dead_note's abort (sched.c:2725) — the §42e "pin-crash wall". The wave-2 SIGABRT agent
  proved this IS the cause (.run/giants/pin_crash_sigabrt.md): pinned families stage 133/133 clean
  once their macros ride along. A propagation wall recorded as a compiler limit for phases = a
  staging-tool artefact.

- _carry_macros: prepend the function-like #define macros the unit body references (file order),
  not already inside the unit. Safe by construction: feeds only the templating path (remap_hseq),
  never make_macro's engine_core.h lift (no #define embedded in a DEFINE_func_*() macro);
  gather_externs only scans func_/D_ so no bogus extern; regression-verified non-GTE exemplars
  carry 0 macros (a no-op where it should be).

THE HONEST PROBE (R14): staging 0 -> 106/112 (6 skip = IMM tier-2). Bounded 8-member gate sample =
0 banked / 8, ALL genuine DIFF (T4 classifier: compiled, wrong bytes — NOT plumbing). 0x8017BEBC is
BYTE-PROVEN NOT TEMPLATABLE: the roadmap's "largest cheap win left" (B2) is REFUTED. The h_seq match
is necessary, not sufficient. This 0% MEANS something because the tool is fixed first.

Carried to T8: harvest the now-unblocked pin families (verify the 133/133 claim + bank).
2026-07-15 19:20:56 -06:00
Drew T 427baba3bf feat(phase-27 T10): completion dashboard (main in the weighted metric) + the resident second oracle
The metrics contract (roadmap §1) wants all three metrics WITH main in the denominators, and the
second, independent boundary oracle (R34) extended beyond the overlays. Both had landmines.

10a — main into the weighted metric, safely:
- weighted_metrics off the func_-only src_stubs regex onto corpus.stubs (R33). THE LANDMINE IS
  REAL: src_stubs("SLUS_007.26") globs src/SLUS_007.26/*.c -> 0 files -> every row "matched" ->
  main 100% + fleet % silently inflates. Routing through corpus.stubs is a PROVEN 0.000pp no-op on
  the existing fleet (overlays are all func_) and closes the curated-name leak.
- a SEPARATE "MAIN game-code weighted" line (0.7%): main's Ghidra sig excludes the LINKED PsyQ
  objects (Ghidra never analysed them), which is exactly right for a game-code metric (LINKED is
  complete, counted in fn-count). Reported un-folded and caveated (month-stale sig, PROVISIONAL) —
  folding a stale/incomplete value into the decomp.dev headline would mislead the flip checkpoint.

10b — the resident second oracle:
- make sig-resident: sig_image on the resident flat blob (byte-derived, not Ghidra). corpus.
  sig_is_independent now covers resident -> audit-corpus checks its boundaries too. Probed clean
  BEFORE wiring (144 fns, all 21 stubs present, 0 phantom), verified 0 phantom + 0 truncated.
- sig-overlays now derives its payload list from config/overlays.mk, not a 0.4.dec glob that
  silently dropped the 4 SC07 index-1 overlays (the audit's own silent-skip class). tools-health
  regenerates sig-overlays + sig-resident first so the audit never crashes on an absent sig.

10c — main's second oracle: docs/second-oracle.md. sig_image can't sign the PS-X EXE yet (0x800
header offset, interleaved data/linked islands, one text range); seeding from splat would destroy
independence for the PHANTOM class specifically. Honest deferral + scoped design, not a fake oracle.

- docs/progress.fleet.md regenerated: 140 binaries · fn-count 82.16% · instr-weighted 67.0%
  (the honest post-T7 drop from 68.9%) · distinct 47.8% · MAIN game-code 0.7% (separate).
- SETUP §6.3 updated (R21).
2026-07-15 18:51:43 -06:00
Drew T 264fe6c115 feat(phase-27 T7): disc-completeness audit — onboard 4 hidden SC07 overlays (136->140) + the type sweep
The whole-binary byte-gate is structurally blind to code nobody onboarded (R34): check-all is
green over the onboarded set no matter what code sits unbuilt on the disc. This reconciles the
onboarded set against every code-bearing PAC payload.

- new_overlay.sh: optional [ENTRY] arg (default 0.4) reaches a non-0.4.dec payload. Onboarded
  ov_SC07_{006,007,010,011} from 1.4.dec (they put graphics at PAC entry 0, the code overlay at
  entry 1 — invisible to the 0.4 hardcode for a month). Each byte-identical (7ca772be / b3b95547 /
  d7b5875d / 9885af74). FLEET 136 -> 140; check-all 140/140 (T2's pass==N re-baselined cleanly).
  difficulty.py NOT in the insertion set anymore (it derives, T6) -> only 3 tool dicts touched.
- tools/disc_code_sweep.py: decode every payload (reusing sig_image.make_insn) and gate code on
  BOTH valid>=0.90 AND jr_$ra density>=0.01. The jr_$ra gate is decisive: isValid() alone flags
  389 false hits (type-0/2 structured data decodes ~100% valid but has ZERO returns); jr_$ra
  separates code (~2.9-3.4%) from data (0.000%), validated on positive+negative controls.
- FINDING (docs/disc-completeness.md): type-4 location overlays are COMPLETE (138/138). All other
  types are data EXCEPT type-1 = 40 code payloads, 1 onboarded (the resident), 39 HIDDEN
  resident-class modules (mostly MAIN.CD/FILE_XXX/1.1). They load at UNKNOWN addresses (not the
  shared overlay slot), so they are NOT mechanically onboardable — byte-verifying a build binary
  needs its load address (P9), knowable only by runtime RE (the Phase-3 method). Deferred with
  evidence, NOT force-onboarded at a guess.
- CONSEQUENCE: game-code TRUE 100% now spans 140 onboarded binaries PLUS ~39 type-1 modules
  pending load-address RE. The roadmap assumed 136 — this is a real re-baselining (the +4 overlays
  also add ~2.45 MB to the denominator; every family propagation is now x138). Flows to T10/T11.
- SETUP §6.3 tool inventory updated (R21).
2026-07-15 18:33:37 -06:00
Drew T 8a1a1a0d79 feat(phase-27 T6): migrate difficulty + exemplar_miner off hand-lists/proxies (R33) — before T7
The 26-A audit named difficulty's 136-entry BINARIES dict as the exact root cause corpus.py:9
describes (a hand-maintained allowlist over a filesystem that already answers the question), and
exemplar_miner's registered_addrs() as a ~60%-wrong proxy for "is this still work?". T7 onboards
new overlays, so these are migrated FIRST or the new binaries silently miss make report.

- exemplar_miner.py: "still a residual?" now = corpus.stubs(source) membership (the INCLUDE_ASM
  invariant, R33), not dp.registered_addrs() (config/dedup.us.yaml — a matched-but-unregistered
  fn, e.g. banked inline or matched-but-local, stayed wrongly in the residual pool).
- difficulty.py: the 136-entry hand-dict -> cfg_for(alias), deriving the mechanical layout
  (src/<a> + asm/<a>/nonmatchings; main/resident the two specials). Validated against the tree
  (src/<a> must exist -> a typo is a clean error, R32), not a hand-list. A newly-onboarded overlay
  now needs zero difficulty registration.
- new_overlay.sh: DROPPED difficulty from the sentinel-insertion set (T6 made it obsolete; leaving
  it would insert a dead dict entry into a file that no longer has a dict). The other 3 tools
  (diff_settings/progress/dup_report) keep their hand-lists — migrated one-at-a-time, byte-gated,
  per the audit's cadence; NOT dup_report.BINARIES, which corpus depends on as the binary list.

VERIFIED:
- difficulty derivation is BYTE-EXACT vs the old dict for all 136 aliases (0 mismatches), and
  old-tool vs new-tool output is byte-identical (.md AND .csv) on the same tree — the diff vs the
  committed docs was pure staleness (committed 2026-06-20, tree at 2026-07-15), NOT my change (R14).
- unknown alias -> clean error, not silent-empty output.
- exemplar_miner runs -> 223 residual stubs (corrected; it's a manual tool, not in make report).
- new_overlay.sh: bash + embedded-python both parse; difficulty absent from the insertion set.
2026-07-15 18:20:55 -06:00
Drew T e0a0becfaa feat(phase-27 T4): one cdecl typedef-strip primitive (was six regexes) + surface cc1 stderr
The plan named two defective regexes; the tree had SIX with complementary holes, each
silently recording the resulting compile failure as "not a match" — a plumbing error
wearing a compiler wall's clothes, the exact class the 26-A audit exists to end (R32).

- cdecl.py: the canonical primitive — typedef_names(tu_path) + strip_provided_typedefs
  (draft, provided). Built on tu_statements (robust) NOT tu_scope (which coverage-asserts
  -> would crash the byte-gate on any unrelated unparseable file-scope statement). Splits
  multi-typedef lines (split_statements, depth-aware); covers scalar AND struct typedefs;
  keeps draft-local types. lru_cached.
- harvest_verify.py: strips PER-TU (cdecl.typedef_names of the draft's real target TU) ->
  unblocks the 39 struct-typedef drafts the scalar-only _TD dropped. And SURFACES cc1
  stderr: build() stashes it; a single-draft failure is classified DIFF / PLUMBING:… /
  CC1-FAIL / SKIP -> .run/harvest_failed.classified.txt. A `redefinition` is no longer
  recorded byte-identically to a codegen miss.
- masked_diff.py: strip_scalar_typedefs() (common.h set derived from the header once, R33,
  cached) replaces SCALAR_TYPEDEF_RE.sub for the ISOLATED compile; wired into match_one +
  p16_permute. Fixes the multi-typedef-LINE skip that discarded 42 masked-MATCH drafts over
  whitespace. Unblocks B4's func_8015C32C (redefinition of 's16').
- canon_sig_reconcile / eval_lora / format_finetune keep their own copies — migrate
  per-bank, byte-gated (the audit-prescribed cadence, not a big-bang swap).

VERIFIED:
- HEADLINE known-answer: func_8015C030 -> MATCH (23 ins) UNEDITED via match_one (was
  CC1-FAIL; the multi-line typedef split alone fixes it — a live x134-family draft that
  was being discarded over whitespace).
- unit: 7/7 scalars stripped; a local struct KEPT; a TU-provided Blk16 stripped.
- classifier unit: DIFF / PLUMBING:… / CC1-FAIL / SKIP all label correctly.
- all 5 edited tools import + AST-parse clean.
- R22 clean-fleet: check-all 136/136; main clean-rebuild 143dbb89. (A mid-test c4546248
  "mismatch" was a stale-incremental artifact from concurrent compiles, cleared by a clean
  rebuild — the R22 lesson; edits touch only tools/, src/ stayed git-clean.)
- SAFETY: a strip bug can only fail-to-bank, never falsely bank (INCLUDE_ASM pastes the
  original asm; a wrong draft always changes bytes -> always fails SHA1).
2026-07-15 18:11:43 -06:00
Drew T ebdef9012b feat(phase-27 T2): make the Makefile fail-closed — the enabling fix for every downstream gate
The roadmap §5 asserted `make report` is fail-closed. It was NOT: .ONESHELL sends each
whole recipe to one `bash -c`, so with no -e only the LAST command's exit survives and
every earlier failure is swallowed. `dedup-check` "gated" purely by being last;
lint_symbol_refs / progress --audit / difficulty / dup_report were non-gates. That is the
26-A audit's own thesis (a loud failure nobody counts is as invisible as a silent one)
biting the audit's infrastructure — and until it's fixed, any R32 assertion added to a
report-invoked tool is swallowed on arrival.

- .SHELLFLAGS := -ec (global fail-closed). ONE documented opt-out: check-env (set +e — its
  contract is accumulate-every-failure-and-report, which -e would truncate at the first
  missing tool).
- check-all:610 grep -c landmine fixed (|| true): grep -c exits 1 on zero matches, which -e
  treats as fatal in a command substitution -> check-all would FAIL exactly when nothing did.
- check-all / extract-all: assert COVERAGE (pass == N), not the absence of a failure marker.
  The old `fail == 0` / `! grep -q` form was a VACUOUS PASS on an empty pipeline (R32).
- new `make tools-health` = audit-corpus + audit-cdecl + report, fail-closed — the deliberate
  pre-matching ritual the roadmap's standing invariant names, and the dependent the two
  derived oracles never had (nothing invoked them). NOT a report/build prereq — audit-cdecl
  cross-compiles every C decl through gcc (~minutes). SETUP §6.3 documents it (R21).

VERIFIED:
- NEGATIVE CONTROL (the proof): a broken lint_symbol_refs makes `make report` exit 0 under
  the old .SHELLFLAGS=-c and exit 2 under -ec. The swallow was real, not theoretical.
- the grep -c landmine + the vacuous-pass both reproduced and fixed in isolation.
- check-env still exits 0 (the opt-out works); recipe sweep found the Makefile already
  -e-aware (set -o pipefail, explicit || true) — line 610 was the only real hazard.
- R22 clean-fleet: make check-all -> 136/136 byte-identical; a forced main re-extract+rebuild
  drove the full splat->cpp->cc1->maspsx->as->ld->objcopy->check pipeline under -e -> 143dbb89.
- audit-corpus 7s / audit-cdecl green / tools-health wired.
2026-07-15 17:56:45 -06:00
Drew T 6e99157bcf chore(phase-27 T3 addendum): widen the .run/giants allowlist — cookbook §45 cited untracked files
Found while reading the seeds for T1: cookbook §45 names
.run/giants/func_80133CD4.fable.c as its worked example and .run/giants/fable_cd4/
as the flagship's gdb oracle — BOTH were untracked. The docs cite artifacts that
were not in the repo.

- .gitignore: widen by FILE TYPE, not directory — .run/giants/*.{c,md,sh} +
  fable_cd4/*.{c,md,sh,gdb,txt}. +49 files / 460K.
- Now preserved: the flagship func_80133CD4 crack + its gdb oracle (§45's cited
  worked example); the byte-verified pf*.c regression ladder (the seeds' own
  Method/reproducibility section cites it: pf2 78, pf_c2 30, pf_d1 35, pf_h1 280);
  the dump.sh/mon*.sh RTL harnesses; the banked giants' drafts (80135480, 80163EC8,
  80166994).
- Still ignored (regenerable via dump.sh, R33): d_pf*.i.*, *.s, dumps_m*/, and the
  ILS/permuter .log files. Negative control re-verified: all 5 probes IGNORED, no
  db.*.gbf staged (R23).

Lesson (R31 candidate): a doc that cites a path is an untested claim about the repo.
The §45 citation and its file were 4 days out of sync; only reading the seed for an
unrelated reason caught it. Candidate lint: cookbook path citations must resolve to
tracked files.
2026-07-15 17:35:29 -06:00
Drew T 2351c43e66 chore(phase-27 T3): preserve the irreplaceable .run/ recon (R20) — 2.2M, not 12.3M
Pulled ahead of T1: the Fable5 sprint's agents work inside .run/, and its Phase-25
seed recons were untracked — an agent overwriting .run/giants/*.opus.c would have
destroyed irreplaceable input. 5 minutes to remove that risk.

- .gitignore: /.run/ -> contents-exclude form (/.run/* + ! exceptions), following the
  /tools/bin/*.sha256 precedent. Resolves R20 (commit irreplaceable RE work) vs R12
  (.run/ is scratch) by splitting the directory on the real axis: what a rerun CANNOT
  reproduce.
- PRESERVED (~2.2M / 31 files): the 6 Phase-25 *.opus.{c,md} giant seed recons (49K);
  the func_80178004 gdb-on-cc1 harness + ORACLE_PROOF.md + the v00-v07 draft ladder +
  the sched/combine .lst evidence (~110K — the distilled output of a 477k-token Fable5
  pass, and the method §52 lever 6 depends on); backlog.jsonl (1.9M) + fuel_manifest.json.
- STILL IGNORED (regenerable, R33): dumps_v00..v07/ and d_pf*.i.* gcc RTL scratch —
  12.3M reproducible via runorc.sh + the .gdb scripts; the MCP log; draft scratch.
  The plan said "track the dirs"; the bytes said the dirs are 96% regenerable.
- VERIFIED both directions: git add --dry-run stages exactly the 30 intended files and
  0 bulk; negative control — ghidra-mcp.log / dumps_v00 / d_pf.i.sched / d_pf.s all
  still IGNORED. No db.*.gbf staged (R23 restart-noise).
2026-07-15 17:32:59 -06:00
Drew T 002f6d7c7b chore(phase-27 T0): Phase Start — plan approved (gate 1) + the P27 verification corrections
- CURRENT_PHASE.md: the approved plan (11 tasks, effort-annotated per R7), the
  Planning-verification section, blockers, and the Task-0 log entry
- harness task list built before work (R28); deps enforced T6->T7, T2->T8,
  T2+T7->T10, T5->T9, all->T11

VERIFICATION (R14 at planning scale — the roadmap's own §0 mandate): 3 read-only
agents checked every P27 premise against the repo. ~28 specifics corrected. The
three that reshaped the plan:
- the 0x8017BEBC probe ("possibly the largest cheap win left") would fail 112/112
  today on a TOOLING defect — extract_unit drops the 8 file-scope gte_* macros the
  banked exemplar needs; 0 of 112 member TUs define them. It would have been logged
  as a 4th 0% exhaustion probe: the 26-A audit's thesis, about to recur.
- `make report` is NOT fail-closed (roadmap §5 asserts it is): .ONESHELL + no -e in
  .SHELLFLAGS => only the last command's exit survives; lint_symbol_refs/progress
  --audit/difficulty/dup_report are swallowed. check-all/extract-all assert fail==0,
  not pass==N => an empty pipeline is a vacuous pass.
- 4 code-bearing SC07 payloads (~2.45 MB, 98.0% plausible-opcode) are invisible to
  every tool: they sit at PAC entry index 1 while new_overlay.sh hardcodes 0.4.dec.
  Zero mentions in docs/ or config/.

DROPPED with reasons: 0x8013C414 (x134 contested by an explicit verified_reach:1;
already drafted) · func_801549F8 (from the SUPERSEDED megaplan; Phase 26 walled it
17/31 and wrote "do NOT hand-grind it"; inside the 0/958 re-gate) · func_8012E364
(the "stale closeness" label is itself stale — close=23) · jtbl_carve "fix first,
load-bearing for B5" (the audit measured it twice: 0 of 5043 jtbl ends differ).
B4 dissolves into T4 — its remedy already ran (A9b, 1/7) and the 1 is banked.

Owner decisions (Drew, 2026-07-15): curated .run/ preservation (R20 vs .gitignore —
every Fable5-sprint input is currently untracked) · full disc audit incl. the
type-sweep, accepting the denominator expansion · Fable5 in 2 waves, distill between.
2026-07-15 17:30:37 -06:00
Drew T 1bbab78c65 feat(phase-26): close — family engine + the tooling-integrity audit + the §52 discovery-flywheel; mechanical harvest byte-proven exhausted; fleet 58.2->68.9% instr (v1.25.0)
- FAMILY ENGINE (Tasks 1-6): family_remap (extended reloc tracker + single-pass subst) +
  family_hseq (the h_seq reframe) + family_sweep (crack-one -> template-x134 -> byte-gate) +
  canon_sig_reconcile v3.2 + rtu_match. Sessions 6-8 cracked 13 cores (58.2->66.5% instr).
- 26-A TOOLING-INTEGRITY AUDIT (inserted half-phase, A0-A11): the tools WERE several of the
  walls. Fixed ~15; DELETED decaying scanners (R33); built corpus.py + cdecl.py (derived,
  coverage-asserted oracles) + make audit-corpus (a SECOND, disagreeing oracle, R34); the
  listCdBuffer 193-slice corpus defect -> 0; masked_diff 150 closeness-lies -> 4; the stale-
  object false-pass closed. Payoff 66.5->68.6% instr. docs/tooling-audit.md AUDIT-CLOSE LEDGER.
- 52 DISCOVERY-FLYWHEEL (Task 7): single Fable5 on func_80178004 = intrinsic 3-integer regalloc
  wall, BUT distilled the walker-family idiom (52); two cheap-Opus waves applied it -> 5 pin-free
  cores banked x134 = 670 instances (68.6->68.9%). 4 named wall classes; 52/52a/52b. pin-guard
  comment false-positive fixed; family_sweep --allow-pins.
- FINDING (R14/P9, 3 probes 0%): the matched-sib mechanical harvest is EXHAUSTED; the manifest's
  ~13k "templatable" members are an h_seq prediction the byte-gate refuses. Phase-26's templating
  thesis is spent -> close, open Phase 27 with byte-gate-honest re-scans.
- R22 clean-fleet 136/136 BYTE-IDENTICAL throughout; dedup 1840/0; 0 NON_MATCHING (G4).
- rules R32 (coverage assertion) / R33 (derive, don't re-derive) / R34 (a second, disagreeing
  oracle). worklog -> phase-ends/logs/Phase26.md (R19). bumps 1.24.0 -> 1.25.0.
2026-07-15 14:54:25 -06:00
Drew T a33c6f85f5 docs(phase-26a): A11 — distill + close the tooling-integrity audit (26-A COMPLETE)
Closed the inserted half-phase. tooling-audit.md: DIAGNOSIS -> AUDIT-CLOSE LEDGER
(A1-A10 outcomes + the payoff 66.5->68.6% instr + remaining/handoff); the "two
rules" -> R32/R33/R34 crisp for P10 ratification at the Phase-26 PhaseEnd.
decision-log: the A10 wall-re-test verdict (R31 -- the broken tools WERE the walls;
the payoff was banked by the fixes; the closeness-0 residual is genuine; the real
deliverable is the 3 rules + the derived-oracle pattern). SETUP: the A9d-A10 tool
changes (R21). Cookbook §51 verified complete; LAW 3 tagged R34.

Observables green: final R22 clean-fleet 136/136 BYTE-IDENTICAL; make report EXIT 0
(dedup 1840/0, C1 227211/227211 signed, lint_symbol_refs wired + passing);
audit-corpus 0 slices; audit-cdecl green. Zero src/config changes this session.

Phase 26 resumes at Task 7 (fresh session).
2026-07-15 00:23:31 -06:00
Drew T e9038a04ea docs(phase-26a): A10 COMPLETE — wall re-test verdict on all 5 walls
The audit thesis is CONFIRMED: the tooling-walls were dissolved by the FIXES and
the payoff banked there (A3f/g/h + A9a/b, 66.5->68.6% instr), while the re-tests
confirm the residual walls are real.
- fuel/closeness-0: CONFIRMED REAL (0/958 bank at fleet scale).
- arity (Phase-15 dead-end): was tooling (A3c order-dependent rule); 13/18 banked.
- def-side loose-typing: was partly tooling; A9b banked func_8017A4AC x134.
- type-heavy: blocking tool build_engine_types was broken (A7 fixed it, now runs);
  the ~1,200-member family harvest is Task-8 integration, not a pure re-test.
- 780 h_seq callee-oracle rejections: was tooling; A3h banked +2,675.
A re-confirmed wall is as valuable as a dissolved one (P9). Next: A11 close.
2026-07-15 00:01:42 -06:00
Drew T b1c58d7668 docs(phase-26a): A10 wave 1 — closeness-0 wall re-test CONFIRMED REAL (0/958 bank)
Re-gated all 958 closeness-0 open-stub backlog drafts through the FIXED gate
across 135 binaries in parallel: banked=0, near=957, failed=71. The closeness-0
backlog is genuine whole-binary near-misses, NOT tooling misses -- match_one's
isolated closeness==0 systematically overstates whole-binary bankability, and the
repaired gate recovers none. P9: a re-confirmed wall is as valuable as a dissolved
one. (The audit's tooling-walls were already banked by A3f/g/h + A9b, +2.1% instr.)

backlog.py: env-gated BACKLOG_NO_RENDER so parallel workers skip the render race
(append is atomic) -- backward-compatible parallel-safety. backlog.md refreshed
with the re-test's whole-binary-informed scores.
2026-07-14 23:54:16 -06:00
Drew T f2b2a9d778 docs(phase-26a): A10 Max-phase — wall re-test scoping + arity measurement (fan-out pending R27)
Scoped the 5 walls. The audit fixes already dissolved the easy wall (A3f/g/h +
A9b, 66.5->68.6% instr). Remaining re-test pool = 2,218 open-stub backlog drafts
(958 at closeness 0). #2 arity: 13/18 Phase-15 arity-conflict fns already banked
(wall largely fell). Confirmed the remaining re-test is breadth (serial gate_stage
timed out). R27 boundary reached: prompting Drew for /effort ultracode before the
parallel re-gate fan-out. Tree clean (HEAD commit:0618 before this).
2026-07-14 22:27:19 -06:00
Drew T ea20bdf9f3 fix(phase-26a): A9g — jr_inventory: retire the ephemeral roster, derive banked from the image (R33)
jr_inventory's `banked` set was filtered by an EPHEMERAL, gitignored
.run/banked_func_*.json roster: a `rm -rf .run` / fresh clone would blind ALL
banked jr at once, cross-address siblings (roster named after the exemplar) were
structurally invisible, and non-leader banked jr were missed. "The purest R33
case in the group" (audit).

FIX (the audit's exact prescription): delete the roster glob + `cand` filter;
`banked` is DERIVED FROM THE IMAGE — a real-C def/define fn is a banked jr iff
family_remap.reloc_targets shows it references a committed .rodata carve offset
(config + image, both durable; cross-address- and non-leader-immune). R32
assertion: every committed carve must resolve to EXACTLY ONE owner or abort (a
stranded/duplicated carve is the §8b func_801734BC incident, never silent).

Also fixed the adjacent finding: the asm_jr scan's func_-fullmatch dropped the
curated-name listCdBuffer jr; now resolved via oss.addr_of(). (The --only path's
own fullmatch is left — it parses user input, not the corpus.)

Perf: read the overlay image ONCE and pass it to reloc_targets(..., data=) — a
new backward-compatible param on family_remap (regression: 0/80 mismatch vs the
re-read path).

Verified: data-param behavior-identical; the R33 win — ov_SC02_000 now finds the
cross-address sibling func_8017FCB0 the roster missed; full-fleet parallel run =
134/134 OK, 0 false aborts, 1336 banked jr == 1336 carves -> 1:1 ownership holds
fleet-wide. Byte-safe: jr_isolate_all is not in the make build/extract path
(R22-neutral); the change makes future isolations strictly more correct.
2026-07-14 22:03:43 -06:00
Drew T 96e025a324 fix(phase-26a): A9f — overlay_src_split swallowed 2 real definitions; the selftest was blind
scan_construct's force_decl latched from the FIRST token and returned at the
first depth-0 `;`, so a definition sharing a physical line with leading externs
(`extern A; extern B; void f(){...}`) was never anchored — absorbed into the
next anchor's preamble. The parser jr_isolate_all rewrites source from was short
two functions in the exemplar overlay. The round-trip selftest is a SERIALISATION
check (a miss lands in a preamble -> round-trip still exact BY CONSTRUCTION), so
it was structurally incapable of seeing this.

FIX (byte-safe): force_decl no longer survives a same-line `;` with trailing
code — re-classify from the remainder and keep scanning so the def anchors (its
leading externs stay in its whole-line item text -> round-trip byte-identical).
Rejected the audit's "split into 3 constructs": round-trip joins whole-line
chunks with `\n`, so sub-line splitting would insert a newline where a space was.
def_name now names the LAST top-level header before `{` (the definition, not the
first same-line extern; byte-identical on every single-def construct).

R32: hidden_definitions() coverage oracle wired into selftest — an independent
detector of `func_XXXX(...){` bodies not anchored. The selftest is now a coverage
check, not just serialisation.

Verified: 2 swallowed -> 0; regression over 1738 overlay .c = 0 round-trip fails,
0 non-monotonic, 0 non-additive changes, +2 anchored defs. Byte-safe: tool not in
the build path (R22-neutral); ov_SC01_077 rebuilds d19c9580; neither def straddles
a committed subseg boundary. Audit ledger line refs were stale (src rewritten);
real cases are ov_SC01_077_after.c:2020 + ov_SC01_077_jr_8015444C.c:1495.
2026-07-14 21:43:52 -06:00
Drew T 68d29ba8af fix(phase-26a): A9e — reconcile_tu already wired into bank_exemplar (A3d); document the ladder
NULL RESULT (P9/R14): the session-12 "wire reconcile_tu into bank_exemplar"
handoff item was stale — A3d (commit:0601) already wired reconcile_tu into
jtbl_family_bank.recover() ("on BOTH banking paths"), and bank_exemplar's
`recovered` stage delegates to fb.recover = cast_call_sites + reconcile_tu.
Proven working by A9b (func_8017A4AC banked at the recovered stage, reconcile_tu
resolving its struct + fn-ptr conflicts). No live tool references the RETIRED
reconcile_decls (only docstrings + the audit-cdecl differential harness).

No code change warranted. Byte-neutral hardening only: document the
raw/scoped/recovered/reconciled fallback-ladder composition inline in
bank_exemplar so a future session does not re-run this "is it wired?" trace.
2026-07-14 21:25:19 -06:00
Drew T 40477281ce fix(phase-26a): A9d — retire the dead Phase-17 canonical-sig chain (R33)
DELETE tools/census_conflict_callees.py + tools/derive_canonical_sigs.py.

- census_conflict_callees: audit-CONFIRMED marked-for-deletion (commit:0593;
  decision-log 836). It re-derives from C text the per-TU "defined/declared/
  stubbed/external?" question that reconcile_tu (Phase 26) answers FROM THE
  BUILD — and does it WRONG in the unsafe direction (unknown -> conflict-free).
- derive_canonical_sigs (census's ONLY consumer): genuinely dead — last touched
  Phase-17 (commit:0140), output .run/canonical_sigs.json read by nothing (no
  Makefile/workflow/import), no-ops on the 2-byte [] input, asm-arity heuristic
  36% wrong vs byte-exact banked C. Its purpose was retired in A3d
  (fleet-majority oracle -> reconcile_tu's per-TU oracle). Deleting census
  orphans it, so the whole dead chain ceases to exist (R33: the best outcome is
  a DELETED SCANNER, not a fixed regex).

Byte-neutral by construction (neither tool is in any build/report path):
module-import smoke over the 13 importable harvest/bank/report/reconcile tools
= all clean; bank_exemplar is a run-only script (indexes sys.argv at module
scope), imports neither deleted module. No src/config change -> no byte moves.

Doc-pointer hygiene: hand-matching-process.md 8a, matching-cookbook.md
(canonical-sig-layer entry), tooling-audit.md (ledger row + derive entry) all
annotated DELETED/historical so nothing points at a nonexistent tool.
2026-07-14 21:21:02 -06:00
Drew T 181191b6af docs(phase-26a): log A9c (lint_symbol_refs green + wired into make report) 2026-07-14 20:46:53 -06:00
Drew T 49004f5b06 docs(phase-26a): log A9a (canon_sig_reconcile fn-ptr fix) + A9b (func_8017A4AC ×134 wall re-test) 2026-07-14 20:39:23 -06:00
Drew T 959cc9b7c1 docs(phase-26a): log A3h — standing-lead harvest (measured, banked +1.3% instr, Bucket X honest-stop) 2026-07-14 19:04:13 -06:00
Drew T 7a3921a3f3 docs(phase-26a): log A3f+A3g (harvest + propagation) + disk-hygiene note 2026-07-14 16:48:23 -06:00
Drew T b89fcc2edc fix(phase-26a): A3e — gate_stage pinned the byte-gate back to 4.9%, OF A3'S OWN FIX
THE WORST DEFECT IN THE AUDIT IS NOT IN A SCANNER. It is one default argument in the CALLER of a
scanner we had already fixed.

    # tools/gate_stage.py:315
    summary = run_gate(a.drafts, binary=b, src=a.src or f"src/{b}/{b}.c", ...)   # ALWAYS the main .c

`src` RESTRICTS the byte-gate to ONE translation unit, and _gate1 does `if src: cmd += ["--src", src]`
-- always truthy. A3 had just taught harvest_verify to DERIVE each draft's home TU *when --src is
omitted*, lifting the byte-gate's reach from 4.9% to 100%. gate_stage NEVER OMITS IT. The fix was
neutralised by its own caller's default, and the PRIMARY BANKING PATH -- every wave, the grinder, the
orchestrator, bulk_harvest -- remained structurally unable to bank 250 of ov_SC01_077's 263 stubs.

WHY IT SURVIVED 26 PHASES: harvest_verify cannot splice a draft whose stub is not in the TU it was
pointed at, so the draft never verifies -- and is then logged as near/failed, i.e. AS A MATCHING
PROBLEM. The wave reports a poor close-rate; the function goes to the backlog as a compiler residual.

    A tool that CANNOT bank a function is indistinguishable, in every log this project keeps,
    from a function that CANNOT BE banked.

PROOF, same draft / same gate / same second: gate_stage rejected func_80129C40; harvest_verify run
directly (no --src) VERIFIED it byte-identical and banked it.

AND A COUNTING BUG THAT HID THE HIDING (gate_stage:261): when match_one says MATCH but the whole-binary
gate rejects, the record is logged status="near" and THE COUNTER IS NEVER INCREMENTED. A 63-draft run
printed `banked 0, near 0, failed 0` -- three zeros that do not sum to 63 -- for phases. Nobody ever
added them up. (The number was not wrong. It was ABSENT.)

ALSO FIXED, sig_unify (the same disease, one level down): it SILENTLY DROPPED 190 of 196 drafts (97%).
`cur_stubs` was read from the main .c (13 of 263 stubs), so any draft whose stub lives in a _jr_ carve
hit `if fn not in cur_stubs: continue` -- dropped BEFORE THE WRITE: never copied to --out, never gated,
never logged, while the summary printed "drafts unified: 6" and read like success. THIS IS GATE_STAGE'S
STAGE-2 RECOVERY -- the pass whose whole job is to rescue the stage-1 failures -- and it has been a
no-op for nearly every draft it was meant to save. Now: TU derived per draft (corpus.stubs), canon
derived from cdecl.tu_scope (cpp -- macro-injected decls finally visible), and _keep() so an
already-acceptable decl is left alone (the §19 "sig_unify regresses canonical drafts" failure mode).
Reach: 6 -> 196 drafts; callee-externs rewritten 2 -> 90; own def-sig 2 -> 86.

MEASURED, all three consumers migrated (196 never-banked drafts):
    near   5 -> 116        failed  190 -> 17
=> 173 of 190 "failures" were PLUMBING, not codegen: now compiling and SCORED instead of invisible.

THE PRIZE (measured, not claimed): the backlog holds 1,588 entries at closeness==0 -- body byte-exact
per match_one, whole-binary gate rejected. 1,215 have been banked since by other paths. 373 ARE STILL
OPEN STUBS WHOSE BODIES ARE ALREADY BYTE-EXACT, sitting in a ledger that calls them unrecoverable.

⚠ THE HARVEST ITSELF IS NOT IN THIS COMMIT, AND IS NOT CLAIMED (P9). Gating the 63 ov_SC01_077 ones
dragged `dedup_propagate --auto-from --recover` behind it; it ran >1h and hit its timeout -- its
first-ever run over the FULL corpus (A6/A7 unblocked the 407 files it could never see). It MUTATES THE
TREE BEFORE IT GATES, so the kill left 859 files + engine_core.h (+544 lines) written and UN-GATED with
the registry never updated. R22 on that tree: 44 passed / 92 FAILED -> `git checkout -- src/ config/`,
fleet restored to 136/136. Nothing lost (H4: the tree was clean, so the revert was one command).
Two real lessons, recorded: dedup_propagate is NOT crash-safe and must never run under a timeout it can
hit; and a 63-draft experiment must not drag an unbounded fleet-wide propagation behind it.

  R22 clean-fleet after revert: 136 passed, 0 failed of 136.  src/ and config/ clean.
  cookbook §51g LAW 11: A FIX IS NOT LANDED UNTIL ITS CALLER STOPS OVERRIDING IT. After fixing a
  scanner, grep every call site and ask whether a caller's default re-disables it. An audit that stops
  at the callee is half an audit.
2026-07-14 15:39:07 -06:00
Drew T 4aae5e7589 fix(phase-26a): A3d — retire the fleet-majority oracle: it was WRONG for the TU 16% of the time, on both banking paths
R33 applied to the worst finding in the audit: this oracle was not fixed, it was RETIRED.

    reconcile_decls asks "what does the FLEET call this symbol?"
    C asks           "what does THIS TRANSLATION UNIT declare?"

The engine is loosely typed -- the same address is legitimately declared with incompatible types in
different overlays -- so a single fleet-wide answer is WRONG FOR SOME TU BY CONSTRUCTION. And it is
worse than a silent skip: it writes an ACTIVELY WRONG declaration into the draft, which then
collides with the very TU it was meant to conform to.

MEASURED across ov_SC01_077's 12 TUs, against what cpp says each TU really declares:

    the fleet oracle AGREES with the TU ................ 2883
    the fleet oracle CONFLICTS with it (cc1 REJECTS)  ..  548    <- 16%
    the TU declares it, the oracle has NO answer ......   357

and it was LIVE ON BOTH BANKING PATHS:
  * gate_stage      -- rewrote 60 of 196 drafts in the current batch
  * jtbl_family_bank -- EVERY SIBLING of the ×134 family sweep, the project's economic engine.
    A poisoned decl means that sibling silently does not bank, and the loss is invisible: the sweep
    simply reports a smaller number. The irony is exact -- that function's own docstring already
    knew the conflicting symbols are PER-OVERLAY, which is precisely why a FLEET oracle could never
    have been right.

reconcile_tu.py (written in Phase 26 but NEVER WIRED) now supersedes it, rebuilt on cdecl:
  * ask cpp what the TU declares (macro-injected DEFINE_func_* externs included -- a raw scan
    cannot see them, §8c / §51g LAW 7);
  * ask cc1 whether the draft's decl can coexist (cdecl.compatible, validated against the real
    gcc-2.7.2 front end on 1,485 live pairs -- NOT the C standard, NOT modern gcc; §51g LAW 9);
  * NOT declared -> leave the draft alone (its extern types are load-bearing: %lo-folding, access
    width, alignment); compatible -> nothing; CONFLICTING -> the TU wins + cast at every USE so the
    draft's intended access survives byte-for-byte;
  * derives WHICH TU from corpus.stubs() rather than a hand-passed --src-file (§51g LAW 10).
  * handles the fn-ptr kind NATIVELY -- which is why it supersedes rather than patches: teaching
    reconcile_decls' parser to see `extern void (*D_x[])(void);` would have ARMED its fn-ptr-blind
    data_access_subs to rewrite a call-through `D_x[i]()` into `((u8 *)D_x)[i]()`. Fixing the regex
    would have detonated a dormant bug.

AND THE NULL RESULT, AGAIN, REPORTED AS SUCH (P9/R14): on the 196 never-banked historical drafts the
new oracle banks EXACTLY AS MANY AS THE OLD ONE -- zero. That tail fails on CODEGEN, not on decl
plumbing. The two disagree on 45 of 196 drafts and the outcome does not move. This is a CORRECTNESS
fix (548 wrong declarations removed from two live pipelines, protecting all FUTURE drafts and every
future family sweep), not a banking win, and it is not being sold as one. Three nulls in one session.

reconcile_decls.py is kept as EVIDENCE, marked RETIRED, with no live caller.

  R22 clean-fleet: make clean + extract-all + check-all -> 136 passed, 0 failed of 136
  src/ untouched (0 changes)   reconcile_tu: 0 coverage defects over 196 drafts
  NOTE: the family-sweep path gets its real exercise at Task 8 -- watch the per-sibling bank rate.
2026-07-14 12:53:14 -06:00
Drew T f4502f11bb fix(phase-26a): A3c — the recovery passes were reconciling 95% of drafts against the WRONG TU
FIRST CONSUMER MIGRATION onto the cdecl oracle — and the compiler taught me two things I had
wrong, one of which reopens a wall that has been closed since Phase 15.

1. cdecl.compatible() — "will cc1 accept these two declarations of one name?"
   The predicate four tools each half-implement and get wrong: norm_sig / _norm_type collapse the
   int family to ONE token, so a SIGNEDNESS change reads as "already compatible" and gets no
   rewrite -- while cc1 REJECTS that redeclaration. Right about codegen, wrong about the front end,
   which never reaches codegen.

2. THE ADJUDICATOR MUST BE THE COMPILER THAT COMPILES YOUR CODE (cookbook §51g LAW 9).
   I wrote the rules from the C standard, then let a compiler judge. It contradicted me -- and then
   the RIGHT compiler contradicted the first one. Three different answers:

       declarations in one TU        | standard | modern gcc | gcc-2.7.2 cc1
       typedef int X;  twice         | error    | ACCEPTS    | ERROR
       extern u16 X; + volatile u16 X| error    | error      | ACCEPTS
       void X(s16);  then  void X(); | error    | error      | ACCEPTS
       void X();     then  void X(s16)| error   | error      | ERROR

   --compat now adjudicates with tools/bin/gcc-2.7.2-psx/cc1, the front end that actually
   arbitrates the build: 1,485/1,485 live corpus pairs agree, 0 disagree, 0 skipped.

3. THE PRIZE: the Phase-15 narrow-param wall rests on a false premise.
   The no-prototype rule is ORDER-DEPENDENT. `void X(s16); void X();` COMPILES; only the reverse
   fails. Phase 15 closed "the 159 arity/narrow-param conflicts" as "no clean deterministic fix --
   it is simply C's default-promotion rule". cc1 does not enforce that rule in the direction the
   wall assumed. Four three-line probes, 90 seconds, zero tokens. -> A10 RE-TEST TARGET.
   Probe the compiler for FACTS; read its source only for LEVERS; byte-validate both. (We read
   gcc-papermario for five phases believing it was 2.7.2. It was 2.8.1.)

4. THE MIGRATION: cast_call_sites canonicalized 95.1% of drafts against a TU that would never
   compile them. `--src-file` is an OPTIONAL HAND-PASSED flag defaulting to src/<ov>/<ov>.c, and no
   caller knows about the Phase-26 _jr_<ADDR> carves: ov_SC01_077 has 263 open stubs across 12 TUs
   and only 13 are in the main .c -- while harvest_verify (A3) correctly splices into the real one.
   Now DERIVED from corpus.stubs() (the INCLUDE_ASM line is self-describing), with the canonical map
   derived from cdecl.tu_scope() (cpp -- so macro-injected DEFINE_func_* decls are finally visible).
   Callee-conflict repair reach: 8 -> 58 of 196 drafts (7x).

5. AND THE NULL RESULT, REPORTED AS SUCH (P9/R14). Those 58 banked ZERO functions. The historical
   draft tail fails on CODEGEN, not plumbing -- func_801387B8, which the audit blames on a single
   unparsed `[4]`, is really 67/100 instructions off with a $s0/$s1 swap (that claim does not
   reproduce on today's tree). The real gain is narrower and still worth having: 52 drafts moved
   from "won't compile" to "compiles, N instructions off" -- from an INVISIBLE failure that reads as
   a compiler wall into a SCORED near-miss the permuter and the §47/§48 dials can act on. That is
   the audit's thesis, not a bank. THREE times in one session a confirmed mechanism produced a null
   consequence.

Also: my own new audit printed "ALL ORACLES GREEN" while silently skipping 100% of its corpus (a
missing -Isrc). The exact bug class, in the tool written to hunt it. An unadjudicable check is not
a passed check.

  R22 clean-fleet: make clean + extract-all + check-all -> 136 passed, 0 failed of 136
  src/ untouched (0 changes)   make audit-cdecl: green   --compat: 1485/1485
  NEXT: sig_unify + reconcile_decls carry the SAME wrong-TU bug (same --src-file flag).
2026-07-14 12:26:01 -06:00
Drew T f9742cf9c0 feat(phase-26a): A3b — cdecl.py, THE C-declaration oracle: one grammar, fifteen deleted models
Fifteen tools each carried their own regex model of "what is a C declaration", and they
disagreed — two tools in ONE pipeline disagree today about whether `extern s32 D_a, D_b;`
is a declaration at all. All fifteen shared one character class,
    extern\s+([A-Za-z_][\w\s\*]*?\bD_[0-9A-Fa-f]+\s*(?:\[\s*\])?)\s*;
which cannot hold '(', ',', or a non-empty [N] — so three whole shapes were invisible to
every one of them: fn-ptr/jump-table arrays, sized arrays (one unparsed `[4]` has blocked
func_801387B8 in 134 TUs), and multi-declarators (the WHOLE line dropped, not just #2..N).

REJECTED the audit's own prescription (a shape-aware alternation per tool, ~15 coordinated
regex edits) on R33 grounds: fifteen hand-maintained models are exactly what diverged, and
an alternation only ever covers the shapes somebody remembered. The thing being scanned HAS
A GRAMMAR. C's declarator grammar is small, closed and TOTAL — it describes fn-ptr arrays,
sized/2-D arrays, multi-declarators, fn-ptr params and K&R identifier-lists without being
told they exist. ~250 lines of recursive descent: LESS code than the regexes it deletes, and
exhaustive by construction rather than by memory. (decision-log 2026-07-14.)

Two statement paths, because the inputs genuinely differ:
  * tu_statements()    - a TU's file scope, derived from cpp. A decl inside a DEFINE_func_*
                         macro body declares NOTHING until the macro is invoked (the §8c law);
                         a raw scan is wrong in both directions. cpp answers it exactly, in
                         54 ms/TU (~20 s for the fleet, cacheable).
  * split_statements() - span-preserving raw split, for drafts (which get rewritten).

THREE ORACLES, whole corpus — a measurement, not a belief:
  * coverage      2,952,246 depth-0 statements -> 2,731,521 declarators, 0 PARSER DEFECTS
  * the real gcc  50,405 distinct declarations compiled beside this parser's reconstruction
                  of each one -> 0 REJECTED
  * differential  0 file-scope symbols the incumbents see that cdecl misses; 26 in
                  engine_core.h they cannot see; 6 they wrongly promote from BLOCK scope

Two ideas worth keeping (cookbook §51g, LAWS 4-8):
  * THE CANDIDATE SET IS DERIVED TOO (R33 applied to R32). At file scope C admits nothing but
    declarations, so R32's over-approximating detector is *every depth-0 statement* — supplied
    by the grammar, with no hand-maintained candidate regex to rot.
  * GCC ADJUDICATES MY OWN COVERAGE GAP. Deciding for myself which failures "don't count" is
    grading my own homework — the habit that wrote the fifteen bugs. A statement gcc ALSO
    rejects is not C (my rejection is correct, the INPUT is corrupt); one gcc ACCEPTS and I do
    not is MY defect. All 33 residual: NOT-C, all dead .run/drafts* scratch, none in src/.

NEW findings (docs/tooling-audit.md):
  * reconcile_decls.DATA_DECL_LINE_RE finds ZERO decls in engine_core.h — it is line-anchored
    and every decl there ends in a '\'. Its "authoritative tier" has ALWAYS been empty.
  * gen_harvest_targets + sig_unify count BLOCK-SCOPE externs (6, byte-proven inside a macro's
    function body) as file-scope canonicals — the §8d `conflicting types` confusion.
  * tu_ambient's func regex ([^()]* params) drops ANY callee with a fn-ptr parameter.
  * R14 near-miss: 33 drafts contain `extern if ((func_80029178(0x119) & 0xFF) != 0);`, written
    by a RECOVERY TOOL — but the source bug was already fixed in Phase 19 (0 garbage / 300 sigs
    today). Mechanism confirmed, consequence nil. Note what it cost while live: a draft that
    cannot compile fails the byte-gate and reads downstream as an INTRINSIC COMPILER WALL.

Bugs the oracles caught in ME (and would otherwise have shipped): `extern s32 (*D_801274D0)(s32);`
parsed the BASE TYPE as the name; a K&R declaration-list flushes as SEVERAL spans, so the body
attached to the wrong one and leaked the K&R parameter names into file scope as fake globals.

SCOPE, deliberate: NO consumer is migrated here, so this cannot move a byte. The audit warns
that making the parser see more ARMS dormant transforms (reconcile_decls.data_access_subs would
mangle `D_1[i]()` -> `((u8 *)D_1)[i]()` the moment fn-ptr decls become visible to it). Migration
is one tool at a time, each byte-gated.

  R22 clean-fleet: make clean + extract-all + check-all -> 136 passed, 0 failed of 136
  make audit-corpus: 0 PHANTOM + 0 TRUNCATED    make audit-cdecl: ALL ORACLES GREEN (new gate)
2026-07-14 11:37:13 -06:00
Drew T c7772bc452 docs(phase-26a): SESSION-9 CLOSE — cookbook §51 (the tooling-integrity laws) + handoff
R30/R16: the context-dependent artifacts, written while the context is live.

cookbook §51 — the SILENT SKIP: the bug class, why the byte-gate cannot see it, the
over-approximating-detector method, and FOUR LAWS:
  1. Derive, don't re-derive — the best outcome is a DELETED SCANNER (28 findings -> one
     derived oracle + ~10 deleted scanners). A derived fact cannot rot; a hand-maintained
     copy of it is a liability that grows with every structural change.
  2. Assert your COVERAGE, not merely your correctness. *** A LOUD FAILURE THAT NOBODY
     COUNTS IS EXACTLY AS INVISIBLE AS A SILENT ONE *** — build_engine_types printed
     '[overlap] handle manually' every single time for four phases while dead on 81% of its
     own corpus. This CORRECTS the first draft of R32 ('fail loud'), which was not enough.
  3. When an oracle is structurally blind to a class of error, add a SECOND ORACLE THAT CAN
     DISAGREE WITH IT — not a better assertion inside it. We had two all along and never made
     them argue. (And scope the comparison to where the second oracle is genuinely independent:
     the same check run outside its domain reports 914 slices when the truth is 193.)
  4. A rule that needs a human to remember it is not a gate. Make it structural.
  + the FALSE-WALL PIPELINE (a silent skip -> a wasted draft -> a backlog 'matching failure'
    -> reserved_walls() PERMANENTLY blacklists a function that was never attempted), and a
    checklist for any new corpus-scanning tool.

CURRENT_PHASE: session-9 handoff — what is done, what remains (each with its spec on disk),
and the R32-corrected / R33 / R34-new rule candidates for P10 ratification.
2026-07-14 10:40:04 -06:00
Drew T a2a507d079 docs(phase-26a): A4/A5 + the stale-object false-pass hole recorded 2026-07-14 10:12:59 -06:00
Drew T 3be734bde8 docs(phase-26a): A3 progress — the oracle is built; 4 of ~10 scanners deleted
corpus.py + build_fuel_manifest + wave_targets + harvest_verify/gate_stage landed.
Remaining: family_manifest/family_hseq (62% of the endgame plan is phantom targets),
DELETE census_conflict_callees (R33), exemplar_miner/difficulty, jr_isolate_all.
2026-07-14 09:30:38 -06:00
Drew T ffb6f1a40f docs(phase-26a): A2 — the full audit; 28 findings survive; the endgame plan was majority-fiction
38 agents / 2.24M tok / 0 err. 32 findings raised -> 28 SURVIVED adversarial verification
(4 REFUTED, 16 downgraded). 40 scanners measured CLEAN. Full write-up: docs/tooling-audit.md ROUND 2.

THE ROOT CAUSE — one bug, ~10 times: a hand-maintained model of the corpus layout (a file
allowlist, a single-.c assumption, a func_-only regex, a REGION_SUB dict) sitting on top of a
filesystem that already answers the question. Every TU split silently widened it.
DECAY PROVEN: .run/fuel_manifest.json (Jul 8) recorded 130 stubs; the same tool today returns 30.
The Phase-26 splits moved ~100 stubs out from under a dict literal last edited in Phase 22 — and
nobody noticed, because an un-nominated target produces SILENCE, not an error.

MEASURED: 91.6% of ALL remaining project gain is invisible to target selection (true 994,633 ins;
the manifest sees 83,305). 117 of 127 reach-134 fns never nominated. harvest_verify cannot see
56,742 of 58,717 (96.6%) open stubs. wave_targets hands 78 of 87 targets a nonexistent asm path.

THREE RESULTS OVERTURN SETTLED CONCLUSIONS:
 1. Phase-22's 'the permuter's fuel is exhausted' is UNSAFE. grinder banks through harvest_verify,
    which sees ONE TU — 1,290 of its own 1,298 queued fns live in another. 99% could never have
    banked. '0 banks since Phase 21' is equally consistent with 'the tool could not bank'.
 2. The Phase-25/26 endgame plan is MAJORITY-FICTION. family-manifest.md advertises 2,758
    multi-member families / 11.0 MB; 1,071 of them / 6.80 MB (62% of the byte-weight) are ALREADY
    FULLY MATCHED. The ranking — the file's whole purpose — is sorted mostly on dead work.
 3. A CORPUS defect the byte-gate is structurally blind to: symbols.us.txt:981 puts a main-EXE DATA
    symbol (listCdBuffer = 0x80180000) into every overlay's symbol stack, but in overlay space that
    address is CODE. splat cuts 97 real functions in half and invents 96 phantom ones = 193 slices
    NOBODY CAN EVER MATCH, in 97 of 134 overlays — and the build stays byte-identical and green,
    because the .s halves are pasted back verbatim. A perfect correctness oracle, a null coverage
    oracle. What saved us: sig_image was RIGHT (58,524/58,621 vs spimdisasm; correct on all 97
    disagreements). A SECOND INDEPENDENT ORACLE is the only reason it was visible at all.

FIX RESTRUCTURED around the root cause: ONE derived corpus oracle (A3) + ~10 DELETED scanners —
not ten fixed regexes. Plus the listCdBuffer corpus fix (A4) and the closeness oracle (A5, which
lies on 155 functions, feeding false walls into reserved_walls()).

decision-log (R31): the why, and the design lesson — a derived fact cannot rot; a hand-maintained
copy of it is a liability that grows with every structural change. We had no instrument that could
report ABSENCE: every gate we owned answered 'is this right?', none answered 'is this all?'
2026-07-14 03:41:43 -06:00
Drew T 92a695ef93 docs(phase-26a): ORDER CORRECTED — the full 18-tool audit runs BEFORE the fix campaign
Drew, mid-session: 'I thought the last session said there were some 15 tools we need to audit.'
He was right, and my ordering was wrong.

I had put the 18-tool audit near the END (as A9). docs/tooling-audit.md prescribes the opposite:
dedup_integrate -> jtbl_family_bank -> the SELECTION tools -> masked_diff/match_one -> THEN the
40 measured findings. The reason is the one that matters:

  A hole in a SELECTION tool makes work invisible to PLANNING — the worst kind, because you
  never know to look.

Fixing on top of unaudited selection tooling means re-running every fix when the audit later
finds the hole. So: A2 is now the full audit; A3-A9 (the fix campaign) are blocked on it.

Tool coverage, stated plainly: A1 (1) + A2 (18) + the fix campaign (~17 already-measured) = ~36
tools — not 82. The filter, from the audit doc: does it PARSE something, and does it GATE or
SELECT work? The remaining ~46 are dead LLM-tier scripts.

A1's result recorded in-file (the three false greens, the causal chain, the null-result blast
radius that confirms R33).
2026-07-14 02:51:59 -06:00
Drew T 978ef703ac docs(phase-26a): A0 — the tooling-integrity audit, as an INSERTED HALF-PHASE (Drew's call)
- Drew (2026-07-14, gate 1): run the audit inside Phase 26, then resume at Task 7.
  Declined the alternative (close Phase 26 early on an unmet milestone -> Phase 27):
  the audit is a PREREQUISITE to structural completion, not a successor to it — the
  tooling that MEASURES the milestone is the thing at fault. Phase-3.5 precedent.
- CURRENT_PHASE.md: the Phase 26-A block (A0-A11), built FROM docs/tooling-audit.md
  (40 measured findings), R33-before-R32 ordering — the best outcome is a DELETED
  scanner, not a fixed regex.
- decision-log (R31): the why, the structural blind spot (a scanner extracts N, the
  true count is M > N, and nobody ever compared N to M — the byte-gate is a perfect
  CORRECTNESS oracle and a NULL COVERAGE oracle), and A1's first finding.
- harness task list built (R28).
2026-07-14 02:39:38 -06:00
Drew T b757de73a8 docs(phase-26): checkpoint now POINTS AT docs/tooling-audit.md as the audit phase's input document
The 40 measured findings were living only in an ephemeral workflow journal outside the repo; the
checkpoint carried my SUMMARY of the audit, not the audit. Now the fresh session is routed to the
evidence, with the priority order (dedup_integrate FIRST — it can print a false green), the R33-before-R32
method (the best outcome is a DELETED scanner), and the real prize: re-test the walls that were diagnosed
on top of the broken 10% callee oracle (the def-side loose-typing wall, the 159 arity conflicts, the
type-heavy tail).
2026-07-14 02:26:14 -06:00
Drew T 6e8c459ade docs(phase-26): SESSION-8 CLOSE — checkpoint for a fresh session; the tooling-integrity audit gates what comes next
RESULTS. Fleet instr-weighted 63.0 -> 66.5%, distinct-code 39.1 -> 46.8%, fn-count 82.61%.
FINAL R22: make clean + extract-all + check-all -> 136/136 BYTE-IDENTICAL, 0 coverage defects.
dedup 1813/0. 0 NON_MATCHING (G4). 31 commits.

13 CORES CRACKED incl. the four heaviest functions in the game (952/890/562/536 ins). The 12-agent
Ultracode wave returned 11/12 first-pass MATCH, each adversarially verified (a skeptic re-ran match_one
+ the §8a jump-table check). Banked x134 this session: func_8015AE2C, func_80178D40, func_8015A3C8,
func_8013FFD8, func_8016AB6C, func_8015444C, func_801380E0 (+ func_8017BEBC x1).

THE TOOLKIT CROSSED A LINE — three ZERO-BYTE DIALS now cover the three passes that produce essentially
every "irreducible" residual, each with a diagnostic signature a cheap agent recognises on sight:
  registers rotated            -> global.c allocno priority -> §47 slider / §48-A pricing dials
  two insns swapped, SAME regs -> sched.c rank_for_schedule LUID tiebreak -> §49 LUID dial
  structure right, count wrong -> loop peel / cross-jump -> §46 / §48-D
That is why 9/12 fell first-pass to ORDINARY agents. Fable5 DISCOVERS a class; everyone else APPLIES it.
New: §46 §47 §48(+A4) §49 §50. Read §50-B before using §48-A1/A4 — it BOUNDS them (the "cross_jump
refunds the bytes" claim is FALSE for a 1-insn tail reached by two jumps; jump.c:1993 minimum=2).

DREW'S DIRECTIVE (binding): the TOOLING-INTEGRITY AUDIT comes BEFORE any further matching work, and is
NOT part of this phase. First act of the fresh session is a Tier-1 phase-boundary call (close Phase 26
early, or run the audit as an inserted phase — Drew decides).

WHY: seven silent-skip tool bugs in one session, and they are a STRUCTURAL blind spot — a scanner
extracts N items, the truth is M > N, and nobody ever compared N to M. The byte-gate is a perfect
CORRECTNESS oracle and a NULL COVERAGE oracle: it has been green since Phase 5 at 0% decompiled (
INCLUDE_ASM pastes the ORIGINAL asm), so a green gate is compatible with ANY decomp %. One 10% hole in
the callee oracle made NINE byte-exact functions look like an intrinsic compiler wall. The real question
the audit answers: how many walls we have already "byte-proven" across 26 phases were lookup misses
wearing a wall's clothes? (The def-side loose-typing wall, the 159 arity conflicts, the type-heavy tail
were ALL diagnosed on top of that hole.) Audit scope so far is 19 of 82 tools (23%), by risk — NOT
comprehensive; dedup_integrate.py is unaudited and can print a FALSE GREEN.

RULE CANDIDATES (P10, Drew ratifies at PhaseEnd):
  R32 Coverage assertion — a corpus scanner must assert its own coverage and fail loud on unparsed input.
  R33 Derive, don't re-derive — where a proven invariant answers the question, derive from it. The best
      audit outcome is not a fixed regex; it is a DELETED scanner.

SELF-CORRECTION ON THE RECORD (P9/R14): I told Drew the headline numbers under-reported by ~190k
instructions. WRONG. weighted_metrics() never calls classify(), so it was structurally immune; the
published numbers were correct all along. I verified the DEFECT but not its BLAST RADIUS. A null result
against a strong prediction is a refutation — chase it.
2026-07-14 02:21:08 -06:00
Drew T 91fd1257ba docs(phase-26): session-8 checkpoint — 3 heaviest cores cracked, 12-core wave 9/12 MATCH, 6 silent-skip bugs fixed 2026-07-14 00:13:50 -06:00
Drew T 1ab9905368 feat(phase-26): §8d scope_data_externs — the ×133 sweep blocker fixed; func_8015AE2C banked ×134
- ROOT CAUSE (R14 — the session-7 diagnosis was half right): the isolated region builds [ OK ]
  WITHOUT the body, so §8b isolation was never implicated. `family_remap.gather_externs` prepends
  carried decls at FILE scope; D_801812A4 is a fn-ptr dispatch table the sibling declares FOUR
  incompatible ways at BLOCK scope inside its own later functions, so the carried file-scope decl
  ESTABLISHES A GLOBAL THE TU NEVER HAD and every later block-scope extern must now agree with it.
  Byte-proven asymmetry: BLOCK(int)->BLOCK(struct*)->FILE(void*) builds; FILE(void*)->BLOCK(int)
  errors. It was the ONLY hard error in the build — all 27 carried function externs were fine raw.

- THE FIX (demote, don't reconcile): tools/scope_data_externs.py emits a carried D_ extern at BLOCK
  scope inside the function body when the TU has no file-scope decl of it above the insertion point.
  Byte-neutral (an extern emits no code; type + access opcodes unchanged) and never worse than raw,
  so it needs no oracle, no type comparator, no fn-ptr parser. Restores fidelity — the original
  declares these symbols at block scope in exactly this way. Wired into jtbl_family_bank as the
  `scoped` stage: raw -> scoped -> recovered -> reconciled (scoped is the base for the later stages).

- reconcile_decls is the WRONG instrument for this class, twice: its oracle answers "what does the
  FLEET call this symbol" when the question is "what can THIS TU see", and its DATA_DECL_LINE_RE
  cannot parse `extern void (*D_x[])(void *);` — silently skipping the very symbols that were
  failing (the phase's third silent-skip bug, after find_site braces + overlay_files splits).

- R17 TRIAGE RULE, first real test, held: `conflicting types` = the compiler REFUSED TO COMPILE =
  a C front-end diagnostic = our Python. Reading cse.c/global.c would have taught nothing.

- RESULT: func_8015AE2C (562 ins, reach 134) swept 133/133 siblings, 0 failures. R22 clean-fleet
  136/136 BYTE-IDENTICAL (534 changed src files); dedup-check 1813 validated / 0 failed; 0
  NON_MATCHING (G4). instr-weighted 63.0 -> 63.6%; distinct-code 39.1 -> 40.5% (+256 unique fns /
  +79,957 ins) — one core, ~0 agent tokens.

- knowledge captured during the producing session (R30/R31/R21): cookbook §8d, decision-log
  2026-07-13 session 8, SETUP tool-inventory row; CURRENT_PHASE session-8 checkpoint.
2026-07-13 20:45:15 -06:00
Drew T 12631df74a docs(phase-26): R17 triage rule — 'wrong bytes' -> read gcc; 'won't compile' -> read our Python
Drew asked whether the x133 sweep blocker warrants a gcc-2.7.2 source read. It does not,
and the distinction is worth pinning down because it routes every future residual:

- The sweep blocker is a C FRONT-END diagnostic (conflicting types: two incompatible
  file-scope decls of one identifier in one TU). gcc is correctly rejecting plain C89.
  The bug is in reconcile_decls (fleet-majority oracle vs the TU's visible decl).
  Reading cse.c/loop.c/global.c would tell you nothing.
- func_8017BEBC (close=2) is the opposite: it compiles fine and emits the wrong bytes, and
  the cause is localized to global.c's allocno-priority tie. THAT is the R17/§45-B target
  (gdb-on-cc1 read of allocno_live_length) — 2 instructions from a 107K-ins bank.

Rule: 'wrong BYTES' -> read the compiler (R17). 'won't COMPILE' -> read our Python.
cookbook §31-triage + the CURRENT_PHASE NEXT block annotated with the routing.
2026-07-13 19:50:51 -06:00
Drew T 9bbfff9b8c docs(phase-26): session-7 checkpoint — heavy-jr waves run; 1 core banked x1, 2 cracks in hand, sweep blocker diagnosed 2026-07-13 18:35:38 -06:00
Drew T f89fd50afb docs(phase-26): session-6 checkpoint — de-risk complete, 3 gate-cap bugs fixed; §41d 2026-07-13 14:44:07 -06:00
Drew T 38ac5659aa feat(phase-26): §8b scoping wall BROKEN — decl-environment reconstruction + lazy per-core isolation
The full 54-jr isolate-all on ov_SC01_077 now builds d19c9580 BYTE-IDENTICAL
(R22 clean-fleet 136/136) — the configuration session 5 could not build. The
heavy-jr harvest (191 cores / 5.53M templatable ins) is unblocked.

- R14 CORRECTION: session-5's "gcc-2.7.2 block-scope-extern TU-persistence" root
  cause was WRONG. There is no gcc quirk — DEFINE_func_* macros expand at FILE
  scope, so their leading externs are genuine file-scope decls that merely live in
  engine_core.h, invisible to any col-0 .c scan (1377 macros / 3929 lines / 1462 syms).
- REJECTED the approved "global symbol->type map + shadow set" design: the engine is
  loosely typed (func_80173544 is DEFINED `s32 f(void*)` yet declared `extern void
  f(void);` inside func_801734BC's body), so declaring every USED symbol hoists that
  block-scope shadow to file scope and CREATES the conflict a shadow-set then dodges.
  Instead reconstruct the original TU's file-scope decl environment and carry it
  strictly FORWARD — conflict-free by construction (every carried decl already
  coexisted with every definition in the one original TU; compatibility is
  order-symmetric; shadows stay in bodies and travel with their item).
- The byte-gate found two MORE lost decl sources, not predicted: (a) a definition is
  itself a declaration for everything below it in its TU (func_8012B2CC undeclared);
  (b) file-local typedefs used by a carried prototype (parse error, Vec3s). K&R defs
  must render `extern T f();` (unprototyped), never f(void).
- LAZY per-core isolation wired into jtbl_family_bank (Drew's call — upfront-x134 =
  ~7,200 region files): jtbl_carve NON-CONTIGUOUS fail-loud -> jr_isolate_all --only
  <core> -> re-extract -> re-carve. Proven on func_80178D40 (890x134, heaviest core):
  carve blocked -> isolated (byte-neutral d19c9580) -> carve in its own subseg.
- TWO LATENT BUGS fixed (both would have corrupted the heavy sweeps):
  * jtbl_carve.func_subseg derived the owning subseg from the ASM TREE, which `make
    extract` never prunes -> after an isolation it returned the STALE owner and
    silently re-created the very collision the isolation removed. Now config-derived.
  * jtbl_family_bank/jtbl_carve revert() DELETED the shared overlays.mk carve var
    unconditionally -> would destroy a COMMITTED carve (all 134 overlays have one) on
    any failed sibling. Now restored to its committed value; only region files created
    by this attempt are removed; dirty-tree preflight refuses to start a sweep.
- docs: cookbook §8b RESOLVED + new §8c "splitting a TU means rebuilding its
  DECLARATION ENVIRONMENT, not moving text"; decision-log 2026-07-13 (R30/R31).
- parser selftest 404/404; R22 clean-fleet 136/136; 0 NON_MATCHING (G4).
2026-07-13 13:25:39 -06:00
Drew T 234788dfc4 feat(phase-26): §8b overlay-src parser (404/404) + jr isolation tool + the gcc-scoping wall finding
- tools/overlay_src_split.py: overlay-.c-aware partition (header = includes + Phase-17
  canonical-sig layer; per-address items = preamble + body; robust def/decl/K&R/DEFINE_func/
  SETTER/RETCONST classification). Fleet-validated 404/404 overlay .c, 341,902 items —
  round-trip exact / 0 unresolved / 0 non-monotonic. The Stage-2 isolation unblock.
- tools/jr_isolate_all.py: multi-cut jr resegment (config split at jr boundaries, source
  repartition + INCLUDE_ASM path repoint, banked-jr carve repoint, -O0 skip, ambient decl
  carry). SINGLE-cut isolation byte-identical (func_8013FFD8 -> d19c9580, R22).
- FINDING (decision-log 2026-07-13): full 54-jr isolation of the dense _after object hits
  gcc-2.7.2 block-scope-extern TU-persistence (func_801734BC/D_80126B3E declared only in
  engine_core.h DEFINE_func macros); mechanical TU-split breaks it. Fix = declaration-
  completion from a global symbol->type map (Drew-approved next step; lazy per-core).
- baseline intact (ov_SC01_077 rebuilds d19c9580); no config/src/binary change committed.
  CURRENT_PHASE session-5 checkpoint + decision-log R31. db.*.gbf = R23 noise, not staged.
2026-07-13 12:00:26 -06:00
Drew T d1dd29d815 feat(phase-26): §8b multi-jtbl same-subseg — contiguous MERGE (built) + isolation scaffold
Session-4 same-subseg handling (the de-risk preamble's harder half; byte-proof of
the merged build + isolation deferred to Stage 2 with concrete cores):

- jtbl_carve.py: MERGE adjacent same-subseg carves into one spanning .rodata piece
  (a code object emits its jtbls contiguous, so two matched jr-fns in one subseg are
  byte-correct iff their jtbls abut). BOUND-FIX: a new jtbl's end is bounded by the
  next raw dlabel OR the next existing carve start (an already-carved adjacent jtbl
  is gone from the data asm -> raw dlabels over-extend it -> false "non-contiguous").
  Config-proven (func_80171B4C 801D8C48 merges with func_801734BC 801D8C68). NO-OP
  for family-1/cross-subseg (single carve per subseg) -> committed configs unaffected.
- jr_isolate.py (scaffold, NOT yet functional): the non-contiguous case — split a fn
  into its own code subseg (whale _o0b precedent) so its jtbl carves independently.
  BLOCKED on split_src_region, which can't partition the overlay .c (global canonical-
  sig extern layer + per-fn callee-externs + DEFINE_func macros + @class annotations,
  ~922 non-address items). Stage-2 build item (overlay-.c-aware source split).
- cookbook §8b (the --order sandwich + the two same-subseg cases + the blocker);
  CURRENT_PHASE session-4 checkpoint updated with the Stage-2 unblock decision.
2026-07-12 22:06:17 -06:00
Drew T ca50ee6978 feat(phase-26): §8 multi-jtbl --order carve + family-1 (func_801734BC ×134)
- ld_interleave.py --order: address-ordered N-piece data->rodata->data sandwich
  for overlays with 2+ matched jr-functions; legacy --front/--tail path is byte-
  untouched (main EXE + the 133 single-carve func_8012ACE0 siblings unaffected)
- jtbl_carve.py rewritten additive/regenerate-from-config: parse the tail data
  region + existing .rodata carves, split the containing data piece for the new
  jtbl, re-emit the address-ordered pieces + the --order arg; same-subseg carve
  collision fails loud (-> jr isolation); idempotent
- jtbl_family_bank.py: `make extract` BEFORE the carve (asm must match the reverted
  committed config; the old error-string retry was fragile) + revert-on-carve-fail
- family-1: func_801734BC (34-ins PURE jr, ov_SC01_077_after) matched in ov077
  (shared-tail switch idiom) + banked 133/133 siblings = x134 — CROSS-subseg
  multi-jtbl (func_8012ACE0 in _a + func_801734BC in _after)
- R22 clean-fleet 136/136 byte-identical (~52s); 0 NON_MATCHING (G4)
2026-07-12 21:51:27 -06:00
Drew T 79ca6b4975 docs(phase-26): refine session-3 checkpoint — small-jr-first as de-risk preamble, then heavy 191
- Drew's sequencing (agreed): do the 45 small jr families FIRST — not for byte-weight (~+1% instr,
  129K ins) but to de-risk + harden the §8 x134 pipeline before the heavy Fable5 cores bet on it.
- decisive technical reason: jtbl_carve only built the single-jtbl carve; func_8012ACE0 is now
  matched in all 133 siblings, so family #2 forces the multi-jtbl address-ordered `ld_interleave
  --order` carve -> build & prove it on cheap 30-ins targets first. Also needs no Fable5.
- guardrail kept explicit: small tier = MEANS (harden pipeline + build multi-jtbl), NOT the
  objective; the 191 heavy jr families (5.53M ins) remain THE byte-weight target -> pivot after.
- CURRENT_PHASE.md SESSION-3 checkpoint updated to Stage 1 (small + build multi-jtbl) -> Stage 2
  (heavy 191, Fable5 un-paused). decision-log addendum with the forcing-function wiki lesson.
2026-07-12 20:24:21 -06:00
Drew T 4e17a7e77b docs(phase-26): session-3 checkpoint — heavy-byte-weight reframe (§8 unlocked the switch cores)
- decision-log (R31): §8 unblocked the SINGLE heaviest byte-weight chunk of the game — 9 of
  the 10 heaviest unmatched family cores are switch (jr) functions (func_80178D40 890x134 =
  477K ins alone); jr substantial = 191 fams / 5.53M templatable ins. My "45 small jr families"
  recommendation (129K ins) was a light-tail trap — Drew caught it against the endgame plan
  (heaviest-byte-weight-first). Corrected next play: Fable5 crack the heavy jr cores -> §8 x134
  bank -> parallel R22 verify; needs Task 7 (Fable5) un-paused (§8 makes that worth it now).
- CURRENT_PHASE.md: SESSION-3 checkpoint as the fresh-session resume point (4 commits this
  session: tiny-band commit:0531, §8 PoC commit:0532, §8 x134 commit:0533, R22 parallel commit:0534;
  distinct-code 30.3->39.1%, instr-weighted 58.2->63.0%, R22 now ~50s)
2026-07-12 19:20:58 -06:00