feat(phase-17): close — canonical-sig layer built; the wall is the compiler, not sigs; pivot to gcc research (v1.16.0)

- canonical-sig layer (session 4): tools/census_conflict_callees.py + derive_canonical_sigs.py
  -> a 20-extern byte-neutral block atop ov_SC01_077.c (LOCAL, not engine_core.h); census
  conflict callees 20->0, blocked targets 24->0; gate pipeline now draft -> sig_unify (MANDATORY)
  -> harvest_verify --chunk 1; fleet 136/136 byte-identical (R22), 55.51% (no regression)
- FINDING (R14/P9): the conflict wall is 7%-reach not ~2x; the 4 reach-134 circular targets are
  ALL gcc-quirk/regalloc/layout-bound (0 banked); the high-reach core IS the quirk tail; struct
  types are byte-neutral for matching (the wall is gcc codegen, not knowledge)
- leverage analysis: fleet % is function-count-weighted (size adds no %); "unblock many" = the
  layer (declaration, not matching); reach is the lever (already reach-sorted); 247 tractable
  reach-134 stubs ~ +3-4% projected
- GO/NO-GO: NO-GO on brute waves at the current ceiling; GO on a compiler-quirk research phase
  (read gcc-2.7.2 source + Xenogears + the §10/regalloc classes, R17) -> then resume the wave
- docs: cookbook §16 corrected + hand-matching-process.md §8 (the layer + the finding + handoff)
- worklog archived -> phase-ends/logs/Phase17.md (R19); bumps 1.15.0 -> 1.16.0
This commit is contained in:
Drew T
2026-06-20 09:35:18 -06:00
parent f0dd9351e4
commit c4b0cef86a
8 changed files with 595 additions and 5 deletions
+3
View File
@@ -1,4 +1,7 @@
{
"worktree": {
"bgIsolation": "none"
},
"hooks": {
"SessionStart": [
{
+54
View File
@@ -332,3 +332,57 @@ Expected to lift whole-binary 33% → ~60% (toward the match_one ceiling) — **
Calibration: 30 fns / 1.89M tokens / 33% whole-binary / +0.47%. Naive scale to 300 ≈ ~4% fleet, token-heavy.
With the canonical-sig layer (33%→~60%) ≈ ~6-7% fleet at ~2× token efficiency. **Build the layer before the
big wave.** Targets: `.run/harvest_targets_s3.json` (300, relocs≤5, reach-sorted; the top-30 are done).
*(Superseded by §8 — the layer was built and the "~2×" did not hold; the wall is the compiler, not sigs.)*
---
## 8. THE CANONICAL-SIG LAYER — BUILT, and the decisive finding (Phase-17 session 4, 2026-06-19/20)
**The layer is BUILT and validated; the "~2× scaling lever" framing was WRONG; the real wall is the
gcc-quirk tail, so the next lever is understanding gcc-2.7.2 (R17 research, Phase 18), NOT more brute waves.**
### 8a. What was built (committed, byte-neutral, reusable)
- `tools/census_conflict_callees.py` — the accurate conflict predicate: an undeclared-`stub` callee with
`decl_sources = n_callers + is_target >= 2` is a sig-conflict risk (a `declared`/`defined`/`extern` callee
is conflict-free; gen_harvest_targets feeds the one sig). Writes `.run/conflict_callees.json`.
- `tools/derive_canonical_sigs.py` — one byte-neutral canonical sig per conflict callee: **`s32` return**
(void→s32 byte-neutral §3a-1; required where `$v0` is used) + **`s32` params** (matched bodies cast int→ptr,
the demo idiom), **arity** from the Ghidra-C cache AND asm read-before-write `$a0–$a3` (agreed on all 14
cached; the 6 non-cached stubs call-site-validated). Writes `.run/canonical_sigs.json`.
- The **20-extern block at the TOP of `src/ov_SC01_077/ov_SC01_077.c`** ("Phase-17 canonical-sig layer").
**LOCAL on purpose** — engine_core.h is shared by all 134 overlays and a reach-1 name (e.g. func_801809BC,
matched differently in ov_SC03_096) would collide. `gen_harvest_targets` + `sig_unify` both already read
the overlay `.c`, so the layer auto-wires with **no tool change**.
- **Pipeline change (mandatory):** harvest_verify accumulates the baseline from the (now block-carrying)
`.c`, so a raw draft's guessed extern clashes with the block even at `--chunk 1`. The wave gate is now
**draft → `sig_unify` (MANDATORY, normalizes drafts to the file-top canonical) → `harvest_verify --chunk 1`
→ propagate.** Census after the layer: **conflict callees 20→0, blocked targets 24→0**, fleet 136/136 (R22).
### 8b. THE FINDING (R14/P9 — this redirects the whole strategy)
- **The conflict wall is small:** for the remaining 270, only **20 callees / 24 targets / 7% of wave reach**.
The "~2×" was the *top-30's in-flight* conflicts, since dissolved by banking those callees.
- **The high-reach core IS the gcc-quirk tail.** Hand-tried the 4 reach-134 *circular* conflict callees
(the §7c "match callees first" move) — **ALL quirk-bound, 0 banked:** func_8012B4B8 = §10 stack-addr
rematerialize-vs-hoist (gcc caches `&mtx`); func_8012B8E4 = `$s0/$s1` regalloc swap, **structurally perfect
75=75** but the permuter probe stalled at base score (external callee `ratan2`, so T6 doesn't even apply —
it's just not in the permuter's search space); func_8016A8FC / func_80169A4C = local-struct-builders
(stack-layout-bound). Drafts in `.run/drafts-s4/`; permuter scratch `.run/permuter/func_8012B8E4/`.
- **Types are byte-neutral for matching (re-confirmed):** matching reads the access *width* off the asm
instruction (`lh`=s16, `lbu`=u8, `lw`=s32), not off any struct def — so emulator-recovered struct types
help *comprehension*, not the byte-close. The wall is the compiler's regalloc/scheduling, which types and
shared-context do not touch.
### 8c. The leverage analysis (answering "do the fewest largest that unlock the most?")
- The fleet % is **function-count-weighted** (`190,949 / 344,010` functions): **every reach-134 match is
+0.039% regardless of size.** Giants bank more *bytes* but the same %. So "fewest largest" gives no % edge.
- "Unblock many" = the canonical-sig layer (declaration removes sig-friction; it does NOT make callers
*matchable* — matching is independent per function). The highest-fan-in callees (func_8012A828 49 callers,
func_80146CA0 46, …) are already `defined`/`declared`/`extern`. Declaring the top-5 *undeclared* keystones
touches only 28 of 900 stubs. So there is no "magic 5 unlocks hundreds."
- **The real lever is reach (size-independent), which we already reach-sort, + the idiom flywheel.** Of the
400 remaining reach-134 stubs, **247 are the tractable shape** (≤80 ins, ≤4 calls); 80 call-heavy (§10
tail), 28 giants. Projected tractable-247 wave ≈ **+3-4% fleet** at the calibration close-rate.
### 8d. The deferred wave (staged, ready to resume after the compiler research)
`.run/harvest_wave_s4.js` = the layer-aware probe (40 tractable reach-134, sig_unify-before-gate). Resume
after Phase 18 lands new gcc-quirk idioms (which raise the close-rate above 33% and so the wave's yield).
+16 -5
View File
@@ -1026,8 +1026,19 @@ the time; the work is byte-closing + sig reconciliation. **Full process: `docs/h
no /mcp) → parallel draft agents (m2c+Ghidra-C+asm+actor-struct+§3a, self-validate `match_one`) → whole-binary
gate (`harvest_verify --chunk 1`) → `sig_unify` recover → `dedup_propagate --auto-from`. Calibration (top-30):
60% match_one MATCH, **33% whole-binary** (+0.47% fleet), 136/136.
**THE CANONICAL-SIG WALL (the ~2× scaling lever):** the entire match_one→whole-binary gap is SIG CONFLICTS
(parallel agents declare shared callees inconsistently; 100% compile-errors, 0 codegen). Fix = a SURGICAL
per-callee canonical-sig layer (match shared callees before callers / seed `engine_core.h`; NOT a blanket
global header — that breaks loose matches, §15). Build it before the big wave. Targets:
`.run/harvest_targets_s3.json`.
**THE CANONICAL-SIG LAYER (built Phase-17 session-4; NOT the ~2× lever it first looked like).** The
match_one→whole-binary gap on the calibration's *top-30* was SIG CONFLICTS (parallel agents declare shared
callees inconsistently → `conflicting types` in the one-big-TU; 100% compile-errors, 0 codegen). Fix = a
SURGICAL per-callee canonical-sig layer: `tools/census_conflict_callees.py` (the conflict predicate:
undeclared-stub callee with `decl_sources = n_callers + is_target >= 2`) + `tools/derive_canonical_sigs.py`
(byte-neutral `s32 func_X(s32...)`, arity from Ghidra-C + asm read-before-write `$a0-$a3`) → a 20-extern
block at the TOP of `ov_SC01_077.c` (LOCAL, not engine_core.h — reach-1 names differ across overlays).
`gen_harvest_targets` + `sig_unify` auto-read it; **the gate pipeline is now draft → `sig_unify` (MANDATORY)
→ `harvest_verify --chunk 1`** (the accumulating baseline now carries the file-top block, so a raw draft's
guessed extern would clash without sig_unify). **SIZING CORRECTION (R14):** for the *remaining 270*, the
conflict wall is only **20 callees / 24 targets / 7% of wave reach** — the "~2×" was the top-30's in-flight
conflicts, since resolved by banking. **The real wall is the gcc-quirk tail, not sig conflicts** — the 4
highest-reach circular targets are ALL §10-hoist / regalloc / layout-bound (0 closed by hand or permuter).
The layer makes a wave *sig-clean*; it does NOT unlock the quirk tail. **→ The match-% lever is understanding
gcc-2.7.2 (R17 compiler-source research, Phase 18), not more brute waves.** Wave deferred; infra staged
(`.run/harvest_wave_s4.js`, 40 tractable reach-134 targets). See `docs/hand-matching-process.md` §8.
+183
View File
@@ -0,0 +1,183 @@
# PhaseEnd — Phase 17: Raise the harness ceiling, then gate the compute run
**Date:** 2026-06-20 · **Project Version:** 1.16.0 · **Phase Status:** Complete (go/no-go decided — NO-GO on brute waves at the current ceiling; PIVOT to compiler-quirk research) · **Generation:** Gen2 (9th phase)
> Gen2 phase 9 of the arc (8→9→10→11→12→13→15→16→**17**; 14 deferred to Gen3+). Ran across **4 sessions**;
> the granular per-session trail (the 5-avenue ceiling test, the demo, the calibration wave, the canonical-sig
> layer build, the 4-circular hand-match attempts, the keystone/leverage analysis) is preserved on-demand at
> **`phase-ends/logs/Phase17.md`** (R19 — NOT auto-loaded; consult only when researching a mechanism). This
> file is the synthesis. Owner decisions (Drew, 2026-06-19/20): pursue all 5 ceiling avenues → on their
> failure, PIVOT to guided hand-matching → build the canonical-sig layer → on the finding that the wall is the
> compiler, **defer the wave and open a fresh compiler-quirk research phase** (goal = match-%, not comprehension).
## Build Log
**Files created/changed and complete — do not recreate:**
*The canonical-sig layer (session 4 — this PhaseEnd commit lands these):*
- `tools/census_conflict_callees.py` — **new.** Censuses the shared callees that block a parallel wave. The
accurate conflict predicate: an undeclared-`stub` callee with `decl_sources = n_callers + is_target >= 2`
(a `defined`/`declared`/`extern` callee is conflict-free). Writes `.run/conflict_callees.json`.
- `tools/derive_canonical_sigs.py` — **new.** One byte-neutral canonical sig per conflict callee: `s32` return
(void→s32 byte-neutral) + `s32` params (bodies cast int→ptr), arity from the Ghidra-C cache AND asm
read-before-write `$a0–$a3` (agreed on all 14 cached; 6 stubs call-site-validated). `.run/canonical_sigs.json`.
- `src/ov_SC01_077/ov_SC01_077.c` — **the 20-extern "Phase-17 canonical-sig layer" block** at file top.
LOCAL on purpose (engine_core.h is shared by 134 overlays; a reach-1 name like func_801809BC is matched
differently in ov_SC03_096 → would collide). `gen_harvest_targets` + `sig_unify` already read the overlay
`.c` → auto-wired, no tool change. **Byte-neutral** (ov_SC01_077 still `d19c9580…`, fleet 136/136).
- `docs/matching-cookbook.md` **§16** (corrected) + `docs/hand-matching-process.md` **§8** (new) — the layer
is BUILT; the "~2× lever" framing was wrong; the gate pipeline is now **draft → `sig_unify` (MANDATORY) →
`harvest_verify --chunk 1`**; the finding + the leverage analysis + the deferred-wave handoff.
*Earlier this phase (sessions 1–3, already committed `commit:0132`→`commit:0139`):*
- `tools/wall_taxonomy.py` + `docs/wall-taxonomy.md` (T1 census); the T3 K&R sig_unify wins (+0.52%);
`tools/ram_probe.py` + `docs/actor-struct.md` + `.run/actor_*` (T5 actor struct, recovered + live-verified,
byte-neutral for matching); the guided-hand-matching demo + idioms; `docs/hand-matching-process.md` (created,
§1 loop / §2 idioms / §3a the 5-move sig playbook / §7 the calibration wave); `tools/ghidra_scripts/
DecompileFunctions.java` (headless Ghidra-C pre-pass); the calibration Ultracode wave matches (→ `engine_core.h`
+ `config/dedup.us.yaml` + `ov_SC01_077.c`); `.run/harvest_wave_s3.js`, `.run/harvest_targets_s3.json`.
**Local artifacts (gitignored / regenerable):** `.run/ghidra_c/*` (300 cached Ghidra-C), `.run/drafts-s3*`,
`.run/drafts-s4*` (this session's 4-circular drafts — none gated), `.run/permuter/func_8012B8E4/` (probe
scratch, reusable: header+macro.inc → target.o), `.run/harvest_wave_s4.js` + `.run/probe_targets_s4.json`
(the staged 40-target tractable probe), `.run/conflict_callees.json`, `.run/canonical_sigs.json`.
**Tools/packages installed:** None — used the existing Phase-4/6/10/11/12/13 toolchain + venv throughout.
**Verification results (literal):**
- **MILESTONE byte-proof:** `make check-all` → **136 passed, 0 failed of 136** (main `143dbb89…`, resident
`8e17e02f…`, all 134 overlays); ov_SC01_077 clean-rebuilds `d19c9580…` (R22). The canonical-sig layer is
byte-neutral fleet-wide.
- **`make report`:** FLEET byte-identical **190,949 / 344,010 = 55.51%** (was 54.48% at phase start; the
+1.03% banked by the demo + calibration wave across sessions 2–3); 0 NON_MATCHING in any default build (G4);
`dedup-check` validated, 0 failed.
- **Canonical-sig layer:** census **conflict callees 20→0, blocked targets 24→0** after the layer; `sig_unify`
rewrites a wrong draft extern to the file-top canonical (verified).
- **The finding (R14/P9):** the conflict wall is **20 callees / 24 targets / 7% of wave reach** (not ~2×). The
4 reach-134 *circular* conflict callees hand-tried are **all gcc-quirk/regalloc/layout-bound, 0 banked**
(func_8012B4B8 §10 hoist; func_8012B8E4 75=75 regalloc-swap, permuter probe stalled at base; func_8016A8FC /
func_80169A4C local-struct-builders). **The high-reach core IS the gcc-quirk tail.**
- `git status`: only `config/`/`tools/`/`src/`/`docs/`/`phase-ends/` tracked; zero ROM-derived/generated bulk
staged. **No Ghidra DB change this phase** (matching used the cached Ghidra-C + asm; no live MCP writes) —
R23 no-op; the `ghidra/ db.*.gbf` churn in `git status` is SessionStart-restart rename noise, do NOT commit it.
**Milestone achieved (confirmed by Drew, gate 2):** the phase set out to raise the harness's per-pass yield to
"eureka" and gate an unattended run. **Measured outcome:** the ceiling did NOT rise to eureka — all 5 planned
avenues were byte-neutral/zero except the guided-hand-matching pivot (fleet 54.48%→55.51%); the canonical-sig
layer was built + validated (sig-conflict wall 20→0, byte-neutral 136/136) BUT the byte-proven finding is that
the **real wall is the gcc-quirk tail, not sig conflicts or missing types** — so the **go/no-go is NO-GO on more
brute waves at the current ceiling, and a GO on a fresh compiler-quirk research phase** that attacks the wall
directly (understand gcc-2.7.2 → raise the close-rate → then resume the built, layer-clean wave).
**Next:** **Phase 18 — Compiler-quirk research (understand gcc-2.7.2 to raise the match-% ceiling).** A
Max-effort, **plan-mode** research phase (Tier-1). See "Notes for Future Phases" for the brief. The harvest
infrastructure is BUILT + staged — resume it after the research lands new cookbook idioms.
## Deviations
| Item | Plan | Actual | Reason |
|---|---|---|---|
| The 5 ceiling avenues (T1–T6) | raise per-pass yield to "eureka" | all byte-neutral/zero except T3 (+0.52%) | the premise was wrong — type recovery is byte-neutral; the permuter doesn't transfer (T6) |
| Whole-phase shape | gate an unattended m2c+permuter run | **PIVOT to guided hand-matching** (Drew), then to **compiler research** | hand-matching is the only proven path on struct fns; but its wall is the compiler, redirecting to R17 research |
| The canonical-sig layer | "the ~2× scaling lever; build before the wave" | built + byte-neutral, but the wall is **7%-reach, not ~2×** | the "~2×" was the top-30's in-flight conflicts, since dissolved by banking (R14 sizing correction) |
| Match the 4 reach-134 circular (validate + bank ×134) | bank the highest-value targets | **0 banked — all quirk-bound** | the high-reach core is the gcc-quirk tail; the layer makes them declarable, not matchable |
| The big wave (task 5) | run the Ultracode wave on the 270 | **staged + DEFERRED** to post-research | understanding the compiler raises the wave's ceiling more than running it now (Drew) |
| Emulator struct types | (raised as a possible silver bullet) | **re-confirmed byte-neutral for matching** | matching reads access widths off the asm opcode, not a struct def; types help comprehension only |
## Commit Message
```
feat(phase-17): close — canonical-sig layer built; the wall is the compiler, not sigs; pivot to gcc research (v1.16.0)
- canonical-sig layer (session 4): tools/census_conflict_callees.py + derive_canonical_sigs.py
-> a 20-extern byte-neutral block atop ov_SC01_077.c (LOCAL, not engine_core.h); census
conflict callees 20->0, blocked targets 24->0; gate pipeline now draft -> sig_unify (MANDATORY)
-> harvest_verify --chunk 1; fleet 136/136 byte-identical (R22), 55.51% (no regression)
- FINDING (R14/P9): the conflict wall is 7%-reach not ~2x; the 4 reach-134 circular targets are
ALL gcc-quirk/regalloc/layout-bound (0 banked); the high-reach core IS the quirk tail; struct
types are byte-neutral for matching (the wall is gcc codegen, not knowledge)
- leverage analysis: fleet % is function-count-weighted (size adds no %); "unblock many" = the
layer (declaration, not matching); reach is the lever (already reach-sorted); 247 tractable
reach-134 stubs ~ +3-4% projected
- GO/NO-GO: NO-GO on brute waves at the current ceiling; GO on a compiler-quirk research phase
(read gcc-2.7.2 source + Xenogears + the §10/regalloc classes, R17) -> then resume the wave
- docs: cookbook §16 corrected + hand-matching-process.md §8 (the layer + the finding + handoff)
- worklog archived -> phase-ends/logs/Phase17.md (R19); bumps 1.15.0 -> 1.16.0
```
## Rules Added This Phase
| Rule | Reason |
|---|---|
| **None (governance).** Phase 17's lessons are *techniques + a strategic finding*, recorded where they belong (the Phase-8/11/13/15/16 precedent): the **canonical-sig layer** (census/derive tools + the local block + sig_unify-before-gate) → cookbook §16 / hand-matching §8; the **finding** (the wall is the gcc-quirk tail, not sigs/types) + the **leverage analysis** (fleet % is function-weighted so size adds no %; reach is the size-independent lever; matching is independent per function) → hand-matching §8c; the **token-economics insight** (for breadth, isolated agents beat serial main-loop because the main loop's accumulating context is re-read every turn → quadratic; the doc-redundancy is secondary) → memory + effort-map. The existing **R17** (web-research compiler internals) is exactly the Phase-18 mandate; **G3/P9/R14/R22/R26/R27** already govern the rest. | A negative-but-decisive result + a tooling layer produces knowledge, not a new norm of conduct. |
## PhaseEnd Changelog
**v1.15.0 → v1.16.0 — Phase 17 complete (Gen2 phase 9; a go/no-go that PIVOTS the strategy).** The phase
tested 5 avenues to raise the harness ceiling to "eureka" (all byte-neutral/zero), then **pivoted to guided
hand-matching** (Drew) — proven on a demo (4/5) and an Ultracode **calibration wave** (+0.47%), driving the
fleet **54.48% → 55.51%** (+1.03% net, 136/136 byte-identical throughout). The calibration exposed a
**canonical-sig wall** (parallel agents declaring shared callees inconsistently → `conflicting types`), and
session 4 **built + validated the canonical-sig layer** (two new tools + a byte-neutral 20-extern block,
sig-conflict surface 20→0). **The decisive finding** (R14/P9): the wall is only **7%-reach, NOT the "~2×
lever"** it first looked like, and the **4 highest-reach targets are all gcc-quirk/regalloc/layout-bound**
(0 banked, even the permuter stalls) — so the **real wall is the compiler's codegen, not signatures or missing
types** (struct types re-confirmed byte-neutral for matching). The leverage analysis settled the "fewest that
unlock the most" question with data (fleet % is function-count-weighted → size adds no %; reach is the
size-independent lever, already maximised; matching is independent per function → no "magic 5"). **Go/no-go
(gate 2, Drew):** NO-GO on more brute waves at the current ceiling; **GO on a fresh compiler-quirk research
phase** (read the gcc-2.7.2 source + Xenogears-decomp + the §10/regalloc quirk classes, R17) to raise the
close-rate — then **resume the built, layer-clean harvest wave** (staged at `.run/harvest_wave_s4.js`). No new
governance rules (techniques → cookbook §16 / hand-matching §8; the strategic finding → the same; token
economics → memory). All 136 binaries byte-identical; 0 NON_MATCHING linked (G4).
## Notes for Future Phases — Phase 18 research brief (the compiler-quirk avenue, match-% goal)
> **Phase 18 = "Understand gcc-2.7.2 to raise the match-% ceiling."** Max-effort, plan-mode (Tier-1). The wall
> is the compiler, so the lever is understanding it. Prioritised (highest match-% leverage first):
1. **[THE LEVER] Read the actual gcc-2.7.2 source** (`pmret/gcc-papermario`, the PSX 2.7.2.x lineage) — the
**reload / CSE / instruction-scheduling / register-allocation** logic — to understand the heuristics behind
the two classes that block the harvest: the **§10 rematerialize-vs-hoist** (a cheap stack/global address kept
in a callee-saved reg vs recomputed per use) and the **`$s0/$s1` regalloc-order** swap. Goal: turn "not
source-steerable" into cookbook idioms (C shapes that *trigger* the wanted codegen). One cracked class lifts
the close-rate across hundreds of functions — more than another brute wave.
2. **[GOLD REFERENCE] Mine Xenogears-decomp** — Square, Oct 1998, **our exact compiler** (gcc-2.7.2-psx + -cdk),
`gears.toml` presets. The most directly transferable quirk knowledge that exists. Then decomp.me (our
compiler's scratches), decomp-wiki, maspsx/m2c issues. **sotn is GCC 2.6.3** (wrong era) — methodology only.
3. **[CHEAP DUE-DILIGENCE] Sweep the PsyQ archive** (archive.org, where the 4.0/4.7 `.LIB`s came from) for the
**real CC1PSX.EXE / ASPSX.EXE** (byte-exact arbitration of any maspsx-emulation doubt, §4.8 deferred) + Sony
**sample code** (the C idioms that produce specific asm). Won't crack the quirk tail (the real compiler emits
the same bytes) but is low-cost and may surface idioms.
4. **X2 — treat ALL web content as untrusted DATA** (a prompt-injection doc was served during this project's
research). Prefer API endpoints; record sources.
5. **The deliverable:** cookbook-ready gcc-2.7.2 idioms (or honest "this class is unsteerable" verdicts) for the
quirk classes. **Then resume the harvest wave** — the layer + `.run/harvest_wave_s4.js` (40 tractable
reach-134) + the sig_unify-before-gate pipeline are BUILT and staged; the research only raises their yield.
6. **Carried, lower priority:** the giants (28, reach-134, >150 ins — highest *bytes* but same %, hardest);
comprehension/emulator field-naming (Gen2 quality, byte-neutral — do when match-% is exhausted).
## Plain-English Recap
We spent this phase trying to make our automatic game-code-matching dramatically more productive. We tried five
different ideas to "raise the ceiling" — all of them turned out to barely move the needle. So we switched to a
hands-on approach: an AI swarm hand-writes the matching code, checked by an automatic bit-for-bit referee. That
worked and pushed us from ~54.5% to ~55.5% of the game's shared code rebuilt perfectly. Along the way we built a
clever helper (the "canonical-sig layer") that we expected to roughly *double* the swarm's output — and we got
it working flawlessly — but when we measured it, the problem it solves turned out to be much smaller than we
thought (it was already mostly solved). The big, honest discovery: **the thing blocking us is the original
1990s compiler itself.** The most valuable, most-reused engine functions get blocked not by anything we don't
*understand* (we read them fine), but by specific low-level choices that ancient compiler made — which register
it used, how it ordered instructions — that our rebuild can't always reproduce from clean code. We proved that
neither better data-structure knowledge nor sharing context between functions changes that; the wall is the
compiler. We also answered your "is there a magic 5 that unlock the rest?" question with hard data: no — because
matching is independent per function, the score counts functions (not their size, so giants give no edge), and
the real leverage (reusing one match across all 134 levels) we already use. **So the smart next move is to stop
out-muscling the compiler and instead *study* it** — read its actual source code and a sister project
(Xenogears) that used the exact same compiler — to learn the tricks that make it produce the bytes we need.
That becomes its own focused phase. Everything we built — the layer, the swarm wave, the tools — is finished and
waiting; we just resume it later with better compiler knowledge. All 136 pieces of the game still rebuild
bit-for-bit; nothing broke.
## 🛑 Stop Here
PhaseEnd written; `CURRENT_PHASE.md` archived → `phase-ends/logs/Phase17.md` (R19). **Drew commits AND pushes**
this PhaseEnd + the Phase-17 session-4 change set (R6/R8 — the 2 new tools, the ov_SC01_077.c layer block, the
cookbook §16 / hand-matching §8 doc updates, this PhaseEnd, the archived log; the `ghidra/ db.*.gbf` churn is
R23 restart-noise — do NOT stage it). No Ghidra DB change this phase (R23 no-op). Gen2 continues — do **NOT**
start Phase 18 here. Start a **fresh session** (effort **Max**, **plan mode**) for **Phase 18 — Compiler-quirk
research** (the brief above). Keep this file forever.
@@ -260,3 +260,42 @@ types — 67% reach incl. the giants) → T6 (validate the 146 permuter candidat
~2× yield/token). 12/30 = genuine gcc-quirk tail. **NEXT (scaling, Max):** build the canonical-sig layer
(establish/enforce shared-callee canonical sigs; match shared callees before callers), then scale the wave
to the remaining ~270 targets (`.run/harvest_targets_s3.json`). Then T7 close.
- 2026-06-19 (session 4, Max): **CANONICAL-SIG LAYER BUILT + a key sizing correction (R14/P9).**
- **Tools (new, committed):** `tools/census_conflict_callees.py` (the accurate conflict predicate:
undeclared-stub callee with `decl_sources = n_callers + is_target >= 2`) + `tools/derive_canonical_sigs.py`
(byte-neutral canonical = `s32` return + `s32`/arity params; arity from Ghidra-C cache AND asm
read-before-write `$a0-$a3`, agreeing on all 14 cached, 6 stubs call-site-validated).
- **THE LAYER:** 20 conflict callees (14 are themselves wave targets / 4 at reach-134; 6 non-target stubs)
declared once as a file-top `extern s32 ...` block in `src/ov_SC01_077/ov_SC01_077.c` (LOCAL, NOT
engine_core.h — reach-1 names like func_801809BC differ across overlays; it's matched in ov_SC03_096).
`gen_harvest_targets` + `sig_unify` both already read ov_SC01_077.c → the layer auto-wires (no tool change).
Census after: **conflict callees 20→0, blocked targets 24→0**. Byte-neutral: ov_SC01_077 clean-rebuilds
`d19c9580` (R22). sig_unify test: rewrites a wrong `extern void func_801758FC(s32)` → canonical
`extern s32 func_801758FC(void)`. ✓
- **SIZING CORRECTION (R14):** the §7c "~2× scaling lever" was measured on the top-30's IN-FLIGHT conflicts;
since those callees got banked the wall shrank. For the remaining 270 it is **only 20 callees / 24 targets /
7% of wave reach** — a modest unblock, NOT the dominant lever. The wave's real ceiling is the gcc-quirk tail
(§2/§10), unchanged by the layer.
- **PIPELINE CHANGE (required):** harvest_verify uses an ACCUMULATING baseline from ov_SC01_077.c (now with
the file-top block) → a raw draft's own guessed extern would clash with the block even at `--chunk 1`. So
the wave gate is now **draft → sig_unify (MANDATORY, normalizes to the file-top canonical) → harvest_verify
--chunk 1 → propagate** (sig_unify was previously a recovery-only pass).
- **NEXT (Drew decision pending — R27 Ultracode prompt + the resized expected value):** scale the wave (task 5,
needs `/effort ultracode`) vs. a focused Max session matching the 4 reach-134 circular targets + tractable
high-reach subset by hand. Layer done (task 2); sig_unify wired (task 3). Tasks 4/5/6/7 open.
- 2026-06-19 (session 4 cont., Max): **TASK 4 — hand-matched the 4 reach-134 circular targets (Drew chose A);
ALL 4 are gcc-quirk/regalloc/layout-bound → 0 banked. Layer validated independently; the finding is the
high-reach core IS the quirk tail (confirms §4).** Drafts in `.run/drafts-s4/` (scratch, none gated).
- `func_8012B4B8` (matrix transform): **§10 stack-addr rematerialize-vs-hoist** — gcc caches `&mtx` in a
callee-saved reg (3 saved); the original re-materializes `addiu $a1,$sp,0x10` per call (2 saved). Not
source-steerable (2 variations tried).
- `func_8012B8E4` (angle-diff + `--expand-div`): **structurally PERFECT (75=75)**, down to a `$s0/$s1`
regalloc swap + a reassociation = 24-mismatch near-miss. **Permuter probe (external callee `ratan2` →
T6 inlined-callee concern does NOT apply): base 530, NO improvement in 90s** → not in the permuter's
randomization space (§3). Scratch `.run/permuter/func_8012B8E4/` (reusable: header+macro.inc → target.o).
- `func_8016A8FC` / `func_80169A4C`: **local-struct-builders** (stack-layout/scheduling-bound). Assessed hard.
- **CONCLUSION:** the layer makes the high-reach circular targets *declarable* but they are the **gcc-quirk
tail, not the sig-conflict wall** — the layer doesn't unlock them. Its value: a **sig-conflict-clean wave**
banking the TRACTABLE (mostly lower-reach) subset. **Wave yield is quirk-tail-limited (~2-4% fleet), NOT
~2×.** This IS the task-7 go/no-go input. **Decision pending (Drew): run the layer-clean wave (task 5,
/effort ultracode) vs. close Phase 17 + defer the wave to a dedicated/unattended run (Phase-16 auto_driver).**
+30
View File
@@ -1,6 +1,36 @@
#include "common.h"
#include "../shared/engine_core.h"
/* ==== Phase-17 canonical-sig layer (tools/derive_canonical_sigs.py) ===================
* ONE byte-neutral canonical signature per undeclared-stub conflict callee, so the parallel
* hand-matching wave declares each shared callee consistently and the one-big-TU build stops
* failing on `conflicting types` (hand-matching-process.md §7c). Form: s32 return (void->s32
* byte-neutral, §3a-1) + s32 params (matched bodies cast int->ptr), arity from Ghidra-C + asm
* read-before-write $a0-$a3 (agree on all 14 cached; 6 stubs call-site-validated). LOCAL to
* this TU on purpose (reach-1 names like func_801809BC differ across overlays, so NOT in the
* shared engine_core.h). Whole-binary harvest_verify byte-gate remains the sole arbiter (G3/P9). */
extern s32 func_8016EC0C(s32 a0, s32 a1); /* match-first, arity 2 */
extern s32 func_8012B4B8(s32 a0); /* match-first, arity 1 */
extern s32 func_801670E4(s32 a0, s32 a1, s32 a2, s32 a3); /* derive-decl, arity 4 */
extern s32 func_80169A4C(s32 a0, s32 a1); /* match-first, arity 2 */
extern s32 func_8016A8FC(s32 a0); /* match-first, arity 1 */
extern s32 func_8012B8E4(s32 a0, s32 a1); /* match-first, arity 2 */
extern s32 func_8015E1B8(s32 a0); /* match-first, arity 1 */
extern s32 func_8015EE08(s32 a0); /* match-first, arity 1 */
extern s32 func_8015F7D4(s32 a0); /* match-first, arity 1 */
extern s32 func_80160B34(s32 a0); /* match-first, arity 1 */
extern s32 func_80165140(s32 a0); /* match-first, arity 1 */
extern s32 func_80161CD0(s32 a0, s32 a1); /* match-first, arity 2 */
extern s32 func_80175268(s32 a0); /* match-first, arity 1 */
extern s32 func_8017EC7C(s32 a0); /* match-first, arity 1 */
extern s32 func_801809BC(s32 a0, s32 a1); /* match-first, arity 2 */
extern s32 func_8012DE2C(s32 a0); /* derive-decl, arity 1 */
extern s32 func_8012DDA4(void); /* derive-decl, arity 0 */
extern s32 func_801759D8(void); /* derive-decl, arity 0 */
extern s32 func_80175820(void); /* derive-decl, arity 0 */
extern s32 func_801758FC(void); /* derive-decl, arity 0 */
/* ==== end canonical-sig layer ==================================================== */
DEFINE_func_80128158() /* dedup: shared engine-core @0x80128158 (src/shared) */
DEFINE_func_80128178() /* dedup: shared engine-core @0x80128178 (src/shared) */
+122
View File
@@ -0,0 +1,122 @@
#!/usr/bin/env python3
"""Census the shared callees that block the parallel hand-matching wave (Phase 17 canonical-sig layer).
The calibration wave (hand-matching-process.md §7c) hit 60% match_one MATCH but only 33% whole-binary;
the entire gap was SIG CONFLICTS — parallel agents each declare an undeclared-stub shared callee (e.g.
func_80131CA8) with a different guessed signature, which then clash in the one-big-TU ov_SC01_077.c
(`conflicting types for func_X`). 100% compile-errors, ZERO codegen mismatches.
A callee that is already `defined` (engine_core.h DEFINE / inline body) or already `declared` (an extern
written somewhere) is conflict-FREE — gen_harvest_targets resolves it and every draft reuses the one sig.
The conflict set is exactly the `stub` callees (no body, no extern anywhere) that >=2 still-stub wave
targets call. Declaring ONE canonical extern for each (the canonical-sig layer) turns them `declared` ->
the wave stops conflicting on them.
This is read-only. Output: the ranked conflict-callee table + the call-edge status totals.
Usage:
tools/census_conflict_callees.py [--source ov_SC01_077] [--targets .run/harvest_targets_s3.json]
"""
import argparse, json, os, importlib.util
from collections import defaultdict
REPO = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
_spec = importlib.util.spec_from_file_location('ght', os.path.join(REPO, 'tools/gen_harvest_targets.py'))
_ght = importlib.util.module_from_spec(_spec)
_spec.loader.exec_module(_ght)
def main():
ap = argparse.ArgumentParser()
ap.add_argument('--source', default='ov_SC01_077')
ap.add_argument('--targets', default='.run/harvest_targets_s3.json')
ap.add_argument('--min-callers', type=int, default=2)
ap.add_argument('--out', default='.run/conflict_callees.json')
args = ap.parse_args()
src = args.source
c_path = os.path.join(REPO, f'src/{src}/{src}.c')
ec = os.path.join(REPO, 'src/shared/engine_core.h')
src_sig = _ght.load_sig(os.path.join(REPO, f'.run/sig.{src}.jsonl')) # addr-int -> {calls,nins,h_exact,reach?}
stubs = _ght.collect_stubs(c_path) # current INCLUDE_ASM set
define_sigs = _ght.collect_define_sigs(ec)
inline_sigs = _ght.collect_inline_sigs(c_path)
defined = {**define_sigs, **inline_sigs}
extern_sigs = _ght.collect_extern_sigs([ec, c_path])
s3 = json.load(open(os.path.join(REPO, args.targets)))
tgt_addrs = [int(t['addr'], 16) for t in s3]
remaining = [a for a in tgt_addrs if a in stubs] # still-stub wave targets
def status(caddr):
if caddr in defined: return 'defined'
if caddr in extern_sigs: return 'declared'
if caddr in stubs: return 'stub'
return 'extern' # resident/EXE, conflict-free
# reach map (how many overlays hold a byte-identical copy) for prioritization
reach = {int(t['addr'], 16): t.get('reach', 1) for t in s3}
callee_callers = defaultdict(set) # callee_addr -> {target addrs calling it (excl. self)}
edge_status = defaultdict(int) # status -> # call edges (over remaining targets)
for a in remaining:
rec = src_sig.get(a)
if not rec:
continue
for ch in set(rec.get('calls', [])):
caddr = int(ch, 16)
edge_status[status(caddr)] += 1
if caddr != a:
callee_callers[caddr].add(a)
# THE conflict predicate (hand-matching-process.md §7c): a sig conflict needs >=2 INDEPENDENT
# declarations of an UNDECLARED stub in the one-big-TU build. Declaration sources for callee X =
# one extern per drafted caller + one definition if X is itself a drafted target (def-vs-extern).
# decl_sources = n_callers + (1 if X is a remaining target else 0)
# An already-`declared`/`defined` callee is conflict-FREE (gen_harvest_targets feeds the one sig).
conflicts = []
for caddr, callers in callee_callers.items():
if status(caddr) != 'stub':
continue
is_t = caddr in remaining
decl_sources = len(callers) + (1 if is_t else 0)
if decl_sources < 2:
continue
crec = src_sig.get(caddr, {})
conflicts.append({
'callee': f'func_{caddr:08X}', 'addr': f'{caddr:08X}',
'n_callers': len(callers), 'is_target': is_t, 'decl_sources': decl_sources,
'kind': 'match-first' if is_t else 'derive-declare',
'callee_nins': crec.get('nins'), 'callee_ncalls': len(crec.get('calls', [])),
'callee_reach': reach.get(caddr, crec.get('reach')),
})
conflicts.sort(key=lambda x: (-x['decl_sources'], -(x['callee_reach'] or 0)))
blocked = {a for a in remaining
if any(int(ch, 16) in {int(c['addr'], 16) for c in conflicts}
for ch in src_sig.get(a, {}).get('calls', []))}
blk_reach = sum(reach.get(a, 1) for a in blocked)
tot_reach = sum(reach.get(a, 1) for a in remaining)
json.dump(conflicts, open(os.path.join(REPO, args.out), 'w'), indent=1)
print(f'wave scope: {len(s3)} s3 targets, {len(remaining)} still-stub (remaining)')
print('call-edge status totals (over remaining targets): '
+ ' '.join(f'{k}:{edge_status[k]}' for k in ('defined', 'declared', 'stub', 'extern')))
mf = sum(1 for c in conflicts if c['is_target'])
print(f'\nCONFLICT CALLEES (undeclared stub, decl_sources>=2): {len(conflicts)} '
f'[{mf} match-first / {len(conflicts)-mf} derive-declare]')
print(f'{"callee":>16} {"callers":>7} {"isTgt":>5} {"srcs":>4} {"nins":>5} {"ncalls":>6} {"reach":>5}')
for c in conflicts:
print(f'{c["callee"]:>16} {c["n_callers"]:>7} {("Y" if c["is_target"] else ""):>5} '
f'{c["decl_sources"]:>4} {str(c["callee_nins"] or "-"):>5} '
f'{str(c["callee_ncalls"] or "-"):>6} {str(c["callee_reach"] or "-"):>5}')
print(f'\nwave targets blocked by >=1 conflict callee: {len(blocked)}/{len(remaining)} '
f'(reach-weighted {blk_reach}/{tot_reach} = {100*blk_reach//max(tot_reach,1)}% of wave reach)')
print(f'wrote {args.out}')
if __name__ == '__main__':
main()
+148
View File
@@ -0,0 +1,148 @@
#!/usr/bin/env python3
"""Derive a byte-neutral canonical signature for each Phase-17 conflict callee (the canonical-sig layer).
Input: .run/conflict_callees.json (from census_conflict_callees.py) — the undeclared-stub callees that
parallel hand-matching agents would declare inconsistently (hand-matching-process.md §7c). For each we
emit ONE canonical `extern s32 func_X(s32 a0, ...);` to seed src/shared/engine_core.h, so gen_harvest_targets
feeds every drafting agent the SAME signature and the one-big-TU build stops conflicting.
Canonical form = WIDEST byte-neutral (hand-matching-process.md §3a):
- return `s32` : void->s32 is byte-neutral (no explicit return => identical epilogue); s32 is REQUIRED
where a caller uses $v0. So s32 is universally safe.
- params `s32` : widest scalar; a matched body casts int->ptr (`*(T*)(a0+off)`, the demo idiom) and a
caller narrows in the call expression. s32 never blocks a match; the byte-gate validates.
- ARITY is the only value that must be exact (a wrong count => "too few/many arguments" at a caller, or a
def/extern arity clash for a circular target). Derived two ways and cross-checked:
(G) Ghidra-C cache .run/ghidra_c/func_<A>.c (FUN_<a>(...) param count) — the static oracle (G1).
(A) asm read-before-write of $a0..$a3 in asm/<src>/nonmatchings/<src>/func_<A>.s — an a-reg whose
FIRST-touching instruction uses it as a SOURCE is an incoming param; dest-only first touch
(lui/lw/move/ALU-dest) = scratch, not a param. Arity = highest param index + 1 (contiguous).
The whole-binary harvest_verify byte-gate remains the sole arbiter (G3/P9): a wrong arity just fails the
gate and is fixed per-callee. This only shapes the drafting / seeds the canonical decls.
Usage:
tools/derive_canonical_sigs.py [--source ov_SC01_077] [--in .run/conflict_callees.json]
"""
import argparse, json, os, re, glob
REPO = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
AREGS = ['a0', 'a1', 'a2', 'a3']
# instruction operand roles restricted to what we need: is the FIRST a-reg touch a source (=> param)?
LOAD = {'lw', 'lh', 'lhu', 'lb', 'lbu', 'lwl', 'lwr', 'll', 'lwc1', 'lwc2', 'ldc1', 'ldc2'} # rt=dest, base=src
STORE = {'sw', 'sh', 'sb', 'swl', 'swr', 'sc', 'swc1', 'swc2', 'sdc1', 'sdc2'} # rt=src, base=src
DEST_FIRST = {'lui', 'li', 'move', 'addu', 'addiu', 'subu', 'and', 'andi', 'or', 'ori', 'xor', # rd/rt=dest, rest=src
'xori', 'nor', 'slt', 'sltu', 'slti', 'sltiu', 'sll', 'srl', 'sra', 'sllv', 'srlv',
'srav', 'mul', 'mult', 'negu', 'neg', 'not', 'mflo', 'mfhi', 'movn', 'movz', 'sub',
'add', 'rotr', 'clz', 'seb', 'seh', 'mfc1', 'mfc2', 'la'}
SRC_ALL = {'beq', 'bne', 'beqz', 'bnez', 'blez', 'bgtz', 'bltz', 'bgez', 'bgezal', 'bltzal', # all regs are src
'jr', 'jalr', 'multu', 'divu', 'div', 'mtlo', 'mthi', 'mtc1', 'mtc2', 'teq', 'tne',
'beql', 'bnel', 'cache'}
def reg_tokens(operand_str):
return re.findall(r'\$([a-z0-9]+)', operand_str)
def a_role(mnem, ops):
"""return (a_sources, a_dest) restricted to a0..a3 for one instruction."""
regs = reg_tokens(ops)
a_in_order = [r for r in regs if r in AREGS]
if not a_in_order:
return set(), set()
if mnem in LOAD:
dest = {regs[0]} & set(AREGS) if regs else set()
src = set(a_in_order) - dest
elif mnem in STORE:
dest, src = set(), set(a_in_order)
elif mnem in DEST_FIRST:
dest = {regs[0]} & set(AREGS) if regs else set()
src = set(a_in_order) - dest
elif mnem in SRC_ALL:
dest, src = set(), set(a_in_order)
else: # unknown: be conservative -> treat as sources (flags a param)
dest, src = set(), set(a_in_order)
return src, dest
INSN_RE = re.compile(r'\*/\s+([a-z][a-z0-9.]*)\s+(.*)$') # after the `... XXXX */` comment
def asm_arity(s_path):
"""highest read-before-write a-reg index +1; returns (arity, note)."""
if not os.path.exists(s_path):
return None, 'no-asm'
first = {} # areg -> 'param' | 'scratch'
for line in open(s_path):
m = INSN_RE.search(line)
if not m:
continue
mnem, ops = m.group(1), m.group(2)
src, dest = a_role(mnem, ops)
for r in AREGS:
if r in first:
continue
if r in src:
first[r] = 'param'
elif r in dest:
first[r] = 'scratch'
arity = 0
for i, r in enumerate(AREGS):
if first.get(r) == 'param':
arity = i + 1
# contiguity note: a param above a scratch/untouched gap (loose-typed) -> flag
gap = any(first.get(AREGS[j]) != 'param' for j in range(arity - 1)) if arity else False
return arity, ('gap' if gap else 'ok')
GHIDRA_SIG_RE = re.compile(r'^\s*[A-Za-z_].*\bFUN_[0-9a-f]+\s*\((.*?)\)\s*$', re.M)
def ghidra_arity(addr):
p = os.path.join(REPO, f'.run/ghidra_c/func_{addr}.c')
if not os.path.exists(p):
return None
m = GHIDRA_SIG_RE.search(open(p).read())
if not m:
return None
params = m.group(1).strip()
if params in ('', 'void'):
return 0
return len([x for x in params.split(',') if x.strip()])
def main():
ap = argparse.ArgumentParser()
ap.add_argument('--source', default='ov_SC01_077')
ap.add_argument('--in', dest='infile', default='.run/conflict_callees.json')
args = ap.parse_args()
asm_dir = os.path.join(REPO, f'asm/{args.source}/nonmatchings/{args.source}')
conf = json.load(open(os.path.join(REPO, args.infile)))
print(f'{"callee":>16} {"kind":>14} {"ghidra":>6} {"asm":>4} {"note":>6} {"->arity":>7} canonical')
rows = []
for c in conf:
addr = c['addr']
g = ghidra_arity(addr)
a, note = asm_arity(os.path.join(asm_dir, f'func_{addr}.s'))
# reconcile: prefer asm (callee's own consumption); if asm gap or asm<ghidra, trust the larger
cand = [x for x in (g, a) if x is not None]
arity = max(cand) if cand else 0
if g is not None and a is not None and g != a:
note = f'G{g}/A{a}'
params = 'void' if arity == 0 else ', '.join('s32 a%d' % i for i in range(arity))
sig = f'extern s32 func_{addr}({params});'
rows.append({'callee': c['callee'], 'addr': addr, 'arity': arity, 'kind': c['kind'],
'ghidra': g, 'asm': a, 'note': note, 'sig': sig})
print(f' func_{addr} {c["kind"]:>14} {str(g):>6} {str(a):>4} {note:>6} {arity:>7} {sig}')
out = os.path.join(REPO, '.run/canonical_sigs.json')
json.dump(rows, open(out, 'w'), indent=1)
print(f'\nwrote {out} ({len(rows)} canonical sigs)')
print('disagreements (G!=A) to eyeball:',
', '.join(r['callee'] for r in rows if r['note'].startswith('G')) or 'none')
if __name__ == '__main__':
main()