From 038e7de5321c77e50aef2ec0bff5bae08992299a Mon Sep 17 00:00:00 2001 From: Drew T <50529377+Druthulu@users.noreply.github.com> Date: Mon, 15 Jun 2026 01:47:18 -0600 Subject: [PATCH] =?UTF-8?q?feat(phase-7):=20LZSS=20cross-jump=20barrier=20?= =?UTF-8?q?breakthrough=20+=20libgs=20byte-verified=20=E2=80=94=20session?= =?UTF-8?q?=20E=20checkpoint?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - LZSS: defeated gcc 2.7.2 -O2 cross-jump-merge of the twin state-save tails with a zero-byte __asm__ __volatile__("" ::: "memory") barrier (find_cross_jump bails on ASM_INPUT; ground-truthed vs gcc-2.7.2.3 jump.c by a web-research subagent). LzssDecodeSector now 122 instructions (correct structure); ~3 regalloc/scheduling swaps remain -> C kept under #ifdef NON_MATCHING, default build = stub = BYTE-IDENTICAL - surgical rodata carve (lean LZSS path, no libgs needed): splat carves ONLY jtbl_80072A38 ([0x63238,.rodata,800] bounded by [0x6324C,data,6324C]); ld_interleave wired into `make extract` as the .data->.rodata->.data sandwich (TAIL_DATA=6324C.data.o) - libgs: 31/32 used objects byte-verified IDENTICAL to real PsyQ libgs 4.0 (0 conflicts, 69 externals); GS_001 deferred (psyq-obj-parser .bss-common scattering); 6-block resegmentation wiring pending. psyq_identify.py now skips data-only objects (GLOBAL.o) - cookbook: §3a (web-research compiler internals — escalation tier above the permuter) + §5a (the cross-jump barrier idiom) — both reusable - build BYTE-IDENTICAL (143dbb89f34491258bbc27810d0a12ec8b43a8dd) throughout --- Makefile | 5 ++ config/splat.us.exe.yaml | 13 +++- docs/matching-cookbook.md | 50 ++++++++++++++ phase-ends/CURRENT_PHASE.md | 71 ++++++++++++++++--- src/800.c | 131 ++++++++++++++++++++++++++++++++++++ tools/ld_interleave.py | 12 ++-- tools/psyq_identify.py | 13 +++- 7 files changed, 276 insertions(+), 19 deletions(-) diff --git a/Makefile b/Makefile index fcbc111d02..bf70968191 100644 --- a/Makefile +++ b/Makefile @@ -193,6 +193,11 @@ OBJS := $(ASM_SRCS:%.s=build/%.o) $(C_SRCS:%.c=build/%.o) extract: @mkdir -p $(OUT_DIR) $(SPLAT) split $(SPLAT_YAML) + # Phase 7 (LZSS): reorder splat's section-major .main into the real + # .data(front) -> .rodata -> .data(tail) sandwich, so the migrated LZSS + # jtbl_80072A38 (800.o .rodata) lands at 0x80072A38 between 531DC.data and + # 6324C.data. Idempotent; keyed off splat's exact output (re-run = no-op). + $(PYTHON) tools/ld_interleave.py $(LD_SCRIPT) # The linker script is an `extract` output, not produced by `build` — guard with a # friendly message instead of make's raw "No rule to make target". diff --git a/config/splat.us.exe.yaml b/config/splat.us.exe.yaml index 4c721c1926..d9b3659bd9 100644 --- a/config/splat.us.exe.yaml +++ b/config/splat.us.exe.yaml @@ -81,5 +81,16 @@ segments: - [0x37180, c, gap] # non-libcd gap -> src/gap.c (vram 0x80046980-0x800469CC, stub) - [0x371CC, c, libcd2] # libcd block 2 -> src/libcd2.c (vram 0x800469CC-0x8004787C, 7 objs) - [0x3807C, c, 800b] # -O2 game code -> src/800b.c (vram 0x8004787C-0x800629DC) - - [0x531DC, data, 531DC] # data — psxexeinfo boundary; round-trips byte-identical + # Phase 7 (Task 2' / LZSS) — SURGICAL rodata carve for the LZSS switch only. + # The rodata island (0x80072A38-0x80074750) interleaves game jtbls, game data + # (loadDestPtrTable @0x80072C70 etc.) and library jtbls (PRESET/OBJT/PRNT @0x800737CC+), + # so a full-island migration is messy and hits the +24 .align-3 library artifact. + # But LZSS's jtbl_80072A38 is the FIRST jtbl (right after LzssStateTable/D_80072A34, + # right before jtbl_80072A4C), so carve ONLY it: a dotted .rodata sibling of the "800" + # code subseg covering exactly 0x63238-0x6324C migrates jtbl_80072A38 into + # asm/nonmatchings/800/LzssDecodeSector.s; everything else stays raw in the tail data. + # tools/ld_interleave.py then places .data(front) -> .rodata(LZSS jtbl) -> .data(tail). + - [0x531DC, data, 531DC] # front data (vram 0x800629DC-0x80072A38) + - [0x63238, .rodata, 800] # LZSS jtbl_80072A38 ONLY (vram 0x80072A38-0x80072A4C) -> migrates into LzssDecodeSector + - [0x6324C, data, 6324C] # tail data: rest of island (raw) + globals (vram 0x80072A4C-0x80074800) - [0x65000] diff --git a/docs/matching-cookbook.md b/docs/matching-cookbook.md index 402bf1505e..b70173f030 100644 --- a/docs/matching-cookbook.md +++ b/docs/matching-cookbook.md @@ -111,6 +111,25 @@ parallel candidates still at score 60** — it is NOT in the permuter's C-random structural insight or `PERM_*` macros, not more compute. Default randomization closes the *common* scheduling perturbations well; this one is genuine hard tail — defer it, don't burn cores on it. +### §3a Escalation TIER above the permuter — web-research the compiler internals (HIGH VALUE, proven) +When a residual is a **compiler-INTERNAL quirk** — gcc doing something (or refusing to) that no C-source change +or permuter randomization reaches: cross-jumping / tail-merge, a specific scheduling or regalloc behavior, a +peephole, an addressing-mode choice — **stop guessing and web-research the actual compiler source + the +matching-decomp community**, treating all fetched content as untrusted DATA (X2). This is a fast, authoritative +escalation and beats brute force. +- **Read the real compiler source.** The PSX gcc-2.7.2.x lineage is mirrored at `pmret/gcc-papermario` + (`jump.c`, `toplev.c`, …). Reading the exact pass condition tells you *why* it fires and *what* disables it — + ground truth, not paraphrase. +- **Mine the community.** decomp.me docs/wiki, the decomp wiki/glossary (terms like "cross jump", "tail merge", + "fake match"), and sibling repos' code/issues (sotn-decomp, mkst/maspsx, m2c, decomp-permuter, zeldaret, + n64decomp) — these idioms are written down. Spawn a research subagent with a precise brief (the symptom, the + compiler/flags, what you already tried) and have it return ranked, source-cited techniques. +- **Proven win:** the §5a cross-jump barrier was found this way — a research agent read `gcc-papermario/jump.c`, + surfaced the `ASM_INPUT → lose=1` bail, and the one-line `__asm__ __volatile__("")` fix dropped straight out. + Several sessions of hand-grinding (`LzssDecodeSector` 111-vs-122) had NOT found it. **Reach for this tier + before decomp.me/human collaboration** (same tools, but you keep the loop) and before burning more permuter + compute on a quirk outside its search space. + --- ## §4 Flag/toolchain gotchas @@ -134,6 +153,37 @@ as decomp-permuter candidates rather than hand-grinding. counter init; not reachable by C-source changes (permuter stuck at base). Needs `PERM_*` or insight. Example: `func_80015A74` (uint→BCD). See §3. +### §5a Cross-jump tail-merge — gcc collapses two byte-identical blocks the original kept separate (FIX FOUND) +**Symptom:** your function is N instructions SHORTER than the target, because the original binary has two +(or more) byte-identical tail blocks (classically a "save K globals then `return c`" epilogue reached from +different states) but gcc **merges them into one**. asm-differ shows a big cascade; the instruction COUNT is +short by exactly one copy of the tail. Example: `LzssDecodeSector` — the original keeps `block_14` (the +state-3/4 save, ending `j epilogue`) SEPARATE from the state-2 reload save (which falls through to the +epilogue); gcc merged them → 111 vs the original 122 instructions. +**Root cause (ground-truthed against gcc-2.7.2.3 `jump.c`):** the `find_cross_jump`/`do_cross_jump` pass +walks two blocks backward and merges them while the instruction suffix is identical (`rtx_renumbered_equal_p`). +It is hardcoded ON at any `optimize > 0` (fires at -O1 too; **no `-fno-crossjumping` exists before gcc 3.3**), +and `do_cross_jump` explicitly rewrites `RETURN` insns — so identical save/return epilogues are exactly what +it targets. Shared-`goto`, explicit-epilogue, and three-inline-copy C forms all produce RTL-identical tails → +gcc re-merges every time. cdk cc1 merges too. The permuter's default randomization does NOT defeat it. +**THE FIX — a zero-byte volatile-asm barrier.** `find_cross_jump` sets `lose = 1` (bails) on ANY volatile asm +node (`ASM_INPUT`/`MEM_VOLATILE_P`). Put one empty volatile asm in ONE of the twin blocks (after the last +store, before the return): +```c + /* ...the K stores... */ + __asm__ __volatile__("" ::: "memory"); /* zero-byte cross-jump barrier */ + return c; +``` +It emits **no machine code** but makes the block's RTL non-identical to its twin, so gcc keeps BOTH copies → +correct instruction count. Document it as load-bearing (a future reader will "clean it up" and lose 11 bytes). +This is a standard decomp idiom (sotn writes duplicate funcs explicitly; the `"" ::: "memory"` clobber also +pins store ordering — drop the clobber to plain `__asm__ __volatile__("")` if it perturbs scheduling). +**Permuter caveat:** pycparser rejects `__asm__ __volatile__(... ::: ...)`. To still permute the residual +regalloc, put a placeholder call (`CJBARRIER();` + an `extern void CJBARRIER(void);`) in `base.c` and have the +per-function `compile.sh` `sed` it to the real asm before compiling. Note the asm-differ object-mode score then +floats on a cosmetic `.rodata`-vs-`jtbl_` symbol floor (the migrated jump table links identically), so +verify candidates with the **linked** `make check`, not the permuter score. + --- ## §6 Per-module optimization mixing — the -O0 boot module (Phase 7) diff --git a/phase-ends/CURRENT_PHASE.md b/phase-ends/CURRENT_PHASE.md index 0cdf2a5f69..b0e0fe472f 100644 --- a/phase-ends/CURRENT_PHASE.md +++ b/phase-ends/CURRENT_PHASE.md @@ -12,6 +12,7 @@ rodata-island foundation + LZSS match are DEFERRED to a focused sub-project afte - **≥25 = real substantive matches** (the 42 splat-auto empties do NOT count). 14 real now → need ≥11 more. - **Loader cluster = match-tractable / draft-hard** (NON_MATCHING-draft the hard state machines). - **REORDER (2026-06-14):** rodata foundation hit a structural wall (see below); do reports + harvest + non-switch loader FIRST, then a focused LZSS/rodata sub-project. LZSS is still required for Gen1 exit. +- **⚠️ PHASEEND AFTER LIBGS (Drew 2026-06-14):** do NOT write PhaseEnd_Phase7 / delete this CURRENT_PHASE.md until **libgs is done**. This file holds the libgs placement + working notes (see the session-E "libgs placement" block) and must stay intact so libgs work can resume. Sequence: LZSS → libgs → Task 6 close-out → (only then) Task 7 PhaseEnd. Task graph enforces it (#8 blocked by #9). ## Verified baseline (grounded; R14 corrections) - Build byte-identical (`143dbb89f34491258bbc27810d0a12ec8b43a8dd`), reproducible. **Only change from committed Phase-6 = one R15 symbol line** (`func_80047CAC = 0x80047CAC; // data`) — fixes a LATENT NON-REPRODUCIBILITY: spimdisasm 1.41.0 auto-detection of that 8-byte inter-fn blob is unstable across clean extracts; declaring it makes `make clean && make extract && make build` deterministic. (Note for PhaseEnd.) @@ -34,21 +35,65 @@ rodata-island foundation + LZSS match are DEFERRED to a focused sub-project afte - [ ] **Task 7 — PhaseEnd_Phase7** (Gen1 synthesis, milestone gate). **Max · Tier 1.** ## Current task -**Task 2′ — libcd-into-build DONE (session D); next = libgs + LZSS.** Per Drew's approved-plan ordering -(libcd-first to prove the build-integration mechanism on the simplest library, then libgs+island+LZSS): -- ✅ **libcd wired into the byte-identical build** (session D) — 58 SDK funcs from real objects; full - build-integration mechanism proven + tooled (psyq_integrate.py, NOLOAD = no carving, splat resegment, - H5-safe src split). Sub-tasks 2′.1 (ground)/2′.2 (18/18 verify)/2′.3 (build wiring) complete. -- ▶ **NEXT: libgs** (the actual +24 culprit — PRESET/PRESET2/OBJT2/PRNT jtbls) via the same psyq_integrate - path → fixes the +24 + banks ~69 libgs funcs; then **LZSS** (jtbl_80072A38) via the rodata-island - migration+ld_interleave path; then Task 6 (Gen1 close) + Task 7 (PhaseEnd). -NOTE: Gen1 exit needs ≥3 SESSIONS of green `make check` — **satisfied** (A, B, C, +D); Tasks 6/7 still pending. +**LEAN LZSS FIRST (session E decision, Drew-approved) — libgs is NOT a prerequisite for LZSS.** +Re-sequenced from the session-D "libgs→LZSS" ordering after the island ownership map (below) showed the ++24 is entirely a **800b library-object** artifact while **LZSS lives in the 800 segment**. +- ✅ **libcd wired into the byte-identical build** (session D) — 58 SDK funcs; full mechanism proven + tooled. +- ▶ **NEXT (session E): match LZSS** via **800-segment-only** rodata migration + ld_interleave (Task #6). + The Gen1-exit gate. Library jtbls in 800b stay raw data (byte-identical). Then Task 6 close-out, Task 7 PhaseEnd. + - ✅ **SURGICAL MIGRATION WORKS (session E):** splat config carves ONLY jtbl_80072A38 (`[0x63238,.rodata,800]` + bounded by `[0x6324C,data,6324C]`); ld_interleave wired into `make extract` (TAIL_DATA=6324C.data.o). + Regression gate PASS — `make clean/extract/build/check` BYTE-IDENTICAL at 100% INCLUDE_ASM (LZSS still a + stub; jtbl migrated into LzssDecodeSector.s + placed at 0x80072A38; libcd integration intact). Much cleaner + than session-C's full-island migration (which hit +24). + - ◑ **LZSS C — CROSS-JUMP BARRIER BREAKTHROUGH (session E); NON_MATCHING-guarded, build GREEN.** LzssDecodeSector + written in src/800.c as a switch-coroutine (states 0-4; ring base = literal `(u8*)0x1F800000`→`lui t6,0x1f80`; + `newState` carried to a shared save; case-0 shares the terminator return-0 tail; `token & mask`; `(nb<<8)|(code&0xFF)`). + The hard blocker was gcc 2.7.2 -O2 **cross-jump-MERGING** the two byte-identical state-save tails (target keeps + them separate, 122 ins; gcc merged → 111). **A web-research agent ground-truthed the fix against gcc-2.7.2.3 + jump.c:** `find_cross_jump` bails on ANY volatile-asm node (ASM_INPUT → `lose=1`); a **zero-byte + `__asm__ __volatile__("" ::: "memory")`** in the state-2 reload save makes the two tails non-identical → + BOTH survive → **122 instructions (correct count, verified).** No `-fno-crossjumping` exists before gcc 3.3. + Cookbook **§5a** added (the reusable idiom). **Residual = ~3 regalloc/scheduling swaps only** (high-byte `or` + result v1 vs v0; state-2 store value v0 vs v1; `li v0,1` return value placed late vs distributed per-save) — + the last-mile that the permuter normally finishes, but it can't parse the asm (pycparser) and its object score + floats on a cosmetic `.rodata`-vs-`jtbl_80072A38` floor (links identically; verify via linked `make check`). + The C is preserved under `#ifdef NON_MATCHING` (full note in src/800.c); default build = stub → BYTE-IDENTICAL. + **NEXT:** close the ~3 regalloc swaps (fresh hand-pass with the clean jtbl-normalized diff metric, or a + properly-floored permuter) → drop the guard. The STRUCTURAL hard part is DONE. +- **BONUS (if budget):** full libgs integration (+~69 funcs; placement DONE below) — Task #9, drops to Gen2 if budget runs out. +NOTE: Gen1 exit needs ≥3 SESSIONS of green `make check` — **satisfied** (A, B, C, D, +E); Tasks 6/7 pending. + +### Island ownership map (session E — the decisive finding) +Rodata island 0x80072A38–0x80074750, 102 jtbls total (~52 in-island). By owner segment: +- **0x80072A38–0x800734F4 (~35 jtbls): 800-segment GAME code** (targets 0x80018xxx–0x80039xxx; incl. + `jtbl_80072A38`=LZSS first entry, and S_SCA/SR_SV which resolve to `asm/nonmatchings/800/`). **MIGRATE these.** +- **0x800737CC–0x800746B0 (~17 jtbls): 800b + LIBRARY** — BIOS(libcd, already integrated), GS_123/PRESET/ + PRESET2/OBJT/OBJT2(libgs), PRNT(libc2), LIBMCRD. **STAY RAW** (untouched flat data → byte-identical). +- **+24 root cause = 6 library jtbls, ALL in 800b** (PRESET_OBJ_744/8FC, PRESET2_OBJ_4D8/A88, OBJT2_OBJ_614, PRNT_OBJ_24C). + LZSS+game jtbls are in 800; the only libgs object in 800 (GS_013, 0x8003D40C) has **no jtbl** → 800-only migration has no library TU → no +24. + +### libgs placement (session E — preserved for the BONUS task #9) +36/201 located; byte-test disambiguation: **PRESET3** (not PRESET2) @0x80055D40, **OBJT3** (not OBJT2) @0x80057094 +(losers have a spurious .data + .text mismatch); GS_131≡RVWUNIT, GS_137≡RVWLUNIT are .text-identical aliases (keep GS_*). +GS_106 @0x80053308 fills a gap (narrow-window). GS_013 @0x8003D40C = far outlier in 800 (6 ins). 4 ambiguous +(GS_101/102/124/125) don't fit gaps → not linked. Main block 0x80051804–0x80057928 (800b), ~4 sub-blocks; gaps 96/48/304 B. +Patched `tools/psyq_identify.py` to skip data-only objects (no .text, e.g. GLOBAL.o). PRESET3/OBJT3 .rdata land in-island (0x80073c98/0x80073ee8). +- **libgs BYTE-VERIFIED (session E):** curated `.run/obj40/libgs_used/` (31 objects). `psyq_link_region … --emit .run/libgs_region` + → **31/32 byte-identical**, 0 conflicts, 69 externals; `.run/libgs_region.{ld,syms}` emitted. Confirms BFM links real + PsyQ libgs 4.0 objects byte-for-byte. **GS_001 EXCLUDED** (known-issue: psyq-obj-parser packs scattered PSD* commons + into `.bss` referenced via `.bss`+offset; 35 words differ at the global-zeroing run — needs per-symbol .bss resolution, + cookbook §9.1 hard case). Block structure: **6 contiguous blocks + 5 gaps** (80/48/1536[GS_001]/48/304 B, all non-libgs + → stay stubs). REMAINING (mechanical): resegment splat 800b into the 6 blocks + 5 gap stubs (~13 subsegs) + split_src + + wire `psyq_integrate` (6 block stubs) → byte-identical → banks ~63 SDK funcs. Simpler first win = block 6 alone (16 objs + incl. PRESET3/OBJT3) as [pre][libgs6][post] → ~35 funcs. Deferred at session-E end (long session); clean continuation. ## Per-session `make check` green log (≥3 sessions needed for the milestone) - 2026-06-14 (session A): `make check` → `143dbb89… BYTE-IDENTICAL` ✓ — baseline restored + reproducibility fix, reports built, **38 real matches** (22 accessor leaves + ResourceGetCdLoc + LoaderResetReadState), build byte-identical throughout. [need ≥2 more sessions] - 2026-06-14 (session B): `make check` → `143dbb89… BYTE-IDENTICAL` ✓ — **per-file -O0 split mechanism** (src/boot.c + Makefile per-file flags); **4 real matches** (GameModeDispatch, DebugMenuHandler, CdQueueBusy, CdReadRequest) → **42 real**; **PsyQ libcd.h infra** (CdlLOC/CdlFILE + 4 named symbols, unlocks the loader cluster); **LoaderInitFileTable + ResourceLoadStateMachine NON_MATCHING-drafted** (→ 4 NM) — **Task 5 non-jtbl loaders COMPLETE** (6 matched + 2 drafted); report tooling fixed (multi-file); cookbook §6/§7/T4. Build byte-identical throughout. [need ≥1 more session] - 2026-06-15 (session C): `make check` → `143dbb89… BYTE-IDENTICAL` ✓ (full `clean && extract && build`, restored after the Task-2′ experiments). **≥3-session bar MET.** This session: fully diagnosed + built the **rodata-island mechanism** (works); root-caused the +24; **proved the PsyQ-library-linking GO** (see below). No new matches (architectural session). Build green at start and after restore. - 2026-06-14 (session D): `make clean && make extract && make build && make check` → `143dbb89… BYTE-IDENTICAL` ✓ — **libcd LINKED INTO THE BUILD** (Drew-approved push-through). The first real PsyQ library is now sourced from real SDK objects in the byte-identical build: **58 libcd SDK functions** linked (not stubs), replacing the libcd-region asm stubs. Idempotent; Makefile-automated; conditional (fresh clone w/o `tools/psyq/` builds via stubs). New committed tooling: `tools/psyq_link.py` (per-object byte-link engine — recovers externals from resolved relocs, weakens psyq-obj-parser's mislabelled `.bss` commons), `psyq_link_lib.py` (whole-lib verify, 18/18 libcd), `psyq_link_region.py` (region link via **NOLOAD** = no data carving), `psyq_integrate.py` (build wiring: splat resegment + .ld swap + external resolution), `split_src_region.py` (H5-safe src split). Cookbook §9.1/9.2/9.3 + R16. **No data carving** (NOLOAD data placement; flat data subseg unchanged). Build green throughout. +- 2026-06-15 (session E): `make clean && make extract && make build && make check` → `143dbb89… BYTE-IDENTICAL` ✓ — start-of-session baseline + after the **LZSS surgical rodata carve** (jtbl_80072A38 migrated + sandwiched) + LZSS C **NON_MATCHING-guarded** (default build = stub). **≥3-session bar already MET (A/B/C/D); E is margin.** Findings: lean LZSS path proven (no libgs needed for it); LZSS C structurally matches (111/122) but blocked on a gcc cross-jump-merge hard-tail (see LZSS block above); libgs placement+disambiguation done (bonus, task #9, deferred per Drew until before PhaseEnd). --- @@ -128,7 +173,13 @@ migration+ld_interleave path above. - LZSS via the proven migration+ld_interleave path (independent of the lib pivot). ## Notes -- Commits accumulate UNCOMMITTED; one phase-end commit by the developer (R8/R6). +- **Rule candidate (PhaseEnd, Drew-flagged session E):** web-research the compiler internals (real compiler + source e.g. `pmret/gcc-papermario`) + decomp community for **compiler-quirk residuals** (cross-jump, + scheduling, regalloc) — a proven escalation tier above the permuter, below decomp.me. Found the LZSS + cross-jump barrier. Captured in cookbook §3a/§5a + memory `web-research-compiler-quirks`. +- **Session-E checkpoint commit:** Drew directed a checkpoint commit (deviates from R8's strict + one-commit-at-phase-end; consistent with the session A–D checkpoint commits in the git log). Commit local in + WSL (no push, no Co-Authored-By per R5); Drew pushes via GitHub Desktop (R6). - `.run/merge_matches.py` = regenerate-800.c + re-apply-matches helper. **H5 caveat:** regen-fresh drops file-level/stub comments; for the permanent 800.c, surgically insert the 70 INCLUDE_RODATA lines instead. - `tools/ld_interleave.py` (committed) = the `.data→.rodata→.data` linker-script interleaver. diff --git a/src/800.c b/src/800.c index 6a1a665100..c6acc3dd77 100644 --- a/src/800.c +++ b/src/800.c @@ -432,7 +432,138 @@ INCLUDE_ASM("asm/nonmatchings/800", func_800184F0); INCLUDE_ASM("asm/nonmatchings/800", func_80018714); +/* LZSS streaming sector decompressor (resumable coroutine state machine). + * Decodes up to 0x800 input bytes per call out of an LZSS stream into a 0x400-byte + * scratchpad ring (lzss_ringBuffer @ 0x1F800000) and to lzss_outPtr. lzss_state is the + * resume point (0=idle/done, 1=fresh, 2=token loop, 3=have code low byte, 4=advance bit). + * Returns 1 if the input sector was consumed mid-stream (resume next call), 0 at the + * stream terminator (back-ref offset 0) or when idle. */ +#ifdef NON_MATCHING +/* CROSS-JUMP BARRIER BREAKTHROUGH (Phase 7): the lone `__asm__ __volatile__("" ::: "memory")` in the + * state-2 reload save (below) is a load-bearing, ZERO-BYTE cross-jump barrier. gcc 2.7.2 -O2's jump.c + * cross-jumping (find_cross_jump) would otherwise MERGE the two byte-identical state-save tails (the + * state-3/4 `save:` block + the state-2 reload save) into one — making the function 111 instructions + * instead of the original 122. A volatile-asm node (ASM_INPUT) makes find_cross_jump bail (lose=1), so + * both saves survive; the empty asm emits no machine code. (No -fno-crossjumping in gcc 2.7.2 — it + * arrived in gcc 3.3.) This defeats the STRUCTURAL blocker; the function is now 122 instructions (the + * correct count, verified vs the original). + * NON_MATCHING residual = ~3 register-allocation / scheduling diffs only: the high-byte `or` result + * lands in v1 vs target v0; the state-2 store value in v0 vs v1; and `li v0,1` (return value) is placed + * late in a merged epilogue vs distributed per-save in the target. These are the last-mile regalloc + * that decomp-permuter normally finishes — but it can't parse the asm barrier (pycparser) and its + * object-mode score has a `.rodata`-vs-`jtbl_80072A38` floor (cosmetic; links identically). Re-enable + * (drop the guard) when the regalloc closes. The surgical rodata carve placing jtbl_80072A38 at + * 0x80072A38 via the .data->.rodata->.data sandwich IS byte-identical and stays live in the stub build. */ +extern u8 lzss_curMask; /* 0x800747A0 */ +extern u8 lzss_curToken; /* 0x800747A4 */ +extern u8 *lzss_outPtr; /* 0x800747AC */ +extern u32 lzss_ringIndex; /* 0x800747B0 */ +extern u16 lzss_partialCode; /* 0x800747B4 */ +extern u32 lzss_state; /* 0x800C7D24 */ + +s32 LzssDecodeSector(u8 *src) { + s32 count = 0x800; + u32 ringIdx = lzss_ringIndex; + u8 mask = lzss_curMask; + u8 token = lzss_curToken; + u8 *out = lzss_outPtr; + u16 code = lzss_partialCode; + u8 *ring = (u8 *)0x1F800000; /* scratchpad ring; original uses the literal (lui 0x1f80), not the symbol */ + u32 readIdx; + s32 len; + u8 b; + u8 nb; + u8 cb; + s32 newState; + + if (lzss_state >= 5) { + goto ret0; + } + switch (lzss_state) { + case 0: + goto term_ret; + case 1: + ringIdx = 1; + mask = 1; + token = *src++; + count--; + case 2: + for (;;) { + if (token & mask) { + b = *src++; + ring[ringIdx] = b; + ringIdx = (ringIdx + 1) & 0x3FF; + count--; + *out++ = b; + } else { + nb = *src++; + count--; + code = (code & 0xFF00) | nb; + if (count != 0) { + goto have_low; + } + newState = 3; + goto save; + have_low: + case 3: + nb = *src++; + count--; + code = (nb << 8) | (code & 0xFF); + readIdx = code & 0x3FF; + if (readIdx == 0) { + lzss_state = 0; + term_ret: + return 0; + } + len = (code >> 10) + 2; + while (len != 0) { + cb = ring[readIdx]; + readIdx = (readIdx + 1) & 0x3FF; + ring[ringIdx] = cb; + ringIdx = (ringIdx + 1) & 0x3FF; + *out++ = cb; + len--; + } + } + if (count != 0) { + goto next_bit; + } + newState = 4; + save: + lzss_state = newState; + lzss_ringIndex = ringIdx; + lzss_curMask = mask; + lzss_curToken = token; + lzss_outPtr = out; + lzss_partialCode = code; + return 1; + next_bit: + case 4: + if (mask == 0x80) { + mask = 1; + token = *src++; + count--; + if (count == 0) { + lzss_state = 2; + lzss_ringIndex = ringIdx; + lzss_curMask = mask; + lzss_curToken = token; + lzss_outPtr = out; + lzss_partialCode = code; + __asm__ __volatile__("" ::: "memory"); /* zero-byte cross-jump barrier — see header */ + return 1; + } + } else { + mask <<= 1; + } + } + } +ret0: + return 0; +} +#else INCLUDE_ASM("asm/nonmatchings/800", LzssDecodeSector); +#endif INCLUDE_ASM("asm/nonmatchings/800", func_80018918); diff --git a/tools/ld_interleave.py b/tools/ld_interleave.py index f230b717fd..2def49c345 100644 --- a/tools/ld_interleave.py +++ b/tools/ld_interleave.py @@ -4,12 +4,14 @@ splat emits one output section (`.main`) section-major in `section_order` (.rodata, .text, .data, .bss), which floats ALL rodata to one place. But this -EXE's real layout is: +EXE's real layout puts .data on BOTH sides of the compiler rodata. For the +SURGICAL LZSS carve (Phase 7), only jtbl_80072A38 is migrated to .rodata; the +rest of the island stays raw inside the tail data object: .text 0x80010000 .. 0x800629DC - .data (front) 0x800629DC .. 0x80072A38 (globals, hand-written ptr tables) - .rodata (island) 0x80072A38 .. 0x80074750 (gcc jtbl_* switch tables + consts) - .data (tail) 0x80074750 .. 0x80074800 (gp base; zero small-data) + .data (front) 0x800629DC .. 0x80072A38 531DC.data.o (globals, ptr tables) + .rodata 0x80072A38 .. 0x80072A4C 800.o (ONLY the migrated LZSS jtbl_80072A38) + .data (tail) 0x80072A4C .. 0x80074800 6324C.data.o (rest of island raw + tail globals) i.e. .data appears on BOTH sides of .rodata, which a single section_order can't express. This script rewrites the `.main {...}` body to the interleaved order: @@ -25,7 +27,7 @@ import re, sys LD = sys.argv[1] if len(sys.argv) > 1 else "build/us/SLUS_007.26.ld" # object basenames whose (.data) belongs to the front / tail region FRONT_DATA = ("531DC.data.o",) -TAIL_DATA = ("64F50.data.o",) +TAIL_DATA = ("6324C.data.o",) src = open(LD).read() diff --git a/tools/psyq_identify.py b/tools/psyq_identify.py index 75e60ce12e..16f2107449 100644 --- a/tools/psyq_identify.py +++ b/tools/psyq_identify.py @@ -24,9 +24,16 @@ text = b[TLO - VRAM_BASE: THI - VRAM_BASE] twords = [struct.unpack_from("