feat(phase-7): PsyQ library-linking proven byte-exact + rodata-island mechanism — session C checkpoint

- rodata-island mechanism PROVEN: migration (dotted .rodata sibling named "800")
  + tools/ld_interleave.py (.data->.rodata->.data linker placement) + data-in-text
  carve; rodata/data sizes come out byte-exact. +24 residual root-caused to the
  .align-3 jumptable file-split padding (6 sites; spim suggests 8 splits)
- PsyQ 4.0 library linking PROVEN end-to-end: libcd SYS.o (483 instrs, full TU + 21
  externals + internal .rdata/.data relocs) links BYTE-IDENTICAL to BFM
  - tools/psyq_lib_split.py: split Sony LIB\x01 archive -> member .OBJ
  - tools/psyq_build_libs.sh: .LIB -> psyq-obj-parser -> ELF .a (14 BFM libs)
  - tools/psyq_identify.py: relocation-masked search -> object link addresses
    (libcd 18/25 used, contiguous from 0x80043088)
  - external symbols recovered from the EXE's own resolved relocations
  - section-alignment 8->4 fix (psyq-obj-parser over-aligns; the +4 mismatch is the tell)
- cookbook §8 (rodata island) + §9 (PsyQ library-linking recipe); CURRENT_PHASE full
  diagnosis, asset inventory, and resume plan
- 3rd consecutive green session (143dbb89 BYTE-IDENTICAL) -> Gen1 >=3-session bar MET
- SDK assets (PsyQ 4.0 USA DTL-S2002 libs + psyq-obj-parser) staged gitignored under
  tools/psyq/; no function matches added (architectural session)
This commit is contained in:
Drew T
2026-06-14 21:24:10 -06:00
parent 8ffb7f0607
commit 26b71c5b2d
6 changed files with 405 additions and 15 deletions
+53
View File
@@ -207,3 +207,56 @@ A function that calls PsyQ library routines needs both the SDK **types** and the
- Done Phase 7: `include/psyq/libcd.h` (CdlLOC 4B, CdlFILE 24B + CdSearchFile/CdPosToInt/CdIntToPos
protos) + the 4 libcd/libetc symbols — unlocks the file-loader cluster. Same pattern for
libgpu/libgte/libspu as they come up.
---
## §8 rodata island (compiler jump tables) — the `.data→.rodata→.data` sandwich (Phase 7)
GCC emits each `switch` jump table into `.rodata`; in this EXE all compiler rodata is ONE island at
0x80072A38–0x80074750, sitting BETWEEN the front `.data` (globals @0x800629DC) and the tail `.data`
(@0x80074750). No single splat `section_order` expresses data→rodata→data. Proven mechanism (session C):
- **Migrate, don't standalone.** A jtbl `.word`s reference function-internal `.L`/`jlabel` targets, so a
separate rodata object can't link — the table MUST co-locate in its function's object. Use a **dotted
`.rodata` subseg whose NAME matches the code subseg** (`[<off>, .rodata, 800]`): `extract=False`, spimdisasm
migrates each single-ref jtbl/const into `asm/nonmatchings/<seg>/<fn>.s` as `.section .rodata`. The
INCLUDE_ASM stub already `.include`s that `.s`, so it flows into the object for free. Multi-ref rodata can't
migrate → splat emits `INCLUDE_RODATA(...)` lines (in a FRESH `.c`). **H5:** don't regen-fresh the curated
`.c` (drops comments) — surgically INSERT just the INCLUDE_RODATA lines.
- **Place explicitly.** splat is section-major (floats all `.rodata` to the front). `tools/ld_interleave.py`
(wired into `make extract`) rewrites the `.main {}` body to text → front-`.data` → `.rodata` → tail-`.data` →
bss, splitting front/tail by object basename. Sizes then land byte-exact.
- **Carve data-in-text.** A trailing non-code table inside the text range (here 0x80062998–0x800629DC) must be
its own `data` subseg, or jumptable analysis mis-extends the last function across it (the +24 `main_TEXT_END`
overrun's first cause).
- **The `.align 3` file-split trap:** GCC 8-aligns jtbls; concatenating many functions into one object injects
padding nops the original (separate TUs) lacked → image grows. spimdisasm PRINTS file-split suggestions at
the misaligned jtbls. Fix = per-file split at those boundaries (sotn-style) — OR link the real library
object (§9) when the owning function is SDK code.
## §9 Link real PsyQ library objects byte-exact (Phase 7 — GO proven)
~350 of BFM's functions are unmodified PsyQ 4.0 SDK code. They are **byte-identical to the real PsyQ library
objects**, so link them directly instead of hand-decompiling — and each library `.o` brings its own correct
alignment (dissolving the library-half of §8's `.align 3` problem). Validated: `CdPosToInt`/`CdIntToPos` EXACT
vs PsyQ libcd; `PRESET_OBJ_*` ∈ `LIBGS.LIB`. Workflow (the decomp-standard psyq-obj-parser path):
- **Tools** (gitignored `tools/psyq/`): `psyq-obj-parser` (decompme prebuilt — `.OBJ`→ELF; rejects `.LIB`),
`lib40/*.LIB` = PsyQ **4.0 USA** libraries (DTL-S2002 R2.0 = BFM's version; extracted from the redump ISO via
`tools/bfm_extract/iso9660.py`). Identify a function's library by searching the `.LIB` for a NON-relocated
instruction run from its EXE bytes (relocated runs false-negative — use leaves or interior runs).
- **Integration:** split `.LIB` (LIB\x01 archive) → `.OBJ` → `psyq-obj-parser` → `ar` per lib → link the `.o`
for each library function and drop its INCLUDE_ASM. BFM mixes 4.0+4.2 library stamps, so a few objects may
need 4.2/4.3 libs — determine per-object by the byte test.
- **Proven full-object link recipe (SYS.o byte-identical to BFM, Phase 7):**
1. **Placement** — `tools/psyq_identify.py <elf_dir>`: relocation-masked search finds each object's `.text`
vram in the EXE. Per library the used objects are CONTIGUOUS in object order → place the first at the
region base, link the rest in order.
2. **Recover externals** — symbols the object references but doesn't define are usually absent from
`symbols.us.txt`; read them straight out of the EXE's RESOLVED relocations: for each reloc, `R_MIPS_26` →
`target = ((word&0x3FFFFFF)<<2)|(pc&0xF0000000)`; an `HI16`+`LO16` pair → `(hi<<16)+signext(lo)`. Feed as
`ld --defsym NAME=0xADDR`.
3. **Alignment** — psyq-obj-parser emits `.text/.rdata/.data` at align 2**3; the original is 4-aligned, so an
8-align bumps the section +4 (the tell: every `LO16` to that section is off by +4). Fix:
`objcopy --set-section-alignment '.rdata=4' --set-section-alignment '.data=4' obj.o obj_a.o` before linking.
4. **Link + verify** — `ld -T <SECTIONS: . = <text vram>; .text:{*(.text)} . = <island>; .rdata:{*(.rodata)
*(.rdata)} . = <data vram>; .data:{*(.data)}> --defsym … obj_a.o` → `objcopy -O binary --only-section
.text` → byte-compare to the EXE. `.rdata`/`.data` vrams are found by searching the EXE for the section
bytes (`objcopy --only-section`). Tools: `tools/psyq_lib_split.py`, `tools/psyq_build_libs.sh`,
`tools/psyq_identify.py`.
+80 -15
View File
@@ -34,30 +34,95 @@ rodata-island foundation + LZSS match are DEFERRED to a focused sub-project afte
- [ ] **Task 7 — PhaseEnd_Phase7** (Gen1 synthesis, milestone gate). **Max · Tier 1.**
## Current task
**Task 5 — Loader cluster** (match-tractable non-switch / draft-hard). Then Task 2′ (LZSS), Task 6, Task 7.
NOTE: Gen1 exit needs ≥3 SESSIONS of green `make check` — cannot complete this session regardless.
**Task 2′ — rodata-island / LZSS gate.** Mechanism now PROVEN end-to-end; a strategic pivot to PsyQ-library
linking is in flight (see the PsyQ spike section below) — that fixes the library-half alignment AND gives
~350 SDK functions byte-exact for free. Then LZSS, Task 6, Task 7.
NOTE: Gen1 exit needs ≥3 SESSIONS of green `make check` — **now satisfied** (A, B, C below); Tasks 6/7 still pending.
## Per-session `make check` green log (≥3 sessions needed for the milestone)
- 2026-06-14 (session A): `make check` → `143dbb89… BYTE-IDENTICAL` ✓ — baseline restored + reproducibility fix, reports built, **38 real matches** (22 accessor leaves + ResourceGetCdLoc + LoaderResetReadState), build byte-identical throughout. [need ≥2 more sessions]
- 2026-06-14 (session B): `make check` → `143dbb89… BYTE-IDENTICAL` ✓ — **per-file -O0 split mechanism** (src/boot.c + Makefile per-file flags); **4 real matches** (GameModeDispatch, DebugMenuHandler, CdQueueBusy, CdReadRequest) → **42 real**; **PsyQ libcd.h infra** (CdlLOC/CdlFILE + 4 named symbols, unlocks the loader cluster); **LoaderInitFileTable + ResourceLoadStateMachine NON_MATCHING-drafted** (→ 4 NM) — **Task 5 non-jtbl loaders COMPLETE** (6 matched + 2 drafted); report tooling fixed (multi-file); cookbook §6/§7/T4. Build byte-identical throughout. [need ≥1 more session]
- 2026-06-15 (session C): `make check` → `143dbb89… BYTE-IDENTICAL` ✓ (full `clean && extract && build`, restored after the Task-2′ experiments). **≥3-session bar MET.** This session: fully diagnosed + built the **rodata-island mechanism** (works); root-caused the +24; **proved the PsyQ-library-linking GO** (see below). No new matches (architectural session). Build green at start and after restore.
---
## Rodata-island foundation — investigation findings (DEFERRED, for Task 2′)
Attempted the 3-way data split (`[data front][.rodata island][data tail]` + `ld_legacy_generation: True`).
**What works:** dotted `.rodata` sibling named `800` → spimdisasm MIGRATES each jump table into its owning
function's `.s` (`jtbl_80072A38` lands inside `LzssDecodeSector.s` with `.section .rodata`/`.section .text`),
references resolve intra-800.o, build LINKS. Non-migrated multi-ref rodata becomes `INCLUDE_RODATA` (70 lines)
in a FRESH-regenerated `800.c`.
**What blocks byte-identity (the structural wall):**
1. Adding any `rodata` subseg turns on global jumptable analysis → merges 56 over-split switch fragments (good) but DROPS an 8-byte inter-fn blob (the `func_80047CAC` issue, now fixed via explicit symbol).
2. A separate rodata object can't link to text-local `.L`/`jlabel` jumptable targets → must migrate (co-locate).
3. splat places sections CONTIGUOUSLY (no explicit `. = addr`), so any size drift shifts the whole image. Observed a **24-byte `.text` overrun** (`main_TEXT_END` 0x800629F4 vs 0x800629DC) → +24B size, 85288 bytes differ. Root cause of the 24B NOT fully pinned.
**Candidate fix for Task 2′:** explicit linker addresses (Makefile post-extract `.ld`-patch placing `.text`@0x80010000, front `.data`@0x800629DC, `800.o(.rodata)`@0x80072A38, tail `.data`@0x80074750) + scope migration so only island tables land in `800.o(.rodata)`. Alternative: sotn-style per-file split (Gen2-scale). Reproduce with the migration config (in git stash / reconstruct from this log).
## Rodata-island foundation — RESOLVED end-to-end (session C, 2026-06-15)
The mechanism now WORKS; only the +24 (a known file-split/alignment artifact) blocks full byte-identity, and
the PsyQ-lib pivot (below) is the chosen fix. Experimental configs saved: `.run/{splat.island.yaml,
symbols.island.txt,Makefile.island,800.c.fresh-throwaway}`; new committed tool `tools/ld_interleave.py`.
**Proven mechanism (3 parts, all validated this session):**
1. **Migration** — dotted `.rodata` sibling named `800` (`[0x63238, .rodata, 800]`) → spimdisasm migrates all
50 jtbls + single-ref consts into their owning `asm/nonmatchings/800/<fn>.s` (99 .s got `.section .rodata`,
`.L`-refs resolve intra-object). Multi-ref consts → 70 `INCLUDE_RODATA` lines (needs a FRESH `800.c`; for the
PERMANENT file, surgically INSERT those 70 lines into the curated 800.c — do NOT regen-fresh, it drops
comments/H5). All 301 jtbl targets ∈ the `800` text seg (none in `boot`), so the single sibling is correct.
2. **Placement** — `tools/ld_interleave.py` (wired into `make extract`) rewrites splat's section-major `.main`
into the real `.data→.rodata→.data` sandwich order (text, front-data@0x800629DC, rodata@0x80072A38,
tail-data@0x80074750). Front/tail split by object basename. rodata + both data sizes came out **byte-exact**.
3. **Data-in-text carve** — the 68-B descriptor table at 0x80062998–0x800629DC (ptrs to start/D_80062998/
D_80074778) must be its own `[0x53198, data, 53198]` subseg or the jumptable analyzer mis-extends the last
code function across it.
**The ONLY residual = +24 (ROOT-CAUSED):** 6× `.align 3` jumptable padding nops injected into `.text` (at
PRESET_OBJ_744, PRESET_OBJ_8FC, PRESET2_OBJ_4D8, PRESET2_OBJ_A88, OBJT2_OBJ_614, PRNT_OBJ_24C). GCC 8-aligns
each switch jtbl, but the original built these as SEPARATE translation units; our single 800.o concatenation
adds padding the original lacked. spimdisasm itself printed **8 file-split suggestions** (rodata 0x6324C,
0x63388, 0x633FC, 0x63920, 0x63C94, 0x64420, 0x64AB4, 0x64CA0). Canonical fix = per-file split — OR the PsyQ-lib
pivot below (most of these are library code).
## PsyQ-library-linking SPIKE — GO PROVEN (session C, the chosen +24 fix + free SDK code)
**Finding:** BFM's PsyQ library functions are **byte-identical to the real PsyQ SDK objects** → link them
directly (byte-exact) instead of hand-decompiling, which ALSO gives each library `.o` correct per-object
alignment (dissolving the library-half of the +24). Validated: `CdPosToInt` (32 instrs) + `CdIntToPos` (65)
EXACT vs PsyQ **4.7** `libcd.a`; `CdPosToInt` also in **4.0** `LIBCD.LIB`; `PRESET_OBJ_108` (a +24 culprit) is
in **4.0 `LIBGS.LIB`** → library code, not game. (Raw-byte lib search has false-negatives on relocated funcs,
e.g. `_spu_FsetPCR`/`OBJT2`/`PRNT` "missed" — needs the real ELF-link test to classify those.)
**Assets staged (gitignored `tools/psyq/`):** `psyq-obj-parser` (decompme prebuilt, works on `.OBJ`; rejects
`.LIB` archives — needs splitting); `psyq4.0/` (4.0 tools: CC1PSX/ASPSX/PSYLIB/…); `conv47/` (4.7 pre-converted
ELF `.a` — quick reference); **`lib40/*.LIB`** = the 20 PsyQ **4.0 USA** libraries (DTL-S2002 R2.0, BFM's exact
version) extracted from the redump via our `tools/bfm_extract/iso9660.py` walker. Footprint in BFM: ~350 funcs
(libsnd 131, libapi/gs 69+, libmcrd 63, libsn 37, libcd 28, libspu 21, …) of 2050 matchable.
**Integration pipeline — BUILT + PROVEN end-to-end (session C):**
- `tools/psyq_lib_split.py` (committed) — splits a `LIB\x01` archive into its member `.OBJ` (locates each
member header by the invariant `u32@(header+12) == LNK_offset − header`). LIBCD → 25 objects ✓.
- `tools/psyq_build_libs.sh` (committed) — `.LIB → .OBJ → psyq-obj-parser → ELF .o → ar` per lib. **Built all
14 BFM libs → `tools/psyq/lib40_elf/*.a`** (gitignored): LIBCD 25, LIBGS 201, LIBSPU 129, LIBSND 163,
LIBMCRD 2, LIBSN 51, LIBAPI 90, LIBETC 7, LIBGTE 381, LIBGPU 12, LIBMATH 48, LIBCARD 18, LIBC 56, LIBC2 46.
- **Byte-match PROVEN at object level:** in libcd `SYS.o`, leaves `CdPosToInt`/`CdIntToPos` are EXACT; relocated
funcs (`CdComstr` …) differ ONLY at their relocation sites → link byte-exact once relocs resolve to BFM
symbol addrs. So every step (split, convert, leaf-match, reloc-resolve) is validated.
**FULL OBJECT LINK — PROVEN BYTE-IDENTICAL (session C):** libcd `SYS.o` (483 instrs, the full TU: leaves +
relocated funcs + 21 externals + internal .rdata/.data) links **byte-for-byte identical to BFM**. The pipeline
+ the 3 last pieces:
- **Identify placement** (`tools/psyq_identify.py`, committed) — relocation-masked search locates each object's
`.text` in BFM. libcd: **18/25 objects found, CONTIGUOUS** at 0x80043088–0x80046D1C in object order (the 7
unused — CDPLAY, C_012–015… — BFM doesn't link). So per-library placement = link the used objects in order at
the region base; addresses are read off, not guessed.
- **Recover externals from BFM** — symbols the object references but doesn't define (e.g. libcd's `CD_pos`,
`CD_com`, `DMACallback`) are NOT in symbols.us.txt, but their addresses are encoded in the EXE's already-
RESOLVED relocations: parse the object's reloc records, read BFM at each site, reconstruct (R_MIPS_26 →
target; HI16/LO16 pair → addr). Recovered all 21 for SYS.o. (Feeds symbols.us.txt over time.)
- **Alignment fix** — psyq-obj-parser sets `.text/.rdata/.data` align=2**3 (8); the original placed them
4-aligned, so an 8-align bumps them +4. `objcopy --set-section-alignment .rdata=4 .data=4` before linking →
exact. (The +4 mismatch is the tell.)
- **Link recipe:** `ld -T <script placing .text@<objaddr> .rdata@<island> .data@<addr>> --defsym <recovered…>
obj.o` → objcopy .text → byte-compare. Proven on SYS.o.
**REMAINING (replication + wiring, NEXT session, Task #5):** generalize the SYS.o recipe to all used objects
per library (identify region → recover externals → set-align → place sections in order → link), then wire into
the build: drop the linked functions' INCLUDE_ASM + carve their raw data, add the lib objects to the link via a
generated `.ld` fragment. ~350 SDK funcs become byte-exact + the library-region jtbl alignment resolves. GAME
switches + **LZSS** (jtbl_80072A38 = island's first entry, before any misalignment) take the proven
migration+ld_interleave path above.
## Blockers / open items
- Task 2′ ld-placement mechanism (above). MCP-mode batching: sig-refresh needs MCP stopped, LZSS needs MCP live.
- `.LIB`→`.OBJ` splitter (LIB\x01 format) — the gate for the lib-linking integration.
- Which 4.x version matches each BFM lib object best (4.0 USA primary; BFM mixes 4.0+4.2 stamps, so some objects
may need 4.2/4.3 libs — determine per-lib during integration via the ELF-link byte test).
- LZSS via the proven migration+ld_interleave path (independent of the lib pivot).
## Notes
- Commits accumulate UNCOMMITTED; one phase-end commit by the developer (R8/R6).
- `.run/merge_matches.py` = the regenerate-800.c + re-apply-matches helper (reusable for Task 2′).
- `.run/merge_matches.py` = regenerate-800.c + re-apply-matches helper. **H5 caveat:** regen-fresh drops
file-level/stub comments; for the permanent 800.c, surgically insert the 70 INCLUDE_RODATA lines instead.
- `tools/ld_interleave.py` (committed) = the `.data→.rodata→.data` linker-script interleaver.
+80
View File
@@ -0,0 +1,80 @@
#!/usr/bin/env python3
"""Reorder splat's generated linker script to honour a .data -> .rodata -> .data
"sandwich" layout (Phase 7 Task 2', the rodata-island problem).
splat emits one output section (`.main`) section-major in `section_order`
(.rodata, .text, .data, .bss), which floats ALL rodata to one place. But this
EXE's real layout is:
.text 0x80010000 .. 0x800629DC
.data (front) 0x800629DC .. 0x80072A38 (globals, hand-written ptr tables)
.rodata (island) 0x80072A38 .. 0x80074750 (gcc jtbl_* switch tables + consts)
.data (tail) 0x80074750 .. 0x80074800 (gp base; zero small-data)
i.e. .data appears on BOTH sides of .rodata, which a single section_order can't
express. This script rewrites the `.main {...}` body to the interleaved order:
text -> front .data -> .rodata -> tail .data -> .bss, keeping splat's START/END/
SIZE symbols. Front vs tail .data is decided by object basename (FRONT_DATA /
TAIL_DATA). All other (empty) .data objects go in the front group.
Idempotent: keyed off splat's exact section-major output; re-running on an
already-patched script is a no-op (the markers won't match). Run post-extract.
"""
import re, sys
LD = sys.argv[1] if len(sys.argv) > 1 else "build/us/SLUS_007.26.ld"
# object basenames whose (.data) belongs to the front / tail region
FRONT_DATA = ("531DC.data.o",)
TAIL_DATA = ("64F50.data.o",)
src = open(LD).read()
# Grab the .main output-section body (between its first '{' and matching '}').
m = re.search(r"(\.main\b.*?\n[ \t]*\{\n)(.*?)(\n[ \t]*\})", src, re.S)
if not m:
sys.exit("ld_interleave: could not find .main { ... } block")
head, body, tail = m.group(1), m.group(2), m.group(3)
# Collect the object input-section lines by linker section, preserving order.
def grab(section):
# lines like: build/src/800.o(.rodata);
return re.findall(rf"^[ \t]*build/\S+\({re.escape(section)}\);", body, re.M)
text_lines = grab(".text")
rodata_lines = grab(".rodata")
data_lines = grab(".data")
bss_lines = grab(".bss")
def is_named(line, names):
return any(n in line for n in names)
front_data = [l for l in data_lines if not is_named(l, TAIL_DATA)]
tail_data = [l for l in data_lines if is_named(l, TAIL_DATA)]
# Sanity: front must contain the FRONT_DATA object.
if not any(is_named(l, FRONT_DATA) for l in front_data):
sys.exit("ld_interleave: front data object not found — config drift?")
if not tail_data:
sys.exit("ld_interleave: tail data object not found — config drift?")
I = " " # 8-space indent matching splat's body
def grp(start, lines, end_sym, size_sym):
out = [f"{I}{start} = .;"]
out += [f"{I}{l.strip()}" for l in lines]
out += [f"{I}. = ALIGN(., 4);", f"{I}{end_sym} = .;"]
if size_sym:
out += [f"{I}{size_sym} = ABSOLUTE({end_sym} - {start});"]
return out
new = [f"{I}FILL(0x00000000);"]
new += grp("main_TEXT_START", text_lines, "main_TEXT_END", "main_TEXT_SIZE")
new += grp("main_DATA_START", front_data, "main_DATA_END", "main_DATA_SIZE")
new += grp("main_RODATA_START", rodata_lines, "main_RODATA_END", "main_RODATA_SIZE")
new += grp("main_DATA2_START", tail_data, "main_DATA2_END", "main_DATA2_SIZE")
new += grp("main_BSS_START", bss_lines, "main_BSS_END", "main_BSS_SIZE")
new_body = "\n".join(new)
out = src[:m.start()] + head + new_body + tail + src[m.end():]
open(LD, "w").write(out)
print(f"ld_interleave: rewrote .main — text={len(text_lines)} "
f"front_data={len(front_data)} rodata={len(rodata_lines)} "
f"tail_data={len(tail_data)} bss={len(bss_lines)}")
+33
View File
@@ -0,0 +1,33 @@
#!/usr/bin/env bash
# Build ELF .a archives from the PsyQ 4.0 .LIB files (Phase 7 — PsyQ-library linking).
# Pipeline per lib: psyq_lib_split.py (.LIB -> .OBJ members) -> psyq-obj-parser (each
# .OBJ -> ELF .o) -> ar (-> <lib>.a). Output (gitignored, SDK-derived):
# tools/psyq/lib40_elf/<LIB>.a + .run/obj40/<lib>/*.{obj,o} scratch
# The .a are then linked into the build for the functions identified as PsyQ SDK code.
# Requires: tools/psyq/lib40/*.LIB (extracted from the DTL-S2002 redump), tools/psyq/
# psyq-obj-parser, mipsel-linux-gnu-ar. SDK assets stay out of git (.gitignore /tools/psyq/).
set -euo pipefail
cd "$(dirname "$0")/.."
LIBDIR=tools/psyq/lib40
OUTDIR=tools/psyq/lib40_elf
SCRATCH=.run/obj40
PARSER=tools/psyq/psyq-obj-parser
AR=mipsel-linux-gnu-ar
mkdir -p "$OUTDIR"
# BFM-relevant libraries (the SDK footprint in the EXE); pass args to override.
LIBS=("${@:-LIBCD LIBGS LIBSPU LIBSND LIBMCRD LIBSN LIBAPI LIBETC LIBGTE LIBGPU LIBMATH LIBCARD LIBC LIBC2}")
for L in ${LIBS[@]}; do
lib="$LIBDIR/$L.LIB"
[ -f "$lib" ] || { echo " skip $L (no $lib)"; continue; }
od="$SCRATCH/$(echo "$L" | tr A-Z a-z)"
rm -rf "$od"; mkdir -p "$od"
n=$(python3 tools/psyq_lib_split.py "$lib" "$od" | head -1 | grep -oE '[0-9]+ objects' | grep -oE '[0-9]+')
ok=0
for o in "$od"/*.obj; do
if "$PARSER" "$o" -o "${o%.obj}.o" >/dev/null 2>&1; then ok=$((ok+1)); fi
done
rm -f "$OUTDIR/$L.a"
$AR rcs "$OUTDIR/$L.a" "$od"/*.o 2>/dev/null || true
printf " %-10s %3s objs -> %3d .o -> %s.a\n" "$L" "${n:-?}" "$ok" "$L"
done
echo "done -> $OUTDIR/"
+84
View File
@@ -0,0 +1,84 @@
#!/usr/bin/env python3
"""Locate where PsyQ library objects are linked in the target EXE.
For each ELF .o (converted from a PsyQ .LIB member), extract its `.text` and the
relocation offsets, build a relocation-masked word pattern (relocated immediate
fields zeroed), and scan the EXE text for the single position where every
NON-relocated word matches. That position is the object's link address in the EXE
(or "absent" if the EXE doesn't link it). This is the placement map the library
linker step consumes.
Usage: psyq_identify.py <elf_dir> [text_lo_vram text_hi_vram]
(defaults to the BFM .text window 0x80010000..0x800629DC)
"""
import struct, subprocess, re, sys, glob, os
EXE = "extracted/retail/SLUS_007.26"
VRAM_BASE = 0x8000F800
ELF_DIR = sys.argv[1] if len(sys.argv) > 1 else ".run/obj40/libcd"
TLO = int(sys.argv[2], 0) if len(sys.argv) > 2 else 0x80010000
THI = int(sys.argv[3], 0) if len(sys.argv) > 3 else 0x800629DC
b = open(EXE, "rb").read()
text = b[TLO - VRAM_BASE: THI - VRAM_BASE]
twords = [struct.unpack_from("<I", text, i)[0] for i in range(0, len(text), 4)]
def obj_text_pattern(o):
"""Return (words, mask) for the object's .text; mask[i]=0 on relocated/jump words."""
d = subprocess.check_output(["mipsel-linux-gnu-objdump", "-dr", "-j", ".text", o],
text=True, stderr=subprocess.DEVNULL)
words, mask = [], []
pending_reloc = False
for line in d.splitlines():
mi = re.match(r"\s+([0-9a-f]+):\s+([0-9a-f]{8})\s", line)
if mi:
w = int(mi.group(2), 16)
words.append(w)
# mask jal/j (opcode 2/3) always (R_MIPS_26 target is link-resolved)
mask.append(0 if (w >> 26) in (2, 3) else 0xFFFFFFFF)
elif "R_MIPS" in line and words:
# relocation annotation follows its instruction line -> mask low 16 (hi/lo/pc16)
if "_26" in line:
mask[-1] = 0
else:
mask[-1] = 0xFFFF0000
return words, mask
def find(words, mask):
n = len(words)
if n < 2:
return None # too short to anchor uniquely
hits = []
for s in range(0, len(twords) - n + 1):
ok = True
for i in range(n):
if (twords[s + i] & mask[i]) != (words[i] & mask[i]):
ok = False; break
if ok:
hits.append(s)
if len(hits) > 1:
break
if len(hits) == 1:
return TLO + hits[0] * 4
return ("ambiguous" if len(hits) > 1 else None)
results = []
for o in sorted(glob.glob(os.path.join(ELF_DIR, "*.o"))):
words, mask = obj_text_pattern(o)
if not words:
results.append((os.path.basename(o), "no-.text", len(words))); continue
r = find(words, mask)
results.append((os.path.basename(o), r, len(words)))
found = [(n, a, l) for n, a, l in results if isinstance(a, int)]
found.sort(key=lambda x: x[1])
print(f"{ELF_DIR}: {len(found)}/{len(results)} objects located in EXE text "
f"[0x{TLO:X}..0x{THI:X}]")
for n, a, l in found:
print(f" 0x{a:08X} {n:18s} ({l} ins)")
absent = [n for n, a, l in results if a is None]
amb = [n for n, a, l in results if a == "ambiguous"]
if absent:
print(f" not linked by EXE ({len(absent)}): {', '.join(absent[:12])}{' …' if len(absent)>12 else ''}")
if amb:
print(f" ambiguous ({len(amb)}): {', '.join(amb)}")
+75
View File
@@ -0,0 +1,75 @@
#!/usr/bin/env python3
"""Split a Sony PSYLINK library archive (LIB\\x01) into its member LNK objects.
Format (derived from PsyQ 4.0 LIBCD.LIB, Phase 7):
"LIB\\x01"
per member:
name[8] space-padded object name (e.g. "CDROM ")
date[4] little-endian timestamp
obj_off[4] little-endian: bytes from THIS member header to its LNK object
<symbol dictionary> (obj_off - 16 bytes: the symbols this object exports)
<LNK object> starts with "LNK\\x02", runs to the next member header
We don't parse the LNK record stream to find each object's end; instead each member
header is located by the invariant u32@(header+12) == (LNK_offset - header) — i.e.
obj_off points exactly at that member's "LNK\\x02". An object therefore spans from its
"LNK\\x02" to the NEXT member header (or EOF for the last). psyq-obj-parser validates
the result downstream (a wrong boundary fails to parse), and the ultimate check is the
byte-match of the linked function against the target EXE.
Usage: psyq_lib_split.py <archive.LIB> <out_dir> -> writes <out_dir>/<NAME>.obj
"""
import struct, sys, os
LNK_MAGIC = b"LNK\x02"
def split(lib_path, out_dir):
d = open(lib_path, "rb").read()
if d[:4] != b"LIB\x01":
sys.exit(f"{lib_path}: not a LIB\\x01 archive (magic {d[:4]!r})")
# 1) all LNK object starts
lnk = []
i = d.find(LNK_MAGIC, 4)
while i != -1:
lnk.append(i)
i = d.find(LNK_MAGIC, i + 4)
if not lnk:
sys.exit(f"{lib_path}: no LNK objects found")
# 2) each LNK's member header: scan back for P with u32@P+12 == lnk-P
headers = []
for L in lnk:
P = None
for cand in range(L - 16, max(L - 0x8000, 0) - 1, -1):
if struct.unpack_from("<I", d, cand + 12)[0] == L - cand:
# sanity: name bytes printable-ish
nm = d[cand:cand + 8]
if all(32 <= c < 127 for c in nm.rstrip(b" \x00") or b" "):
P = cand; break
if P is None:
sys.exit(f"{lib_path}: could not locate member header for LNK@0x{L:x}")
headers.append(P)
# 3) extract: object k = [lnk[k] : headers[k+1]] (EOF for last)
os.makedirs(out_dir, exist_ok=True)
members = []
for k, L in enumerate(lnk):
end = headers[k + 1] if k + 1 < len(headers) else len(d)
name = d[headers[k]:headers[k] + 8].rstrip(b" \x00").decode("ascii", "replace")
# uniquify (PSYLIB names are unique, but be safe)
obj = d[L:end]
out = os.path.join(out_dir, f"{name}.obj")
n = 1
while os.path.exists(out):
out = os.path.join(out_dir, f"{name}_{n}.obj"); n += 1
open(out, "wb").write(obj)
members.append((name, len(obj)))
return members
if __name__ == "__main__":
if len(sys.argv) != 3:
sys.exit(__doc__)
ms = split(sys.argv[1], sys.argv[2])
print(f"{os.path.basename(sys.argv[1])}: {len(ms)} objects -> {sys.argv[2]}")
for nm, sz in ms[:8]:
print(f" {nm}.obj ({sz} B)")
if len(ms) > 8:
print(f" … +{len(ms)-8} more")