mirror of
https://github.com/Druthulu/BFM-decomp
synced 2026-10-05 17:03:23 -04:00
feat(phase-7): PsyQ library-linking proven byte-exact + rodata-island mechanism — session C checkpoint
- rodata-island mechanism PROVEN: migration (dotted .rodata sibling named "800")
+ tools/ld_interleave.py (.data->.rodata->.data linker placement) + data-in-text
carve; rodata/data sizes come out byte-exact. +24 residual root-caused to the
.align-3 jumptable file-split padding (6 sites; spim suggests 8 splits)
- PsyQ 4.0 library linking PROVEN end-to-end: libcd SYS.o (483 instrs, full TU + 21
externals + internal .rdata/.data relocs) links BYTE-IDENTICAL to BFM
- tools/psyq_lib_split.py: split Sony LIB\x01 archive -> member .OBJ
- tools/psyq_build_libs.sh: .LIB -> psyq-obj-parser -> ELF .a (14 BFM libs)
- tools/psyq_identify.py: relocation-masked search -> object link addresses
(libcd 18/25 used, contiguous from 0x80043088)
- external symbols recovered from the EXE's own resolved relocations
- section-alignment 8->4 fix (psyq-obj-parser over-aligns; the +4 mismatch is the tell)
- cookbook §8 (rodata island) + §9 (PsyQ library-linking recipe); CURRENT_PHASE full
diagnosis, asset inventory, and resume plan
- 3rd consecutive green session (143dbb89 BYTE-IDENTICAL) -> Gen1 >=3-session bar MET
- SDK assets (PsyQ 4.0 USA DTL-S2002 libs + psyq-obj-parser) staged gitignored under
tools/psyq/; no function matches added (architectural session)
This commit is contained in:
@@ -207,3 +207,56 @@ A function that calls PsyQ library routines needs both the SDK **types** and the
|
||||
- Done Phase 7: `include/psyq/libcd.h` (CdlLOC 4B, CdlFILE 24B + CdSearchFile/CdPosToInt/CdIntToPos
|
||||
protos) + the 4 libcd/libetc symbols — unlocks the file-loader cluster. Same pattern for
|
||||
libgpu/libgte/libspu as they come up.
|
||||
|
||||
---
|
||||
|
||||
## §8 rodata island (compiler jump tables) — the `.data→.rodata→.data` sandwich (Phase 7)
|
||||
GCC emits each `switch` jump table into `.rodata`; in this EXE all compiler rodata is ONE island at
|
||||
0x80072A38–0x80074750, sitting BETWEEN the front `.data` (globals @0x800629DC) and the tail `.data`
|
||||
(@0x80074750). No single splat `section_order` expresses data→rodata→data. Proven mechanism (session C):
|
||||
- **Migrate, don't standalone.** A jtbl `.word`s reference function-internal `.L`/`jlabel` targets, so a
|
||||
separate rodata object can't link — the table MUST co-locate in its function's object. Use a **dotted
|
||||
`.rodata` subseg whose NAME matches the code subseg** (`[<off>, .rodata, 800]`): `extract=False`, spimdisasm
|
||||
migrates each single-ref jtbl/const into `asm/nonmatchings/<seg>/<fn>.s` as `.section .rodata`. The
|
||||
INCLUDE_ASM stub already `.include`s that `.s`, so it flows into the object for free. Multi-ref rodata can't
|
||||
migrate → splat emits `INCLUDE_RODATA(...)` lines (in a FRESH `.c`). **H5:** don't regen-fresh the curated
|
||||
`.c` (drops comments) — surgically INSERT just the INCLUDE_RODATA lines.
|
||||
- **Place explicitly.** splat is section-major (floats all `.rodata` to the front). `tools/ld_interleave.py`
|
||||
(wired into `make extract`) rewrites the `.main {}` body to text → front-`.data` → `.rodata` → tail-`.data` →
|
||||
bss, splitting front/tail by object basename. Sizes then land byte-exact.
|
||||
- **Carve data-in-text.** A trailing non-code table inside the text range (here 0x80062998–0x800629DC) must be
|
||||
its own `data` subseg, or jumptable analysis mis-extends the last function across it (the +24 `main_TEXT_END`
|
||||
overrun's first cause).
|
||||
- **The `.align 3` file-split trap:** GCC 8-aligns jtbls; concatenating many functions into one object injects
|
||||
padding nops the original (separate TUs) lacked → image grows. spimdisasm PRINTS file-split suggestions at
|
||||
the misaligned jtbls. Fix = per-file split at those boundaries (sotn-style) — OR link the real library
|
||||
object (§9) when the owning function is SDK code.
|
||||
|
||||
## §9 Link real PsyQ library objects byte-exact (Phase 7 — GO proven)
|
||||
~350 of BFM's functions are unmodified PsyQ 4.0 SDK code. They are **byte-identical to the real PsyQ library
|
||||
objects**, so link them directly instead of hand-decompiling — and each library `.o` brings its own correct
|
||||
alignment (dissolving the library-half of §8's `.align 3` problem). Validated: `CdPosToInt`/`CdIntToPos` EXACT
|
||||
vs PsyQ libcd; `PRESET_OBJ_*` ∈ `LIBGS.LIB`. Workflow (the decomp-standard psyq-obj-parser path):
|
||||
- **Tools** (gitignored `tools/psyq/`): `psyq-obj-parser` (decompme prebuilt — `.OBJ`→ELF; rejects `.LIB`),
|
||||
`lib40/*.LIB` = PsyQ **4.0 USA** libraries (DTL-S2002 R2.0 = BFM's version; extracted from the redump ISO via
|
||||
`tools/bfm_extract/iso9660.py`). Identify a function's library by searching the `.LIB` for a NON-relocated
|
||||
instruction run from its EXE bytes (relocated runs false-negative — use leaves or interior runs).
|
||||
- **Integration:** split `.LIB` (LIB\x01 archive) → `.OBJ` → `psyq-obj-parser` → `ar` per lib → link the `.o`
|
||||
for each library function and drop its INCLUDE_ASM. BFM mixes 4.0+4.2 library stamps, so a few objects may
|
||||
need 4.2/4.3 libs — determine per-object by the byte test.
|
||||
- **Proven full-object link recipe (SYS.o byte-identical to BFM, Phase 7):**
|
||||
1. **Placement** — `tools/psyq_identify.py <elf_dir>`: relocation-masked search finds each object's `.text`
|
||||
vram in the EXE. Per library the used objects are CONTIGUOUS in object order → place the first at the
|
||||
region base, link the rest in order.
|
||||
2. **Recover externals** — symbols the object references but doesn't define are usually absent from
|
||||
`symbols.us.txt`; read them straight out of the EXE's RESOLVED relocations: for each reloc, `R_MIPS_26` →
|
||||
`target = ((word&0x3FFFFFF)<<2)|(pc&0xF0000000)`; an `HI16`+`LO16` pair → `(hi<<16)+signext(lo)`. Feed as
|
||||
`ld --defsym NAME=0xADDR`.
|
||||
3. **Alignment** — psyq-obj-parser emits `.text/.rdata/.data` at align 2**3; the original is 4-aligned, so an
|
||||
8-align bumps the section +4 (the tell: every `LO16` to that section is off by +4). Fix:
|
||||
`objcopy --set-section-alignment '.rdata=4' --set-section-alignment '.data=4' obj.o obj_a.o` before linking.
|
||||
4. **Link + verify** — `ld -T <SECTIONS: . = <text vram>; .text:{*(.text)} . = <island>; .rdata:{*(.rodata)
|
||||
*(.rdata)} . = <data vram>; .data:{*(.data)}> --defsym … obj_a.o` → `objcopy -O binary --only-section
|
||||
.text` → byte-compare to the EXE. `.rdata`/`.data` vrams are found by searching the EXE for the section
|
||||
bytes (`objcopy --only-section`). Tools: `tools/psyq_lib_split.py`, `tools/psyq_build_libs.sh`,
|
||||
`tools/psyq_identify.py`.
|
||||
|
||||
+80
-15
@@ -34,30 +34,95 @@ rodata-island foundation + LZSS match are DEFERRED to a focused sub-project afte
|
||||
- [ ] **Task 7 — PhaseEnd_Phase7** (Gen1 synthesis, milestone gate). **Max · Tier 1.**
|
||||
|
||||
## Current task
|
||||
**Task 5 — Loader cluster** (match-tractable non-switch / draft-hard). Then Task 2′ (LZSS), Task 6, Task 7.
|
||||
NOTE: Gen1 exit needs ≥3 SESSIONS of green `make check` — cannot complete this session regardless.
|
||||
**Task 2′ — rodata-island / LZSS gate.** Mechanism now PROVEN end-to-end; a strategic pivot to PsyQ-library
|
||||
linking is in flight (see the PsyQ spike section below) — that fixes the library-half alignment AND gives
|
||||
~350 SDK functions byte-exact for free. Then LZSS, Task 6, Task 7.
|
||||
NOTE: Gen1 exit needs ≥3 SESSIONS of green `make check` — **now satisfied** (A, B, C below); Tasks 6/7 still pending.
|
||||
|
||||
## Per-session `make check` green log (≥3 sessions needed for the milestone)
|
||||
- 2026-06-14 (session A): `make check` → `143dbb89… BYTE-IDENTICAL` ✓ — baseline restored + reproducibility fix, reports built, **38 real matches** (22 accessor leaves + ResourceGetCdLoc + LoaderResetReadState), build byte-identical throughout. [need ≥2 more sessions]
|
||||
- 2026-06-14 (session B): `make check` → `143dbb89… BYTE-IDENTICAL` ✓ — **per-file -O0 split mechanism** (src/boot.c + Makefile per-file flags); **4 real matches** (GameModeDispatch, DebugMenuHandler, CdQueueBusy, CdReadRequest) → **42 real**; **PsyQ libcd.h infra** (CdlLOC/CdlFILE + 4 named symbols, unlocks the loader cluster); **LoaderInitFileTable + ResourceLoadStateMachine NON_MATCHING-drafted** (→ 4 NM) — **Task 5 non-jtbl loaders COMPLETE** (6 matched + 2 drafted); report tooling fixed (multi-file); cookbook §6/§7/T4. Build byte-identical throughout. [need ≥1 more session]
|
||||
- 2026-06-15 (session C): `make check` → `143dbb89… BYTE-IDENTICAL` ✓ (full `clean && extract && build`, restored after the Task-2′ experiments). **≥3-session bar MET.** This session: fully diagnosed + built the **rodata-island mechanism** (works); root-caused the +24; **proved the PsyQ-library-linking GO** (see below). No new matches (architectural session). Build green at start and after restore.
|
||||
|
||||
---
|
||||
|
||||
## Rodata-island foundation — investigation findings (DEFERRED, for Task 2′)
|
||||
Attempted the 3-way data split (`[data front][.rodata island][data tail]` + `ld_legacy_generation: True`).
|
||||
**What works:** dotted `.rodata` sibling named `800` → spimdisasm MIGRATES each jump table into its owning
|
||||
function's `.s` (`jtbl_80072A38` lands inside `LzssDecodeSector.s` with `.section .rodata`/`.section .text`),
|
||||
references resolve intra-800.o, build LINKS. Non-migrated multi-ref rodata becomes `INCLUDE_RODATA` (70 lines)
|
||||
in a FRESH-regenerated `800.c`.
|
||||
**What blocks byte-identity (the structural wall):**
|
||||
1. Adding any `rodata` subseg turns on global jumptable analysis → merges 56 over-split switch fragments (good) but DROPS an 8-byte inter-fn blob (the `func_80047CAC` issue, now fixed via explicit symbol).
|
||||
2. A separate rodata object can't link to text-local `.L`/`jlabel` jumptable targets → must migrate (co-locate).
|
||||
3. splat places sections CONTIGUOUSLY (no explicit `. = addr`), so any size drift shifts the whole image. Observed a **24-byte `.text` overrun** (`main_TEXT_END` 0x800629F4 vs 0x800629DC) → +24B size, 85288 bytes differ. Root cause of the 24B NOT fully pinned.
|
||||
**Candidate fix for Task 2′:** explicit linker addresses (Makefile post-extract `.ld`-patch placing `.text`@0x80010000, front `.data`@0x800629DC, `800.o(.rodata)`@0x80072A38, tail `.data`@0x80074750) + scope migration so only island tables land in `800.o(.rodata)`. Alternative: sotn-style per-file split (Gen2-scale). Reproduce with the migration config (in git stash / reconstruct from this log).
|
||||
## Rodata-island foundation — RESOLVED end-to-end (session C, 2026-06-15)
|
||||
The mechanism now WORKS; only the +24 (a known file-split/alignment artifact) blocks full byte-identity, and
|
||||
the PsyQ-lib pivot (below) is the chosen fix. Experimental configs saved: `.run/{splat.island.yaml,
|
||||
symbols.island.txt,Makefile.island,800.c.fresh-throwaway}`; new committed tool `tools/ld_interleave.py`.
|
||||
**Proven mechanism (3 parts, all validated this session):**
|
||||
1. **Migration** — dotted `.rodata` sibling named `800` (`[0x63238, .rodata, 800]`) → spimdisasm migrates all
|
||||
50 jtbls + single-ref consts into their owning `asm/nonmatchings/800/<fn>.s` (99 .s got `.section .rodata`,
|
||||
`.L`-refs resolve intra-object). Multi-ref consts → 70 `INCLUDE_RODATA` lines (needs a FRESH `800.c`; for the
|
||||
PERMANENT file, surgically INSERT those 70 lines into the curated 800.c — do NOT regen-fresh, it drops
|
||||
comments/H5). All 301 jtbl targets ∈ the `800` text seg (none in `boot`), so the single sibling is correct.
|
||||
2. **Placement** — `tools/ld_interleave.py` (wired into `make extract`) rewrites splat's section-major `.main`
|
||||
into the real `.data→.rodata→.data` sandwich order (text, front-data@0x800629DC, rodata@0x80072A38,
|
||||
tail-data@0x80074750). Front/tail split by object basename. rodata + both data sizes came out **byte-exact**.
|
||||
3. **Data-in-text carve** — the 68-B descriptor table at 0x80062998–0x800629DC (ptrs to start/D_80062998/
|
||||
D_80074778) must be its own `[0x53198, data, 53198]` subseg or the jumptable analyzer mis-extends the last
|
||||
code function across it.
|
||||
**The ONLY residual = +24 (ROOT-CAUSED):** 6× `.align 3` jumptable padding nops injected into `.text` (at
|
||||
PRESET_OBJ_744, PRESET_OBJ_8FC, PRESET2_OBJ_4D8, PRESET2_OBJ_A88, OBJT2_OBJ_614, PRNT_OBJ_24C). GCC 8-aligns
|
||||
each switch jtbl, but the original built these as SEPARATE translation units; our single 800.o concatenation
|
||||
adds padding the original lacked. spimdisasm itself printed **8 file-split suggestions** (rodata 0x6324C,
|
||||
0x63388, 0x633FC, 0x63920, 0x63C94, 0x64420, 0x64AB4, 0x64CA0). Canonical fix = per-file split — OR the PsyQ-lib
|
||||
pivot below (most of these are library code).
|
||||
|
||||
## PsyQ-library-linking SPIKE — GO PROVEN (session C, the chosen +24 fix + free SDK code)
|
||||
**Finding:** BFM's PsyQ library functions are **byte-identical to the real PsyQ SDK objects** → link them
|
||||
directly (byte-exact) instead of hand-decompiling, which ALSO gives each library `.o` correct per-object
|
||||
alignment (dissolving the library-half of the +24). Validated: `CdPosToInt` (32 instrs) + `CdIntToPos` (65)
|
||||
EXACT vs PsyQ **4.7** `libcd.a`; `CdPosToInt` also in **4.0** `LIBCD.LIB`; `PRESET_OBJ_108` (a +24 culprit) is
|
||||
in **4.0 `LIBGS.LIB`** → library code, not game. (Raw-byte lib search has false-negatives on relocated funcs,
|
||||
e.g. `_spu_FsetPCR`/`OBJT2`/`PRNT` "missed" — needs the real ELF-link test to classify those.)
|
||||
**Assets staged (gitignored `tools/psyq/`):** `psyq-obj-parser` (decompme prebuilt, works on `.OBJ`; rejects
|
||||
`.LIB` archives — needs splitting); `psyq4.0/` (4.0 tools: CC1PSX/ASPSX/PSYLIB/…); `conv47/` (4.7 pre-converted
|
||||
ELF `.a` — quick reference); **`lib40/*.LIB`** = the 20 PsyQ **4.0 USA** libraries (DTL-S2002 R2.0, BFM's exact
|
||||
version) extracted from the redump via our `tools/bfm_extract/iso9660.py` walker. Footprint in BFM: ~350 funcs
|
||||
(libsnd 131, libapi/gs 69+, libmcrd 63, libsn 37, libcd 28, libspu 21, …) of 2050 matchable.
|
||||
**Integration pipeline — BUILT + PROVEN end-to-end (session C):**
|
||||
- `tools/psyq_lib_split.py` (committed) — splits a `LIB\x01` archive into its member `.OBJ` (locates each
|
||||
member header by the invariant `u32@(header+12) == LNK_offset − header`). LIBCD → 25 objects ✓.
|
||||
- `tools/psyq_build_libs.sh` (committed) — `.LIB → .OBJ → psyq-obj-parser → ELF .o → ar` per lib. **Built all
|
||||
14 BFM libs → `tools/psyq/lib40_elf/*.a`** (gitignored): LIBCD 25, LIBGS 201, LIBSPU 129, LIBSND 163,
|
||||
LIBMCRD 2, LIBSN 51, LIBAPI 90, LIBETC 7, LIBGTE 381, LIBGPU 12, LIBMATH 48, LIBCARD 18, LIBC 56, LIBC2 46.
|
||||
- **Byte-match PROVEN at object level:** in libcd `SYS.o`, leaves `CdPosToInt`/`CdIntToPos` are EXACT; relocated
|
||||
funcs (`CdComstr` …) differ ONLY at their relocation sites → link byte-exact once relocs resolve to BFM
|
||||
symbol addrs. So every step (split, convert, leaf-match, reloc-resolve) is validated.
|
||||
|
||||
**FULL OBJECT LINK — PROVEN BYTE-IDENTICAL (session C):** libcd `SYS.o` (483 instrs, the full TU: leaves +
|
||||
relocated funcs + 21 externals + internal .rdata/.data) links **byte-for-byte identical to BFM**. The pipeline
|
||||
+ the 3 last pieces:
|
||||
- **Identify placement** (`tools/psyq_identify.py`, committed) — relocation-masked search locates each object's
|
||||
`.text` in BFM. libcd: **18/25 objects found, CONTIGUOUS** at 0x80043088–0x80046D1C in object order (the 7
|
||||
unused — CDPLAY, C_012–015… — BFM doesn't link). So per-library placement = link the used objects in order at
|
||||
the region base; addresses are read off, not guessed.
|
||||
- **Recover externals from BFM** — symbols the object references but doesn't define (e.g. libcd's `CD_pos`,
|
||||
`CD_com`, `DMACallback`) are NOT in symbols.us.txt, but their addresses are encoded in the EXE's already-
|
||||
RESOLVED relocations: parse the object's reloc records, read BFM at each site, reconstruct (R_MIPS_26 →
|
||||
target; HI16/LO16 pair → addr). Recovered all 21 for SYS.o. (Feeds symbols.us.txt over time.)
|
||||
- **Alignment fix** — psyq-obj-parser sets `.text/.rdata/.data` align=2**3 (8); the original placed them
|
||||
4-aligned, so an 8-align bumps them +4. `objcopy --set-section-alignment .rdata=4 .data=4` before linking →
|
||||
exact. (The +4 mismatch is the tell.)
|
||||
- **Link recipe:** `ld -T <script placing .text@<objaddr> .rdata@<island> .data@<addr>> --defsym <recovered…>
|
||||
obj.o` → objcopy .text → byte-compare. Proven on SYS.o.
|
||||
|
||||
**REMAINING (replication + wiring, NEXT session, Task #5):** generalize the SYS.o recipe to all used objects
|
||||
per library (identify region → recover externals → set-align → place sections in order → link), then wire into
|
||||
the build: drop the linked functions' INCLUDE_ASM + carve their raw data, add the lib objects to the link via a
|
||||
generated `.ld` fragment. ~350 SDK funcs become byte-exact + the library-region jtbl alignment resolves. GAME
|
||||
switches + **LZSS** (jtbl_80072A38 = island's first entry, before any misalignment) take the proven
|
||||
migration+ld_interleave path above.
|
||||
|
||||
## Blockers / open items
|
||||
- Task 2′ ld-placement mechanism (above). MCP-mode batching: sig-refresh needs MCP stopped, LZSS needs MCP live.
|
||||
- `.LIB`→`.OBJ` splitter (LIB\x01 format) — the gate for the lib-linking integration.
|
||||
- Which 4.x version matches each BFM lib object best (4.0 USA primary; BFM mixes 4.0+4.2 stamps, so some objects
|
||||
may need 4.2/4.3 libs — determine per-lib during integration via the ELF-link byte test).
|
||||
- LZSS via the proven migration+ld_interleave path (independent of the lib pivot).
|
||||
|
||||
## Notes
|
||||
- Commits accumulate UNCOMMITTED; one phase-end commit by the developer (R8/R6).
|
||||
- `.run/merge_matches.py` = the regenerate-800.c + re-apply-matches helper (reusable for Task 2′).
|
||||
- `.run/merge_matches.py` = regenerate-800.c + re-apply-matches helper. **H5 caveat:** regen-fresh drops
|
||||
file-level/stub comments; for the permanent 800.c, surgically insert the 70 INCLUDE_RODATA lines instead.
|
||||
- `tools/ld_interleave.py` (committed) = the `.data→.rodata→.data` linker-script interleaver.
|
||||
|
||||
@@ -0,0 +1,80 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Reorder splat's generated linker script to honour a .data -> .rodata -> .data
|
||||
"sandwich" layout (Phase 7 Task 2', the rodata-island problem).
|
||||
|
||||
splat emits one output section (`.main`) section-major in `section_order`
|
||||
(.rodata, .text, .data, .bss), which floats ALL rodata to one place. But this
|
||||
EXE's real layout is:
|
||||
|
||||
.text 0x80010000 .. 0x800629DC
|
||||
.data (front) 0x800629DC .. 0x80072A38 (globals, hand-written ptr tables)
|
||||
.rodata (island) 0x80072A38 .. 0x80074750 (gcc jtbl_* switch tables + consts)
|
||||
.data (tail) 0x80074750 .. 0x80074800 (gp base; zero small-data)
|
||||
|
||||
i.e. .data appears on BOTH sides of .rodata, which a single section_order can't
|
||||
express. This script rewrites the `.main {...}` body to the interleaved order:
|
||||
text -> front .data -> .rodata -> tail .data -> .bss, keeping splat's START/END/
|
||||
SIZE symbols. Front vs tail .data is decided by object basename (FRONT_DATA /
|
||||
TAIL_DATA). All other (empty) .data objects go in the front group.
|
||||
|
||||
Idempotent: keyed off splat's exact section-major output; re-running on an
|
||||
already-patched script is a no-op (the markers won't match). Run post-extract.
|
||||
"""
|
||||
import re, sys
|
||||
|
||||
LD = sys.argv[1] if len(sys.argv) > 1 else "build/us/SLUS_007.26.ld"
|
||||
# object basenames whose (.data) belongs to the front / tail region
|
||||
FRONT_DATA = ("531DC.data.o",)
|
||||
TAIL_DATA = ("64F50.data.o",)
|
||||
|
||||
src = open(LD).read()
|
||||
|
||||
# Grab the .main output-section body (between its first '{' and matching '}').
|
||||
m = re.search(r"(\.main\b.*?\n[ \t]*\{\n)(.*?)(\n[ \t]*\})", src, re.S)
|
||||
if not m:
|
||||
sys.exit("ld_interleave: could not find .main { ... } block")
|
||||
head, body, tail = m.group(1), m.group(2), m.group(3)
|
||||
|
||||
# Collect the object input-section lines by linker section, preserving order.
|
||||
def grab(section):
|
||||
# lines like: build/src/800.o(.rodata);
|
||||
return re.findall(rf"^[ \t]*build/\S+\({re.escape(section)}\);", body, re.M)
|
||||
|
||||
text_lines = grab(".text")
|
||||
rodata_lines = grab(".rodata")
|
||||
data_lines = grab(".data")
|
||||
bss_lines = grab(".bss")
|
||||
|
||||
def is_named(line, names):
|
||||
return any(n in line for n in names)
|
||||
|
||||
front_data = [l for l in data_lines if not is_named(l, TAIL_DATA)]
|
||||
tail_data = [l for l in data_lines if is_named(l, TAIL_DATA)]
|
||||
# Sanity: front must contain the FRONT_DATA object.
|
||||
if not any(is_named(l, FRONT_DATA) for l in front_data):
|
||||
sys.exit("ld_interleave: front data object not found — config drift?")
|
||||
if not tail_data:
|
||||
sys.exit("ld_interleave: tail data object not found — config drift?")
|
||||
|
||||
I = " " # 8-space indent matching splat's body
|
||||
def grp(start, lines, end_sym, size_sym):
|
||||
out = [f"{I}{start} = .;"]
|
||||
out += [f"{I}{l.strip()}" for l in lines]
|
||||
out += [f"{I}. = ALIGN(., 4);", f"{I}{end_sym} = .;"]
|
||||
if size_sym:
|
||||
out += [f"{I}{size_sym} = ABSOLUTE({end_sym} - {start});"]
|
||||
return out
|
||||
|
||||
new = [f"{I}FILL(0x00000000);"]
|
||||
new += grp("main_TEXT_START", text_lines, "main_TEXT_END", "main_TEXT_SIZE")
|
||||
new += grp("main_DATA_START", front_data, "main_DATA_END", "main_DATA_SIZE")
|
||||
new += grp("main_RODATA_START", rodata_lines, "main_RODATA_END", "main_RODATA_SIZE")
|
||||
new += grp("main_DATA2_START", tail_data, "main_DATA2_END", "main_DATA2_SIZE")
|
||||
new += grp("main_BSS_START", bss_lines, "main_BSS_END", "main_BSS_SIZE")
|
||||
new_body = "\n".join(new)
|
||||
|
||||
out = src[:m.start()] + head + new_body + tail + src[m.end():]
|
||||
open(LD, "w").write(out)
|
||||
print(f"ld_interleave: rewrote .main — text={len(text_lines)} "
|
||||
f"front_data={len(front_data)} rodata={len(rodata_lines)} "
|
||||
f"tail_data={len(tail_data)} bss={len(bss_lines)}")
|
||||
@@ -0,0 +1,33 @@
|
||||
#!/usr/bin/env bash
|
||||
# Build ELF .a archives from the PsyQ 4.0 .LIB files (Phase 7 — PsyQ-library linking).
|
||||
# Pipeline per lib: psyq_lib_split.py (.LIB -> .OBJ members) -> psyq-obj-parser (each
|
||||
# .OBJ -> ELF .o) -> ar (-> <lib>.a). Output (gitignored, SDK-derived):
|
||||
# tools/psyq/lib40_elf/<LIB>.a + .run/obj40/<lib>/*.{obj,o} scratch
|
||||
# The .a are then linked into the build for the functions identified as PsyQ SDK code.
|
||||
# Requires: tools/psyq/lib40/*.LIB (extracted from the DTL-S2002 redump), tools/psyq/
|
||||
# psyq-obj-parser, mipsel-linux-gnu-ar. SDK assets stay out of git (.gitignore /tools/psyq/).
|
||||
set -euo pipefail
|
||||
cd "$(dirname "$0")/.."
|
||||
LIBDIR=tools/psyq/lib40
|
||||
OUTDIR=tools/psyq/lib40_elf
|
||||
SCRATCH=.run/obj40
|
||||
PARSER=tools/psyq/psyq-obj-parser
|
||||
AR=mipsel-linux-gnu-ar
|
||||
mkdir -p "$OUTDIR"
|
||||
# BFM-relevant libraries (the SDK footprint in the EXE); pass args to override.
|
||||
LIBS=("${@:-LIBCD LIBGS LIBSPU LIBSND LIBMCRD LIBSN LIBAPI LIBETC LIBGTE LIBGPU LIBMATH LIBCARD LIBC LIBC2}")
|
||||
for L in ${LIBS[@]}; do
|
||||
lib="$LIBDIR/$L.LIB"
|
||||
[ -f "$lib" ] || { echo " skip $L (no $lib)"; continue; }
|
||||
od="$SCRATCH/$(echo "$L" | tr A-Z a-z)"
|
||||
rm -rf "$od"; mkdir -p "$od"
|
||||
n=$(python3 tools/psyq_lib_split.py "$lib" "$od" | head -1 | grep -oE '[0-9]+ objects' | grep -oE '[0-9]+')
|
||||
ok=0
|
||||
for o in "$od"/*.obj; do
|
||||
if "$PARSER" "$o" -o "${o%.obj}.o" >/dev/null 2>&1; then ok=$((ok+1)); fi
|
||||
done
|
||||
rm -f "$OUTDIR/$L.a"
|
||||
$AR rcs "$OUTDIR/$L.a" "$od"/*.o 2>/dev/null || true
|
||||
printf " %-10s %3s objs -> %3d .o -> %s.a\n" "$L" "${n:-?}" "$ok" "$L"
|
||||
done
|
||||
echo "done -> $OUTDIR/"
|
||||
@@ -0,0 +1,84 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Locate where PsyQ library objects are linked in the target EXE.
|
||||
|
||||
For each ELF .o (converted from a PsyQ .LIB member), extract its `.text` and the
|
||||
relocation offsets, build a relocation-masked word pattern (relocated immediate
|
||||
fields zeroed), and scan the EXE text for the single position where every
|
||||
NON-relocated word matches. That position is the object's link address in the EXE
|
||||
(or "absent" if the EXE doesn't link it). This is the placement map the library
|
||||
linker step consumes.
|
||||
|
||||
Usage: psyq_identify.py <elf_dir> [text_lo_vram text_hi_vram]
|
||||
(defaults to the BFM .text window 0x80010000..0x800629DC)
|
||||
"""
|
||||
import struct, subprocess, re, sys, glob, os
|
||||
|
||||
EXE = "extracted/retail/SLUS_007.26"
|
||||
VRAM_BASE = 0x8000F800
|
||||
ELF_DIR = sys.argv[1] if len(sys.argv) > 1 else ".run/obj40/libcd"
|
||||
TLO = int(sys.argv[2], 0) if len(sys.argv) > 2 else 0x80010000
|
||||
THI = int(sys.argv[3], 0) if len(sys.argv) > 3 else 0x800629DC
|
||||
|
||||
b = open(EXE, "rb").read()
|
||||
text = b[TLO - VRAM_BASE: THI - VRAM_BASE]
|
||||
twords = [struct.unpack_from("<I", text, i)[0] for i in range(0, len(text), 4)]
|
||||
|
||||
def obj_text_pattern(o):
|
||||
"""Return (words, mask) for the object's .text; mask[i]=0 on relocated/jump words."""
|
||||
d = subprocess.check_output(["mipsel-linux-gnu-objdump", "-dr", "-j", ".text", o],
|
||||
text=True, stderr=subprocess.DEVNULL)
|
||||
words, mask = [], []
|
||||
pending_reloc = False
|
||||
for line in d.splitlines():
|
||||
mi = re.match(r"\s+([0-9a-f]+):\s+([0-9a-f]{8})\s", line)
|
||||
if mi:
|
||||
w = int(mi.group(2), 16)
|
||||
words.append(w)
|
||||
# mask jal/j (opcode 2/3) always (R_MIPS_26 target is link-resolved)
|
||||
mask.append(0 if (w >> 26) in (2, 3) else 0xFFFFFFFF)
|
||||
elif "R_MIPS" in line and words:
|
||||
# relocation annotation follows its instruction line -> mask low 16 (hi/lo/pc16)
|
||||
if "_26" in line:
|
||||
mask[-1] = 0
|
||||
else:
|
||||
mask[-1] = 0xFFFF0000
|
||||
return words, mask
|
||||
|
||||
def find(words, mask):
|
||||
n = len(words)
|
||||
if n < 2:
|
||||
return None # too short to anchor uniquely
|
||||
hits = []
|
||||
for s in range(0, len(twords) - n + 1):
|
||||
ok = True
|
||||
for i in range(n):
|
||||
if (twords[s + i] & mask[i]) != (words[i] & mask[i]):
|
||||
ok = False; break
|
||||
if ok:
|
||||
hits.append(s)
|
||||
if len(hits) > 1:
|
||||
break
|
||||
if len(hits) == 1:
|
||||
return TLO + hits[0] * 4
|
||||
return ("ambiguous" if len(hits) > 1 else None)
|
||||
|
||||
results = []
|
||||
for o in sorted(glob.glob(os.path.join(ELF_DIR, "*.o"))):
|
||||
words, mask = obj_text_pattern(o)
|
||||
if not words:
|
||||
results.append((os.path.basename(o), "no-.text", len(words))); continue
|
||||
r = find(words, mask)
|
||||
results.append((os.path.basename(o), r, len(words)))
|
||||
|
||||
found = [(n, a, l) for n, a, l in results if isinstance(a, int)]
|
||||
found.sort(key=lambda x: x[1])
|
||||
print(f"{ELF_DIR}: {len(found)}/{len(results)} objects located in EXE text "
|
||||
f"[0x{TLO:X}..0x{THI:X}]")
|
||||
for n, a, l in found:
|
||||
print(f" 0x{a:08X} {n:18s} ({l} ins)")
|
||||
absent = [n for n, a, l in results if a is None]
|
||||
amb = [n for n, a, l in results if a == "ambiguous"]
|
||||
if absent:
|
||||
print(f" not linked by EXE ({len(absent)}): {', '.join(absent[:12])}{' …' if len(absent)>12 else ''}")
|
||||
if amb:
|
||||
print(f" ambiguous ({len(amb)}): {', '.join(amb)}")
|
||||
@@ -0,0 +1,75 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Split a Sony PSYLINK library archive (LIB\\x01) into its member LNK objects.
|
||||
|
||||
Format (derived from PsyQ 4.0 LIBCD.LIB, Phase 7):
|
||||
"LIB\\x01"
|
||||
per member:
|
||||
name[8] space-padded object name (e.g. "CDROM ")
|
||||
date[4] little-endian timestamp
|
||||
obj_off[4] little-endian: bytes from THIS member header to its LNK object
|
||||
<symbol dictionary> (obj_off - 16 bytes: the symbols this object exports)
|
||||
<LNK object> starts with "LNK\\x02", runs to the next member header
|
||||
|
||||
We don't parse the LNK record stream to find each object's end; instead each member
|
||||
header is located by the invariant u32@(header+12) == (LNK_offset - header) — i.e.
|
||||
obj_off points exactly at that member's "LNK\\x02". An object therefore spans from its
|
||||
"LNK\\x02" to the NEXT member header (or EOF for the last). psyq-obj-parser validates
|
||||
the result downstream (a wrong boundary fails to parse), and the ultimate check is the
|
||||
byte-match of the linked function against the target EXE.
|
||||
|
||||
Usage: psyq_lib_split.py <archive.LIB> <out_dir> -> writes <out_dir>/<NAME>.obj
|
||||
"""
|
||||
import struct, sys, os
|
||||
|
||||
LNK_MAGIC = b"LNK\x02"
|
||||
|
||||
def split(lib_path, out_dir):
|
||||
d = open(lib_path, "rb").read()
|
||||
if d[:4] != b"LIB\x01":
|
||||
sys.exit(f"{lib_path}: not a LIB\\x01 archive (magic {d[:4]!r})")
|
||||
# 1) all LNK object starts
|
||||
lnk = []
|
||||
i = d.find(LNK_MAGIC, 4)
|
||||
while i != -1:
|
||||
lnk.append(i)
|
||||
i = d.find(LNK_MAGIC, i + 4)
|
||||
if not lnk:
|
||||
sys.exit(f"{lib_path}: no LNK objects found")
|
||||
# 2) each LNK's member header: scan back for P with u32@P+12 == lnk-P
|
||||
headers = []
|
||||
for L in lnk:
|
||||
P = None
|
||||
for cand in range(L - 16, max(L - 0x8000, 0) - 1, -1):
|
||||
if struct.unpack_from("<I", d, cand + 12)[0] == L - cand:
|
||||
# sanity: name bytes printable-ish
|
||||
nm = d[cand:cand + 8]
|
||||
if all(32 <= c < 127 for c in nm.rstrip(b" \x00") or b" "):
|
||||
P = cand; break
|
||||
if P is None:
|
||||
sys.exit(f"{lib_path}: could not locate member header for LNK@0x{L:x}")
|
||||
headers.append(P)
|
||||
# 3) extract: object k = [lnk[k] : headers[k+1]] (EOF for last)
|
||||
os.makedirs(out_dir, exist_ok=True)
|
||||
members = []
|
||||
for k, L in enumerate(lnk):
|
||||
end = headers[k + 1] if k + 1 < len(headers) else len(d)
|
||||
name = d[headers[k]:headers[k] + 8].rstrip(b" \x00").decode("ascii", "replace")
|
||||
# uniquify (PSYLIB names are unique, but be safe)
|
||||
obj = d[L:end]
|
||||
out = os.path.join(out_dir, f"{name}.obj")
|
||||
n = 1
|
||||
while os.path.exists(out):
|
||||
out = os.path.join(out_dir, f"{name}_{n}.obj"); n += 1
|
||||
open(out, "wb").write(obj)
|
||||
members.append((name, len(obj)))
|
||||
return members
|
||||
|
||||
if __name__ == "__main__":
|
||||
if len(sys.argv) != 3:
|
||||
sys.exit(__doc__)
|
||||
ms = split(sys.argv[1], sys.argv[2])
|
||||
print(f"{os.path.basename(sys.argv[1])}: {len(ms)} objects -> {sys.argv[2]}")
|
||||
for nm, sz in ms[:8]:
|
||||
print(f" {nm}.obj ({sz} B)")
|
||||
if len(ms) > 8:
|
||||
print(f" … +{len(ms)-8} more")
|
||||
Reference in New Issue
Block a user