mirror of
https://github.com/Druthulu/BFM-decomp
synced 2026-09-29 15:18:24 -04:00
feat(phase-19): T1 -O0 split lever — infra proven + 6 -O0 fns matched (ov_SC01_077)
- config/splat.ov_SC01_077.yaml: 3-object code split (ov_SC01_077_a before / ov_SC01_077_o0 -O0 cluster / ov_SC01_077 after). A single object's .text cannot be split around a middle object, so before/after are distinct objects; the after-region keeps the ov_SC01_077 name (bulk matched C + asm paths unchanged). - Makefile: target-specific CC1FLAGS:=-O0 for ov_SC01_077_o0.o (src/boot.c precedent). - src/ov_SC01_077/ov_SC01_077_a.c (new before-region) + ov_SC01_077.c (after-region) + ov_SC01_077_o0.c (new -O0 cluster): 6/16 -O0 fns matched byte-perfect (func_8013B568/B7F4/BD34/C360/C938/C964). - DEFERRED to Phase 20 (Drew, on the discovered difficulty): the 10 remaining -O0 fns hit an indexed-global %lo-folding codegen quirk (gcc-source research, R17) + the x134 per-overlay rollout (engine_core.h compiles -O2, cannot carry -O0 fns). Corrects the Phase-18 backlog premise (afternoon/free-x134) per R14. Documented in ov_SC01_077_o0.c. - verified clean: main 143dbb89, resident 8e17e02f, ov_SC01_077 d19c9580 + 2 overlays.
This commit is contained in:
@@ -439,6 +439,11 @@ build/src/%.o: src/%.c
|
||||
# pattern recipe reads $(CC1FLAGS), so this overrides it for just build/src/boot.o):
|
||||
build/src/boot.o: CC1FLAGS := -quiet -O0 -G0 -mips1 -mcpu=3000 -mgas -msoft-float -fgnu-linker
|
||||
|
||||
# Phase-19 T1: same per-file -O0 mechanism for the ov_SC01_077 -O0 cluster (16 contiguous fns
|
||||
# vram 0x8013B568..0x8013C98C, prologue sig 21F0A003). Split into its own .c by the splat config
|
||||
# (config/splat.ov_SC01_077.yaml) so this override reaches just that .o.
|
||||
build/src/ov_SC01_077/ov_SC01_077_o0.o: CC1FLAGS := -quiet -O0 -G0 -mips1 -mcpu=3000 -mgas -msoft-float -fgnu-linker
|
||||
|
||||
# link (the .ld pulls in the .o by path) + objcopy to the raw PS-X EXE image.
|
||||
$(OUT): $(OBJS) $(ASSET_OBJS) $(LD_SCRIPT)
|
||||
@set -e
|
||||
|
||||
@@ -51,7 +51,15 @@ segments:
|
||||
vram: 0x80128158
|
||||
align: 4
|
||||
subsegments:
|
||||
- [0x0, c, ov_SC01_077] # code: vram 0x80128158..0x80186AD0 (file 0x0..0x5E978)
|
||||
# Phase-19 T1: per-file -O0 split (src/boot.c precedent). The 16 contiguous functions
|
||||
# vram 0x8013B568..0x8013C98C were compiled -O0 (prologue sig 21F0A003); gcc-2.7.2 has no
|
||||
# per-function optimize pragma, so they need their own -O0 .o (the Makefile forces
|
||||
# CC1FLAGS=-O0 on ov_SC01_077_o0.o). A single object's .text can't be split around a middle
|
||||
# object, so before/after are DISTINCT objects: before -> ov_SC01_077_a (smaller, asm paths
|
||||
# rewritten), after KEEPS the ov_SC01_077 name (the bulk matched C + asm paths unchanged).
|
||||
- [0x0, c, ov_SC01_077_a] # -O2 before: vram 0x80128158..0x8013B568 (file 0x0..0x13410)
|
||||
- [0x13410, c, ov_SC01_077_o0] # -O0 cluster: 16 fns vram 0x8013B568..0x8013C98C (file 0x13410..0x14834)
|
||||
- [0x14834, c, ov_SC01_077] # -O2 after: vram 0x8013C98C..0x80186AD0 (file 0x14834..0x5E978)
|
||||
- [0x5E978, data, tail] # data tail (word-aligned bulk): file 0x5E978..0xB29D4
|
||||
- [0xB29D4, bin, trailing] # final 3 bytes (EOF 0xB29D7 not word-aligned; spimdisasm drops
|
||||
# a trailing partial word and a <4-byte `data` carve emits nothing,
|
||||
|
||||
@@ -0,0 +1,36 @@
|
||||
# CURRENT PHASE — Phase 19: Scale the toolkit (wave-harvest at scale + cheap front-loaded levers)
|
||||
|
||||
**Started:** 2026-06-20 · **Generation:** Gen2 (11th phase) · **Plan approved (gate 1):** Drew, 2026-06-20
|
||||
**Effort:** Max (planning/synthesis) · xHigh (T1/T2 execution) · **Ultracode for T3 wave batches** (prompt at the transition, R26/R27)
|
||||
**Plan file:** `~/.claude/plans/plan-mode-enabled-max-golden-truffle.md`
|
||||
|
||||
## Scope decisions (Drew, gate 1)
|
||||
- **Bounded flywheel** — T1 + T2 + **2–3 wave batches of 50**, bank gains, close at a clean checkpoint with a Phase-20 backlog. (Not open-ended; not infra-only.)
|
||||
- **gcc-research spike (T4) = CONDITIONAL in-phase** — runs ONLY if the waves plateau on a residual class blocking meaningful reach.
|
||||
|
||||
## Baseline at phase start (verified this session)
|
||||
- Fleet **56.64%** byte-identical-from-source (194,839 / 344,010); 136/136 binaries byte-identical; 1450 dedup groups (0 failed); 0 NON_MATCHING.
|
||||
- `ov_SC01_077`: 1719/2586 = 66.5%, 867 INCLUDE_ASM stubs remaining; clean-rebuild SHA `d19c9580…`.
|
||||
- Staged: `.run/harvest_wave_p18s1.js` (120 tractable reach-134 targets; P18 ran 31). Ghidra-C cache `.run/ghidra_c/` = 300 fns (fresh).
|
||||
|
||||
## Tasks
|
||||
- [x] **T1 — Per-file -O0 split-file lever** — DONE (re-scoped by Drew: bank infra + 6, defer rest). The 3-object split works (`ov_SC01_077_a` before / `ov_SC01_077_o0` -O0 / `ov_SC01_077` after — a single object's .text can't be split around a middle object). **6/16 -O0 fns matched** byte-perfect (`func_8013B568/B7F4/BD34/C360/C938/C964`). **DEFERRED to Phase 20:** the 10 remaining -O0 fns hit an indexed-global `%lo`-folding codegen quirk (gcc-source research, R17 — our cc1 materializes the address, the original folds %lo); AND the ×134 per-overlay rollout (engine_core.h compiles -O2, can't carry -O0 fns → each overlay needs its own -O0 split). Two findings corrected the Phase-18 backlog's "afternoon, free ×134" premise (R14).
|
||||
- [ ] **T2 — Recovery tooling.** `tools/cast_call_sites.py` (call-site fn-ptr casts) + implicit-int propagate-first. Build/validate on `.run/drafts-p18s1/` near-misses.
|
||||
- [ ] **T3 — Scale waves (2–3 batches of 50, Ultracode).** Regen Ghidra-C for fresh targets, `gen_harvest_targets.py` manifest, Step-1 wave prompt → recovery (T2) → gate → `dedup_propagate` ×134 → checkpoint → improve. Each batch = a progress report.
|
||||
- [ ] **T4 — (CONDITIONAL) gcc-research spike.** Only on a blocking residual class (loop-guard order `func_8012C2D0`; §10 schedule `func_8014F2E0`).
|
||||
- [ ] **PhaseEnd** — verify all checkboxes (P7), demonstrate milestone, write `PhaseEnd_Phase19.md` (+ `## Plain-English Recap`, R25; Phase-20 backlog), archive worklog (R19).
|
||||
|
||||
## Milestone (gate 2, Drew confirms)
|
||||
-O0 lever banked + recovery tooling proven + toolkit scaled across 2–3 batches of 50, fleet up materially (target +3–5%, ~56.6% → ~60%), 136/136 byte-identical from clean (R22), 0 NON_MATCHING (G4), closed at a clean checkpoint with a Phase-20 backlog.
|
||||
|
||||
## Progress log
|
||||
*(append one line per task as completed)*
|
||||
- 2026-06-20 — Phase started; plan approved (gate 1); task list built (R28); CURRENT_PHASE.md created. Beginning T1.
|
||||
- 2026-06-20 — T1: built + proved the 3-object -O0 split infra; matched 6/16 -O0 fns (ov_SC01_077 d19c9580). Two consults with Drew on T1 scope: (1) full -O0 rollout chosen, then (2) on discovering the remaining fns are research-grade (the %lo quirk), Drew chose "bank 6 + infra, defer rest to Phase 20." Fleet unchanged (the 6 are ov_SC01_077-local until rollout). Verified main/resident/ov_SC01_077/+2 byte-identical. Committed checkpoint. → T2.
|
||||
|
||||
## Blockers
|
||||
*(none)*
|
||||
|
||||
## Notes / carried to Phase 20
|
||||
- **The -O0 lever completion (NEW, from T1):** infra is built + proven in ov_SC01_077 (the 3-object split). Remaining: (a) **match the 10 indexed-global -O0 fns** — blocked by the `%lo`-folding quirk (our cc1 materializes `lui;addiu;addu;sw 0(reg)`; original folds `lui;addu;sw %lo(sym)(reg)`); needs gcc-2.7.2 source research (R17/§17), maybe unsteerable from C. Stubs: func_8013B598/B6A0/B7AC/B83C/BC7C/BCDC/BD74/C08C/C0F8/C414. (b) **The ×134 rollout** — each overlay needs its own -O0 split (uniform offsets 0x13410/0x14834, shared bodies header); engine_core.h can't carry -O0 fns. ~+0.6% fleet when complete. (c) the 2 -O0 outliers (func_80144B9C giant + func_801457A4).
|
||||
- Most giants (28 reach-134 >150-ins); per-overlay unique remainder (×1); comprehension/emulator actor-field naming; Wine/CC1PSX (parked); public flip (Phase 14, Gen3+).
|
||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,73 @@
|
||||
#include "common.h"
|
||||
|
||||
/* Phase-19 T1: the -O0 cluster (16 contiguous fns vram 0x8013B568..0x8013C98C, prologue sig
|
||||
* 21F0A003 = -O0 frame pointer). Split into its own .c so the Makefile forces CC1FLAGS=-O0 on
|
||||
* ov_SC01_077_o0.o (src/boot.c precedent). A single object's .text can't be split around a middle
|
||||
* object, so the overlay code is 3 objects: ov_SC01_077_a (before) / ov_SC01_077_o0 (this, -O0) /
|
||||
* ov_SC01_077 (after). Data externs hoisted + deconflicted; intra-cluster callees forward-declared.
|
||||
*
|
||||
* 6/16 matched here (reach-134, but the ×134 fleet rollout is DEFERRED to Phase 20: -O0 functions
|
||||
* can't propagate via engine_core.h — that header compiles -O2 in other overlays — so each overlay
|
||||
* needs its own -O0 split).
|
||||
*
|
||||
* 10 stubs DEFERRED to Phase 20: they hit an -O0 codegen quirk — for indexed global access
|
||||
* `arr[i]=x`, our cc1 materializes the address (lui;addiu;addu;sw 0(reg)) while the original folds
|
||||
* %lo (lui;addu idx;sw %lo(sym)(reg), 1 ins shorter). That's gcc address-splitting behavior our cc1
|
||||
* reproduces differently at -O0; cracking it is gcc-source research (Phase-18 §17 / R17 territory),
|
||||
* not a C-form fix. The 6 matched below are scalar-store / pointer-loop / simple-call (no index). */
|
||||
extern s32 D_80187270;
|
||||
extern s32 D_801DAB24;
|
||||
extern s32 D_801DAAC0;
|
||||
|
||||
void func_8013B83C(s32 a0, s32 a1, s32 a2);
|
||||
void func_8013BD74(void *a0, s32 a1);
|
||||
|
||||
void func_8013B568(s32 arg0) {
|
||||
D_80187270 = arg0;
|
||||
}
|
||||
|
||||
INCLUDE_ASM("asm/ov_SC01_077/nonmatchings/ov_SC01_077_o0", func_8013B598);
|
||||
|
||||
INCLUDE_ASM("asm/ov_SC01_077/nonmatchings/ov_SC01_077_o0", func_8013B6A0);
|
||||
|
||||
INCLUDE_ASM("asm/ov_SC01_077/nonmatchings/ov_SC01_077_o0", func_8013B7AC);
|
||||
|
||||
void func_8013B7F4(s32 a0, s32 a1) {
|
||||
func_8013B83C(a0, a1, D_801DAB24);
|
||||
}
|
||||
|
||||
INCLUDE_ASM("asm/ov_SC01_077/nonmatchings/ov_SC01_077_o0", func_8013B83C);
|
||||
|
||||
INCLUDE_ASM("asm/ov_SC01_077/nonmatchings/ov_SC01_077_o0", func_8013BC7C);
|
||||
|
||||
INCLUDE_ASM("asm/ov_SC01_077/nonmatchings/ov_SC01_077_o0", func_8013BCDC);
|
||||
|
||||
void func_8013BD34(s32 a0) {
|
||||
func_8013BD74(&D_801DAAC0, a0);
|
||||
}
|
||||
|
||||
INCLUDE_ASM("asm/ov_SC01_077/nonmatchings/ov_SC01_077_o0", func_8013BD74);
|
||||
|
||||
INCLUDE_ASM("asm/ov_SC01_077/nonmatchings/ov_SC01_077_o0", func_8013C08C);
|
||||
|
||||
INCLUDE_ASM("asm/ov_SC01_077/nonmatchings/ov_SC01_077_o0", func_8013C0F8);
|
||||
|
||||
void func_8013C360(s32 a0) {
|
||||
s32 *p;
|
||||
u32 i;
|
||||
p = (s32 *)(a0 + 0x10);
|
||||
for (i = 0; i < *(u32 *)(a0 + 8); i++) {
|
||||
*(s32 *)(*(s32 *)p) = *(s32 *)((s32)p + 4);
|
||||
p = (s32 *)((s32)p + 0xC);
|
||||
}
|
||||
}
|
||||
|
||||
INCLUDE_ASM("asm/ov_SC01_077/nonmatchings/ov_SC01_077_o0", func_8013C414);
|
||||
|
||||
void func_8013C938(void) {
|
||||
D_801DAAC0 = 1;
|
||||
}
|
||||
|
||||
void func_8013C964(void) {
|
||||
D_801DAAC0 = 0;
|
||||
}
|
||||
Reference in New Issue
Block a user