From 19c676f474fa60b8ea2c1bf9f53baf1c0da7e410 Mon Sep 17 00:00:00 2001 From: Drew T <50529377+Druthulu@users.noreply.github.com> Date: Tue, 8 Sep 2026 13:05:16 -0600 Subject: [PATCH] =?UTF-8?q?docs(phase-34):=20outreach=20=E2=80=94=20the=20?= =?UTF-8?q?decomp.me=20#ai=20channel=20reply=20drafted=20from=20the=20reco?= =?UTF-8?q?rd=20(Drew's=20words=20to=20edit)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- docs/outreach/tools-announcement.md | 40 ++++++++++++++++++++++------- 1 file changed, 31 insertions(+), 9 deletions(-) diff --git a/docs/outreach/tools-announcement.md b/docs/outreach/tools-announcement.md index d2a2bb376..e409c72c4 100644 --- a/docs/outreach/tools-announcement.md +++ b/docs/outreach/tools-announcement.md @@ -8,16 +8,9 @@ **Discord post:** ``` -Made a small tool while doing the Brave Fencer Musashi decomp that might be useful to other PS1 projects: xsig -(https://github.com/Druthulu/xsig). It hashes functions with the relocated fields masked out, so the same function at -two different link addresses gets the same signature. We used it to find shared code across overlays (2,220 groups in -BFM) and to check other games: against Xenogears, Vagrant Story and Tomba it found 103 shared functions, all PsyQ -library code, so mostly it's good for spotting library functions and shared overlay code. Stdlib only, MIT, has tests. -The full decomp is at https://github.com/Druthulu/BFM-decomp if anyone's curious, everything's matched. +I made a small tool while doing the Brave Fencer Musashi decomp that might be useful to other PS1 projects: xsig (https://github.com/Druthulu/xsig). It hashes functions with the relocated fields masked out, so the same function at two different link addresses gets the same signature. I used it to find shared code across overlays (2,220 groups in BFM) and to check other games: against Xenogears, Vagrant Story and Tomba it found 103 shared functions, all PsyQ library code, so mostly it's good for spotting library functions and shared overlay code. Stdlib only, MIT, has tests. The full decomp is at https://github.com/Druthulu/BFM-decomp if anyone's curious, everything's matched. -Next step for it is going past exact matches: try the 1:1 first, then widen to same-shape functions (registers and -immediates masked, then same structure) and report which tier matched and what differs, so it also finds near-copies -across overlays and across games, not just library code. That's the v2 I'm starting on. +Next step for it is going past exact matches: try the 1:1 first, then widen to same-shape functions (registers and immediates masked, then same structure) and eport which tier matched and what differs, so it also finds near-copies across overlays and across games, not just library code. That's the v2 I'm starting on. ``` **Decompedia tools-page row (if the page takes rows):** name xsig, language Python (stdlib), license MIT, platform MIPS / @@ -30,3 +23,32 @@ No date is promised in the post. **Facts the post rests on:** `tools/xsig/README.md` (the worked example: 103 hits across three games, all library; 126 against Tomba with one 19-instruction non-library HIGH), `config/dedup.us.yaml` (2,220 groups), the E4 log entry (8/8 tests). + +## The decomp.me Discord `#ai` channel reply (Drew, 2026-09-08 — drafted from the record, his words to edit) + +> Context: people in that channel were describing the trouble of keeping an AI assistant on track. Facts the draft rests on: +> `docs/verification.md` (218/218 from a clean rebuild), `docs/gen3-standards.md` (the definition of matched; names only with +> evidence; sotn's style guide adopted), `docs/how-to-ai-decomp/` ch.01 (governance), ch.04 (oracles), ch.12 (the failure museum), +> `docs/retrospective.md`. No number in it is typed from memory. Linked: the how-to index, the governance chapter, the kit page. + +``` +I ran into the same thing doing Brave Fencer Musashi (PS1). Ended up at 218 binaries rebuilding byte for byte +from C, and honestly the model was never the hard part, the setup around it was. What kept it on track: + +The only definition of done is the community one: instruction-identical including register allocation, and +the whole-binary hash equal to the original, checked on every build. Nothing "functionally equivalent" ever +counted, so the assistant can't talk its way to a match. What isn't ours is said plainly (1,256 Sony library +objects linked, five hand-written asm routines kept verbatim). + +Names only with evidence. If nothing proves what a function is, it stays func_XXXXXXXX. Wrong names are +worse than none, so we adopted the sotn style guide for the cleanup rather than inventing one. + +The rules grew from actual failures, not from a prompt. Every session starts by re-reading them, every +session ends with a checkpoint file the next one replays word for word, and one task closes before the +next opens. The compiler quirks got solved by reading the gcc 2.7.2 source, not by guessing. + +I wrote all of it up, including what went wrong, which is the more useful half: +https://github.com/Druthulu/BFM-decomp/wiki/How-to-AI-decomp +Chapter 1 is the on-track part, chapter 12 is the failure list. If you're starting a project from zero +there's a starter kit page too: https://github.com/Druthulu/BFM-decomp/wiki/Start-a-new-decomp-project +```