From 3f51101ca9d8cc7acb9d1cce63e65136d0581a7f Mon Sep 17 00:00:00 2001 From: Drew T <50529377+Druthulu@users.noreply.github.com> Date: Tue, 4 Aug 2026 13:29:51 -0600 Subject: [PATCH] =?UTF-8?q?docs(phase-30):=20SESSION-33..37=20checkpoint?= =?UTF-8?q?=20=E2=80=94=20PAUSED;=2016=20ungated=20wave-5=20drafts=20prese?= =?UTF-8?q?rved=20in=20git?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Paused at Drew's request for a Windows restart. Nothing running, tree lock free, tree clean (R23 db churn aside). R22 run 19x this session, 140/140 every time. THE ONE THING OWED: `.run/s37/*/func_*.c` — 16 wave-5 drafts, all claiming MATCH, NONE gated. Force-added to git (.run/ is gitignored) because they cost ~2.7M agent tokens and Workflow's resumeFromRunId cache is SAME-SESSION-ONLY, so it does not survive the restart. Resume by GATING them, not by re-running the wave. Also preserved: the wave scripts (.run/s36w.js, .run/s37w.js), manifests, the hardened gate driver, and the four capture drivers. MEASURED THIS SESSION (both answer questions Drew asked): - The wave PROMPT is the lever. Bank rate 76% -> 77% -> 100% -> 100% on the same models and the same gate, prompt the only variable. The jump was STEP 0 (a magic-literal grep of src/, ahead of engine_core.h) — and that step came from a wave-2 agent's index_gap report, i.e. the agents write the next prompt. - The pipeline() fix, before/after: wave 4 parallel() 14 targets / 136 min / 2.5x parallelism; wave 5 pipeline() 16 targets / 82 min / 3.8x. 40% faster on 14% more targets. The two-batch design was a hard barrier with 46-min dead gaps at each boundary; the harness already caps at 16 so the batching bought nothing. Housekeeping: removed 3 stray cc1 intermediates (t.i, t.i.greg, t.s) that an agent left at the repo ROOT — scratch belongs under .run/ (R12), same class as the gccdump.lreg noted at the Phase-24 close. --- .run/s33_capture.py | 33 ++++ .run/s34_capture.py | 37 +++++ .run/s35_capture.py | 37 +++++ .run/s36_capture.py | 37 +++++ .run/s36_wave.json | 156 ++++++++++++++++++ .run/s36w.js | 153 ++++++++++++++++++ .run/s37/ov_SC02_026/func_801825A4.c | 76 +++++++++ .run/s37/ov_SC02_028/func_8018B13C.c | 16 ++ .run/s37/ov_SC02_028/func_8018B354.c | 16 ++ .run/s37/ov_SC02_035/func_80184328.c | 16 ++ .run/s37/ov_SC02_039/func_8017D930.c | 37 +++++ .run/s37/ov_SC02_041/func_801837B0.c | 231 +++++++++++++++++++++++++++ .run/s37/ov_SC03_024/func_80181914.c | 107 +++++++++++++ .run/s37/ov_SC03_092/func_80184370.c | 189 ++++++++++++++++++++++ .run/s37/ov_SC03_092/func_8018478C.c | 164 +++++++++++++++++++ .run/s37/ov_SC03_108/func_8017F2E8.c | 171 ++++++++++++++++++++ .run/s37/ov_SC03_124/func_8018893C.c | 149 +++++++++++++++++ .run/s37/ov_SC06_008/func_8017ED80.c | 54 +++++++ .run/s37/ov_SC06_008/func_8017F100.c | 60 +++++++ .run/s37/ov_SC06_008/func_80181B20.c | 69 ++++++++ .run/s37/ov_SC06_015/func_8017D8AC.c | 222 +++++++++++++++++++++++++ .run/s37/ov_SC06_032/func_8018EDB0.c | 141 ++++++++++++++++ .run/s37_wave.json | 178 +++++++++++++++++++++ .run/s37w.js | 152 ++++++++++++++++++ .run/s6f_gate.py | 11 +- phase-ends/CURRENT_PHASE.md | 109 ++++++++++++- 26 files changed, 2618 insertions(+), 3 deletions(-) create mode 100644 .run/s33_capture.py create mode 100644 .run/s34_capture.py create mode 100644 .run/s35_capture.py create mode 100644 .run/s36_capture.py create mode 100644 .run/s36_wave.json create mode 100644 .run/s36w.js create mode 100644 .run/s37/ov_SC02_026/func_801825A4.c create mode 100644 .run/s37/ov_SC02_028/func_8018B13C.c create mode 100644 .run/s37/ov_SC02_028/func_8018B354.c create mode 100644 .run/s37/ov_SC02_035/func_80184328.c create mode 100644 .run/s37/ov_SC02_039/func_8017D930.c create mode 100644 .run/s37/ov_SC02_041/func_801837B0.c create mode 100644 .run/s37/ov_SC03_024/func_80181914.c create mode 100644 .run/s37/ov_SC03_092/func_80184370.c create mode 100644 .run/s37/ov_SC03_092/func_8018478C.c create mode 100644 .run/s37/ov_SC03_108/func_8017F2E8.c create mode 100644 .run/s37/ov_SC03_124/func_8018893C.c create mode 100644 .run/s37/ov_SC06_008/func_8017ED80.c create mode 100644 .run/s37/ov_SC06_008/func_8017F100.c create mode 100644 .run/s37/ov_SC06_008/func_80181B20.c create mode 100644 .run/s37/ov_SC06_015/func_8017D8AC.c create mode 100644 .run/s37/ov_SC06_032/func_8018EDB0.c create mode 100644 .run/s37_wave.json create mode 100644 .run/s37w.js diff --git a/.run/s33_capture.py b/.run/s33_capture.py new file mode 100644 index 000000000..1c43f7687 --- /dev/null +++ b/.run/s33_capture.py @@ -0,0 +1,33 @@ +#!/usr/bin/env python3 +"""Capture the whole-binary blocker for a gate-refused draft: splice -> build -> read the +compiler's OWN error -> revert. Agents cannot run the gate, so a TU-level declaration conflict is +invisible to them (§136a: classify on the build's OUTPUT, never its exit status).""" +import sys, os, re, subprocess, shutil +sys.path.insert(0, 'tools'); import corpus +for spec in sys.argv[1:]: + ov, fn = spec.split(':') + addr = int(fn.split('_')[1], 16) + st = corpus.stubs(ov); rec = st.get(addr) + if not rec: + print(f"== {ov} {fn}: NOT A LIVE STUB (already banked?)"); continue + tu = rec.path + draft = f".run/s34/{ov}/{fn}.c" + stub = f'INCLUDE_ASM("{rec.asm_dir}", {fn});' + t = open(tu).read() + if t.count(stub) != 1: + print(f"== {ov} {fn}: stub line not found verbatim in {tu}"); continue + bak = t + open(tu, 'w').write(t.replace(stub, open(draft).read())) + r = subprocess.run(['make', 'build', f'BINARY={ov}'], capture_output=True, text=True) + open(tu, 'w').write(bak) # revert ALWAYS + out = r.stdout + r.stderr + # §136a, applied to MYSELF: a narrow keyword filter is exactly how a real error goes unseen. + # Keep every cc1/ld diagnostic line that names a source position, minus the known SHB noise. + errs = [l for l in out.splitlines() + if re.search(r'\.[ch]:\d+|undefined reference|\bError \d+', l) + and 'built-in function' not in l and 'Makefile:' not in l] + print(f"== {ov} {fn} (build rc={r.returncode})") + if errs: + for l in errs[:6]: print(" ", l.strip()) + else: + print(" NO COMPILE ERROR -> the draft builds clean; the miss is a byte DIFF at whole-binary") diff --git a/.run/s34_capture.py b/.run/s34_capture.py new file mode 100644 index 000000000..951132f2e --- /dev/null +++ b/.run/s34_capture.py @@ -0,0 +1,37 @@ +#!/usr/bin/env python3 +"""Capture the whole-binary blocker for a gate-refused draft: splice -> build -> read the +compiler's OWN error -> revert. Agents cannot run the gate, so a TU-level declaration conflict is +invisible to them (§136a: classify on the build's OUTPUT, never its exit status).""" +import sys, os, re, subprocess, shutil +sys.path.insert(0, 'tools'); import corpus +for spec in sys.argv[1:]: + ov, fn = spec.split(':') + addr = int(fn.split('_')[1], 16) + st = corpus.stubs(ov); rec = st.get(addr) + if not rec: + print(f"== {ov} {fn}: NOT A LIVE STUB (already banked?)"); continue + tu = rec.path + draft = f".run/s35/{ov}/{fn}.c" + stub = f'INCLUDE_ASM("{rec.asm_dir}", {fn});' + t = open(tu).read() + if t.count(stub) != 1: + print(f"== {ov} {fn}: stub line not found verbatim in {tu}"); continue + bak = t + open(tu, 'w').write(t.replace(stub, open(draft).read())) + r = subprocess.run(['make', 'build', f'BINARY={ov}'], capture_output=True, text=True) + open(tu, 'w').write(bak) # revert ALWAYS + out = r.stdout + r.stderr + # §136a, applied to MYSELF: a narrow keyword filter is exactly how a real error goes unseen. + # Keep every cc1/ld diagnostic line that names a source position, minus the known SHB noise. + HARD = re.compile(r'undeclared|redefinition|conflicting types|parse error|undefined reference' + r'|syntax error|too (many|few) arguments|invalid|incompatible') + lines_ = [l for l in out.splitlines() + if re.search(r'\.[ch]:\d+|undefined reference', l) + and 'built-in function' not in l and 'Makefile:' not in l] + hard = [l for l in lines_ if HARD.search(l) and 'warning' not in l.lower()] + errs = hard or [l for l in lines_ if 'warning' not in l.lower()] or lines_ + print(f"== {ov} {fn} (build rc={r.returncode})") + if errs: + for l in errs[:6]: print(" ", l.strip()) + else: + print(" NO COMPILE ERROR -> the draft builds clean; the miss is a byte DIFF at whole-binary") diff --git a/.run/s35_capture.py b/.run/s35_capture.py new file mode 100644 index 000000000..4c0b0cc25 --- /dev/null +++ b/.run/s35_capture.py @@ -0,0 +1,37 @@ +#!/usr/bin/env python3 +"""Capture the whole-binary blocker for a gate-refused draft: splice -> build -> read the +compiler's OWN error -> revert. Agents cannot run the gate, so a TU-level declaration conflict is +invisible to them (§136a: classify on the build's OUTPUT, never its exit status).""" +import sys, os, re, subprocess, shutil +sys.path.insert(0, 'tools'); import corpus +for spec in sys.argv[1:]: + ov, fn = spec.split(':') + addr = int(fn.split('_')[1], 16) + st = corpus.stubs(ov); rec = st.get(addr) + if not rec: + print(f"== {ov} {fn}: NOT A LIVE STUB (already banked?)"); continue + tu = rec.path + draft = f".run/s36/{ov}/{fn}.c" + stub = f'INCLUDE_ASM("{rec.asm_dir}", {fn});' + t = open(tu).read() + if t.count(stub) != 1: + print(f"== {ov} {fn}: stub line not found verbatim in {tu}"); continue + bak = t + open(tu, 'w').write(t.replace(stub, open(draft).read())) + r = subprocess.run(['make', 'build', f'BINARY={ov}'], capture_output=True, text=True) + open(tu, 'w').write(bak) # revert ALWAYS + out = r.stdout + r.stderr + # §136a, applied to MYSELF: a narrow keyword filter is exactly how a real error goes unseen. + # Keep every cc1/ld diagnostic line that names a source position, minus the known SHB noise. + HARD = re.compile(r'undeclared|redefinition|conflicting types|parse error|undefined reference' + r'|syntax error|too (many|few) arguments|invalid|incompatible') + lines_ = [l for l in out.splitlines() + if re.search(r'\.[ch]:\d+|undefined reference', l) + and 'built-in function' not in l and 'Makefile:' not in l] + hard = [l for l in lines_ if HARD.search(l) and 'warning' not in l.lower()] + errs = hard or [l for l in lines_ if 'warning' not in l.lower()] or lines_ + print(f"== {ov} {fn} (build rc={r.returncode})") + if errs: + for l in errs[:6]: print(" ", l.strip()) + else: + print(" NO COMPILE ERROR -> the draft builds clean; the miss is a byte DIFF at whole-binary") diff --git a/.run/s36_capture.py b/.run/s36_capture.py new file mode 100644 index 000000000..4c0b0cc25 --- /dev/null +++ b/.run/s36_capture.py @@ -0,0 +1,37 @@ +#!/usr/bin/env python3 +"""Capture the whole-binary blocker for a gate-refused draft: splice -> build -> read the +compiler's OWN error -> revert. Agents cannot run the gate, so a TU-level declaration conflict is +invisible to them (§136a: classify on the build's OUTPUT, never its exit status).""" +import sys, os, re, subprocess, shutil +sys.path.insert(0, 'tools'); import corpus +for spec in sys.argv[1:]: + ov, fn = spec.split(':') + addr = int(fn.split('_')[1], 16) + st = corpus.stubs(ov); rec = st.get(addr) + if not rec: + print(f"== {ov} {fn}: NOT A LIVE STUB (already banked?)"); continue + tu = rec.path + draft = f".run/s36/{ov}/{fn}.c" + stub = f'INCLUDE_ASM("{rec.asm_dir}", {fn});' + t = open(tu).read() + if t.count(stub) != 1: + print(f"== {ov} {fn}: stub line not found verbatim in {tu}"); continue + bak = t + open(tu, 'w').write(t.replace(stub, open(draft).read())) + r = subprocess.run(['make', 'build', f'BINARY={ov}'], capture_output=True, text=True) + open(tu, 'w').write(bak) # revert ALWAYS + out = r.stdout + r.stderr + # §136a, applied to MYSELF: a narrow keyword filter is exactly how a real error goes unseen. + # Keep every cc1/ld diagnostic line that names a source position, minus the known SHB noise. + HARD = re.compile(r'undeclared|redefinition|conflicting types|parse error|undefined reference' + r'|syntax error|too (many|few) arguments|invalid|incompatible') + lines_ = [l for l in out.splitlines() + if re.search(r'\.[ch]:\d+|undefined reference', l) + and 'built-in function' not in l and 'Makefile:' not in l] + hard = [l for l in lines_ if HARD.search(l) and 'warning' not in l.lower()] + errs = hard or [l for l in lines_ if 'warning' not in l.lower()] or lines_ + print(f"== {ov} {fn} (build rc={r.returncode})") + if errs: + for l in errs[:6]: print(" ", l.strip()) + else: + print(" NO COMPILE ERROR -> the draft builds clean; the miss is a byte DIFF at whole-binary") diff --git a/.run/s36_wave.json b/.run/s36_wave.json new file mode 100644 index 000000000..047a3a73c --- /dev/null +++ b/.run/s36_wave.json @@ -0,0 +1,156 @@ +[ + { + "fn": "func_8018BED0", + "ov": "ov_SC03_091", + "sub": "asm/ov_SC03_091/nonmatchings/ov_SC03_091_jr_80182268", + "tu": "src/ov_SC03_091/ov_SC03_091_jr_80182268.c", + "n": 328, + "m": 8, + "ti": 2624, + "model": "sonnet", + "seed": 0 + }, + { + "fn": "func_8018BAB4", + "ov": "ov_SC03_091", + "sub": "asm/ov_SC03_091/nonmatchings/ov_SC03_091_jr_80182268", + "tu": "src/ov_SC03_091/ov_SC03_091_jr_80182268.c", + "n": 263, + "m": 8, + "ti": 2104, + "model": "sonnet", + "seed": 0 + }, + { + "fn": "func_8018A808", + "ov": "ov_SC02_027", + "sub": "asm/ov_SC02_027/nonmatchings/ov_SC02_027_jr_8017D898", + "tu": "src/ov_SC02_027/ov_SC02_027_jr_8017D898.c", + "n": 44, + "m": 41, + "ti": 1804, + "model": "sonnet", + "seed": 0 + }, + { + "fn": "func_80181D38", + "ov": "ov_SC03_006", + "sub": "asm/ov_SC03_006/nonmatchings/ov_SC03_006_jr_8017AE2C", + "tu": "src/ov_SC03_006/ov_SC03_006_jr_8017AE2C.c", + "n": 143, + "m": 9, + "ti": 1287, + "model": "sonnet", + "seed": 1 + }, + { + "fn": "func_8017ED54", + "ov": "ov_SC06_014", + "sub": "asm/ov_SC06_014/nonmatchings/ov_SC06_014_jr_8017BEBC", + "tu": "src/ov_SC06_014/ov_SC06_014_jr_8017BEBC.c", + "n": 265, + "m": 4, + "ti": 1060, + "model": "sonnet", + "seed": 0 + }, + { + "fn": "func_8017C290", + "ov": "ov_SC03_024", + "sub": "asm/ov_SC03_024/nonmatchings/ov_SC03_024_jr_8017AE2C", + "tu": "src/ov_SC03_024/ov_SC03_024_jr_8017AE2C.c", + "n": 229, + "m": 4, + "ti": 916, + "model": "sonnet", + "seed": 0 + }, + { + "fn": "func_80182A78", + "ov": "ov_SC03_002", + "sub": "asm/ov_SC03_002/nonmatchings/ov_SC03_002_jr_8017D604", + "tu": "src/ov_SC03_002/ov_SC03_002_jr_8017D604.c", + "n": 180, + "m": 5, + "ti": 900, + "model": "sonnet", + "seed": 1 + }, + { + "fn": "func_801803E0", + "ov": "ov_SC03_108", + "sub": "asm/ov_SC03_108/nonmatchings/ov_SC03_108_jr_8017BEBC", + "tu": "src/ov_SC03_108/ov_SC03_108_jr_8017BEBC.c", + "n": 225, + "m": 4, + "ti": 900, + "model": "sonnet", + "seed": 1 + }, + { + "fn": "func_8018829C", + "ov": "ov_SC06_018", + "sub": "asm/ov_SC06_018/nonmatchings/ov_SC06_018_jr_8017C24C", + "tu": "src/ov_SC06_018/ov_SC06_018_jr_8017C24C.c", + "n": 150, + "m": 6, + "ti": 900, + "model": "sonnet", + "seed": 1 + }, + { + "fn": "func_8018B29C", + "ov": "ov_SC02_028", + "sub": "asm/ov_SC02_028/nonmatchings/ov_SC02_028_jr_8017D898", + "tu": "src/ov_SC02_028/ov_SC02_028_jr_8017D898.c", + "n": 44, + "m": 20, + "ti": 880, + "model": "sonnet", + "seed": 0 + }, + { + "fn": "func_8017DF18", + "ov": "ov_SC01_004", + "sub": "asm/ov_SC01_004/nonmatchings/ov_SC01_004_jr_8017BE9C", + "tu": "src/ov_SC01_004/ov_SC01_004_jr_8017BE9C.c", + "n": 175, + "m": 5, + "ti": 875, + "model": "sonnet", + "seed": 1 + }, + { + "fn": "func_80181CF0", + "ov": "ov_SC06_008", + "sub": "asm/ov_SC06_008/nonmatchings/ov_SC06_008_jr_8017C294", + "tu": "src/ov_SC06_008/ov_SC06_008_jr_8017C294.c", + "n": 124, + "m": 7, + "ti": 868, + "model": "sonnet", + "seed": 1 + }, + { + "fn": "func_8017D4A4", + "ov": "ov_SC01_004", + "sub": "asm/ov_SC01_004/nonmatchings/ov_SC01_004_jr_8017BE9C", + "tu": "src/ov_SC01_004/ov_SC01_004_jr_8017BE9C.c", + "n": 173, + "m": 5, + "ti": 865, + "model": "sonnet", + "seed": 1 + }, + { + "fn": "func_8018A970", + "ov": "ov_SC02_027", + "sub": "asm/ov_SC02_027/nonmatchings/ov_SC02_027_jr_8017D898", + "tu": "src/ov_SC02_027/ov_SC02_027_jr_8017D898.c", + "n": 41, + "m": 21, + "ti": 861, + "model": "sonnet", + "seed": 0 + } +] \ No newline at end of file diff --git a/.run/s36w.js b/.run/s36w.js new file mode 100644 index 000000000..bd7790f2e --- /dev/null +++ b/.run/s36w.js @@ -0,0 +1,153 @@ +export const meta = { + name: 'p30-s36-SONNET', + description: 'P30 S34: fresh h_seq families, open sites DERIVED from corpus.stubs', + phases: [ + { title: 'Draft', detail: 'one agent per target, size-routed per the §136i ladder (haiku/sonnet/opus)' }, + { title: 'Escalate', detail: 'next rung up (haiku->sonnet, sonnet->opus) on any non-MATCH' }, + ], +} + +// args = { targets: [...compact records...], extra: "" } +// Accept a JSON string too — an invocation can deliver args stringified and pipeline() then dies. +const A = {targets: [{"fn": "func_8018BED0", "ov": "ov_SC03_091", "sub": "asm/ov_SC03_091/nonmatchings/ov_SC03_091_jr_80182268", "tu": "src/ov_SC03_091/ov_SC03_091_jr_80182268.c", "n": 328, "m": 8, "ti": 2624, "model": "sonnet", "seed": 0}, {"fn": "func_8018BAB4", "ov": "ov_SC03_091", "sub": "asm/ov_SC03_091/nonmatchings/ov_SC03_091_jr_80182268", "tu": "src/ov_SC03_091/ov_SC03_091_jr_80182268.c", "n": 263, "m": 8, "ti": 2104, "model": "sonnet", "seed": 0}, {"fn": "func_8018A808", "ov": "ov_SC02_027", "sub": "asm/ov_SC02_027/nonmatchings/ov_SC02_027_jr_8017D898", "tu": "src/ov_SC02_027/ov_SC02_027_jr_8017D898.c", "n": 44, "m": 41, "ti": 1804, "model": "sonnet", "seed": 0}, {"fn": "func_80181D38", "ov": "ov_SC03_006", "sub": "asm/ov_SC03_006/nonmatchings/ov_SC03_006_jr_8017AE2C", "tu": "src/ov_SC03_006/ov_SC03_006_jr_8017AE2C.c", "n": 143, "m": 9, "ti": 1287, "model": "sonnet", "seed": 1}, {"fn": "func_8017ED54", "ov": "ov_SC06_014", "sub": "asm/ov_SC06_014/nonmatchings/ov_SC06_014_jr_8017BEBC", "tu": "src/ov_SC06_014/ov_SC06_014_jr_8017BEBC.c", "n": 265, "m": 4, "ti": 1060, "model": "sonnet", "seed": 0}, {"fn": "func_8017C290", "ov": "ov_SC03_024", "sub": "asm/ov_SC03_024/nonmatchings/ov_SC03_024_jr_8017AE2C", "tu": "src/ov_SC03_024/ov_SC03_024_jr_8017AE2C.c", "n": 229, "m": 4, "ti": 916, "model": "sonnet", "seed": 0}, {"fn": "func_80182A78", "ov": "ov_SC03_002", "sub": "asm/ov_SC03_002/nonmatchings/ov_SC03_002_jr_8017D604", "tu": "src/ov_SC03_002/ov_SC03_002_jr_8017D604.c", "n": 180, "m": 5, "ti": 900, "model": "sonnet", "seed": 1}, {"fn": "func_801803E0", "ov": "ov_SC03_108", "sub": "asm/ov_SC03_108/nonmatchings/ov_SC03_108_jr_8017BEBC", "tu": "src/ov_SC03_108/ov_SC03_108_jr_8017BEBC.c", "n": 225, "m": 4, "ti": 900, "model": "sonnet", "seed": 1}, {"fn": "func_8018829C", "ov": "ov_SC06_018", "sub": "asm/ov_SC06_018/nonmatchings/ov_SC06_018_jr_8017C24C", "tu": "src/ov_SC06_018/ov_SC06_018_jr_8017C24C.c", "n": 150, "m": 6, "ti": 900, "model": "sonnet", "seed": 1}, {"fn": "func_8018B29C", "ov": "ov_SC02_028", "sub": "asm/ov_SC02_028/nonmatchings/ov_SC02_028_jr_8017D898", "tu": "src/ov_SC02_028/ov_SC02_028_jr_8017D898.c", "n": 44, "m": 20, "ti": 880, "model": "sonnet", "seed": 0}, {"fn": "func_8017DF18", "ov": "ov_SC01_004", "sub": "asm/ov_SC01_004/nonmatchings/ov_SC01_004_jr_8017BE9C", "tu": "src/ov_SC01_004/ov_SC01_004_jr_8017BE9C.c", "n": 175, "m": 5, "ti": 875, "model": "sonnet", "seed": 1}, {"fn": "func_80181CF0", "ov": "ov_SC06_008", "sub": "asm/ov_SC06_008/nonmatchings/ov_SC06_008_jr_8017C294", "tu": "src/ov_SC06_008/ov_SC06_008_jr_8017C294.c", "n": 124, "m": 7, "ti": 868, "model": "sonnet", "seed": 1}, {"fn": "func_8017D4A4", "ov": "ov_SC01_004", "sub": "asm/ov_SC01_004/nonmatchings/ov_SC01_004_jr_8017BE9C", "tu": "src/ov_SC01_004/ov_SC01_004_jr_8017BE9C.c", "n": 173, "m": 5, "ti": 865, "model": "sonnet", "seed": 1}, {"fn": "func_8018A970", "ov": "ov_SC02_027", "sub": "asm/ov_SC02_027/nonmatchings/ov_SC02_027_jr_8017D898", "tu": "src/ov_SC02_027/ov_SC02_027_jr_8017D898.c", "n": 41, "m": 21, "ti": 861, "model": "sonnet", "seed": 0}], extra: "START HERE \u2014 SIBLING-FIRST IS THE FASTEST ROUTE (\u00a7136c, measured this session: it produced several\nFIRST-DRAFT matches). Before you derive anything from the .s:\n 1. grep src/shared/engine_core.h for a `DEFINE_func_*` macro body that is a NEAR-TWIN of your\n target (same struct-offset chain, same shape, differing only in constants/one call). This is a\n family wave \u2014 the twin usually EXISTS, because that is what a family is.\n 2. grep your own TU for an already-BANKED sibling (a real function definition, not an INCLUDE_ASM).\n 3. Reuse its EXPRESSION FORMS and its DECLARATION FORMS verbatim. They are already byte-proven to\n produce the gcc-2.7.2 schedule and register assignment you need.\nSearch order: engine_core.h near-twin -> same-TU banked sibling -> the .s -> the Ghidra seed LAST\n(the seed was byte-proven to be an ENTIRELY DIFFERENT body twice this session).\n\nTHE LOCAL-VARIABLE LEVER (\u00a7136 \u2014 highest-yield finding, 25 banks). gcc-2.7.2 allocates ONE PSEUDO\nPER C LOCAL, and local-alloc.c:472 REFUSES a local allocno whose REG_N_DEATHS > 1 \u2014 promoting it to\na GLOBAL allocno that loses the low register. The number and scope of your locals moves whole\nregister assignments. REACH FOR THIS BEFORE register __asm__ PINS.\n L1. Same $v0/$v1 pair swapped in ONE arm only => you reused ONE local across N arms. SPLIT it into\n per-arm block-scoped locals.\n L2. One extra `sw $sN` in the prologue, frame otherwise identical => SPLIT a compound initializer:\n `x = *(u8*)p << k;` makes TWO pseudos; `x = *(u8*)p; x = x << k;` reuses one.\n L3. An extra `addu $vX,$v0,$zero` after a `jal` AND a later copy of the same value => the source\n had TWO variables with the first PINNED (an unpinned pseudo coalesces the pair away).\n L4. LENGTH-DRIFT short by `addiu $sN,$sp,K` + a save/restore pair => write a POINTER local assigned\n before the loop and used only inside it.\n L5. Local stack slots are assigned in DECLARATION order ascending from 0x10, independent of use\n order. Frame-offset drift with correct code is a declaration-ORDER problem.\n\nTYPE-FORM RULES:\n T1. A real `mult $rX,$rY` with a small constant => the multiplier is a NON-CONST LOCAL, not a\n literal (a literal goes through synth_mult's sll/addu chain). `s32 r = K;` as its own statement.\n T2. `li $sN,0xfff0` + `addu` where the target has `addiu $vN,$vN,-0x10` => gcc narrowed to HImode.\n Hoist the call to its own statement; put the load+subtract in a BLOCK-SCOPED s32 temp.\n T3. `andi $vN,0xffff` after a `jal` the target lacks => the TU declares that callee u16/s16-\n returning. Do NOT change the decl \u2014 cast at the call (idiom 9 on the RETURN axis).\n T4. Unexplained `addu $vX,$aY,$zero` near a conditional branch (including in its DELAY SLOT,\n consumed only by the fall-through arm) + the feeding `lh` loads out of order => the value is an\n s16 LOCAL, not s32; LOAD_EXTEND_OP folded the widening extend into a plain move.\n T5. LENGTH-DRIFT +1 with a narrow load of the SAME stack slot => gcc narrowed a memory-operand\n `local >> 16`. Bind the local to an s32 temp used twice.\n T6. A symbol read once at a constant offset but built into $s1 by lui/addiu => cache it in a\n POINTER LOCAL. A direct D_xxx[k] folds %lo per use and shrinks the frame.\n T7. BRANCH POLARITY: if match_one prints BRANCH-POLARITY, invert the source condition \u2014 gcc-2.7.2\n flips the branch to place the longer block as the fall-through. (Closed 8 mismatches in one\n edit this wave; the index keys the literal string \"BRANCH-POLARITY\" straight to \u00a73-T4.)\n\nSCHEDULING:\n S1. The MEM_IN_STRUCT_P escape does NOT apply when the blocking store has a VARYING (register-base)\n address \u2014 true_dependence() only drops the edge for a CONSTANT-address store. Then the lever is\n SOURCE ORDER: assign the load to a temp ABOVE the stores. (But same-base `reg+const` addresses\n ARE disambiguated by memrefs_conflict_p, so those hoists are free.)\n S2. An unfilled load-delay nop where the target fills it with a trailing call's arg setup => hoist a\n LOAD: split `*p = *p + 1` into `v = *p + 1; ... *p = v;`.\n S3. NEVER pin an incoming PARAMETER \u2014 it turns the param's `move` into a schedulable body insn and\n reshuffles the whole prologue. Pin loop variables only.\n S4. If no C lever moves a 3-6 instruction schedule/regalloc residual, a \u00a721 ZERO-BYTE RE-TIE\n barrier can anchor it: `__asm__ __volatile__(\"\" : \"=r\"(v) : \"0\"(v));` between the loads and the\n use. Try the source-level levers first; this closed one case after six other variants failed.\n\nDECLARATION SURFACE (decides whether a byte-correct draft BANKS):\n D1. `conflicting types` for a D_ symbol you cannot find declared in the split .c => the decl lives\n inside a DEFINE_func_*() MACRO BODY in src/shared/engine_core.h. Reuse its canonical type.\n An 8-byte-stride table declared `s32 D_x[][2]` must be indexed [i][0]/[i][1].\n D2. cc1 reports only the FIRST conflict. EVERY reconciled draft this session had a SECOND hidden\n conflict, sometimes BELOW the splice point. grep the WHOLE TU in ONE pass for every symbol.\n D3. A DECLARATION CONFLICT ABORTS THE COMPILE, so it hides the byte question entirely. If you are\n handed a draft as \"byte-correct, only declaration-blocked\", RUN match_one ON IT FIRST \u2014 two of\n three such drafts this session also had a real codegen residual behind the conflict.\n D4. If you compile anything, use a PROCESS-UNIQUE scratch path. A shared one silently compiled\n another agent's file and returned a meaningless success this session.\n\n\nx2-9 BAND. \u00a7136c sibling-first still applies but CHECK the twin EXISTS first (\u00a7136e) \u2014 if the whole\nfamily is in nonmatchings, go straight to the .s. \u00a7136j: sub-120-ins targets tend to fail on\nDECLARATIONS, 120+ on real codegen; do the whole-TU one-pass symbol grep (D2) either way.\nIf match_one reports REGALLOC-PERM (a clean 2-register swap), see \u00a7137: it is a TWO-COMPILE\nARITHMETIC problem \u2014 read R and L from `cc1 -dl -dg`, evaluate floor_log2(R)*R/L*1e4*size for both\ncontenders AND their ranked neighbours to get the admissible window, then place a zero-byte\n`__asm__ __volatile__(\"\" ::\"r\"(v))` so L lands inside it. Do NOT reach for the permuter first.\n\n\nS33 MEASURED (promote into your first pass):\n - \u00a7138 RECONCILE DIRECTION: if the TU declares your symbol ABOVE the splice point, DELETE your\n duplicate; if BELOW, KEEP a decl in the TU's EXACT shape and cast at the use. Picking the wrong\n direction CREATES the next error. grep the TU for the symbol and compare line numbers with the\n stub line BEFORE editing either way.\n - A repeated typedef OR a repeated bare `struct Tag {...}` is a C89 ERROR even when textually\n identical. If the TU already defines the type trio you copied from a sibling, do not redeclare\n it \u2014 and check for the bare struct tag too, not just the typedefs (that one cost an extra round).\n - func_8017D174 (793 ins) closed a \u00a7137 allocno tie AND a sched2 loop-head rotation JOINTLY with\n four zero-byte asms: two `\"=r\"`/`\"0\"` re-ties splitting a live range, plus two volatile sliders\n placed in a DIFFERENT basic block, so they lift the live-length count without perturbing the head\n schedule. A slider inside the block you are trying to fix will regress it.\n - Rank/trust nothing from a map's `exemplar` field: derive the open sites from corpus.stubs. The\n map's exemplar can point at an instance that is ALREADY banked, which hides the whole family.\n\n\nS34 MEASURED \u2014 DO THIS FIRST:\n - STEP 0 of sibling-first, ahead of engine_core.h: `grep -rn \"\" src/` with a distinctive\n literal from YOUR .s (a magic word, an unusual mask, an odd immediate). \u00a7136c's first two steps\n are same-TU/shared-header scoped and CANNOT reach a banked twin in another overlay's TU \u2014 and the\n big template classes live cross-overlay. Measured: func_80188C04 (328 ins) was byte-identical to\n an already-banked func_801833F0 in ov_SC02_028; one grep found it and the body was reused\n verbatim with only file-local type/macro suffixes renamed.\n - If a sibling of yours banked into the SAME TU recently, the TU now carries ITS type names for the\n shared data symbols. Use the TU's names; do not introduce your own suffixed duplicates.\n - \"normalized distance 0\" is NOT an h_exact guarantee. Relocs are masked by normalization.\n\n\nS35 (wave 3 banked 13/13 \u2014 the magic-grep STEP 0 is doing the work; lead with it):\n - Write ONLY your deliverable `.run///func_.c` into the drafts directory. Scratch\n files there break the gate. Put experiments anywhere else under .run/.\n - If the TU declares your function itself incompatibly (`extern void f(void);` vs a def that takes\n an argument), that is the SELF axis: define under a private C name bound to the real symbol \u2014\n `extern aF() __asm__(\"func_\"); aF() { ... }`. Three of\n this session's reconciles were exactly this, and it has zero blast radius on callers.\n"} +const T = Array.isArray(A) ? A : A.targets +const EXTRA = (Array.isArray(A) ? '' : A.extra) || '' +if (!Array.isArray(T)) throw new Error('args.targets must be an array') + +const VERDICT = { + type: 'object', + additionalProperties: false, + required: ['fn', 'status', 'summary'], + properties: { + fn: { type: 'string' }, + status: { type: 'string', enum: ['MATCH', 'DIFF', 'BLOCKED'] }, + closeness: { type: 'number', description: 'mismatching instructions remaining; 0 for MATCH' }, + klass: { type: 'string', description: 'residual class if not MATCH' }, + summary: { type: 'string', description: 'what you did and what the residual is, <=4 sentences' }, + levers: { type: 'string', description: 'cookbook sections / idioms that CLOSED the residual' }, + index_hit: { type: 'boolean' }, + index_gap: { type: 'string' }, + }, +} + +function prompt(t, escalated) { + const asm = `${t.sub}/${t.fn}.s` + return `You are matching ONE PS1 function to byte-identical gcc-2.7.2 output for the Brave Fencer +Musashi decompilation. Your ONLY deliverable is a C file at \`.run/s36/${t.ov}/${t.fn}.c\`. + +TARGET + function ${t.fn} + binary ${t.ov} + target asm ${asm} <- THE GROUND TRUTH. Read this FIRST and in full. + TU it lands in ${t.tu} + asm-subdir ${t.sub} + size ${t.n} instructions + leverage family of ${t.m} members / ${t.ti} templatable instructions — a byte-match here + propagates ${t.m}x across the fleet. +${t.seed ? ` ghidra seed .run/ghidra_c/${t.fn}.c <- A HINT ONLY. It is sometimes an ENTIRELY + DIFFERENT body (byte-proven this phase). If it disagrees with the .s, THE .s WINS.` : ` ghidra seed (none cached — work from the .s)`} +${t.retry ? ` ** RETRY ** A previous wave recorded: "${t.retry}". That is a data point, not a + verdict. Re-derive from the .s; do not assume the earlier verdict was right.` : ''} + +HOW TO WORK (this order is the measured-fastest) +1. \`docs/cookbook-index.md\` is a SYMPTOM-KEYED index of 364 byte-verified idioms. Grep it for your + residual's symptom BEFORE deriving anything. Measured: index-first took a wave's bank rate from + 57% to 100%. Then read the section it names in \`docs/matching-cookbook.md\`. + \`docs/gcc-2.7.2-map/{sched,regalloc,loop,cse_expr}.md\` is the compiler-source-derived map for + scheduling / register-allocation residuals. +2. Read the target \`.s\` completely: frame size, callee-saved registers, jal targets, every + \`%hi/%lo\` symbol. +3. Read the TU (${t.tu}) for EVERY symbol your draft will name. cc1 reports only the FIRST conflict, + so a draft can look one edit from done and hold three more. grep the whole TU in ONE pass. + Match its existing declarations EXACTLY; push any type disagreement to a CAST AT THE USE SITE + rather than redeclaring the symbol. +4. Write the draft, then verify: + .venv/bin/python tools/match_one.py ${t.fn} --c .run/s36/${t.ov}/${t.fn}.c --asm-subdir ${t.sub} + Iterate until it prints MATCH; it names the exact mismatching instructions. + +IDIOMS THAT CLOSED RESIDUALS IN THE LAST WAVES (cookbook §135 — all byte-verified) + 1. UNSIGNED switch index => pure equality chain, NO range test. No \`slti\` bound check in the + target's switch means the index is u32, not s32. + 2. \`a0[0x46]\` (ARRAY_REF) sets MEM_IN_STRUCT_P and lets a load HOIST past a constant-address + store; \`*(s16 *)((s32)a0 + 0x8C)\` (INDIRECT_REF) keeps the dependence. Many 4-instruction + "scheduling residuals" are just this type-form choice. + 3. A constant store whose top bit is set in the STORED width needs an UNSIGNED destination: + \`*(u16 *)p = 0x8C00\` emits \`ori\`; through \`s16\` it folds negative and emits \`addiu\`. + 4. The list scheduler PRESERVES the relative order of disambiguable stores. A store written late + in source SINKS. If a store lands too late, move it EARLIER IN SOURCE (not a permuter job). + 5. A \`short\` loop counter blocks strength reduction; walking explicit pointers (\`p++\`) + reproduces the original biv/giv set. + 6. Frame size off by a constant => DEAD LOCALS. If ALL diffs are \`sp\`-relative immediates off by + one constant delta, add the padding declaration. + 7. An INTERIOR address has no symbol — a \`lui/addiu\` pair can build an offset INTO a symbol. + Find the containing symbol in the data \`.s\` and index into it; declaring the interior address + as its own extern link-fails. + 8. NEVER redeclare a C-library name (\`memcpy\` etc.). + 9. Loose typing is pervasive: if the TU declares \`void f(void)\` but the asm passes \`$a0\`, call + through a cast — \`((void(*)(s32))f)(a0)\` — do NOT change the declaration. +${EXTRA ? `\nPROMOTED FROM THE PREVIOUS BATCH (fresh, byte-verified this session)\n${EXTRA}\n` : ''} +HARD RULES + * Write ONLY \`.run/s36/${t.ov}/${t.fn}.c\`. NEVER edit \`src/\`, \`asm/\`, \`config/\`, \`include/\`, + the Makefile, or any tracked file. + * Do NOT run \`make\`, \`make build\`, \`make extract\`, or \`tools/harvest_verify.py\`. The + whole-binary gate is the orchestrator's job and the sole arbiter of a match. + * \`match_one\` MATCH is NECESSARY BUT NOT SUFFICIENT — it compiles standalone and cannot see the + TU's other declarations. Step 3 is what makes a MATCH actually BANK. + * Report honestly. A DIFF with a precise residual class routes the next attempt; a false MATCH + just gets caught by the byte-gate and wastes a cycle. + * ${escalated ? 'A cheaper rung of the model ladder already attempted this and did not reach MATCH. Read its draft at the path above, but re-derive from the .s rather than trusting it.' : 'Work economically — most functions this size close from the .s plus one or two index lookups.'} + +Return the structured verdict.` +} + +phase('Draft') + +// S10 measured a SERVER-side throttle at 30-wide: 14 of 30 agents never ran. The limiter is +// CAPACITY, not capability (§136i: Sonnet banked 13/16 = 81% on 125-793 ins targets, at least as +// well as Opus). So deal TWO batches of ~9 rather than one wide fan-out. +const HALF = Math.ceil(T.length / 2) +const BATCHES = [T.slice(0, HALF), T.slice(HALF)] + +async function runOne(t) { + const v = await agent(prompt(t, false), { + label: `draft:${t.fn}(${t.n}i,x${t.m})`, + phase: 'Draft', + model: t.model, + schema: VERDICT, + }) + if (!v) return { t, v: { fn: t.fn, status: 'BLOCKED', summary: 'agent returned no verdict' }, tier: t.model } + if (v.status === 'MATCH' || t.model === 'opus') return { t, v, tier: t.model } + const nextTier = t.model === 'haiku' ? 'sonnet' : 'opus' + const v2 = await agent(prompt(t, true), { + label: `escalate:${t.fn}`, + phase: 'Escalate', + model: nextTier, + schema: VERDICT, + }) + return { t, v: v2 && v2.status === 'MATCH' ? v2 : (v2 || v), tier: nextTier + '-escalated' } +} + +const results = [] +for (let b = 0; b < BATCHES.length; b++) { + log(`batch ${b + 1}/${BATCHES.length}: ${BATCHES[b].length} targets`) + const r = await parallel(BATCHES[b].map((t) => () => runOne(t))) + results.push(...r.filter(Boolean)) +} + +const ok = results.filter(Boolean) +const matched = ok.filter((r) => r.v && r.v.status === 'MATCH') +log(`s36-wave: ${matched.length}/${T.length} claim MATCH (the gate is the arbiter)`) + +return { + claimed_match: matched.map((r) => r.t.fn), + verdicts: ok.map((r) => ({ + fn: r.t.fn, ov: r.t.ov, nins: r.t.n, members: r.t.m, tier: r.tier, + status: r.v ? r.v.status : 'NONE', + closeness: r.v ? r.v.closeness : null, + klass: r.v ? r.v.klass : null, + levers: r.v ? r.v.levers : null, + index_hit: r.v ? r.v.index_hit : null, + index_gap: r.v ? r.v.index_gap : null, + summary: r.v ? r.v.summary : null, + })), +} diff --git a/.run/s37/ov_SC02_026/func_801825A4.c b/.run/s37/ov_SC02_026/func_801825A4.c new file mode 100644 index 000000000..99a6f3256 --- /dev/null +++ b/.run/s37/ov_SC02_026/func_801825A4.c @@ -0,0 +1,76 @@ +extern u8 *func_801290DC(s32 a0, u8 *a1); +extern void func_80049CAC(s32 a0, s32 a1); +extern void func_800484EC(s32 a0, s32 a1, s32 a2); +extern s32 rand(void); +extern u16 D_801AC2F0[][4]; +extern s32 D_801AC2E4[]; + +s32 func_801825A4(void *a0, s32 a1, s32 a2, s32 a3) { + u8 *obj; + s32 sub; + s32 rvA; + s16 rvB; + s32 p1; + + p1 = a1; + obj = func_801290DC(0x41, (u8 *)a0); + if (obj == 0) { + return 0; + } + { + s32 idx; + + rvA = rand(); + rvB = rvA; + sub = *(s32 *)(obj + 0x20); + *(s32 *)(sub + 0x20) = (s32) D_801AC2E4; + idx = a3 & 0xF; + { + u16 *tbl = D_801AC2F0[idx]; + *(u8 *)(sub + 0x27) = (u8) tbl[2]; + *(u16 *)(sub + 0x28) = tbl[0]; + *(u16 *)(sub + 0x2A) = tbl[1]; + } + + if ((a3 & 0x8000) != 0) { + s16 spd = (rvA * 0x10000 >> 0x10) % 0x200 + 0x400; + *(u16 *)(sub + 0x1A) = spd; + *(u16 *)(sub + 0x18) = spd; + *(s32 *)(obj + 0x34) = spd; + } else { + s16 spd = (rvA * 0x10000 >> 0x10) % 0xC00 + 0x400; + *(u16 *)(sub + 0x1A) = spd; + *(u16 *)(sub + 0x18) = spd; + *(s32 *)(obj + 0x34) = spd; + } + + { + s32 sign; + s32 h; + s32 mod128; + s16 srcvec[4]; + s32 buf[8]; + s32 trailing[4]; + + sign = -1; + if ((rvB & 1) != 0) { + sign = 1; + } + h = rvB; + mod128 = h % 128; + + srcvec[2] = 0; + trailing[1] = 0; + trailing[0] = 0; + trailing[2] = a2; + + srcvec[0] = (s16) (sign * mod128 - 0x300); + srcvec[1] = (s16) (p1 + sign * (h % 0x300)); + + func_80049CAC((s32) srcvec, (s32) buf); + func_800484EC((s32) buf, (s32) trailing, (s32) (obj + 0x10)); + *(s32 *)(obj + 0x1C) = 0x2D; + } + } + return (s32) obj; +} diff --git a/.run/s37/ov_SC02_028/func_8018B13C.c b/.run/s37/ov_SC02_028/func_8018B13C.c new file mode 100644 index 000000000..70e698fcd --- /dev/null +++ b/.run/s37/ov_SC02_028/func_8018B13C.c @@ -0,0 +1,16 @@ +typedef struct { s32 w[4]; } Blk16_8018B13C; + +extern s32 func_8004787C(s32 a0); +extern void func_80028620(s32, void *); +extern Blk16_8018B13C D_800A5EA8; +extern Blk16_8018B13C D_801D0320; +extern s32 D_800A5EB0; + +void func_8018B13C(s32 a0) { + Blk16_8018B13C *s1 = &D_800A5EA8; + + *s1 = D_801D0320; + D_800A5EB0 = func_8004787C(*(s16 *)(a0 + 0xFE)) * 10 / 4096 - 5; + *(s16 *)(a0 + 0xFE) = (*(u16 *)(a0 + 0xFE) + 0x71) & 0xFFF; + func_80028620(2, s1); +} diff --git a/.run/s37/ov_SC02_028/func_8018B354.c b/.run/s37/ov_SC02_028/func_8018B354.c new file mode 100644 index 000000000..0e82b8a78 --- /dev/null +++ b/.run/s37/ov_SC02_028/func_8018B354.c @@ -0,0 +1,16 @@ +typedef struct { s32 w[4]; } Blk16_8018B29C; + +extern s32 func_8004787C(s32 a0); +extern void func_80028620(s32, void *); +extern Blk16_8018B29C D_800A5E88; +extern Blk16_8018B29C D_801D03A8; +extern s32 D_800A5E90; + +void func_8018B354(s32 a0) { + Blk16_8018B29C *s1 = &D_800A5E88; + + *s1 = D_801D03A8; + D_800A5E90 = func_8004787C(*(s16 *)(a0 + 0xFE)) / 512 + 8; + *(s16 *)(a0 + 0xFE) = (*(u16 *)(a0 + 0xFE) + 0x71) & 0xFFF; + func_80028620(0, s1); +} diff --git a/.run/s37/ov_SC02_035/func_80184328.c b/.run/s37/ov_SC02_035/func_80184328.c new file mode 100644 index 000000000..d9546f2f5 --- /dev/null +++ b/.run/s37/ov_SC02_035/func_80184328.c @@ -0,0 +1,16 @@ +typedef struct { s32 w[4]; } Blk16_80184328; + +extern s32 func_8004787C(s32 a0); +extern void func_80028620(s32, void *); +extern s32 D_800A5E90; +extern Blk16_80184328 D_800A5E88; +extern Blk16_80184328 D_801BBE34; + +void func_80184328(s32 a0) { + Blk16_80184328 *s1 = &D_800A5E88; + + *s1 = D_801BBE34; + D_800A5E90 = func_8004787C(*(s16 *)(a0 + 0xFE)) * 20 / 4096 - 10; + *(s16 *)(a0 + 0xFE) = (*(u16 *)(a0 + 0xFE) + 0x71) & 0xFFF; + func_80028620(0, s1); +} diff --git a/.run/s37/ov_SC02_039/func_8017D930.c b/.run/s37/ov_SC02_039/func_8017D930.c new file mode 100644 index 000000000..2978eb90d --- /dev/null +++ b/.run/s37/ov_SC02_039/func_8017D930.c @@ -0,0 +1,37 @@ +extern s32 func_8017DBD0(u32 a0v); +extern s16 func_8017DB14(u32 a0); + +void func_8017D930(void *a0, void *a1) { + u16 *r = (u16 *)a0; + s16 *m = (s16 *)a1; + s16 cx; + s16 sx; + s16 cy; + s16 sy; + s16 cz; + s16 sz; + s32 sxsy; + s32 cxcz; + s32 cxsz; + + cx = func_8017DBD0(r[0] & 0xFFF); + sx = func_8017DB14(r[0] & 0xFFF); + cy = func_8017DBD0(r[1] & 0xFFF); + sy = func_8017DB14(r[1] & 0xFFF); + cz = func_8017DBD0(r[2] & 0xFFF); + sz = func_8017DB14(r[2] & 0xFFF); + + cxsz = (cx * sz) >> 15; + cxcz = (cx * cz) >> 15; + sxsy = (sx * sy) >> 15; + + m[0] = (cz * cy) >> 15; + m[1] = ((sxsy * cz) >> 15) - cxsz; + m[2] = ((cxcz * sy) >> 15) + ((sx * sz) >> 15); + m[3] = (sz * cy) >> 15; + m[4] = ((sxsy * sz) >> 15) + cxcz; + m[5] = ((cxsz * sy) >> 15) - ((sx * cz) >> 15); + m[6] = -sy; + m[7] = (cy * sx) >> 15; + m[8] = (cy * cx) >> 15; +} diff --git a/.run/s37/ov_SC02_041/func_801837B0.c b/.run/s37/ov_SC02_041/func_801837B0.c new file mode 100644 index 000000000..98d680145 --- /dev/null +++ b/.run/s37/ov_SC02_041/func_801837B0.c @@ -0,0 +1,231 @@ +/* func_801837B0 — ov_SC02_041, TU ov_SC02_041_jr_8017BEBC.c. + * + * Family of 4 (824 templatable ins): func_801837B0 / func_80186DD4 (ov_SC04_002) / + * func_801861A4 (ov_SC04_005) / func_80182810 (ov_SC04_007). + * + * Draws the two triangles of an entity's ground marker: builds a 4-vertex gouraud packet on + * the stack (verts 0x10, colours 0x30, code 0x40), whose colour ramp is derived from the s32 + * at a0+0x1C (*15, then >>1 for the second component, then -0x10 clamped-at-0 for the far + * pair). Bit 0 of the u16 at a0+0x70 picks which channel carries the bright value. Then the + * camera matrix cached at D_800AF630+0x18 is loaded and two rtpt+rtps+avsz4 groups run over a + * sliding window of the SVECTORs at a0+0xDC..0x104 (stride 8), queueing the packet through + * func_80017714 whenever the GTE flag masked with ~0x1000 is clear. + * + * DECLARATION SURFACE (whole-TU one-pass grep, D2): + * * D_800AF630 — TU line 53 `extern u8 D_800AF630[];` (identical extern repeated, legal) + * * func_80017714— TU line 4324 `extern void func_80017714(void *);` (identical extern repeated) + * * func_801837B0— TU line 5235 `extern void func_801837B0(s32 a0);` — the definition below + * uses that exact prototype, so no SELF-axis rename is needed. + * * every type tag and every gte_ macro carries the _801837B0 suffix: this TU already defines + * gte_ldv0 twice (lines 2687 and 4332, different bodies) and gte_SetRotMatrix at 5033, all + * ABOVE the splice point at 5246, so unsuffixed names would collide/redefine. + * + * CODEGEN NOTES (what the .s pins): + * * ONE aggregate `pkt` holds verts + colours + code. The code word MUST live inside the + * object handed to func_80017714 or gcc dead-stores it (a separate `s32 tag` local is + * eliminated and the frame comes out 8 short: 0x80 instead of 0x88). + * * `c` is s16 but `d` is s32 — that asymmetry IS in the target: after the -0x10 the `c < 0` + * test needs `sll $v1,$v1,16; bgez` (HImode) while the `d < 0` test is a bare `bgez $a3` + * (an `sra` by 17 already leaves 18 sign bits, so no re-extension is emitted). + * * `p`/`q` are a SECOND pair of s16 locals holding the value actually stored. They produce + * the `addu $a1,$v1,$zero` / `addu $a2,$a3,$zero` copies in the head and after each -0x10, + * and they occupy $a1/$a2 — which is what pushes the D_800AF630 base out to $t0 and the + * reload register out to $t1, exactly as the target has them. `q` must be s16, not s32: an + * s32 `q` is copy-propagated into `d` and the three `addu $a2,...` copies vanish. + * * The colour writes are CHAINED assignments; gcc expands `A = B = C = D = q` innermost-first, + * so the store order is D,C,B,A and exactly one QImode conversion temp is materialised per + * chain ($a0 for the first, $v0 reused for the rest). + * * The -0x10 on the second channel goes through a BLOCK-SCOPED `s32 e`, not `d -= 0x10`. + * Two separate effects, both required: + * - splitting `d` into `d` + `e` cuts REG_N_REFS(d) from 10 to 4, which drops d below p + * and q in global.c's allocno priority. With `d -= 0x10` the ranking is d,c,p,q and the + * three registers come out permuted ($a1/$a2/$a3 = d/p/q instead of p/q/d) — §137. + * - scoping `e` INSIDE each arm (one pseudo per arm instead of one shared) is what fixes + * the last residual: a function-scope `e` leaves `addiu $a3,$a3,-0x10` scheduled one slot + * too early, ahead of the `addu $v0,$a2,$zero` chain temp. That reorder was INVARIANT + * under every legal statement permutation (8 tried), i.e. it was never a source-order + * problem — it is the allocno/live-range split. + */ + +extern u8 D_800AF630[]; +extern void func_80017714(void *); + +typedef struct { s16 vx, vy, vz, pad; } SV_801837B0; /* 8 bytes */ +typedef struct { u8 r, g, b, cd; } CV_801837B0; /* 4 bytes */ + +typedef struct { + SV_801837B0 v[4]; /* +0x00 : sxy0..3 (v[0].vz doubles as the otz slot) */ + CV_801837B0 rgb[4]; /* +0x20 */ + u32 code; /* +0x30 */ + u32 pad; /* +0x34 */ +} PKT_801837B0; /* 0x38 */ + +#define gte_ldv0_801837B0(r0) __asm__ volatile ( \ + "lwc2 $0, 0( %0 );" \ + "lwc2 $1, 4( %0 )" \ + : \ + : "r"( r0 ) ) + +#define gte_ldv3_801837B0(r0, r1, r2) __asm__ volatile ( \ + "lwc2 $0, 0( %0 );" \ + "lwc2 $1, 4( %0 );" \ + "lwc2 $2, 0( %1 );" \ + "lwc2 $3, 4( %1 );" \ + "lwc2 $4, 0( %2 );" \ + "lwc2 $5, 4( %2 )" \ + : \ + : "r"( r0 ), "r"( r1 ), "r"( r2 ) ) + +#define gte_rtps_801837B0() __asm__ volatile ("nop;nop;rtps") +#define gte_rtpt_801837B0() __asm__ volatile ("nop;nop;rtpt") +#define gte_avsz4_801837B0() __asm__ volatile ("nop;nop;avsz4") + +#define gte_stsxy_801837B0(r0) __asm__ volatile ( \ + "swc2 $14, 0( %0 )" \ + : \ + : "r"( r0 ) \ + : "memory" ) + +#define gte_stsxy3_801837B0(r0, r1, r2) __asm__ volatile ( \ + "swc2 $12, 0( %0 );" \ + "swc2 $13, 0( %1 );" \ + "swc2 $14, 0( %2 )" \ + : \ + : "r"( r0 ), "r"( r1 ), "r"( r2 ) \ + : "memory" ) + +#define gte_stotz_801837B0(r0) __asm__ volatile ( \ + "swc2 $7, 0( %0 )" \ + : \ + : "r"( r0 ) \ + : "memory" ) + +#define gte_stflg_801837B0(r0) __asm__ volatile ( \ + "cfc2 $12, $31;" \ + "nop;" \ + "sw $12, 0( %0 )" \ + : \ + : "r"( r0 ) \ + : "$12", "memory" ) + +#define gte_SetRotMatrix_801837B0(r0) __asm__ volatile ( \ + "lw $12, 0( %0 );" \ + "lw $13, 4( %0 );" \ + "ctc2 $12, $0;" \ + "ctc2 $13, $1;" \ + "lw $12, 8( %0 );" \ + "lw $13, 12( %0 );" \ + "lw $14, 16( %0 );" \ + "ctc2 $12, $2;" \ + "ctc2 $13, $3;" \ + "ctc2 $14, $4" \ + : \ + : "r"( r0 ) \ + : "$12", "$13", "$14" ) + +#define gte_SetTransMatrix_801837B0(r0) __asm__ volatile ( \ + "lw $12, 20( %0 );" \ + "lw $13, 24( %0 );" \ + "ctc2 $12, $5;" \ + "lw $14, 28( %0 );" \ + "ctc2 $13, $6;" \ + "ctc2 $14, $7" \ + : \ + : "r"( r0 ) \ + : "$12", "$13", "$14" ) + +void func_801837B0(s32 a0) +{ + PKT_801837B0 pkt; /* 0x10 */ + s32 flag1; /* 0x48 */ + s32 flag2; /* 0x4C */ + s32 otz; /* 0x50 */ + u8 *base; + s32 *m; + s16 c, p, q; + s32 d; + + c = *(s32 *)(a0 + 0x1C) * 15; + base = D_800AF630; + p = c; + d = c >> 1; + q = d; + + if ((*(u16 *)(a0 + 0x70) & 1) == 0) { + pkt.rgb[0].b = pkt.rgb[1].b = p; + pkt.rgb[0].r = pkt.rgb[1].r = pkt.rgb[0].g = pkt.rgb[1].g = q; + c -= 0x10; + p = c; + { + s32 e; /* block-scoped: one pseudo per arm (§136/L1) */ + e = d - 0x10; + q = e; + if (c < 0) { + p = 0; + } + if (e < 0) { + q = 0; + } + } + pkt.rgb[2].b = pkt.rgb[3].b = p; + pkt.rgb[2].r = pkt.rgb[3].r = pkt.rgb[2].g = pkt.rgb[3].g = q; + } else { + pkt.rgb[0].r = pkt.rgb[1].r = p; + pkt.rgb[0].g = pkt.rgb[1].g = pkt.rgb[0].b = pkt.rgb[1].b = q; + c -= 0x10; + p = c; + { + s32 e; /* block-scoped: one pseudo per arm (§136/L1) */ + e = d - 0x10; + q = e; + if (c < 0) { + p = 0; + } + if (e < 0) { + q = 0; + } + } + pkt.rgb[2].r = pkt.rgb[3].r = p; + pkt.rgb[2].g = pkt.rgb[3].g = pkt.rgb[2].b = pkt.rgb[3].b = q; + } + + pkt.code = 0x50000000; + + m = (s32 *)(base + 0x18); + gte_SetRotMatrix_801837B0(m); + gte_SetTransMatrix_801837B0(m); + + gte_ldv3_801837B0((SV_801837B0 *)(a0 + 0xFC), (SV_801837B0 *)(a0 + 0x104), + (SV_801837B0 *)(a0 + 0xEC)); + gte_rtpt_801837B0(); + gte_stflg_801837B0(&flag1); + gte_stsxy3_801837B0(&pkt.v[0], &pkt.v[1], &pkt.v[2]); + gte_ldv0_801837B0((SV_801837B0 *)(a0 + 0xF4)); + gte_rtps_801837B0(); + gte_stflg_801837B0(&flag2); + flag1 = flag1 | flag2; + gte_stsxy_801837B0(&pkt.v[3]); + gte_avsz4_801837B0(); + gte_stotz_801837B0(&otz); + if ((flag1 & 0xFFFFEFFF) == 0) { + pkt.v[0].vz = (s16)otz; + func_80017714(&pkt.v[0]); + } + + gte_ldv3_801837B0((SV_801837B0 *)(a0 + 0xEC), (SV_801837B0 *)(a0 + 0xF4), + (SV_801837B0 *)(a0 + 0xDC)); + gte_rtpt_801837B0(); + gte_stflg_801837B0(&flag1); + gte_stsxy3_801837B0(&pkt.v[0], &pkt.v[1], &pkt.v[2]); + gte_ldv0_801837B0((SV_801837B0 *)(a0 + 0xE4)); + gte_rtps_801837B0(); + gte_stflg_801837B0(&flag2); + flag1 = flag1 | flag2; + gte_stsxy_801837B0(&pkt.v[3]); + gte_avsz4_801837B0(); + gte_stotz_801837B0(&otz); + if ((flag1 & 0xFFFFEFFF) == 0) { + pkt.v[0].vz = (s16)otz; + func_80017714(&pkt.v[0]); + } +} diff --git a/.run/s37/ov_SC03_024/func_80181914.c b/.run/s37/ov_SC03_024/func_80181914.c new file mode 100644 index 000000000..1ef85a1ca --- /dev/null +++ b/.run/s37/ov_SC03_024/func_80181914.c @@ -0,0 +1,107 @@ +/* func_80181914 — byte-verified twin of func_8017D77C (src/ov_SC04_018/ov_SC04_018_jr_8017AE2C.c:4271), + * found via §136e/S34 magic-word grep on the `AAA32294` instruction bytes (the + * D_800AF630+0xA3AA far-offset lhu). func_8017D77C is the banked member of this + * family; this body was derived by remapping its symbols: + * func_8017D77C -> func_80181914 D_801E6F58 -> D_801C12C0 + * func_8017D9B8 -> func_80181B50 func_8017DAC4 -> func_80181C5C + * Ent_8017D6EC -> Ent_80181914 (own function-suffixed typedef -- same 0x24-stride + * shape as this TU's own Ent_8017D6EC_80181884 at TU:4633-4639, + * given a fresh name/no shared tag so splicing this in below that + * file-scope typedef can't hit the C89 duplicate-typedef error, S33). + * func_80181C5C is already banked in THIS TU (defined further down as + * `void func_80181C5C(void *arg0)`), matching the direct call below with no cast + * needed. func_80181B50 has no declaration anywhere in this TU but its own + * INCLUDE_ASM stub, so it is declared directly as `void (void *)` (no idiom-9 cast + * needed, unlike the twin's func_8017D9B8 which was `void (void)`). + */ + +extern u8 D_80078EB1; +extern u8 D_80078E78[]; +extern u8 D_800AF630[]; +extern s32 func_8004787C(s32 a0); +extern void func_8012AD44(s32 *a0, s16 a1); +extern void func_80181B50(void *arg0); +extern void func_80181C5C(void *arg0); + +typedef struct { + u8 unk00[0x16]; + s16 unk16; + u8 unk18[4]; + s32 unk1C; + u8 unk20[4]; +} Ent_80181914; + +extern Ent_80181914 D_801C12C0[]; + +void func_80181914(s32 *arg0) +{ + u8 *m = D_800AF630; + u8 *q = D_80078E78; + Ent_80181914 *p; + Ent_80181914 *r; + s32 i; + s32 j; + s32 n; + s32 h; + s32 c; + s32 e; + + if (D_80078EB1 >= 9) { + n = 0; + for (j = 0; j < 4; j++) { + r = &D_801C12C0[j]; + if (r->unk1C == 0) { + n++; + } + } + if (n == 0) { + func_8012AD44(arg0, 0); + return; + } + } + + i = 0; + do { + p = (Ent_80181914 *)((s32)D_801C12C0 + i * 0x24); + if (p->unk1C == 0) { + if (*(s16 *)((s32)p + 0xC) > 0x400) { + if (*(s32 *)((s32)p + 0x18) != 0) { + *(s32 *)((s32)p + 0x18) = *(s32 *)((s32)p + 0x18) - 1; + } else { + *(s16 *)((s32)p + 0xC) = *(s16 *)((s32)p + 0xC) + 11; + } + } else { + *(s16 *)((s32)p + 0xC) = *(s16 *)((s32)p + 0xC) + 11; + } + if (*(s16 *)((s32)p + 0xC) > 0x800) { + *(s16 *)((s32)p + 0xC) = 0; + } + c = (func_8004787C(*(s16 *)((s32)p + 0xC)) / 64) & 0xFF; + c = c | ((c << 16) | (c << 8)); + *(s32 *)((s32)p + 0x4) = c; + c = (func_8004787C(*(s16 *)((s32)p + 0xC)) / 256) & 0xFF; + c = c | ((c << 16) | (c << 8)); + *(s32 *)((s32)p + 0x8) = c; + if ((*(u16 *)(m + 0xA3AA) & 1) == 0) { + e = *(u16 *)((s32)p + 0xE) + 1; + *(u16 *)((s32)p + 0xE) = e; + if ((s16)e >= 0x40) { + *(s16 *)((s32)p + 0xE) = 0; + } + } + h = *(s16 *)((s32)p + 0xC); + if (h == 0) { + p->unk1C = 1; + } else if (h > 0x555) { + if (*(s32 *)((s32)p + 0x20) == 0) { + *(s32 *)((s32)p + 0x20) = 1; + if (q[0x39] < 9) { + func_80181B50(p); + } + } + } + func_80181C5C(p); + } + i++; + } while (i < 4); +} diff --git a/.run/s37/ov_SC03_092/func_80184370.c b/.run/s37/ov_SC03_092/func_80184370.c new file mode 100644 index 000000000..8260fc63a --- /dev/null +++ b/.run/s37/ov_SC03_092/func_80184370.c @@ -0,0 +1,189 @@ +/* func_80184370 — ov_SC03_092 (263 ins, jr-function). + * Family twin of the BANKED exemplar func_80182FD4 (ov_SC02_028, + * src/ov_SC02_028/ov_SC02_028_jr_8017D898.c:4184) — same h_seq template + * (family_hseq.json, diff_class PURE, 7 addr instances), body reused verbatim + * with only the tail-call target (func_8018478C here vs func_801833F0 there, + * the per-overlay sibling in the same wave) and local type-tag suffixes + * renamed. See docs/matching-cookbook.md S34/S35/S36 (magic-grep / sibling + * reuse) — this is the STEP 0 win: the whole body is a byte-proven template. + * + * 12-point (6-segment) ribbon/trail projector. + * a0 = entity, a1 = ctx. Projects the entity origin (RTPS) into buf[0], then + * the 12 offset points at (*(a0+0x20))->pt[0..11] into buf[1..12]. + * + * Gate 1: RTPS flag & ~0x1000 must be clear. + * Gate 2: z = otz + 1, biased by the 12-bit field of the u16 @0x2C according + * to its top two bits (0xC000 = subtract & clamp at 0, else add), and + * the whole draw is dropped unless z < 0x1000. + * Bit 0x8000 of the s16 @0x1E picks the "per-segment OT" variant: the six EVEN + * points go through RTPT three-at-a-time (screen xy only) and the six ODD ones + * go through RTPS one at a time, each also depositing its own otz+1 into + * buf[13..18]; func_8018478C then gets the OT *base*. Otherwise all 12 points + * go through RTPT three-at-a-time and func_8018478C gets the single slot &ot[z]. + * + * buf is ONE flat 19-word packet (76 bytes; gcc rounds the BLKmode stack slot + * up to 80) — the OT indices are written as buf[13 + (i >> 1)]. + */ + +#define gte_ldv0(r0) __asm__ volatile ( \ + "lwc2 $0, 0( %0 );" \ + "lwc2 $1, 4( %0 )" \ + : \ + : "r"( r0 ) ) + +#define gte_ldv3(r0, r1, r2) __asm__ volatile ( \ + "lwc2 $0, 0( %0 );" \ + "lwc2 $1, 4( %0 );" \ + "lwc2 $2, 0( %1 );" \ + "lwc2 $3, 4( %1 );" \ + "lwc2 $4, 0( %2 );" \ + "lwc2 $5, 4( %2 )" \ + : \ + : "r"( r0 ), "r"( r1 ), "r"( r2 ) ) + +#define gte_rtps() __asm__ volatile ("nop;nop;rtps") +#define gte_rtpt() __asm__ volatile ("nop;nop;rtpt") + +#define gte_stsxy(r0) __asm__ volatile ( \ + "swc2 $14, 0( %0 )" \ + : \ + : "r"( r0 ) \ + : "memory" ) + +#define gte_stsxy3(r0, r1, r2) __asm__ volatile ( \ + "swc2 $12, 0( %0 );" \ + "swc2 $13, 0( %1 );" \ + "swc2 $14, 0( %2 )" \ + : \ + : "r"( r0 ), "r"( r1 ), "r"( r2 ) \ + : "memory" ) + +#define gte_stszotz(r0) __asm__ volatile ( \ + "mfc2 $12, $19;" \ + "nop;" \ + "sra $12, $12, 2;" \ + "sw $12, 0( %0 )" \ + : \ + : "r"( r0 ) \ + : "$12", "memory" ) + +#define gte_stflg(r0) __asm__ volatile ( \ + "cfc2 $12, $31;" \ + "nop;" \ + "sw $12, 0( %0 )" \ + : \ + : "r"( r0 ) \ + : "$12", "memory" ) + +void func_80184370(void *a0, void *a1) +{ + typedef struct { u16 vx, vy; } Pt2_80184370; + typedef struct { u8 pad[0x10]; Pt2_80184370 pt[12]; } Src_80184370; + typedef struct { u16 vx, vy, vz, pad; } Vec8_80184370; + + extern void func_8001E094(void); + extern void func_8001E378(void *a0); + extern void func_8018478C(void *a0, void *a1, u32 *a2, u32 *a3); + extern u8 D_800A6610[]; + extern short D_800B9A02; + + Vec8_80184370 base; + Vec8_80184370 v[3]; + u32 buf[19]; + long flag; + long otz; + long flag2; + + Src_80184370 *s; + u32 *ot; + s32 z; + s32 i; + u16 w; + + ot = (u32 *)&D_800A6610[(*(u16 *)&D_800B9A02) << 14]; + s = *(Src_80184370 **)((s32)a0 + 0x20); + if (s == 0) { + return; + } + + if (*(s32 *)((s32)a0 + 0x34) != 0) { + func_8001E094(); + } else { + func_8001E378(a0); + } + + base.vx = *(u16 *)((s32)a0 + 0x2E); + base.vy = *(u16 *)((s32)a0 + 0x30); + base.vz = *(u16 *)((s32)a0 + 0x32); + + gte_ldv0(&base); + gte_rtps(); + gte_stsxy(&buf[0]); + gte_stflg(&flag); + gte_stszotz(&otz); + + if (flag & ~0x1000) { + return; + } + + w = *(u16 *)((s32)a0 + 0x2C); + z = otz + 1; + if ((w & 0xC000) != 0) { + if ((w & 0xC000) == 0xC000) { + z -= (w & 0xFFF); + if (z < 0) { + z = 0; + } + } else { + z += (w & 0xFFF); + } + } + if (z >= 0x1000) { + return; + } + + if (*(s16 *)((s32)a0 + 0x1E) & 0x8000) { + for (i = 0; i < 12; i += 6) { + v[0].vx = base.vx + s->pt[i].vx; + v[0].vy = base.vy + s->pt[i].vy; + v[0].vz = base.vz; + v[1].vx = base.vx + s->pt[i + 2].vx; + v[1].vy = base.vy + s->pt[i + 2].vy; + v[1].vz = base.vz; + v[2].vx = base.vx + s->pt[i + 4].vx; + v[2].vy = base.vy + s->pt[i + 4].vy; + v[2].vz = base.vz; + gte_ldv3(&v[0], &v[1], &v[2]); + gte_rtpt(); + gte_stsxy3(&buf[i + 1], &buf[i + 3], &buf[i + 5]); + } + for (i = 1; i < 12; i += 2) { + v[0].vx = base.vx + s->pt[i].vx; + v[0].vy = base.vy + s->pt[i].vy; + v[0].vz = base.vz; + gte_ldv0(&v[0]); + gte_rtps(); + gte_stsxy(&buf[i + 1]); + gte_stflg(&flag2); + gte_stszotz(&otz); + buf[13 + (i >> 1)] = otz + 1; + } + func_8018478C(a0, a1, &buf[0], ot); + } else { + for (i = 0; i < 12; i += 3) { + v[0].vx = base.vx + s->pt[i].vx; + v[0].vy = base.vy + s->pt[i].vy; + v[0].vz = base.vz; + v[1].vx = base.vx + s->pt[i + 1].vx; + v[1].vy = base.vy + s->pt[i + 1].vy; + v[1].vz = base.vz; + v[2].vx = base.vx + s->pt[i + 2].vx; + v[2].vy = base.vy + s->pt[i + 2].vy; + v[2].vz = base.vz; + gte_ldv3(&v[0], &v[1], &v[2]); + gte_rtpt(); + gte_stsxy3(&buf[i + 1], &buf[i + 2], &buf[i + 3]); + } + func_8018478C(a0, a1, &buf[0], ot + z); + } +} diff --git a/.run/s37/ov_SC03_092/func_8018478C.c b/.run/s37/ov_SC03_092/func_8018478C.c new file mode 100644 index 000000000..6691f5533 --- /dev/null +++ b/.run/s37/ov_SC03_092/func_8018478C.c @@ -0,0 +1,164 @@ +#include "common.h" + +typedef struct { + u32 addr : 24; + u32 len : 8; +} PTag_8018478C; + +typedef struct { + PTag_8018478C tag; + u8 r0, g0, b0, code; + u16 x0, y0; + u8 u0, v0; + u16 clut; + u16 x1, y1; + u8 u1, v1; + u16 tpage; + u16 x2, y2; + u8 u2, v2; + u16 pad2; + u16 x3, y3; + u8 u3, v3; + u16 pad3; +} Ft4_8018478C; + +typedef struct { + PTag_8018478C tag; + u32 code0; +} Drm_8018478C; + +#define ADDPRIM_8018478C(o, p) \ + (((PTag_8018478C *)(p))->addr = ((PTag_8018478C *)(o))->addr, \ + ((PTag_8018478C *)(o))->addr = (u32)(p)) + +/* lhu / sll 16 / sra 19 : signed 13-bit field held at bit 3 of a u16 */ +#define SR3_8018478C(a) (((s32)(*(u16 *)(a) << 16)) >> 19) + +void func_8018478C(void *ent, void *spr, u32 *q, u32 *ot) +{ + extern u8 *D_800A5E60; + register s32 zr __asm__("$0"); + + Ft4_8018478C *poly; + Drm_8018478C *dm; + s32 vtx; + u32 flags; + s32 tp, abr, code, shift; + u16 x; + u32 y; + s32 tpage, clut, cy; + s32 su, sv, uu, vv; + u16 fl; + s32 i, m; + u32 b; + + poly = (Ft4_8018478C *)D_800A5E60; + fl = *(u16 *)((s32)ent + 0x1E) & 0x8000; + flags = *(u32 *)((s32)ent + 4); + tp = (flags >> 24) & 3; + shift = 2 - tp; + x = *(u16 *)((s32)ent + 0x28) + (*(s16 *)((s32)spr + 4) >> shift); + y = *(u16 *)((s32)ent + 0x2A) + *(u16 *)((s32)spr + 6); + vtx = *(s32 *)((s32)ent + 0x20); + D_800A5E60 += 0xF0; + if (flags & 0x40000000) { + code = 0x2E; + abr = (flags >> 28) & 3; + } else { + code = 0x2C; + abr = 1; + } + tpage = (tp << 7) | (abr << 5) | ((y & 0x100) >> 4) | ((x & 0x3C0) >> 6) | + ((y & 0x200) << 2); + b = *(u8 *)((s32)ent + 0x27); + cy = (b + 0x100) << 6; + if (b < 0xE0) { + clut = cy | 0x16; + } else { + clut = cy | 0x10; + } + { + s32 t = ((x - ((tpage & 0xF) << 6)) << shift) + + (*(u16 *)((s32)spr + 4) & ((1 << shift) - 1)); + su = t + zr; + __asm__ __volatile__("" ::"r"(t + *(u8 *)((s32)spr + 2) - 1)); + __asm__ __volatile__("" ::"r"(t)); + } + y = y & 0xFFFF; + if (tpage & 0x10) { + sv = y - 0x100; + } else { + sv = y + zr; + } + __asm__ __volatile__("" ::"r"(sv + *(u8 *)((s32)spr + 3) - 1)); + + i = 1; + uu = su; + vv = sv; + for (; i < 13; i += 2, poly++) { + s32 k = (i - 1) * 4; + __asm__ __volatile__("" ::"r"(i)); + poly->tag.len = 9; + poly->code = code; + poly->tpage = tpage; + poly->u2 = su; + poly->v2 = sv; + poly->u0 = SR3_8018478C(vtx + k + 0x10) + uu; + poly->v0 = SR3_8018478C(vtx + k + 0x12) + vv; + poly->u1 = SR3_8018478C(vtx + i * 4 + 0x10) + uu; + poly->v1 = SR3_8018478C(vtx + i * 4 + 0x12) + vv; + /* k2 is recorded as a giv HERE, AFTER the u1/v1 mem giv. */ + { + s32 k2 = (i + 1) * 4; + poly->u3 = SR3_8018478C(vtx + k2 + 0x10) + uu; + poly->v3 = SR3_8018478C(vtx + k2 + 0x12) + vv; + } + poly->clut = clut; + __asm__ __volatile__("" ::"r"(i), "r"(i)); + poly->r0 = *(u8 *)((s32)ent + 0x24); + poly->g0 = *(u8 *)((s32)ent + 0x25); + poly->b0 = *(u8 *)((s32)ent + 0x26); + poly->x0 = ((u16 *)q)[i * 2]; + poly->y0 = ((u16 *)q)[i * 2 + 1]; + poly->x1 = ((u16 *)q)[i * 2 + 2]; + poly->y1 = ((u16 *)q)[i * 2 + 3]; + poly->x2 = ((u16 *)q)[0]; + poly->y2 = ((u16 *)q)[1]; + poly->x3 = ((u16 *)q)[i * 2 + 4]; + poly->y3 = ((u16 *)q)[i * 2 + 5]; + if (fl != 0) { + ADDPRIM_8018478C(&ot[*(s32 *)((s32)q + 0x34 + (i >> 1) * 4)], poly); + } else { + ADDPRIM_8018478C(ot, poly); + } + } + poly[-1].x3 = ((u16 *)q)[2]; + poly[-1].y3 = ((u16 *)q)[3]; + poly[-1].u3 = su + SR3_8018478C(vtx + 0x10); + poly[-1].v3 = sv + SR3_8018478C(vtx + 0x12); + + if (flags & 0x40000000) { + dm = (Drm_8018478C *)poly; + if (fl != 0) { + m = 0; + D_800A5E60 += 0x30; + do { + m++; + dm->tag.len = 1; + dm->code0 = (abr << 5) | 0xE100000A; + ADDPRIM_8018478C(&ot[*(volatile s32 *)((s32)q + 0x34)], dm); + q = (u32 *)((s32)q + 4); + dm++; + } while (m < 6); + } else { + dm->code0 = (abr << 5) | 0xE100000A; + D_800A5E60 += 8; + dm->tag.len = 1; + ADDPRIM_8018478C(ot, dm); + /* Zero-byte live-range stretch: puts vtx's allocno priority + * inside the only admissible window. Placement is + * load-bearing; do not move this statement. */ + __asm__ __volatile__("" ::"r"(vtx)); + } + } +} diff --git a/.run/s37/ov_SC03_108/func_8017F2E8.c b/.run/s37/ov_SC03_108/func_8017F2E8.c new file mode 100644 index 000000000..931ff62d1 --- /dev/null +++ b/.run/s37/ov_SC03_108/func_8017F2E8.c @@ -0,0 +1,171 @@ +void func_8017F2E8(s32 param_1) +{ + extern void func_801808B4(void *a0); + extern void func_8012B178(s32 a0, s32 a1); + extern s32 func_8012CBA4(s32 a0); + extern s32 func_8012B608(s32 a0, s32 a1, s32 a2); + extern s32 func_8012BEE8(s32 a0); + extern void func_8012ADE4(u8 *a0); + extern s32 func_8012BD3C(s32 a0, s32 a1, s32 a2); + extern void func_8012B23C(s32 a0); + extern s32 func_80143B6C(s32 a0, s32 a1); + extern void func_80131E00(s32 a0, s32 a1); + + s32 sVar1; + s32 iVar1; + u16 state; + s32 uVar2; + s32 pad[4]; + + if (*(s16 *)(param_1 + 0xa) >= 0x10) { + func_801808B4((void *)param_1); + return; + } + + sVar1 = func_8004787C((*(s32 *)(param_1 + 0xe4) << 6) & 0x7c0); + *(s16 *)(*(s32 *)(param_1 + 0x20) + 0x1a) = (sVar1 >> 1) + 0x800; + + state = *(u16 *)(param_1 + 0x34); + switch (state) { + case 0: { + s32 iVar3 = 0x1000 - sVar1; + s32 a1; + if (iVar3 >= 0) { + a1 = -0x8000 - (iVar3 << 4); + } else { + a1 = -0x8000 - ((sVar1 - 0x1000) << 4); + } + func_8012B178(param_1, a1); + iVar1 = func_8012CBA4(param_1); + if ((iVar1 & 0xff) == 0x1a) { + goto EXIT_808B4; + } + if ((iVar1 & 0x6000) == 0) { + goto MERGE_4C0; + } + if (--*(s32 *)(param_1 + 0xe8) != 0) { + goto TAIL_5D8; + } + { + s32 iVar2; + s32 nVar; + s32 sVar2; + + *(u16 *)(param_1 + 0x34) = 1; + *(s32 *)(param_1 + 0x1c) = 0x20; + iVar2 = rand(); + nVar = iVar2 % 1024; + sVar2 = *(s16 *)(*(s32 *)(param_1 + 0x20) + 0x12); + if (rand() & 1) { + uVar2 = sVar2 + nVar; + } else { + uVar2 = sVar2 - nVar; + } + } + goto MERGE_4E0; + } + + case 1: { + s32 ptr0 = *(s32 *)(param_1 + 0x20); + s32 arg0 = *(s16 *)(ptr0 + 0x12); + s32 ret = func_8012B608(arg0, *(s32 *)(param_1 + 0xe0), 0x14); + s32 ptr1 = *(s32 *)(param_1 + 0x20); + s32 sum; + + sum = *(u16 *)(ptr1 + 0x12) + ret; + __asm__ __volatile__("" : "=r"(sum) : "0"(sum)); + ret = 0x1000 - sVar1; + *(s16 *)(ptr1 + 0x12) = sum; + if (ret < 0) { + ret = sVar1 - 0x1000; + } + ret = ret << 4; + ret = -ret; + { + s32 arg1 = ret - 0x4000; + __asm__ __volatile__("" : "=r"(arg1) : "0"(arg1)); + func_8012B178(param_1, arg1); + } + iVar1 = func_8012CBA4(param_1); + if ((iVar1 & 0xff) == 0x1a) { + goto EXIT_808B4; + } + if ((iVar1 & 0x6000) != 0) { + goto MERGE_4E8; + } + } + MERGE_4C0: + func_8012ADE4((u8 *)param_1); + { + /* §17 pin: the merge-block pointer must land in $v1 so reload's + * scratch for the constant store takes $v0 (see .L8017F4C0). */ + register s32 ptr2 __asm__("$3"); + ptr2 = *(s32 *)(param_1 + 0x20); + *(u16 *)(param_1 + 0x34) = 2; + uVar2 = *(s16 *)(ptr2 + 0x12) + 0x800; + } + MERGE_4E0: + *(s32 *)(param_1 + 0xe0) = uVar2; + goto TAIL_5D8; + MERGE_4E8: + if (func_8012BEE8(param_1) != 0) { + *(u16 *)(param_1 + 0x34) = 0; + *(s32 *)(param_1 + 0xe8) = 0x40; + } + goto TAIL_5D8; + + case 2: { + s32 ptrA = *(s32 *)(param_1 + 0x20); + s32 arg1 = *(s32 *)(param_1 + 0xe0); + s32 arg0 = *(s16 *)(ptrA + 0x12); + s32 ret = func_8012B608(arg0, arg1, 0xa); + + if (ret == 0) { + *(u16 *)(param_1 + 0x34) = 0; + } + { + s32 ptrB = *(s32 *)(param_1 + 0x20); + *(s16 *)(ptrB + 0x12) = *(u16 *)(ptrB + 0x12) + ret; + } + goto TAIL_5D8; + } + + case 3: + iVar1 = func_8012CBA4(param_1); + if ((iVar1 & 0xff) != 0x1a) { + goto NOT_1A; + } + EXIT_808B4: + func_801808B4((void *)param_1); + return; + NOT_1A: + if ((iVar1 & 0x2000) == 0) { + goto L8017F598; + } + func_8012B23C(param_1); + func_8012B178(param_1, 0xFFFD8000); + *(u16 *)(param_1 + 0x34) = 0; + goto TAIL_5D8; + L8017F598: + if ((*(s32 *)(param_1 + 0x1c) & 3) != 0) { + goto L8017F5B4; + } + func_80143B6C(param_1, 1); + L8017F5B4: + *(s32 *)(param_1 + 0x1c) += 1; + if (*(s32 *)(param_1 + 0x1c) < 0x3c) { + goto TAIL_5DC; + } + func_80131E00(param_1, 0xd); + } + +TAIL_5D8: +TAIL_5DC: + if (func_8012BD3C(param_1, 0x400, 0x24000) == 1) { + *(s16 *)(param_1 + 2) = 3; + } + if (0x100000 < *(s32 *)(param_1 + 0x14)) { + *(s32 *)(param_1 + 0x14) = 0x100000; + } + *(s32 *)(param_1 + 0xe4) += 1; +} diff --git a/.run/s37/ov_SC03_124/func_8018893C.c b/.run/s37/ov_SC03_124/func_8018893C.c new file mode 100644 index 000000000..10c0cced6 --- /dev/null +++ b/.run/s37/ov_SC03_124/func_8018893C.c @@ -0,0 +1,149 @@ +/* func_8018893C — twin of the byte-matched func_8018BCD4 in ov_SC03_001/ov_SC03_001_jr_8017AE2C.c + * (family of 4: this fn + ov_SC04_018/func_80189214 + ov_SC04_019/func_80189214 + + * ov_SC05_017/func_80188F14; all normalize byte-identical to this shape). Only the two + * overlay-local data symbols differ (D_801BC522/D_801BC536 here vs D_801C11BA/D_801C11CE + * there). Caller in this TU (func_801880E0) already declares the matching prototype: + * extern s32 *func_8018893C(s32 *, void *, s32, void *, s32); + */ + +extern s32 *func_8018893C(s32 *, void *, s32, void *, s32); + +s32 *func_8018893C(out, src, idx, w, col) + s32 *out; + void *src; + s16 idx; + void *w; + s32 col; +{ + extern short D_800B9A02; /* TU-visible spelling (file scope) */ + extern s16 func_8014168C(s16 a0); + extern u16 D_8011511A; + extern u16 D_80115116; + extern u8 D_80115138[]; + extern u8 D_80115140[]; + extern u8 D_80115148[]; + extern u8 D_80115158[]; + extern u16 D_801BC522; + extern u16 D_801BC536; + + typedef struct { + u32 *ot; /* 0x00 */ + u32 pad[4]; /* 0x04..0x13 -> 0x14 stride */ + } Env_8018893C; + extern Env_8018893C D_800AE7BC[]; + + typedef struct { u32 addr : 24; u32 len : 8; } PTag_8018893C; + + /* §37 /s-DEP LATTICE: this store MUST be MEM_IN_STRUCT_P. `out[0] = k` folds `out + 0` + * away and expands as a NON-/s mem, which keeps the true_dependence edge to the incoming + * 5th-arg slot and pins `lw $v1,0x38($sp)` after it. The target has the arg load AFTER + * the out[0] store, i.e. the edge must be DROPPED — sched.c's drop clause needs the store + * /s + varying and the load non-/s + fixed-address. A COMPONENT_REF at offset 0 supplies + * the /s. */ + typedef struct { u32 w; } W_8018893C; + + s32 c; + s16 t; + u16 *q; + s32 m; + volatile u16 *pbh; + + c = D_80115138[idx]; + + ((W_8018893C *)out)->w = 0x04000000; + *((u8 *)out + 0xC) = 0x30; + *((u8 *)out + 0xD) = 0x48; + *(s16 *)((u8 *)out + 0xE) = 0x4056; + out[1] = col | 0x64000000; + + if (c < 10) { + /* §135 idiom 9/T3: the fleet decl returns s16; cast at the CALL so the return feeds + * `sll $v0,$v0,1` with no `andi`/re-extend. */ + t = ((s32 (*)(s16)) func_8014168C)(idx) * 2; + } else { + t = (D_80115148[idx * 2] - D_80115140[idx]) * 2; + } + + q = (u16 *)(t * 2 + (s32)src); + pbh = (volatile u16 *)&D_800B9A02; + *(s16 *)((u8 *)out + 0x8) = q[0] - 8; + *(s16 *)((u8 *)out + 0xA) = q[1]; + *(s16 *)((u8 *)out + 0x12) = 8; + *(s16 *)((u8 *)out + 0x10) = 8; + + /* §36 BITFIELD STORE = THE MASK-ORDER DECOUPLER. The head materializes 0x00FFFFFF + * (lui+ori) BEFORE 0xFF000000 (lui) while the body ANDs the dest first. `volatile` on the + * index read is what keeps the SECOND *pbh load alive (a plain read cse-folds and the + * fn loses 5 ins). */ + ((PTag_8018893C *)out)->addr = + ((PTag_8018893C *)(D_800AE7BC[*pbh].ot + 2))->addr; + ((PTag_8018893C *)(D_800AE7BC[*pbh].ot + 2))->addr = (u32)out; + + out += 5; + if (c < 10) { + return out; + } + + m = D_8011511A; + if (m == idx) { + if ((D_80115116 & 8) != 0) { + s16 j; + s32 k; + s16 y; + s16 eight; + u16 *pb; + /* §137/§36: m24 is a 2-insn constant, but the natural allocno order puts j/k ahead + * of it and rotates $a2/$a3/$t0. Pinning m24 alone restores the whole rotation. */ + register u32 m24 __asm__("$6"); + u32 mhi; + s32 vsum; + + j = 0; + k = m; + eight = 8; + pb = (u16 *)&D_800B9A02; + m24 = 0xFFFFFF; + mhi = 0xFF000000; + /* NOTE: the loop deliberately has NO source pointer of its own. gcc's combine_givs + * folds every `out + K` reference into ONE address giv based at out+0x12; adding a + * `u8 *p = (u8*)out + 0x12` biv makes `*(s16*)p` use the biv directly, which + * survives as a SECOND register (the +1-instruction LENGTH-DRIFT). The giv's base + * is the LAST such reference in source order — hence the +8 (`vsum`) store sits + * before the 0x10/0x12 stores here even though the scheduler sinks it back. */ + for (; j < 2; j++) { + if (j == 0) { + if (D_80115140[k] == 0) { + continue; + } + *((u8 *)out + 0xD) = 0x30; + y = D_801BC522 - 2; + } else { + s32 k2 = k * 2; + if ((((s8 *)D_80115158)[k2] - ((s8 *)D_80115140)[k]) < 7) { + continue; + } + *((u8 *)out + 0xD) = 0x38; + y = D_801BC536 + 1; + } + *(s16 *)((u8 *)out + 0xA) = y; + __asm__("" ::: "memory"); /* keeps the y store ahead of the 0x64808080 pair */ + *(u32 *)out = 0x04000000; + *((u8 *)out + 0xC) = 0x78; + *(u32 *)((u8 *)out + 0x4) = 0x64808080; + *(s16 *)((u8 *)out + 0xE) = 0x4056; + vsum = *(s32 *)((u8 *)w + 8) + *(s32 *)((u8 *)w + 0xC) - 0xC; + *(s16 *)((u8 *)out + 0x8) = vsum; + *(s16 *)((u8 *)out + 0x10) = eight; + *(s16 *)((u8 *)out + 0x12) = eight; + *(u32 *)out = ((*(u32 *)out) & mhi) | (D_800AE7BC[*pb].ot[2] & m24); + { + register u32 *op __asm__("$4"); + op = D_800AE7BC[*pb].ot; + op[2] = (op[2] & mhi) | (((u32)out) & m24); + } + out += 5; + } + } + } + return out; +} diff --git a/.run/s37/ov_SC06_008/func_8017ED80.c b/.run/s37/ov_SC06_008/func_8017ED80.c new file mode 100644 index 000000000..f7c0c71e2 --- /dev/null +++ b/.run/s37/ov_SC06_008/func_8017ED80.c @@ -0,0 +1,54 @@ +#include "common.h" + +/* TU declares func_8017ED80 as void(void) (S35 self-axis) but the asm takes + * $a0 as a pointer parameter — bind through a private C name. */ +extern void aF8017ED80(void *param_1) __asm__("func_8017ED80"); + +void aF8017ED80(void *param_1) { + u8 *a0 = (u8 *)param_1; + s32 pad_[4]; + s32 iVar2; + u32 uVar1; + + if (*(u16 *)(a0 + 0x0) != 0) { + iVar2 = *(s32 *)(a0 + 0xCC); + *(u16 *)(iVar2 + 0x8) = *(u16 *)(a0 + 0x6); + *(u16 *)(iVar2 + 0xA) = *(u16 *)(a0 + 0xA); + *(u16 *)(iVar2 + 0xC) = *(u16 *)(a0 + 0xE); + *(u16 *)(iVar2 + 0x10) = *(u16 *)(*(s32 *)(a0 + 0x20) + 0x10); + *(u16 *)(iVar2 + 0x12) = *(u16 *)(*(s32 *)(a0 + 0x20) + 0x12) + *(u16 *)(a0 + 0xFC); + *(u16 *)(iVar2 + 0x14) = *(u16 *)(*(s32 *)(a0 + 0x20) + 0x14); + *(u16 *)(iVar2 + 0x18) = *(u16 *)(*(s32 *)(a0 + 0x20) + 0x18); + *(u16 *)(iVar2 + 0x1A) = *(u16 *)(*(s32 *)(a0 + 0x20) + 0x1A); + *(u16 *)(iVar2 + 0x1C) = *(u16 *)(*(s32 *)(a0 + 0x20) + 0x1C); + if (*(s32 *)(*(s32 *)(a0 + 0x20) + 0x4) < 0) { + register u32 val __asm__("$2"); + val = *(u32 *)(iVar2 + 0x4); + uVar1 = val | 0x80000000; + } else { + __asm__ __volatile__(""); + uVar1 = *(u32 *)(iVar2 + 0x4) & 0x7FFFFFFF; + } + *(u32 *)(iVar2 + 0x4) = uVar1; + + iVar2 = *(s32 *)(a0 + 0xD0); + *(u16 *)(iVar2 + 0x8) = *(u16 *)(a0 + 0x6); + *(u16 *)(iVar2 + 0xA) = *(u16 *)(a0 + 0xA); + *(u16 *)(iVar2 + 0xC) = *(u16 *)(a0 + 0xE); + *(u16 *)(iVar2 + 0x10) = *(u16 *)(*(s32 *)(a0 + 0x20) + 0x10); + *(u16 *)(iVar2 + 0x12) = *(u16 *)(*(s32 *)(a0 + 0x20) + 0x12) + *(u16 *)(a0 + 0xFE); + *(u16 *)(iVar2 + 0x14) = *(u16 *)(*(s32 *)(a0 + 0x20) + 0x14); + *(u16 *)(iVar2 + 0x18) = *(u16 *)(*(s32 *)(a0 + 0x20) + 0x18); + *(u16 *)(iVar2 + 0x1A) = *(u16 *)(*(s32 *)(a0 + 0x20) + 0x1A); + *(u16 *)(iVar2 + 0x1C) = *(u16 *)(*(s32 *)(a0 + 0x20) + 0x1C); + if (*(s32 *)(*(s32 *)(a0 + 0x20) + 0x4) < 0) { + register u32 val __asm__("$2"); + val = *(u32 *)(iVar2 + 0x4); + uVar1 = val | 0x80000000; + } else { + __asm__ __volatile__(""); + uVar1 = *(u32 *)(iVar2 + 0x4) & 0x7FFFFFFF; + } + *(u32 *)(iVar2 + 0x4) = uVar1; + } +} diff --git a/.run/s37/ov_SC06_008/func_8017F100.c b/.run/s37/ov_SC06_008/func_8017F100.c new file mode 100644 index 000000000..ae639f9af --- /dev/null +++ b/.run/s37/ov_SC06_008/func_8017F100.c @@ -0,0 +1,60 @@ +#include "common.h" + +extern s32 D_801151D4; +extern s32 rand(void); +extern s32 VectorNormalSS(void *a0, void *a1); +extern s32 func_8012C588(s32 a0, s32 a1); +extern u8 *func_8012913C(); +extern void func_8017ED80(void); + +void func_8017F100(s32 a0) +{ + u16 nv[4]; + s32 g; + s32 s0; + s32 idx; + + g = D_801151D4; + + if (*(u16 *)a0 != 0) { + idx = *(s32 *)(a0 + 0x1C) & 3; + switch (idx) { + case 0: + s0 = func_8012C588(0x281, a0); + if (s0 != 0) { + *(s32 *)(s0 + 0x1C) = 2; + *(s16 *)(s0 + 0x12) = (rand() & 0x1F) - 0x10; + *(s16 *)(s0 + 0x16) = -((rand() & 0xF) + 0x10); + *(s16 *)(s0 + 0x1A) = (rand() & 0x1F) - 0x10; + } + break; + case 1: + break; + case 2: + case 3: + s0 = (s32)func_8012913C(0x23); + if (s0 != 0) { + *(s16 *)(s0 + 0x6) = *(u16 *)(a0 + 0x6) + (rand() & 0x3F) - 0x20; + *(s16 *)(s0 + 0xA) = *(u16 *)(a0 + 0xA) + (rand() & 0x3F) - 0x30; + { + s32 r = rand(); + s32 t = *(u16 *)(a0 + 0xE); + *(s32 *)(s0 + 0x18) = 0; + *(s32 *)(s0 + 0x14) = 0; + *(s32 *)(s0 + 0x10) = 0; + *(s16 *)(s0 + 0xE) = t + (r & 0x3F) - 0x20; + } + *(s16 *)(s0 + 0x34) = (rand() & 0x17FF) + 0x800; + nv[0] = *(s32 *)(g + 0x5C) - *(u16 *)(s0 + 0x6); + nv[1] = *(s32 *)(g + 0x60) - *(u16 *)(s0 + 0xA); + nv[2] = *(s32 *)(g + 0x64) - *(u16 *)(s0 + 0xE); + VectorNormalSS(nv, nv); + *(s16 *)(s0 + 0x6) = *(u16 *)(s0 + 0x6) + ((s16)nv[0] >> 6); + *(s16 *)(s0 + 0xA) = *(u16 *)(s0 + 0xA) + ((s16)nv[1] >> 6); + *(s16 *)(s0 + 0xE) = *(u16 *)(s0 + 0xE) + ((s16)nv[2] >> 6); + } + break; + } + ((void (*)(s32))func_8017ED80)(a0); + } +} diff --git a/.run/s37/ov_SC06_008/func_80181B20.c b/.run/s37/ov_SC06_008/func_80181B20.c new file mode 100644 index 000000000..1fae41487 --- /dev/null +++ b/.run/s37/ov_SC06_008/func_80181B20.c @@ -0,0 +1,69 @@ +#include "common.h" + +extern void func_8012B370(int a0); +extern void func_8004914C(void *a0); +extern void func_800491AC(void *a0); +extern void func_8002D4C8(s32 a0, s32 a1); +extern s32 RotTransPers(s32 a0, s32 a1, s32 *a2, s32 *a3); +extern u8 D_800AF648; + +void func_80181B20(s32 a0) { + + struct { + s16 v[3]; /* sp+0x10 */ + s16 pad1; /* sp+0x16 */ + u16 sxy[2]; /* sp+0x18 */ + s32 z; /* sp+0x1C */ + s32 flag; /* sp+0x20 */ + } L; + + s16 hp; + s32 p; + + hp = *(s16 *)(*(s32 *)(a0 + 0x20) + 0x14); + if (hp < 0x301) { + s32 count; + s32 total; + + count = *(u16 *)(a0 + 0xFC); + total = *(u16 *)(a0 + 0xFE); + count = count + 1; + total = total + count; + *(u16 *)(a0 + 0xFC) = count; + *(u16 *)(a0 + 0xFE) = total; + *(u16 *)(*(s32 *)(a0 + 0x20) + 0x14) = + *(u16 *)(*(s32 *)(a0 + 0x20) + 0x14) + total; + } else { + p = *(s32 *)(a0 + 0xD0); + *(s16 *)(a0 + 2) = 3; + *(s32 *)(a0 + 0x1C) = 0x1E; + *(s16 *)(a0 + 0x98) = 0; + *(s16 *)(p + 2) = 3; + *(s32 *)(p + 0x1C) = 0x1E; + *(s16 *)(p + 0x98) = 0; + + L.v[0] = *(s32 *)(*(s32 *)(a0 + 0x20) + 0x48); + L.v[1] = *(s32 *)(*(s32 *)(a0 + 0x20) + 0x4C); + L.v[2] = *(s32 *)(*(s32 *)(a0 + 0x20) + 0x50); + { register void *r4 __asm__("$4"); r4 = &D_800AF648; func_8004914C(r4); } + { register void *r4 __asm__("$4"); r4 = &D_800AF648; func_800491AC(r4); } + RotTransPers((s32)L.v, (s32)L.sxy, &L.z, &L.flag); + if (L.flag >= 0 && (u32)((L.sxy[0] + 0xEF) & 0xFFFF) < 0x1DF + && (u32)((L.sxy[1] + 0xB3) & 0xFFFF) < 0x167) { + s32 x = (s16)L.sxy[0]; + s32 ax; + ax = x; + if (x < 0) { + ax = -x; + } + ax = ((0xF0 - ax) * 0x7F) / 0xF0; + x = (x + 0xF0) / 0x1E; + if (x == 0x10) { + x = 0xF; + } + x = x << 8; + func_8002D4C8(0x961, (ax | (0x3000 | x)) & 0xFFFF); + } + } + func_8012B370(a0); +} diff --git a/.run/s37/ov_SC06_015/func_8017D8AC.c b/.run/s37/ov_SC06_015/func_8017D8AC.c new file mode 100644 index 000000000..c93c1ca87 --- /dev/null +++ b/.run/s37/ov_SC06_015/func_8017D8AC.c @@ -0,0 +1,222 @@ +/* func_8017D8AC -- ov_SC06_015 (family of 5: also func_8017ED54/ov_SC06_014, + * func_8017D9F0/ov_SC06_013, func_801830B4/ov_SC06_016, func_80181164/ov_SC06_030 -- all + * byte-identical modulo per-overlay local static data addresses). + * + * Builds a 5-pointed symmetric star/cross set of screen coordinates around the projected + * position of the object, then emits 4 Gouraud-shaded quads (POLY_G4, len=8, code=0x3A) whose + * per-vertex colour/x/y are selected from that coordinate set via small per-overlay index + * tables, and finally adds the sprite-header primitive itself to the OT. + */ + +extern s32 D_80126950; +extern s16 D_800B9A02; +extern s32 D_800A651C; + +extern void RotTransSV(void *a0, void *a1, void *a2); +extern void func_8004914C(void *a0); +extern void func_800491AC(void *a0); +extern s32 RotTransPers(s32 a0, s32 a1, s32 *a2, s32 *a3); + +extern void *func_80010A08(s32 a0); +extern s32 GetTPage(s32 a0, s32 a1, s32 a2, s32 a3); +extern s32 func_8005A600(s32 a0, s32 a1, s32 a2, s32 a3, s32 a4); +extern s32 AddPrim(s32 a0, void *a1); + +extern u8 D_800AF648; + +extern s32 D_8018F81C; +extern s32 D_8018F824[]; + +extern u8 D_8018F830[]; +extern u8 D_8018F831[]; +extern u8 D_8018F832[]; +extern u8 D_8018F833[]; +extern u8 D_8018F840[]; +extern u8 D_8018F841[]; +extern u8 D_8018F842[]; +extern u8 D_8018F843[]; +extern u8 D_8018F850[]; +extern u8 D_8018F851[]; +extern u8 D_8018F852[]; +extern u8 D_8018F853[]; + +/* §37/§124 SELF-axis: the TU declares `extern void func_8017D8AC(void);` (called with zero + args at func_8017E51C) while the byte-true definition takes an s32. Define under a private + C name bound to the real symbol. */ +extern void aF8017D8AC(s32 param_1) __asm__("func_8017D8AC"); +void aF8017D8AC(s32 param_1) { + /* single raw local block; sp+0x18 .. sp+0x5B (verified via the outgoing-arg-area rounded + * to 8 for func_8005A600's 5th argument, matching the sibling family's identical frame). */ + u8 buf[0x44]; + + /* Loop-carried locals pinned to the exact callee-saved registers the target uses -- + * each is reused for TWO logical roles (early-phase / loop-phase), matching the target's + * own register reuse across the two roles (§ "don't conclude unsteerable, try register + * pins" -- these are loop/local values, never the incoming parameter, so S3 is respected). */ + register s32 vVec __asm__("$17"); /* s1: RotTransSV/Pers vector ptr, then loop counter i */ + register s32 vFlag __asm__("$18"); /* s2: RotTransSV/Pers flag ptr, then grid anchor */ + register u8 *p0 __asm__("$16"); /* s0: D_800AF648 ptr, then prim-write pointer */ + register u8 *cur __asm__("$19"); /* s3: walking AddPrim arg (s4object+0xC, +=0x24) */ + register void *s4o __asm__("$20"); /* s4: func_80010A08(0x9C) result */ + register s32 *ctab __asm__("$21"); /* s5: &D_8018F824 colour table */ + register s32 s6 __asm__("$22"); /* s6: OT-base + zcount, the AddPrim OT arg */ + + s32 zcount; + s32 radius, inner; + u16 cx, cy; + s32 tpage; + + { + register s32 *p __asm__("$2") = + (s32 *)(*(s32 *)(param_1 + 0x20) + 0x34); + __asm__ __volatile__( + "lw $12, 0(%0)\n" + "lw $13, 4(%0)\n" + "ctc2 $12, $0\n" + "ctc2 $13, $1\n" + "lw $12, 8(%0)\n" + "lw $13, 12(%0)\n" + "lw $14, 16(%0)\n" + "ctc2 $12, $2\n" + "ctc2 $13, $3\n" + "ctc2 $14, $4\n" + "lw $12, 20(%0)\n" + "lw $13, 24(%0)\n" + "ctc2 $12, $5\n" + "lw $14, 28(%0)\n" + "ctc2 $13, $6\n" + "ctc2 $14, $7\n" + : : "r"(p) : "$12", "$13", "$14", "memory"); + } + + { + /* §gcc-2.7.2-map/sched.md rule 7: a single-SET pseudo gets a "birthing boost" that + * sinks it to just before its first consumer, regardless of source order. A 2nd SET + * that survives CSE (a zero-byte re-tie) disables the boost so it schedules at its + * natural (early) priority instead. */ + void *vecAddr = &D_8018F81C; + void *arg1; + __asm__ __volatile__("" : "=r"(vecAddr) : "0"(vecAddr)); + vVec = (s32)(buf + 0x10); + arg1 = (void *)vVec; + __asm__ __volatile__("" : "=r"(arg1) : "0"(arg1)); + vFlag = (s32)(buf + 0x38); + RotTransSV(vecAddr, arg1, (void *)vFlag); + } + + p0 = &D_800AF648; + func_8004914C(p0); + func_800491AC(p0); + + zcount = RotTransPers(vVec, (s32)(buf + 0x3C), (s32 *)(buf + 0x40), (s32 *)vFlag); + if (zcount <= 0) { + return; + } + if (*(s32 *)vFlag < 0) { + return; + } + zcount = zcount << 2; + + radius = ((D_80126950 + 0x1F4) * 48) / zcount; + + cx = *(u16 *)(buf + 0x3C); + cy = *(u16 *)(buf + 0x3E); + + { + register s32 t1 __asm__("$9") = + *(s32 *)((s8 *)&D_800A651C + (u16)D_800B9A02 * 0x14); + s6 = t1 + zcount; + } + + *(u16 *)(buf + 0x1C) = cx; + *(u16 *)(buf + 0x2C) = cy; + + *(u16 *)(buf + 0x18) = cx - radius; + + inner = (radius * 179) >> 8; + + *(u16 *)(buf + 0x1A) = cx - inner; + *(u16 *)(buf + 0x1E) = cx + inner; + *(u16 *)(buf + 0x20) = cx + radius; + *(u16 *)(buf + 0x28) = cy - radius; + *(u16 *)(buf + 0x2A) = cy - inner; + *(u16 *)(buf + 0x2E) = cy + inner; + *(u16 *)(buf + 0x30) = cy + radius; + + s4o = func_80010A08(0x9C); + if (s4o == 0) { + return; + } + + tpage = GetTPage(0, 1, 0, 0); + func_8005A600((s32)s4o, 0, 0, (u16)tpage, 0); + + { + register s32 t __asm__("$2") = (s32)((u8 *)s4o + 0xC); + cur = (u8 *)t; + ctab = D_8018F824; + vFlag = (s32)buf; + p0 = (u8 *)s4o + 0x2E; + vVec = 0; + + *(s32 *)(buf + 0x04) = (s32)((u8 *)s4o + 0x30); + *(s32 *)(buf + 0x08) = (s32)((u8 *)s4o + 0x54); + *(s32 *)(buf + 0x00) = t; + *(s32 *)(buf + 0x0C) = (s32)((u8 *)s4o + 0x78); + } + + for (; vVec < 0x10;) { + s32 color0, color1, color2, color3; + u16 x0, x1, x2, x3; + u16 y0, y1, y2, y3; + + color0 = ctab[D_8018F850[vVec]]; + *(s32 *)(p0 - 0x1E) = color0; + color1 = ctab[D_8018F851[vVec]]; + *(s32 *)(p0 - 0x16) = color1; + color2 = ctab[D_8018F852[vVec]]; + *(s32 *)(p0 - 0xE) = color2; + color3 = ctab[D_8018F853[vVec]]; + *(u8 *)(p0 - 0x1F) = 8; + *(u8 *)(p0 - 0x1B) = 0x3A; + *(s32 *)(p0 - 0x6) = color3; + + x0 = *(u16 *)(vFlag + D_8018F830[vVec] * 2 + 0x18); + *(u16 *)(p0 - 0x1A) = x0; + x1 = *(u16 *)(vFlag + D_8018F831[vVec] * 2 + 0x18); + *(u16 *)(p0 - 0x12) = x1; + x2 = *(u16 *)(vFlag + D_8018F832[vVec] * 2 + 0x18); + *(u16 *)(p0 - 0xA) = x2; + x3 = *(u16 *)(vFlag + D_8018F833[vVec] * 2 + 0x18); + *(u16 *)(p0 - 0x2) = x3; + + y0 = *(u16 *)(vFlag + D_8018F840[vVec] * 2 + 0x28); + *(u16 *)(p0 - 0x18) = y0; + y1 = *(u16 *)(vFlag + D_8018F841[vVec] * 2 + 0x28); + *(u16 *)(p0 - 0x10) = y1; + { + s32 argA0; + u8 *argA1; + + y2 = *(u16 *)(vFlag + D_8018F842[vVec] * 2 + 0x28); + argA0 = s6; + __asm__ __volatile__("" : "=r"(argA0) : "0"(argA0)); + *(u16 *)(p0 - 0x8) = y2; + argA1 = cur; + __asm__ __volatile__("" : "=r"(argA1) : "0"(argA1)); + { + u8 idx3 = D_8018F843[vVec]; + cur += 0x24; + y3 = *(u16 *)(vFlag + idx3 * 2 + 0x28); + vVec += 4; + } + + *(u16 *)(p0) = y3; + + AddPrim(argA0, argA1); + } + p0 += 0x24; + } + + AddPrim(s6, s4o); +} diff --git a/.run/s37/ov_SC06_032/func_8018EDB0.c b/.run/s37/ov_SC06_032/func_8018EDB0.c new file mode 100644 index 000000000..03a639dd4 --- /dev/null +++ b/.run/s37/ov_SC06_032/func_8018EDB0.c @@ -0,0 +1,141 @@ +// func_8018EDB0 — ov_SC06_032 / ov_SC06_032_jr_8017C24C +// Sibling of func_8018F060 (ov_SC06_018 / ov_SC06_018_jr_8017C24C, banked MATCH, same jr group +// "jr_8017C24C"). The .s is byte-identical in structure to the banked sibling: same prologue +// (func_8004914C/func_800491AC on ((s32*)param_1)[8]+0x34), same two 3-iteration RotTransSV loops +// (src[i].vx/vy from a per-overlay SVEC array selected by (param_2&1)*3, vz = -param_3[0] then +// -param_3[1]), same B4 spill-copy of param_4 into pkt[8] + pkt[9]=0x50000000, same DRAW() GTE +// macro (gte_ldv3/rtpt/stflg/stsxy3/ldv0/rtps/stflg/or-flags/stsxy/avsz4/stotz + OT-range guarded +// func_80017254 insert) repeated 3x with the shared D_800AF630+0x18 draw-mode calls. Only the +// per-overlay source-vector symbol differs (D_801CC7FC here vs D_801D1228 there); all other data +// symbols (D_800AF630, D_800B9A02, D_800A651C, D_800AE610) and callees are declared identically to +// the banked sibling, verbatim (S34/S36: reuse the byte-proven declaration + expression forms). +// +// DECLARATION SURFACE (whole-TU one-pass grep, D2): +// * func_8004914C(void *a0) / func_800491AC(void *a0) / RotTransSV(void*,void*,void*) are ALL +// already `extern` at this TU's file scope (lines 2637-2639) — do NOT redeclare here. +// * D_800B9A02 is `extern short D_800B9A02;` at file scope (line 2451/2453) — do NOT redeclare; +// unsigned access is forced at the USE with `*(u16 *)&D_800B9A02` (TU's own canonical spelling, +// already used at TU lines 2842/6033). +// * D_800AF630, D_800A651C, D_800AE610, func_80017254, D_801CC7FC appear nowhere else in this TU +// — declare fresh, matching the banked sibling's exact canonical shapes. +// * func_8018EDB0 has no caller anywhere in src/ (INCLUDE_ASM stub only) — signature unconstrained. + +typedef struct { u16 vx, vy, vz, pad; } SVEC_8018EDB0; /* u16 source vec -> lhu */ +typedef struct { s16 vx, vy, vz, pad; } SVECTOR_8018EDB0; /* 8 bytes, align 2 */ +typedef struct { u8 d[8]; } __attribute__((packed, aligned(1))) B8_8018EDB0; +typedef struct { u8 d[4]; } __attribute__((packed, aligned(1))) B4_8018EDB0; + +extern void func_80017254(void *a0); +extern SVEC_8018EDB0 D_801CC7FC[]; +extern short D_800B9A02; +extern u8 D_800A651C[]; +extern u8 D_800AE610[]; +extern u8 D_800AF630[]; + +#define gte_ldv0(r0) __asm__ volatile ( \ + "lwc2 $0, 0( %0 );" "lwc2 $1, 4( %0 )" : : "r"( r0 ) ) +#define gte_ldv3(r0, r1, r2) __asm__ volatile ( \ + "lwc2 $0, 0( %0 );" "lwc2 $1, 4( %0 );" "lwc2 $2, 0( %1 );" \ + "lwc2 $3, 4( %1 );" "lwc2 $4, 0( %2 );" "lwc2 $5, 4( %2 )" \ + : : "r"( r0 ), "r"( r1 ), "r"( r2 ) ) +#define gte_rtps() __asm__ volatile ("nop;nop;rtps") +#define gte_rtpt() __asm__ volatile ("nop;nop;rtpt") +#define gte_stsxy(r0) __asm__ volatile ( \ + "swc2 $14, 0( %0 )" : : "r"( r0 ) : "memory" ) +#define gte_stsxy3(r0, r1, r2) __asm__ volatile ( \ + "swc2 $12, 0( %0 );" "swc2 $13, 0( %1 );" "swc2 $14, 0( %2 )" \ + : : "r"( r0 ), "r"( r1 ), "r"( r2 ) : "memory" ) +#define gte_avsz4() __asm__ volatile ("nop;nop;avsz4") +#define gte_stotz(r0) __asm__ volatile ( \ + "swc2 $7, 0( %0 )" : : "r"( r0 ) : "memory" ) +#define gte_stflg(r0) __asm__ volatile ( \ + "cfc2 $12, $31;" "nop;" "sw $12, 0( %0 )" : : "r"( r0 ) : "$12", "memory" ) + +#define DRAW() \ + mb = (u8 *)&D_800AF630; \ + __asm__("" : "=r"(mb) : "0"(mb)); \ + func_8004914C(mb + 0x18); \ + __asm__("" : "=r"(mb) : "0"(mb)); \ + func_800491AC(mb + 0x18); \ + pc = (long *)&pkt[0]; \ + p1 = (long *)&pkt[2]; \ + p2 = (long *)&pkt[4]; \ + p3 = (long *)&pkt[6]; \ + pfl = &flag1; \ + p0 = pc; \ + gte_ldv3(p0, p1, p2); \ + gte_rtpt(); \ + gte_stflg(pfl); \ + __asm__ __volatile__("" :: "r"(pfl)); \ + gte_stsxy3(p0, p1, p2); \ + gte_ldv0(p3); \ + gte_rtps(); \ + gte_stflg(&flag2); \ + flag1 |= flag2; \ + gte_stsxy(p3); \ + gte_avsz4(); \ + gte_stotz(&otz); \ + oz = otz; \ + *(s16 *)((u8 *)pkt + 4) = (s16)oz; \ + if (oz > 0 && flag1 >= 0 && \ + !((u32)&D_800AE610 < (u32)(*(s32 *)((u8 *)&D_800A651C \ + + (*(u16 *)&D_800B9A02) * 0x14) + oz * 4))) { \ + func_80017254(pc); \ + } + +void func_8018EDB0(s32 param_1, u32 param_2, u16 *param_3, u32 param_4) { + u8 dead[0x20]; /* 0x10 : reserved, never referenced (holds the frame at 0xC8) */ + SVECTOR_8018EDB0 scratch; /* 0x30 */ + B8_8018EDB0 out[6]; /* 0x38 */ + u32 pkt[10]; /* 0x68 : v0..v3 + color(0x88) + code(0x8C) */ + SVECTOR_8018EDB0 rtflag; /* 0x90 (8 bytes -> 0x94 gap) */ + s32 flag1; /* 0x98 */ + s32 flag2; /* 0x9C */ + s32 otz; /* 0xA0 */ + SVEC_8018EDB0 *src; + s32 i, oz; + long *p0, *p1, *p2, *p3, *pfl; + u8 *mb; + register long *pc __asm__("$16"); + + func_8004914C((void *)(((s32 *)param_1)[8] + 0x34)); + func_800491AC((void *)(((s32 *)param_1)[8] + 0x34)); + + src = &D_801CC7FC[(param_2 & 1) * 3]; + for (i = 0; i < 3; i++) { + scratch.vx = src[i].vx; + scratch.vy = src[i].vy; + scratch.vz = -param_3[0]; + RotTransSV(&scratch, &out[i], &rtflag); + } + for (i = 0; i < 3; i++) { + scratch.vx = src[i].vx; + scratch.vy = src[i].vy; + scratch.vz = -param_3[1]; + RotTransSV(&scratch, &out[3 + i], &rtflag); + } + + *(B4_8018EDB0 *)&pkt[8] = *(B4_8018EDB0 *)¶m_4; + pkt[9] = 0x50000000; + + /* face 0: verts 0,1,3,4 */ + *(B8_8018EDB0 *)&pkt[0] = out[0]; + *(B8_8018EDB0 *)&pkt[2] = out[1]; + *(B8_8018EDB0 *)&pkt[4] = out[3]; + *(B8_8018EDB0 *)&pkt[6] = out[4]; + DRAW(); + + /* face 1: verts 1,2,4,5 */ + *(B8_8018EDB0 *)&pkt[0] = out[1]; + *(B8_8018EDB0 *)&pkt[2] = out[2]; + *(B8_8018EDB0 *)&pkt[4] = out[4]; + *(B8_8018EDB0 *)&pkt[6] = out[5]; + DRAW(); + + /* face 2: verts 2,0,5,3 */ + *(B8_8018EDB0 *)&pkt[0] = out[2]; + *(B8_8018EDB0 *)&pkt[2] = out[0]; + *(B8_8018EDB0 *)&pkt[4] = out[5]; + *(B8_8018EDB0 *)&pkt[6] = out[3]; + DRAW(); +} diff --git a/.run/s37_wave.json b/.run/s37_wave.json new file mode 100644 index 000000000..61627b8d3 --- /dev/null +++ b/.run/s37_wave.json @@ -0,0 +1,178 @@ +[ + { + "fn": "func_8018478C", + "ov": "ov_SC03_092", + "sub": "asm/ov_SC03_092/nonmatchings/ov_SC03_092_jr_8017AE2C", + "tu": "src/ov_SC03_092/ov_SC03_092_jr_8017AE2C.c", + "n": 328, + "m": 7, + "ti": 2296, + "model": "sonnet", + "seed": 0 + }, + { + "fn": "func_80184370", + "ov": "ov_SC03_092", + "sub": "asm/ov_SC03_092/nonmatchings/ov_SC03_092_jr_8017AE2C", + "tu": "src/ov_SC03_092/ov_SC03_092_jr_8017AE2C.c", + "n": 263, + "m": 7, + "ti": 1841, + "model": "sonnet", + "seed": 1 + }, + { + "fn": "func_8018B13C", + "ov": "ov_SC02_028", + "sub": "asm/ov_SC02_028/nonmatchings/ov_SC02_028_jr_8017D898", + "tu": "src/ov_SC02_028/ov_SC02_028_jr_8017D898.c", + "n": 44, + "m": 40, + "ti": 1760, + "model": "sonnet", + "seed": 0 + }, + { + "fn": "func_80181914", + "ov": "ov_SC03_024", + "sub": "asm/ov_SC03_024/nonmatchings/ov_SC03_024_jr_8017DF84", + "tu": "src/ov_SC03_024/ov_SC03_024_jr_8017DF84.c", + "n": 143, + "m": 8, + "ti": 1144, + "model": "sonnet", + "seed": 0 + }, + { + "fn": "func_8017F2E8", + "ov": "ov_SC03_108", + "sub": "asm/ov_SC03_108/nonmatchings/ov_SC03_108_jr_8017BEBC", + "tu": "src/ov_SC03_108/ov_SC03_108_jr_8017BEBC.c", + "n": 214, + "m": 4, + "ti": 856, + "model": "sonnet", + "seed": 1 + }, + { + "fn": "func_8017F100", + "ov": "ov_SC06_008", + "sub": "asm/ov_SC06_008/nonmatchings/ov_SC06_008_jr_8017C294", + "tu": "src/ov_SC06_008/ov_SC06_008_jr_8017C294.c", + "n": 122, + "m": 7, + "ti": 854, + "model": "sonnet", + "seed": 1 + }, + { + "fn": "func_8017D930", + "ov": "ov_SC02_039", + "sub": "asm/ov_SC02_039/nonmatchings/ov_SC02_039_jr_8017BEBC", + "tu": "src/ov_SC02_039/ov_SC02_039_jr_8017BEBC.c", + "n": 121, + "m": 7, + "ti": 847, + "model": "sonnet", + "seed": 1 + }, + { + "fn": "func_80184328", + "ov": "ov_SC02_035", + "sub": "asm/ov_SC02_035/nonmatchings/ov_SC02_035_jr_8017BEBC", + "tu": "src/ov_SC02_035/ov_SC02_035_jr_8017BEBC.c", + "n": 44, + "m": 19, + "ti": 836, + "model": "sonnet", + "seed": 0 + }, + { + "fn": "func_801837B0", + "ov": "ov_SC02_041", + "sub": "asm/ov_SC02_041/nonmatchings/ov_SC02_041_jr_8017BEBC", + "tu": "src/ov_SC02_041/ov_SC02_041_jr_8017BEBC.c", + "n": 206, + "m": 4, + "ti": 824, + "model": "sonnet", + "seed": 1 + }, + { + "fn": "func_8018B354", + "ov": "ov_SC02_028", + "sub": "asm/ov_SC02_028/nonmatchings/ov_SC02_028_jr_8017D898", + "tu": "src/ov_SC02_028/ov_SC02_028_jr_8017D898.c", + "n": 41, + "m": 20, + "ti": 820, + "model": "sonnet", + "seed": 0 + }, + { + "fn": "func_8017ED80", + "ov": "ov_SC06_008", + "sub": "asm/ov_SC06_008/nonmatchings/ov_SC06_008_jr_8017C294", + "tu": "src/ov_SC06_008/ov_SC06_008_jr_8017C294.c", + "n": 117, + "m": 7, + "ti": 819, + "model": "sonnet", + "seed": 1 + }, + { + "fn": "func_8018893C", + "ov": "ov_SC03_124", + "sub": "asm/ov_SC03_124/nonmatchings/ov_SC03_124_jr_8017AE2C", + "tu": "src/ov_SC03_124/ov_SC03_124_jr_8017AE2C.c", + "n": 203, + "m": 4, + "ti": 812, + "model": "sonnet", + "seed": 0 + }, + { + "fn": "func_80181B20", + "ov": "ov_SC06_008", + "sub": "asm/ov_SC06_008/nonmatchings/ov_SC06_008_jr_8017C294", + "tu": "src/ov_SC06_008/ov_SC06_008_jr_8017C294.c", + "n": 116, + "m": 7, + "ti": 812, + "model": "sonnet", + "seed": 1 + }, + { + "fn": "func_8017D8AC", + "ov": "ov_SC06_015", + "sub": "asm/ov_SC06_015/nonmatchings/ov_SC06_015_jr_8017BEBC", + "tu": "src/ov_SC06_015/ov_SC06_015_jr_8017BEBC.c", + "n": 265, + "m": 3, + "ti": 795, + "model": "sonnet", + "seed": 0 + }, + { + "fn": "func_8018EDB0", + "ov": "ov_SC06_032", + "sub": "asm/ov_SC06_032/nonmatchings/ov_SC06_032_jr_8017C24C", + "tu": "src/ov_SC06_032/ov_SC06_032_jr_8017C24C.c", + "n": 397, + "m": 2, + "ti": 794, + "model": "sonnet", + "seed": 0 + }, + { + "fn": "func_801825A4", + "ov": "ov_SC02_026", + "sub": "asm/ov_SC02_026/nonmatchings/ov_SC02_026_jr_8017C180", + "tu": "src/ov_SC02_026/ov_SC02_026_jr_8017C180.c", + "n": 129, + "m": 6, + "ti": 774, + "model": "sonnet", + "seed": 1 + } +] \ No newline at end of file diff --git a/.run/s37w.js b/.run/s37w.js new file mode 100644 index 000000000..d2627da96 --- /dev/null +++ b/.run/s37w.js @@ -0,0 +1,152 @@ +export const meta = { + name: 'p30-s37-SONNET-pipeline', + description: 'P30 S34: fresh h_seq families, open sites DERIVED from corpus.stubs', + phases: [ + { title: 'Draft', detail: 'one agent per target, size-routed per the §136i ladder (haiku/sonnet/opus)' }, + { title: 'Escalate', detail: 'next rung up (haiku->sonnet, sonnet->opus) on any non-MATCH' }, + ], +} + +// args = { targets: [...compact records...], extra: "" } +// Accept a JSON string too — an invocation can deliver args stringified and pipeline() then dies. +const A = {targets: [{"fn": "func_8018478C", "ov": "ov_SC03_092", "sub": "asm/ov_SC03_092/nonmatchings/ov_SC03_092_jr_8017AE2C", "tu": "src/ov_SC03_092/ov_SC03_092_jr_8017AE2C.c", "n": 328, "m": 7, "ti": 2296, "model": "sonnet", "seed": 0}, {"fn": "func_80184370", "ov": "ov_SC03_092", "sub": "asm/ov_SC03_092/nonmatchings/ov_SC03_092_jr_8017AE2C", "tu": "src/ov_SC03_092/ov_SC03_092_jr_8017AE2C.c", "n": 263, "m": 7, "ti": 1841, "model": "sonnet", "seed": 1}, {"fn": "func_8018B13C", "ov": "ov_SC02_028", "sub": "asm/ov_SC02_028/nonmatchings/ov_SC02_028_jr_8017D898", "tu": "src/ov_SC02_028/ov_SC02_028_jr_8017D898.c", "n": 44, "m": 40, "ti": 1760, "model": "sonnet", "seed": 0}, {"fn": "func_80181914", "ov": "ov_SC03_024", "sub": "asm/ov_SC03_024/nonmatchings/ov_SC03_024_jr_8017DF84", "tu": "src/ov_SC03_024/ov_SC03_024_jr_8017DF84.c", "n": 143, "m": 8, "ti": 1144, "model": "sonnet", "seed": 0}, {"fn": "func_8017F2E8", "ov": "ov_SC03_108", "sub": "asm/ov_SC03_108/nonmatchings/ov_SC03_108_jr_8017BEBC", "tu": "src/ov_SC03_108/ov_SC03_108_jr_8017BEBC.c", "n": 214, "m": 4, "ti": 856, "model": "sonnet", "seed": 1}, {"fn": "func_8017F100", "ov": "ov_SC06_008", "sub": "asm/ov_SC06_008/nonmatchings/ov_SC06_008_jr_8017C294", "tu": "src/ov_SC06_008/ov_SC06_008_jr_8017C294.c", "n": 122, "m": 7, "ti": 854, "model": "sonnet", "seed": 1}, {"fn": "func_8017D930", "ov": "ov_SC02_039", "sub": "asm/ov_SC02_039/nonmatchings/ov_SC02_039_jr_8017BEBC", "tu": "src/ov_SC02_039/ov_SC02_039_jr_8017BEBC.c", "n": 121, "m": 7, "ti": 847, "model": "sonnet", "seed": 1}, {"fn": "func_80184328", "ov": "ov_SC02_035", "sub": "asm/ov_SC02_035/nonmatchings/ov_SC02_035_jr_8017BEBC", "tu": "src/ov_SC02_035/ov_SC02_035_jr_8017BEBC.c", "n": 44, "m": 19, "ti": 836, "model": "sonnet", "seed": 0}, {"fn": "func_801837B0", "ov": "ov_SC02_041", "sub": "asm/ov_SC02_041/nonmatchings/ov_SC02_041_jr_8017BEBC", "tu": "src/ov_SC02_041/ov_SC02_041_jr_8017BEBC.c", "n": 206, "m": 4, "ti": 824, "model": "sonnet", "seed": 1}, {"fn": "func_8018B354", "ov": "ov_SC02_028", "sub": "asm/ov_SC02_028/nonmatchings/ov_SC02_028_jr_8017D898", "tu": "src/ov_SC02_028/ov_SC02_028_jr_8017D898.c", "n": 41, "m": 20, "ti": 820, "model": "sonnet", "seed": 0}, {"fn": "func_8017ED80", "ov": "ov_SC06_008", "sub": "asm/ov_SC06_008/nonmatchings/ov_SC06_008_jr_8017C294", "tu": "src/ov_SC06_008/ov_SC06_008_jr_8017C294.c", "n": 117, "m": 7, "ti": 819, "model": "sonnet", "seed": 1}, {"fn": "func_8018893C", "ov": "ov_SC03_124", "sub": "asm/ov_SC03_124/nonmatchings/ov_SC03_124_jr_8017AE2C", "tu": "src/ov_SC03_124/ov_SC03_124_jr_8017AE2C.c", "n": 203, "m": 4, "ti": 812, "model": "sonnet", "seed": 0}, {"fn": "func_80181B20", "ov": "ov_SC06_008", "sub": "asm/ov_SC06_008/nonmatchings/ov_SC06_008_jr_8017C294", "tu": "src/ov_SC06_008/ov_SC06_008_jr_8017C294.c", "n": 116, "m": 7, "ti": 812, "model": "sonnet", "seed": 1}, {"fn": "func_8017D8AC", "ov": "ov_SC06_015", "sub": "asm/ov_SC06_015/nonmatchings/ov_SC06_015_jr_8017BEBC", "tu": "src/ov_SC06_015/ov_SC06_015_jr_8017BEBC.c", "n": 265, "m": 3, "ti": 795, "model": "sonnet", "seed": 0}, {"fn": "func_8018EDB0", "ov": "ov_SC06_032", "sub": "asm/ov_SC06_032/nonmatchings/ov_SC06_032_jr_8017C24C", "tu": "src/ov_SC06_032/ov_SC06_032_jr_8017C24C.c", "n": 397, "m": 2, "ti": 794, "model": "sonnet", "seed": 0}, {"fn": "func_801825A4", "ov": "ov_SC02_026", "sub": "asm/ov_SC02_026/nonmatchings/ov_SC02_026_jr_8017C180", "tu": "src/ov_SC02_026/ov_SC02_026_jr_8017C180.c", "n": 129, "m": 6, "ti": 774, "model": "sonnet", "seed": 1}], extra: "START HERE \u2014 SIBLING-FIRST IS THE FASTEST ROUTE (\u00a7136c, measured this session: it produced several\nFIRST-DRAFT matches). Before you derive anything from the .s:\n 1. grep src/shared/engine_core.h for a `DEFINE_func_*` macro body that is a NEAR-TWIN of your\n target (same struct-offset chain, same shape, differing only in constants/one call). This is a\n family wave \u2014 the twin usually EXISTS, because that is what a family is.\n 2. grep your own TU for an already-BANKED sibling (a real function definition, not an INCLUDE_ASM).\n 3. Reuse its EXPRESSION FORMS and its DECLARATION FORMS verbatim. They are already byte-proven to\n produce the gcc-2.7.2 schedule and register assignment you need.\nSearch order: engine_core.h near-twin -> same-TU banked sibling -> the .s -> the Ghidra seed LAST\n(the seed was byte-proven to be an ENTIRELY DIFFERENT body twice this session).\n\nTHE LOCAL-VARIABLE LEVER (\u00a7136 \u2014 highest-yield finding, 25 banks). gcc-2.7.2 allocates ONE PSEUDO\nPER C LOCAL, and local-alloc.c:472 REFUSES a local allocno whose REG_N_DEATHS > 1 \u2014 promoting it to\na GLOBAL allocno that loses the low register. The number and scope of your locals moves whole\nregister assignments. REACH FOR THIS BEFORE register __asm__ PINS.\n L1. Same $v0/$v1 pair swapped in ONE arm only => you reused ONE local across N arms. SPLIT it into\n per-arm block-scoped locals.\n L2. One extra `sw $sN` in the prologue, frame otherwise identical => SPLIT a compound initializer:\n `x = *(u8*)p << k;` makes TWO pseudos; `x = *(u8*)p; x = x << k;` reuses one.\n L3. An extra `addu $vX,$v0,$zero` after a `jal` AND a later copy of the same value => the source\n had TWO variables with the first PINNED (an unpinned pseudo coalesces the pair away).\n L4. LENGTH-DRIFT short by `addiu $sN,$sp,K` + a save/restore pair => write a POINTER local assigned\n before the loop and used only inside it.\n L5. Local stack slots are assigned in DECLARATION order ascending from 0x10, independent of use\n order. Frame-offset drift with correct code is a declaration-ORDER problem.\n\nTYPE-FORM RULES:\n T1. A real `mult $rX,$rY` with a small constant => the multiplier is a NON-CONST LOCAL, not a\n literal (a literal goes through synth_mult's sll/addu chain). `s32 r = K;` as its own statement.\n T2. `li $sN,0xfff0` + `addu` where the target has `addiu $vN,$vN,-0x10` => gcc narrowed to HImode.\n Hoist the call to its own statement; put the load+subtract in a BLOCK-SCOPED s32 temp.\n T3. `andi $vN,0xffff` after a `jal` the target lacks => the TU declares that callee u16/s16-\n returning. Do NOT change the decl \u2014 cast at the call (idiom 9 on the RETURN axis).\n T4. Unexplained `addu $vX,$aY,$zero` near a conditional branch (including in its DELAY SLOT,\n consumed only by the fall-through arm) + the feeding `lh` loads out of order => the value is an\n s16 LOCAL, not s32; LOAD_EXTEND_OP folded the widening extend into a plain move.\n T5. LENGTH-DRIFT +1 with a narrow load of the SAME stack slot => gcc narrowed a memory-operand\n `local >> 16`. Bind the local to an s32 temp used twice.\n T6. A symbol read once at a constant offset but built into $s1 by lui/addiu => cache it in a\n POINTER LOCAL. A direct D_xxx[k] folds %lo per use and shrinks the frame.\n T7. BRANCH POLARITY: if match_one prints BRANCH-POLARITY, invert the source condition \u2014 gcc-2.7.2\n flips the branch to place the longer block as the fall-through. (Closed 8 mismatches in one\n edit this wave; the index keys the literal string \"BRANCH-POLARITY\" straight to \u00a73-T4.)\n\nSCHEDULING:\n S1. The MEM_IN_STRUCT_P escape does NOT apply when the blocking store has a VARYING (register-base)\n address \u2014 true_dependence() only drops the edge for a CONSTANT-address store. Then the lever is\n SOURCE ORDER: assign the load to a temp ABOVE the stores. (But same-base `reg+const` addresses\n ARE disambiguated by memrefs_conflict_p, so those hoists are free.)\n S2. An unfilled load-delay nop where the target fills it with a trailing call's arg setup => hoist a\n LOAD: split `*p = *p + 1` into `v = *p + 1; ... *p = v;`.\n S3. NEVER pin an incoming PARAMETER \u2014 it turns the param's `move` into a schedulable body insn and\n reshuffles the whole prologue. Pin loop variables only.\n S4. If no C lever moves a 3-6 instruction schedule/regalloc residual, a \u00a721 ZERO-BYTE RE-TIE\n barrier can anchor it: `__asm__ __volatile__(\"\" : \"=r\"(v) : \"0\"(v));` between the loads and the\n use. Try the source-level levers first; this closed one case after six other variants failed.\n\nDECLARATION SURFACE (decides whether a byte-correct draft BANKS):\n D1. `conflicting types` for a D_ symbol you cannot find declared in the split .c => the decl lives\n inside a DEFINE_func_*() MACRO BODY in src/shared/engine_core.h. Reuse its canonical type.\n An 8-byte-stride table declared `s32 D_x[][2]` must be indexed [i][0]/[i][1].\n D2. cc1 reports only the FIRST conflict. EVERY reconciled draft this session had a SECOND hidden\n conflict, sometimes BELOW the splice point. grep the WHOLE TU in ONE pass for every symbol.\n D3. A DECLARATION CONFLICT ABORTS THE COMPILE, so it hides the byte question entirely. If you are\n handed a draft as \"byte-correct, only declaration-blocked\", RUN match_one ON IT FIRST \u2014 two of\n three such drafts this session also had a real codegen residual behind the conflict.\n D4. If you compile anything, use a PROCESS-UNIQUE scratch path. A shared one silently compiled\n another agent's file and returned a meaningless success this session.\n\n\nx2-9 BAND. \u00a7136c sibling-first still applies but CHECK the twin EXISTS first (\u00a7136e) \u2014 if the whole\nfamily is in nonmatchings, go straight to the .s. \u00a7136j: sub-120-ins targets tend to fail on\nDECLARATIONS, 120+ on real codegen; do the whole-TU one-pass symbol grep (D2) either way.\nIf match_one reports REGALLOC-PERM (a clean 2-register swap), see \u00a7137: it is a TWO-COMPILE\nARITHMETIC problem \u2014 read R and L from `cc1 -dl -dg`, evaluate floor_log2(R)*R/L*1e4*size for both\ncontenders AND their ranked neighbours to get the admissible window, then place a zero-byte\n`__asm__ __volatile__(\"\" ::\"r\"(v))` so L lands inside it. Do NOT reach for the permuter first.\n\n\nS33 MEASURED (promote into your first pass):\n - \u00a7138 RECONCILE DIRECTION: if the TU declares your symbol ABOVE the splice point, DELETE your\n duplicate; if BELOW, KEEP a decl in the TU's EXACT shape and cast at the use. Picking the wrong\n direction CREATES the next error. grep the TU for the symbol and compare line numbers with the\n stub line BEFORE editing either way.\n - A repeated typedef OR a repeated bare `struct Tag {...}` is a C89 ERROR even when textually\n identical. If the TU already defines the type trio you copied from a sibling, do not redeclare\n it \u2014 and check for the bare struct tag too, not just the typedefs (that one cost an extra round).\n - func_8017D174 (793 ins) closed a \u00a7137 allocno tie AND a sched2 loop-head rotation JOINTLY with\n four zero-byte asms: two `\"=r\"`/`\"0\"` re-ties splitting a live range, plus two volatile sliders\n placed in a DIFFERENT basic block, so they lift the live-length count without perturbing the head\n schedule. A slider inside the block you are trying to fix will regress it.\n - Rank/trust nothing from a map's `exemplar` field: derive the open sites from corpus.stubs. The\n map's exemplar can point at an instance that is ALREADY banked, which hides the whole family.\n\n\nS34 MEASURED \u2014 DO THIS FIRST:\n - STEP 0 of sibling-first, ahead of engine_core.h: `grep -rn \"\" src/` with a distinctive\n literal from YOUR .s (a magic word, an unusual mask, an odd immediate). \u00a7136c's first two steps\n are same-TU/shared-header scoped and CANNOT reach a banked twin in another overlay's TU \u2014 and the\n big template classes live cross-overlay. Measured: func_80188C04 (328 ins) was byte-identical to\n an already-banked func_801833F0 in ov_SC02_028; one grep found it and the body was reused\n verbatim with only file-local type/macro suffixes renamed.\n - If a sibling of yours banked into the SAME TU recently, the TU now carries ITS type names for the\n shared data symbols. Use the TU's names; do not introduce your own suffixed duplicates.\n - \"normalized distance 0\" is NOT an h_exact guarantee. Relocs are masked by normalization.\n\n\nS35 (wave 3 banked 13/13 \u2014 the magic-grep STEP 0 is doing the work; lead with it):\n - Write ONLY your deliverable `.run///func_.c` into the drafts directory. Scratch\n files there break the gate. Put experiments anywhere else under .run/.\n - If the TU declares your function itself incompatibly (`extern void f(void);` vs a def that takes\n an argument), that is the SELF axis: define under a private C name bound to the real symbol \u2014\n `extern aF() __asm__(\"func_\"); aF() { ... }`. Three of\n this session's reconciles were exactly this, and it has zero blast radius on callers.\n\n\nS36: STEP 0 is now the highest-yield move in this prompt \u2014 waves 3 and 4 both banked 100% and most\nagents reported `index_gap: none` because the grep resolved the target in one pass. LEAD WITH IT.\n - If two targets in your wave share a TU, the FIRST one banked puts its file-local type tags into\n that TU. If you copy a sibling's body, do NOT also copy its `struct X_ {...}` definition \u2014\n the TU may already have it ABOVE your splice point, and a repeated tag is a C89 error even when\n identical. Use the TU's tag; delete your duplicate (\u00a7138 reconcile direction).\n"} +const T = Array.isArray(A) ? A : A.targets +const EXTRA = (Array.isArray(A) ? '' : A.extra) || '' +if (!Array.isArray(T)) throw new Error('args.targets must be an array') + +const VERDICT = { + type: 'object', + additionalProperties: false, + required: ['fn', 'status', 'summary'], + properties: { + fn: { type: 'string' }, + status: { type: 'string', enum: ['MATCH', 'DIFF', 'BLOCKED'] }, + closeness: { type: 'number', description: 'mismatching instructions remaining; 0 for MATCH' }, + klass: { type: 'string', description: 'residual class if not MATCH' }, + summary: { type: 'string', description: 'what you did and what the residual is, <=4 sentences' }, + levers: { type: 'string', description: 'cookbook sections / idioms that CLOSED the residual' }, + index_hit: { type: 'boolean' }, + index_gap: { type: 'string' }, + }, +} + +function prompt(t, escalated) { + const asm = `${t.sub}/${t.fn}.s` + return `You are matching ONE PS1 function to byte-identical gcc-2.7.2 output for the Brave Fencer +Musashi decompilation. Your ONLY deliverable is a C file at \`.run/s37/${t.ov}/${t.fn}.c\`. + +TARGET + function ${t.fn} + binary ${t.ov} + target asm ${asm} <- THE GROUND TRUTH. Read this FIRST and in full. + TU it lands in ${t.tu} + asm-subdir ${t.sub} + size ${t.n} instructions + leverage family of ${t.m} members / ${t.ti} templatable instructions — a byte-match here + propagates ${t.m}x across the fleet. +${t.seed ? ` ghidra seed .run/ghidra_c/${t.fn}.c <- A HINT ONLY. It is sometimes an ENTIRELY + DIFFERENT body (byte-proven this phase). If it disagrees with the .s, THE .s WINS.` : ` ghidra seed (none cached — work from the .s)`} +${t.retry ? ` ** RETRY ** A previous wave recorded: "${t.retry}". That is a data point, not a + verdict. Re-derive from the .s; do not assume the earlier verdict was right.` : ''} + +HOW TO WORK (this order is the measured-fastest) +1. \`docs/cookbook-index.md\` is a SYMPTOM-KEYED index of 364 byte-verified idioms. Grep it for your + residual's symptom BEFORE deriving anything. Measured: index-first took a wave's bank rate from + 57% to 100%. Then read the section it names in \`docs/matching-cookbook.md\`. + \`docs/gcc-2.7.2-map/{sched,regalloc,loop,cse_expr}.md\` is the compiler-source-derived map for + scheduling / register-allocation residuals. +2. Read the target \`.s\` completely: frame size, callee-saved registers, jal targets, every + \`%hi/%lo\` symbol. +3. Read the TU (${t.tu}) for EVERY symbol your draft will name. cc1 reports only the FIRST conflict, + so a draft can look one edit from done and hold three more. grep the whole TU in ONE pass. + Match its existing declarations EXACTLY; push any type disagreement to a CAST AT THE USE SITE + rather than redeclaring the symbol. +4. Write the draft, then verify: + .venv/bin/python tools/match_one.py ${t.fn} --c .run/s37/${t.ov}/${t.fn}.c --asm-subdir ${t.sub} + Iterate until it prints MATCH; it names the exact mismatching instructions. + +IDIOMS THAT CLOSED RESIDUALS IN THE LAST WAVES (cookbook §135 — all byte-verified) + 1. UNSIGNED switch index => pure equality chain, NO range test. No \`slti\` bound check in the + target's switch means the index is u32, not s32. + 2. \`a0[0x46]\` (ARRAY_REF) sets MEM_IN_STRUCT_P and lets a load HOIST past a constant-address + store; \`*(s16 *)((s32)a0 + 0x8C)\` (INDIRECT_REF) keeps the dependence. Many 4-instruction + "scheduling residuals" are just this type-form choice. + 3. A constant store whose top bit is set in the STORED width needs an UNSIGNED destination: + \`*(u16 *)p = 0x8C00\` emits \`ori\`; through \`s16\` it folds negative and emits \`addiu\`. + 4. The list scheduler PRESERVES the relative order of disambiguable stores. A store written late + in source SINKS. If a store lands too late, move it EARLIER IN SOURCE (not a permuter job). + 5. A \`short\` loop counter blocks strength reduction; walking explicit pointers (\`p++\`) + reproduces the original biv/giv set. + 6. Frame size off by a constant => DEAD LOCALS. If ALL diffs are \`sp\`-relative immediates off by + one constant delta, add the padding declaration. + 7. An INTERIOR address has no symbol — a \`lui/addiu\` pair can build an offset INTO a symbol. + Find the containing symbol in the data \`.s\` and index into it; declaring the interior address + as its own extern link-fails. + 8. NEVER redeclare a C-library name (\`memcpy\` etc.). + 9. Loose typing is pervasive: if the TU declares \`void f(void)\` but the asm passes \`$a0\`, call + through a cast — \`((void(*)(s32))f)(a0)\` — do NOT change the declaration. +${EXTRA ? `\nPROMOTED FROM THE PREVIOUS BATCH (fresh, byte-verified this session)\n${EXTRA}\n` : ''} +HARD RULES + * Write ONLY \`.run/s37/${t.ov}/${t.fn}.c\`. NEVER edit \`src/\`, \`asm/\`, \`config/\`, \`include/\`, + the Makefile, or any tracked file. + * Do NOT run \`make\`, \`make build\`, \`make extract\`, or \`tools/harvest_verify.py\`. The + whole-binary gate is the orchestrator's job and the sole arbiter of a match. + * \`match_one\` MATCH is NECESSARY BUT NOT SUFFICIENT — it compiles standalone and cannot see the + TU's other declarations. Step 3 is what makes a MATCH actually BANK. + * Report honestly. A DIFF with a precise residual class routes the next attempt; a false MATCH + just gets caught by the byte-gate and wastes a cycle. + * ${escalated ? 'A cheaper rung of the model ladder already attempted this and did not reach MATCH. Read its draft at the path above, but re-derive from the .s rather than trusting it.' : 'Work economically — most functions this size close from the .s plus one or two index lookups.'} + +Return the structured verdict.` +} + +phase('Draft') + +// S36 MEASUREMENT: the two-batch `parallel()` design was a HARD BARRIER — batch 2 could not start +// until batch 1's slowest agent finished. Measured across waves 1-4 from agent transcript +// timestamps: median agent 9-33 min but wall-clock 125-208 min, i.e. only 2.1-4.8x effective +// parallelism, with 37-50 min dead gaps visible at every batch boundary. The batching existed to +// dodge S10's 30-wide server throttle, but the harness already caps workflow agents at +// min(16, cores-2) and these waves are 13-14 targets — so it bought nothing and cost a barrier. +// pipeline() has NO barrier: each target flows draft -> escalate independently, so wall-clock is +// the slowest SINGLE chain rather than the sum of two batch maxima. +const results = await pipeline( + T, + (t) => agent(prompt(t, false), { + label: `draft:${t.fn}(${t.n}i,x${t.m})`, + phase: 'Draft', + model: t.model, + schema: VERDICT, + }).then((v) => ({ t, v })), + + async ({ t, v }) => { + if (!v) return { t, v: { fn: t.fn, status: 'BLOCKED', summary: 'agent returned no verdict' }, tier: t.model } + if (v.status === 'MATCH' || t.model === 'opus') return { t, v, tier: t.model } + const nextTier = t.model === 'haiku' ? 'sonnet' : 'opus' + const v2 = await agent(prompt(t, true), { + label: `escalate:${t.fn}`, + phase: 'Escalate', + model: nextTier, + schema: VERDICT, + }) + return { t, v: v2 && v2.status === 'MATCH' ? v2 : (v2 || v), tier: nextTier + '-escalated' } + }, +) + +const ok = results.filter(Boolean) +const matched = ok.filter((r) => r.v && r.v.status === 'MATCH') +log(`s37-wave: ${matched.length}/${T.length} claim MATCH (the gate is the arbiter)`) + +return { + claimed_match: matched.map((r) => r.t.fn), + verdicts: ok.map((r) => ({ + fn: r.t.fn, ov: r.t.ov, nins: r.t.n, members: r.t.m, tier: r.tier, + status: r.v ? r.v.status : 'NONE', + closeness: r.v ? r.v.closeness : null, + klass: r.v ? r.v.klass : null, + levers: r.v ? r.v.levers : null, + index_hit: r.v ? r.v.index_hit : null, + index_gap: r.v ? r.v.index_gap : null, + summary: r.v ? r.v.summary : null, + })), +} diff --git a/.run/s6f_gate.py b/.run/s6f_gate.py index c86c50144..1002e8a34 100644 --- a/.run/s6f_gate.py +++ b/.run/s6f_gate.py @@ -6,7 +6,7 @@ caught it). So this driver never trusts a recorded path: for every draft on disk where that function's INCLUDE_ASM actually lives, groups by (binary, split), and runs the whole-binary byte-gate once per group. The gate is the sole arbiter (G3/P9). """ -import sys, os, glob, subprocess, collections, shutil +import sys, os, re, glob, subprocess, collections, shutil sys.path.insert(0, 'tools') import corpus @@ -18,7 +18,14 @@ skipped = [] for d in sorted(glob.glob(sys.argv[1] if len(sys.argv)>1 else '.run/s6f/*/*.c')): ov = os.path.basename(os.path.dirname(d)) fn = os.path.basename(d)[:-2] - addr = int(fn.split('_')[1], 16) + # S35: an agent left scratch files (test_licm*.c) in the drafts dir and this line died on + # int('full', 16), taking the whole gate with it. A drafts dir is agent-writable, so treat a + # non-conforming name as a NAMED, COUNTED skip — never a crash (R32). + m_ = re.fullmatch(r'func_([0-9A-Fa-f]{8})', fn) + if not m_: + skipped.append((ov, fn, 'not a func_.c deliverable — agent scratch?')) + continue + addr = int(m_.group(1), 16) st = corpus.stubs(ov) if addr not in st: skipped.append((ov, fn, 'not a live stub (already banked?)')) diff --git a/phase-ends/CURRENT_PHASE.md b/phase-ends/CURRENT_PHASE.md index bb4f0c54b..b0268f32b 100644 --- a/phase-ends/CURRENT_PHASE.md +++ b/phase-ends/CURRENT_PHASE.md @@ -147,7 +147,114 @@ stub on a named wall/behemoth/queue ledger** — 140/140 byte-identical througho --- -# 🛑 SESSION-33/34/35 CHECKPOINT (2026-08-04) — FRESH SESSION SAFE HERE +# 🛑 SESSION-33..37 CHECKPOINT (2026-08-04) — FRESH SESSION SAFE · PAUSED FOR A WINDOWS RESTART +> **NOTHING IS RUNNING. Tree lock FREE. Tree CLEAN** but for the R23 `db.*.gbf` churn — never stage. +> Effort: **ultracode**. **R22 clean-fleet run NINETEEN times this session, 140/140 every time.** +> Drew paused here to restart Windows. **Drew's standing decision: NO phase close — keep grinding.** + +## ⚠️ THE ONE THING THAT IS OWED: 16 UNGATED WAVE-5 DRAFTS +`.run/s37/*/func_*.c` — **16 drafts, all claiming MATCH, NONE gated, NONE banked.** They are +**force-added to git** (`.run/` is otherwise gitignored) because they cost ~2.7M agent tokens and +the Workflow `resumeFromRunId` cache is SAME-SESSION-ONLY, so it does not survive the restart. +**Resume by gating them — do not re-run the wave:** +``` +tools/treelock.sh g .venv/bin/python .run/s6f_gate.py '.run/s37/*/func_*.c' +``` +then: capture any failure (`.venv/bin/python .run/s36_capture.py :` — it ranks HARD errors +above warnings) → reconcile per the §138 direction rule → `make sig-overlays` + +`tools/family_hseq.py` + `family_sweep --hseq --band all --only ` → **R22** → commit. +Manifest: `.run/s37_wave.json` (16 targets / 16,884 templatable ins). Script: `.run/s37w.js`. + +## FLEET (at the last commit, before the ungated drafts) +**96.27% fn-count · 94.1% instr-weighted · 88.6% distinct-code** (77,723 uniq) · dedup **1910/0** · +C1 241216/241216 · **0 NON_MATCHING** (G4). HEAD `commit:1391` + this checkpoint. +Session opened 96.01 / 93.6 / 88.0. Phase opened 92.00 / 87.5 / 78.0 ⇒ **+4.27 / +6.6 / +10.6pp.** + +## WHAT THIS SESSION DID +| lane | result | +|---|---| +| SC07 EXTEND | **0/36 → 31/36** | +| PROPAGATE head | **0 → 18,545/18,545 ins** (5/5 classes) | +| wave 1 (17 tgt) | 13 heads + 65 members + 2 reconciled | +| wave 2 (13 tgt) | 9 heads + 1 reconciled + 18 members | +| wave 3 (13 tgt) | **13/13** + 21 members | +| wave 4 (14 tgt) | **14/14** + 26 members | +| wave 5 (16 tgt) | **16/16 claimed — UNGATED, see above** | +| reconcile lane | **21/22 lifetime** | + +## 🔑 THE RESULT WORTH KEEPING: the wave prompt is the lever, and agents write it +Bank rate, **same models and same gate — the prompt was the only variable**: +**76% → 77% → 100% → 100%** (wave 5 pending its gate). +The jump came from **STEP 0: `grep -rn "" src/`** with a distinctive literal from the target +`.s`, placed AHEAD of `engine_core.h` in §136c's search order. §136c's first two steps are +same-TU/shared-header scoped and structurally CANNOT reach a banked twin in another overlay's TU — +where the big template classes live. **That step came from a wave-2 agent's `index_gap` report.** +By wave 4 most agents cited it by name and reported `index_gap: none`. → memory +`wave-prompt-seed-step0-and-gaps`; cookbook §138. + +## ⚡ THE `pipeline()` FIX — measured, not assumed (Drew asked why waves were slow) +The two-batch `parallel()` design was a **hard barrier**: batch 2 could not start until batch 1's +slowest agent finished. +| | wave 4 `parallel()` | wave 5 `pipeline()` | +|---|---|---| +| targets | 14 | **16** | +| wall-clock | 136 min | **82 min** | +| parallelism | 2.5× | **3.8×** | +| median agent | 14 min | 9 min | +**40% faster on 14% more targets.** `.run/s37w.js` is the pipeline version — copy its execution +block for every future wave. (Slowest single agent is still ~50 min; that is real `match_one` +iteration, 200-300 turns, and is the irreducible floor now.) + +## 📌 THE DISTILLED RULES (cookbook §138 + index; all measured this session) +1. **Bucket a gate refusal by the (macro-shape, TU-shape) PAIR, not the SYMBOL** — the same symbol + conflicts in BOTH directions across the fleet. +2. **Reconcile direction depends on WHERE the TU's decl is.** ABOVE the splice ⇒ DELETE your + duplicate (§100); BELOW ⇒ KEEP a decl in the TU's EXACT shape and cast at the use (§17a-1 D2). + Picking wrong CREATES the next error. **A wave can manufacture this for itself**: two targets in + one TU means the first to bank puts its type tags in the second's way (wave 4's only failure). +3. **STEP 0 magic-literal grep** (above). +4. **Never rank off the family map's `exemplar` field** — it is IN-FAMILY and can name an + ALREADY-BANKED instance. Derive open sites from `corpus.stubs`: 16,696 ins → **41,023** on the + same map. +5. **Three carry variants hide in one "CARRY-FIXABLE" bucket** — multi-line comment (fix the tool) · + draft-local `struct Tag` (use the shared type) · file-scope `static inline` helper (hand-author + + EXCLUDE the source overlay; gcc-2.7.2 accepts implicit decls so it passes `compiles_standalone` + and only fails 137 gates later). + +## 🔬 SETTLED — do not re-litigate +- **`func_801758FC` / `func_80132018` are NOT templatable.** Probed with `.run/s6_diag.py` (2 builds): + the remapped member is **BUILD OK + byte diff** ⇒ genuine per-member codegen. The h_seq refusal + ceiling, confirmed. Same for `func_8018A808`'s family (0/14). +- **`_alias_decl_for` was ONE function, not a class** (91 alias decls fleet-wide, 90 already matched). +- **"normalized distance 0" is NOT an h_exact guarantee** — relocs are masked. + +## ▶ RESUME ORDER +1. **Gate `.run/s37/*/func_*.c`** (above) → reconcile → propagate → R22 → commit. +2. **Wave 6**: derive the corpus.stubs way (rule 4) from a REGENERATED map. Pool at pause: **852 + families / 225,217 ins**. Exclude the 3 walls (`0x801412a8`, `0x80178004`, whale `0x80144b9c`) + and the ledgered residuals (`0x8017c294` close=12, `0x8017f7b4`, `0x801898e4`, `0x80186e24` + close=187, `0x801758fc`, `0x80132018`, `0x8018a808`). Use `.run/s37w.js`'s pipeline block. +3. **EXTEND's last 5**: `func_80144B9C` ×4 needs the §38 `-O0` shared-header route (and + `dedup_extend` should refuse-and-name that class per R32); `func_80149954` ×1 sits behind + `func_80147364`'s u16 params. + +## 🧰 MY PROCESS ERRORS — one mechanism: the signal sampled was not the thing measured +1. `nohup CMD &` inside a backgrounded call ⇒ the harness signalled the WRAPPER; the fleet check + stood at **63/140** and I nearly read it as a pass. +2. **`pgrep -x make` is right for ONE make, WRONG for a campaign** of sequential makes (fires in a + gap); **`pgrep -f ` SELF-MATCHES** so that waiter never exits. Use the campaign's real argv + or `treelock.sh --status`. +3. A `corpus.stubs` probe mid-rebuild returned garbage; R32's assertion refused to answer. +4. **I predicted a fix without reading the macro in front of me** — relaxed 42 `(void)` decls and + asserted it unblocked the lane; it banked 0/1 in all 134 because that macro declares `(u8*)`. + → the PAIR rule. +5. **I violated §136a in my own capture tool** — a narrow keyword filter reported "NO COMPILE ERROR" + on a build failing with `redefinition of struct PW8017C290`. +No bad bytes from any of them — the byte-gate and R22 caught everything. + +--- + +# 🛑 (superseded) SESSION-33/34/35 CHECKPOINT (2026-08-04) > **Tree CLEAN** but for the R23 `db.*.gbf` churn — never stage. Effort: **ultracode**. > **R22 clean-fleet run FIFTEEN times, 140/140 every time.** HEAD `commit:1388`. > **Drew's standing decision: NO phase close — keep grinding, run waves all night.**