phase12: merges 7-8 — 625 bodies / 634 regions (from 621 / 630)

Merge 7: worker D's 0x8006AD5C (84 B, a GOAL B row, first spelling) and 0x8006D1C4 (140 B).
Merge 8: worker C's GTE pair 0x80010810 / 0x8009C69C (60 B each, on the DEFAULT toolchain, via
the inline-asm hatch). Every row verified from a fresh --work dir against the exact md5, merged to
a candidate, gated whole-binary, promoted only on result=MATCH.

  sf3_match gate     c_regions=634  differing_bytes=0  result=MATCH
                     sha1 e173426c157384ebf1b6caf8c6fea18a85a14af9  (unchanged)
  make check         exit 0     extents-verify  regions=634 disagreements=0 AGREE
  registry audit     634 rows ordered, non-overlapping, 625 distinct sources, 0 missing

INTEGRITY CHECK THAT PASSED, ON A REAL HAZARD. Worker C edited the headers of the three
ALREADY-MERGED cc1bin sources after I merged them, changing their md5s (6afb89a7 / 51ab9083 /
c30c5bfa). The registry points at those paths, so a content change would have made it stale. I
re-ran the full gate on the current registry: MATCH. Comments only. This is exactly what the
md5-in-claim-row guard exists for, and here the whole-binary gate is what settled it.

C'S INLINE-ASM HATCH WAS JUSTIFIED TWICE, AND THE SECOND REASON IS A NEW FINDING.
The GTE pair is the d=0/15 identical pair, so one solve meant two bodies. The hatch was needed
because (1) ~20 spellings and ALL TEN vendored cc1 builds fold the dead `move t0,a1` into the
negation, and `((y ^ -1) + 1)` is the only spelling reaching the right LENGTH while lowering to
nor+addiu -- so length alone was never evidence; and (2) THE ORIGINAL'S NEGATION IS THE TRAPPING
`sub` (funct 0x22), not `subu` (0x23). objdump prints `neg` for the original and `negu` for the
candidate, so a mnemonic comparison cannot see it: a one-byte funct-field difference, cookbook 6's
class. With the copy fixed but the C negation kept, the row sits at differing_bytes=1.

AND C CLOSED THE ALT-CC1 QUESTION AGAINST ITS OWN INTEREST. On the 0x80050674 pair the residual is
differing_bytes=1 on the default cc1 and differing_bytes=22 on BOTH 2.8.1 and 2.91.66-psx, so it is
not a compiler-revision artifact. C then declined to spend the inline-asm hatch on it, on the
grounds that the residual is a mundane operand order inside an otherwise all-C body and the hatch
would be doing COSMETIC work. Ruling: endorsed. The hatch is for shapes plain C provably cannot
express; pinning three registers to win one operand order is not that. The row is classified with
its mechanism and an exhausted dimension instead.

OPEN, RECORDED WITH THE EXPERIMENTS RATHER THAN A GUESS: the workflow's checklist item
`excluded_already_registered == registry size` held exactly through merges 5, 6 and 7 and diverges
at merge 8 (634 registry rows, counter 632). Two hypotheses were tested and BOTH REFUTED: the two
new rows are not extent-graded non-exact (both are `exact`/`term=jr_ra`), and it is not "negatives
rows are counted under the negatives filter instead" (11 registry rows have negatives addresses,
but removing only the 2 newest from the registry reproduces the count exactly). Two failed
root-cause attempts, so per AGENTS.md rule 10 it is recorded with the evidence instead of a third
guess. No correctness risk, and that is measured: the whole-binary gate proves all 634 rows
byte-exact simultaneously, the registry audit is clean, and extents-verify agrees. It is a
diagnostic, not a gate. To be reconciled at the close, where the negatives-index reconciliation
happens anyway and is probably the same question.
This commit is contained in:
Christopher Williams
2026-09-24 17:30:17 -04:00
parent 677d982398
commit b8502c5693
7 changed files with 1146 additions and 769 deletions
+768 -769
View File
File diff suppressed because it is too large Load Diff
+4
View File
@@ -10,6 +10,7 @@
# the build is the all-payload data baseline and contains no C.
#
# Matched so far:
0x80010810 0x8001084C src/func_80010810.c
0x80012780 0x8001278C src/func_80012780.c
0x800127F0 0x8001281C src/func_800127F0.c
0x8001281C 0x80012834 src/func_8001281C.c
@@ -290,6 +291,7 @@
0x800697C4 0x800697FC src/func_800697C4.c
0x8006A98C 0x8006AA10 src/func_8006A98C.c
0x8006AA10 0x8006AA88 src/func_8006AA10.c
0x8006AD5C 0x8006ADB0 src/func_8006AD5C.c
0x8006ADB0 0x8006AE04 src/func_8006ADB0.c
0x8006AE04 0x8006AE54 src/func_8006AE04.c
0x8006B0A8 0x8006B184 src/func_8006B0A8.c
@@ -311,6 +313,7 @@
0x8006BC08 0x8006BC34 src/func_8006BC08.c
0x8006BC34 0x8006BC74 src/func_8006BC34.c
0x8006BC74 0x8006BF30 src/func_8006BC74.c
0x8006D1C4 0x8006D250 src/func_8006D1C4.c
0x8006EC94 0x8006ECD4 src/func_8006EC94.c
0x8006F28C 0x8006F2F4 src/func_8006F28C.c
0x8006F6BC 0x8006F6F4 src/func_8006F6BC.c
@@ -416,6 +419,7 @@
0x8009AC08 0x8009AC28 src/func_8009AC08.c
0x8009B464 0x8009B4A0 src/func_8009B464.c
0x8009B56C 0x8009B638 src/func_8009B56C.c
0x8009C69C 0x8009C6D8 src/func_8009C69C.c
0x8009C904 0x8009CB28 src/func_8009C904.c
0x8009D798 0x8009D8A0 src/func_8009D798.c
0x8009D8A0 0x8009D8E0 src/func_8009D8A0.c
1 # Code-region registry: one C region per matched function.
10 # the build is the all-payload data baseline and contains no C.
11 #
12 # Matched so far:
13 0x80010810
14 0x80012780
15 0x800127F0
16 0x8001281C
291 0x800697C4
292 0x8006A98C
293 0x8006AA10
294 0x8006AD5C
295 0x8006ADB0
296 0x8006AE04
297 0x8006B0A8
313 0x8006BC08
314 0x8006BC34
315 0x8006BC74
316 0x8006D1C4
317 0x8006EC94
318 0x8006F28C
319 0x8006F6BC
419 0x8009AC08
420 0x8009B464
421 0x8009B56C
422 0x8009C69C
423 0x8009C904
424 0x8009D798
425 0x8009D8A0
+71
View File
@@ -480,3 +480,74 @@ Now `crc32("band:address") % 3` — membership depends only on the row's own ide
removing 10 rows moves 0 assignments.** Buckets came out 310/330/353, which is a 12% spread; the
partition is a **default, not a wall**, and abutting neighbours are explicitly claimable outside it.
`crc32`, not Python's `hash()`, which is salted per process and would re-partition every run.
### Merges 7-8 — **625 bodies / 634 regions**
Merge 7: worker D's `0x8006AD5C` (84 B, a **Goal B** row, first spelling) and `0x8006D1C4` (140 B).
Merge 8: worker C's **GTE pair** `0x80010810`/`0x8009C69C` (60 B each, **on the DEFAULT toolchain**,
via the inline-asm hatch). Every row verified from a fresh `--work` dir against the exact md5,
merged to a candidate, gated whole-binary, promoted only on `result=MATCH`, `make check` exit 0.
**Integrity check on a shared-namespace hazard, and it passed.** Worker C edited the *headers* of
the three already-merged `cc1bin` sources after I had merged them, which changed their md5s
(`6afb89a7…`, `51ab9083…`, `c30c5bfa…` — C reported all three). The registry points at those paths,
so a content change would have made the registry stale. I re-ran the **full gate on the current
registry: `c_regions=632 differing_bytes=0 result=MATCH`** — comments only, nothing broke. Recorded
because this is precisely the hazard the md5-in-claim-row guard exists for, and here the whole-binary
gate was the thing that settled it.
**C's inline-asm hatch was justified twice over, and the second reason is a new finding.** The GTE
pair is the d=0/15 identical pair, so one solve was two bodies. C used the hatch because:
1. ~20 spellings and **all ten vendored cc1 builds** fold the dead `move t0,a1` into the negation,
and `((y ^ -1) + 1)` is the only spelling reaching the right LENGTH (60) while lowering to
`nor`+`addiu` (17 differing bytes) — so **length alone was never evidence**;
2. **the original's negation is the TRAPPING `sub` (funct 0x22), not `subu` (0x23).** `objdump`
prints `neg` for the original and `negu` for the candidate, so a mnemonic comparison cannot see
it — it is a **one-byte funct-field difference**, cookbook 6's class. With the copy fixed but the
C negation kept, the row sits at `differing_bytes=1` (0x00084023 vs 0x00084022).
C explicitly noted this does **not** put the row in the blocked `trapping-arithmetic` class: only
the negate traps, and the row's class is `register-tiebreak`. The escape hatch is one statement
(`move` + `sub`), everything else in C, `t0` pinned to `$8`, and the file header carries the full
auditable proof.
**C also CLOSED the alt-cc1 question on the 0x80050674 pair, against its own interest**: that
residual is `differing_bytes=1` on the default cc1 and **`differing_bytes=22` on both 2.8.1 and
2.91.66-psx** (5 spellings each). So it is not a compiler-revision artifact. C declined to spend the
inline-asm hatch on it, reasoning that the residual is a mundane operand order inside an otherwise
all-C body and the hatch would be doing **cosmetic** work. **Orchestrator ruling: the decline is
endorsed.** The hatch exists for shapes plain C provably cannot express; pinning three registers to
win one operand order is not that, and the row is classified with its mechanism and an exhausted
dimension instead. That is the right use of a decline.
### ⚠ OPEN: the `excluded_already_registered == registry size` invariant no longer reconciles
The workflow's checklist requires `excluded_already_registered == registry size` after every
worklist regeneration. It held exactly through merges 5, 6 and 7 and **diverges at merge 8**:
| | registry | excluded_already_registered |
|---|---|---|
| merge 5 | 629 | 629 |
| merge 6 | 630 | 630 |
| merge 7 | 632 | 632 |
| **merge 8** | **634** | **632** |
**Two hypotheses, both REFUTED by experiment:**
1. *"The two new rows are extent-graded non-exact and are filtered earlier."* **No** — both are
graded `exact` with `term=jr_ra`, same as the rows that ARE counted.
2. *"Rows still present in the negatives index are excluded by the negatives filter instead."*
**No** — 11 registry rows have negatives-index addresses, but removing only the 2 newest from the
registry reproduces the count exactly (`registry=632 rows → 632`; a 634-row registry → also 632).
So the counter is lagging by exactly the 2 rows added in merge 8 (`0x80010810`, `0x8009C69C`), and
**I do not know why.** I am recording it rather than guessing a third mechanism, per AGENTS.md
rule 10: two distinct root-cause attempts have failed, so stop and report the evidence.
**It carries no correctness risk and that is measured, not assumed:** the whole-binary gate proves
every one of the 634 registry rows is byte-exact simultaneously, the registry audit shows 634 rows
ordered, non-overlapping, 625 distinct sources, 0 missing, and `extents-verify` agrees at 634. The
invariant is a *diagnostic*, not a gate. **To be reconciled at the phase close**, where the full
census is available and the negatives-index reconciliation is done anyway — the two are probably the
same question, since the divergence appeared on the first merge whose additions were all
negatives-index rows that no earlier merge had registered.
+93
View File
@@ -0,0 +1,93 @@
/*
* func_80010810 — 60 bytes at 0x80010810..0x8001084C
*
* PHASE 12 WORKER C. A GTE (COP2) row, and one that needed the INLINE-ASM ESCAPE HATCH.
* The hatch was taken with the developer's authorisation and the proof is below.
*
* WHAT THE BODY DOES, read off the instruction stream (only the bytes are evidence):
* move t0,a1 t0 = y
* sub t0,zero,t0 t0 = -t0 <- FUNCT 0x22, the TRAPPING `sub`
* sll t0,t0,0x10 t0 <<= 16
* andi a0,a0,0xffff x &= 0xffff
* or t0,t0,a0 t0 = (VY << 16) | VX
* mtc2 t0,$0 VXY0
* mtc2 a2,$1 VZ0
* nop
* nop
* cop2 0x180001 nRTPS
* nop
* mfc2 t0,$19 t0 = SZ3
* nop
* jr ra
* move v0,t0
*
* `include/gtemac.h` already records VXY0/VZ0/SZ3 and `gte_nRTPS()` as PROVED BY THIS ROW
* (its comments cite 0x80010810 by address). The twin 0x8009C69C is byte-identical (raw
* words d=0/15, see src/func_8009C69C.c).
*
* --- WHY INLINE ASM: the two instructions plain C cannot emit (the proof) ---
*
* The residual is TWO facts, and both had to be established separately.
*
* (1) A DEAD ARGUMENT COPY. The original's pack begins with `move t0,a1` and only then
* `neg t0,t0`. Writing the negation in C -- `t0 = y; t0 = -t0;`, or `-(y)` directly, or
* any of the named-temporary / cast / two's-complement-manual spellings -- makes cc1's
* `combine` pass fuse the copy into the negation and emit the ONE-instruction
* `subu t0,zero,a1`. Measured: 20 spellings, ALL giving candidate_bytes=56 where the
* original has 60. The spellings tried:
* `-(y)` inline; `int ny; ny = y; ny = -ny;`; `ny = y; ny = 0 - ny;`;
* `short vy = -y;`; `unsigned short vx = x, vy = -y;`; `int ny = y;` then `-ny`;
* `register int t0 asm("$8")` with `t0 = y; t0 = -t0;`;
* `register int t0 asm("$8")` with `t0 = y;` then `-t0` inside the pack expression;
* `register int t0 asm("$8")` with an explicit `pack` temporary;
* `register int y` / `register`-qualified parameters; `(int)(short)y`;
* `y * -65536`; `(char *)`/`(int)` casts;
* and `((y ^ -1) + 1)`, the ONLY spelling that reaches the correct LENGTH (60) --
* it lowers to `nor`+`addiu` (17 differing bytes), which proves the length alone is
* not evidence of the right sequence.
* And it is not a compiler-revision artifact either: ALL TEN cc1 builds vendored in
* `tools/old-gcc/` (2.5.7-psx, 2.6.0-psx, 2.6.3-psx, 2.7.2, 2.7.2-cdk, 2.7.2-psx,
* 2.8.0-psx, 2.8.1-psx, 2.91.66-psx, 2.95.2-psx) give 56 on this source. There is no C
* spelling and no vendored cc1 that emits a dead copy; the copy is required by the
* original's bytes, so it is emitted explicitly.
*
* (2) THE NEGATION IS THE TRAPPING `sub`, NOT `subu`. The original's word 1 is 0x00084022:
* funct 0x22 = `sub`, which traps on overflow. cc1's `-t0` always lowers to `negu`
* (0x00084023, funct 0x23 = `subu`). A mnemonic comparison calls both of them `neg`
* -- objdump prints `neg` for the original and `negu` for the candidate -- and the
* difference is ONE byte in the funct field. This is cookbook finding 6's class
* ("count right + residual 1-8 bytes: diff the ENCODING, not the mnemonic") with a
* 1-byte residual that no C negation can reach. NOTE this is a local fact about this
* row, not an argument that the row is in the blocked `trapping-arithmetic` class:
* 0x80010810 is classed `register-tiebreak`, and only the negate here traps.
* With (1) alone the row sits at 1 differing byte (0x00084023 vs 0x00084022); with
* both, it is byte-exact.
*
* The two instructions therefore go through one asm statement that emits exactly them:
* __asm__ volatile ("move %0,%1\n\tsub %0,$zero,%0" : "=r"(t0) : "r"(y));
* Everything else in the body stays in C, and `t0` is a register variable pinned to $8 so
* the COP2 operand registers are the original's (the idiom established for 0x800F3E18).
* The `nop`s are written at the call site, per the gtemac.h convention for this file.
*
* LIMITS. The parameter names and the meaning of the packed coordinates are hypotheses
* read off the instruction shape; the function's purpose (an nRTPS projection returning
* SZ3) is inferred from the macro set that this row itself proved, not independently.
* The command field is named by its VALUE, per the gtemac.h convention.
*/
#include "../include/gtemac.h"
int func_80010810(int x, int y, int z)
{
register int t0 __asm__("$8");
__asm__ volatile ("move %0,%1\n\tsub %0,$zero,%0" : "=r"(t0) : "r"(y));
gte_ldVXY0(((unsigned int)t0 << 16) | ((unsigned int)x & 0xFFFF));
gte_ldVZ0(z);
__asm__ volatile ("nop");
__asm__ volatile ("nop");
gte_nRTPS();
__asm__ volatile ("nop");
gte_stSZ3(t0);
__asm__ volatile ("nop");
return t0;
}
+21
View File
@@ -0,0 +1,21 @@
/*
* func_8006AD5C — 84 bytes at 0x8006AD5C..0x8006ADB0 [DRAFT header]
*/
extern char *D_80121E88;
extern unsigned char D_80122374;
void func_800C2D88(char *);
void func_8006AD5C(void)
{
char *p = D_80121E88;
while (p != 0) {
char *n = *(char **)(p + 12);
char *next = *(char **)(n + 396);
if (D_80122374 == 0)
func_800C2D88(p);
p = next;
}
}
+94
View File
@@ -0,0 +1,94 @@
/*
* func_8006D1C4 — 140 bytes at 0x8006D1C4..0x8006D250
*
* Byte-identical reconstruction. Initialises a 64-byte stack record (two 16-byte
* vectors copied from the third argument when it is non-null, plus a flag word
* whose bit 0 records whether that argument was non-null) and then calls the
* FIRST ARGUMENT as a function pointer, passing the second argument and the
* record's address.
*
* The observed instructions are:
* addiu sp,sp,-88
* move t0,a0 ; fn — must survive the call's a0 setup
* move a3,a1 ; the value passed on as the call's first argument
* sw ra,80(sp)
* beqz a2,common ; if (src == 0) skip the copies
* sw zero,20(sp) ; (delay slot) s.f1 = 0
* lw v0,0(a2) / lw v1,4(a2) / lw a0,8(a2) / lw a1,12(a2)
* sw v0,24(sp) / sw v1,28(sp) / sw a0,32(sp) / sw a1,36(sp) ; s.a
* lw v0,16(a2) / lw v1,20(a2) / lw a0,24(a2) / lw a1,28(a2)
* sw v0,40(sp) / sw v1,44(sp) / sw a0,48(sp) / sw a1,52(sp) ; s.b
* common:
* li v1,-2 ; mask for bit 0
* move a0,a3 ; the call's first argument
* lw v0,56(sp) ; s.flag
* addiu a1,sp,16 ; the record's address
* and v0,v0,v1
* sltu v1,zero,a2 ; (src != 0)
* or v0,v0,v1
* jalr t0 ; fn(second argument, &s)
* sw v0,56(sp) ; (delay slot) s.flag = ...
* epilogue:
* lw ra,80(sp) / addiu sp,sp,88 / jr ra / nop
* (the return value is whatever the indirect call left in v0)
*
* FOUR things the bytes fix:
*
* 1. `s.f1 = 0` IS AN UNCONDITIONAL INITIALISATION BEFORE THE `if`, NOT A
* STATEMENT INSIDE IT. Written inside the `if (src != 0)` block the row comes
* out at the correct 140 bytes with **13 differing bytes**: the store loses
* the `beqz` delay slot to the `move a3,a1` parameter copy, which then sits in
* the fall-through block, and the whole entry-block order changes
* (`sw ra` moves first). Moving the initialisation before the branch gives it
* the delay slot and keeps BOTH parameter copies in the entry block — the
* original exactly. The delay slot is what makes an otherwise-invisible
* source-order choice load-bearing (cookbook 149/163/184).
* 2. THE TWO VECTOR COPIES ARE 16-BYTE STRUCT ASSIGNMENTS. Each is 4 batched
* loads then 4 batched stores (cookbook 96/102), so the source copies two
* `int[4]` structs, not eight element stores; element stores would serialise
* behind load-delay nops.
* 3. THE FIRST ARGUMENT IS CALLED INDIRECTLY AND MUST BE COPIED OUT OF `a0`
* FIRST, because the call's own argument setup overwrites `a0` with the
* second argument (`move a0,a3`). Likewise the second argument is copied out
* of `a1` because `a1` receives `&s`. Both copies land in caller-saved
* registers (`t0`, `a3`) — no saved register is needed, only `ra` is stacked.
* 4. THE RECORD IS 64 BYTES AND ONLY +4, +8..+39 AND +40 ARE TOUCHED. The frame
* arithmetic fixes the size: 88 = 16 (outgoing argument area) + 64 (the
* record at sp+16..sp+79) + 4 (`ra` at 80) + 4 padding. The bytes at +0
* (sp+16) and +44..+63 are never written here, so the trailing `unknown[5]` is
* a declaration sized to reproduce the frame, not a recovered field: the
* address of the record escapes into the callee, which is what makes the
* whole size load-bearing (cookbook 59/131).
*
* LIMITS: every displacement, the frame and the mask are read from the bytes.
* Field names and types are hypotheses; in particular `f0` at +0 and the
* trailing fields are unnamed because nothing writes them, and the flag word's
* meaning is only "bit 0 = (src != 0), and the low bit is preserved otherwise".
* The indirect call's signature is inferred from the argument registers that are
* set, and its return value is whatever the callee leaves in v0.
*/
struct func_8006D1C4_V {
int v[4];
};
struct func_8006D1C4_S {
int f0; /* +0 — never written here */
int f1; /* +4 — zeroed unconditionally */
struct func_8006D1C4_V a; /* +8 — copied from src[0..3] */
struct func_8006D1C4_V b; /* +24 — copied from src[4..7] */
int flag; /* +40 — bit 0 = (src != 0), low bit cleared */
int unknown[5]; /* +44..+63 — sized to reproduce the frame */
};
int func_8006D1C4(int (*fn)(), int a1, int *src) {
struct func_8006D1C4_S s;
s.f1 = 0;
if (src != 0) {
s.a = *(struct func_8006D1C4_V *)src;
s.b = *(struct func_8006D1C4_V *)(src + 4);
}
s.flag = (s.flag & -2) | (src != 0);
return fn(a1, &s);
}
+95
View File
@@ -0,0 +1,95 @@
/*
* func_8009C69C — 60 bytes at 0x8009C69C..0x8009C6D8
*
* PHASE 12 WORKER C. A GTE (COP2) row, and one that needed the INLINE-ASM ESCAPE HATCH.
* The hatch was taken with the developer's authorisation and the proof is below.
*
* WHAT THE BODY DOES, read off the instruction stream (only the bytes are evidence):
* move t0,a1 t0 = y
* sub t0,zero,t0 t0 = -t0 <- FUNCT 0x22, the TRAPPING `sub`
* sll t0,t0,0x10 t0 <<= 16
* andi a0,a0,0xffff x &= 0xffff
* or t0,t0,a0 t0 = (VY << 16) | VX
* mtc2 t0,$0 VXY0
* mtc2 a2,$1 VZ0
* nop
* nop
* cop2 0x180001 nRTPS
* nop
* mfc2 t0,$19 t0 = SZ3
* nop
* jr ra
* move v0,t0
*
* `include/gtemac.h` already records VXY0/VZ0/SZ3 and `gte_nRTPS()` as PROVED BY THIS ROW
* (its comments cite 0x80010810 by address). The twin 0x80010810 is byte-identical (raw
* words d=0/15, see src/func_80010810.c): the two 15-word bodies differ in ZERO words, so this
* file is the twin's body with the symbol renamed, and the whole proof below applies
* unchanged. Verified independently at its own address.
*
* --- WHY INLINE ASM: the two instructions plain C cannot emit (the proof) ---
*
* The residual is TWO facts, and both had to be established separately.
*
* (1) A DEAD ARGUMENT COPY. The original's pack begins with `move t0,a1` and only then
* `neg t0,t0`. Writing the negation in C -- `t0 = y; t0 = -t0;`, or `-(y)` directly, or
* any of the named-temporary / cast / two's-complement-manual spellings -- makes cc1's
* `combine` pass fuse the copy into the negation and emit the ONE-instruction
* `subu t0,zero,a1`. Measured: 20 spellings, ALL giving candidate_bytes=56 where the
* original has 60. The spellings tried:
* `-(y)` inline; `int ny; ny = y; ny = -ny;`; `ny = y; ny = 0 - ny;`;
* `short vy = -y;`; `unsigned short vx = x, vy = -y;`; `int ny = y;` then `-ny`;
* `register int t0 asm("$8")` with `t0 = y; t0 = -t0;`;
* `register int t0 asm("$8")` with `t0 = y;` then `-t0` inside the pack expression;
* `register int t0 asm("$8")` with an explicit `pack` temporary;
* `register int y` / `register`-qualified parameters; `(int)(short)y`;
* `y * -65536`; `(char *)`/`(int)` casts;
* and `((y ^ -1) + 1)`, the ONLY spelling that reaches the correct LENGTH (60) --
* it lowers to `nor`+`addiu` (17 differing bytes), which proves the length alone is
* not evidence of the right sequence.
* And it is not a compiler-revision artifact either: ALL TEN cc1 builds vendored in
* `tools/old-gcc/` (2.5.7-psx, 2.6.0-psx, 2.6.3-psx, 2.7.2, 2.7.2-cdk, 2.7.2-psx,
* 2.8.0-psx, 2.8.1-psx, 2.91.66-psx, 2.95.2-psx) give 56 on this source. There is no C
* spelling and no vendored cc1 that emits a dead copy; the copy is required by the
* original's bytes, so it is emitted explicitly.
*
* (2) THE NEGATION IS THE TRAPPING `sub`, NOT `subu`. The original's word 1 is 0x00084022:
* funct 0x22 = `sub`, which traps on overflow. cc1's `-t0` always lowers to `negu`
* (0x00084023, funct 0x23 = `subu`). A mnemonic comparison calls both of them `neg`
* -- objdump prints `neg` for the original and `negu` for the candidate -- and the
* difference is ONE byte in the funct field. This is cookbook finding 6's class
* ("count right + residual 1-8 bytes: diff the ENCODING, not the mnemonic") with a
* 1-byte residual that no C negation can reach. NOTE this is a local fact about this
* row, not an argument that the row is in the blocked `trapping-arithmetic` class:
* 0x80010810 is classed `register-tiebreak`, and only the negate here traps.
* With (1) alone the row sits at 1 differing byte (0x00084023 vs 0x00084022); with
* both, it is byte-exact.
*
* The two instructions therefore go through one asm statement that emits exactly them:
* __asm__ volatile ("move %0,%1\n\tsub %0,$zero,%0" : "=r"(t0) : "r"(y));
* Everything else in the body stays in C, and `t0` is a register variable pinned to $8 so
* the COP2 operand registers are the original's (the idiom established for 0x800F3E18).
* The `nop`s are written at the call site, per the gtemac.h convention for this file.
*
* LIMITS. The parameter names and the meaning of the packed coordinates are hypotheses
* read off the instruction shape; the function's purpose (an nRTPS projection returning
* SZ3) is inferred from the macro set that this row itself proved, not independently.
* The command field is named by its VALUE, per the gtemac.h convention.
*/
#include "../include/gtemac.h"
int func_8009C69C(int x, int y, int z)
{
register int t0 __asm__("$8");
__asm__ volatile ("move %0,%1\n\tsub %0,$zero,%0" : "=r"(t0) : "r"(y));
gte_ldVXY0(((unsigned int)t0 << 16) | ((unsigned int)x & 0xFFFF));
gte_ldVZ0(z);
__asm__ volatile ("nop");
__asm__ volatile ("nop");
gte_nRTPS();
__asm__ volatile ("nop");
gte_stSZ3(t0);
__asm__ volatile ("nop");
return t0;
}