phase12: merges 7-8 — 625 bodies / 634 regions (from 621 / 630)
Merge 7: worker D's 0x8006AD5C (84 B, a GOAL B row, first spelling) and 0x8006D1C4 (140 B).
Merge 8: worker C's GTE pair 0x80010810 / 0x8009C69C (60 B each, on the DEFAULT toolchain, via
the inline-asm hatch). Every row verified from a fresh --work dir against the exact md5, merged to
a candidate, gated whole-binary, promoted only on result=MATCH.
sf3_match gate c_regions=634 differing_bytes=0 result=MATCH
sha1 e173426c157384ebf1b6caf8c6fea18a85a14af9 (unchanged)
make check exit 0 extents-verify regions=634 disagreements=0 AGREE
registry audit 634 rows ordered, non-overlapping, 625 distinct sources, 0 missing
INTEGRITY CHECK THAT PASSED, ON A REAL HAZARD. Worker C edited the headers of the three
ALREADY-MERGED cc1bin sources after I merged them, changing their md5s (6afb89a7 / 51ab9083 /
c30c5bfa). The registry points at those paths, so a content change would have made it stale. I
re-ran the full gate on the current registry: MATCH. Comments only. This is exactly what the
md5-in-claim-row guard exists for, and here the whole-binary gate is what settled it.
C'S INLINE-ASM HATCH WAS JUSTIFIED TWICE, AND THE SECOND REASON IS A NEW FINDING.
The GTE pair is the d=0/15 identical pair, so one solve meant two bodies. The hatch was needed
because (1) ~20 spellings and ALL TEN vendored cc1 builds fold the dead `move t0,a1` into the
negation, and `((y ^ -1) + 1)` is the only spelling reaching the right LENGTH while lowering to
nor+addiu -- so length alone was never evidence; and (2) THE ORIGINAL'S NEGATION IS THE TRAPPING
`sub` (funct 0x22), not `subu` (0x23). objdump prints `neg` for the original and `negu` for the
candidate, so a mnemonic comparison cannot see it: a one-byte funct-field difference, cookbook 6's
class. With the copy fixed but the C negation kept, the row sits at differing_bytes=1.
AND C CLOSED THE ALT-CC1 QUESTION AGAINST ITS OWN INTEREST. On the 0x80050674 pair the residual is
differing_bytes=1 on the default cc1 and differing_bytes=22 on BOTH 2.8.1 and 2.91.66-psx, so it is
not a compiler-revision artifact. C then declined to spend the inline-asm hatch on it, on the
grounds that the residual is a mundane operand order inside an otherwise all-C body and the hatch
would be doing COSMETIC work. Ruling: endorsed. The hatch is for shapes plain C provably cannot
express; pinning three registers to win one operand order is not that. The row is classified with
its mechanism and an exhausted dimension instead.
OPEN, RECORDED WITH THE EXPERIMENTS RATHER THAN A GUESS: the workflow's checklist item
`excluded_already_registered == registry size` held exactly through merges 5, 6 and 7 and diverges
at merge 8 (634 registry rows, counter 632). Two hypotheses were tested and BOTH REFUTED: the two
new rows are not extent-graded non-exact (both are `exact`/`term=jr_ra`), and it is not "negatives
rows are counted under the negatives filter instead" (11 registry rows have negatives addresses,
but removing only the 2 newest from the registry reproduces the count exactly). Two failed
root-cause attempts, so per AGENTS.md rule 10 it is recorded with the evidence instead of a third
guess. No correctness risk, and that is measured: the whole-binary gate proves all 634 rows
byte-exact simultaneously, the registry audit is clean, and extents-verify agrees. It is a
diagnostic, not a gate. To be reconciled at the close, where the negatives-index reconciliation
happens anyway and is probably the same question.
This commit is contained in:
+768
-769
File diff suppressed because it is too large
Load Diff
@@ -10,6 +10,7 @@
|
||||
# the build is the all-payload data baseline and contains no C.
|
||||
#
|
||||
# Matched so far:
|
||||
0x80010810 0x8001084C src/func_80010810.c
|
||||
0x80012780 0x8001278C src/func_80012780.c
|
||||
0x800127F0 0x8001281C src/func_800127F0.c
|
||||
0x8001281C 0x80012834 src/func_8001281C.c
|
||||
@@ -290,6 +291,7 @@
|
||||
0x800697C4 0x800697FC src/func_800697C4.c
|
||||
0x8006A98C 0x8006AA10 src/func_8006A98C.c
|
||||
0x8006AA10 0x8006AA88 src/func_8006AA10.c
|
||||
0x8006AD5C 0x8006ADB0 src/func_8006AD5C.c
|
||||
0x8006ADB0 0x8006AE04 src/func_8006ADB0.c
|
||||
0x8006AE04 0x8006AE54 src/func_8006AE04.c
|
||||
0x8006B0A8 0x8006B184 src/func_8006B0A8.c
|
||||
@@ -311,6 +313,7 @@
|
||||
0x8006BC08 0x8006BC34 src/func_8006BC08.c
|
||||
0x8006BC34 0x8006BC74 src/func_8006BC34.c
|
||||
0x8006BC74 0x8006BF30 src/func_8006BC74.c
|
||||
0x8006D1C4 0x8006D250 src/func_8006D1C4.c
|
||||
0x8006EC94 0x8006ECD4 src/func_8006EC94.c
|
||||
0x8006F28C 0x8006F2F4 src/func_8006F28C.c
|
||||
0x8006F6BC 0x8006F6F4 src/func_8006F6BC.c
|
||||
@@ -416,6 +419,7 @@
|
||||
0x8009AC08 0x8009AC28 src/func_8009AC08.c
|
||||
0x8009B464 0x8009B4A0 src/func_8009B464.c
|
||||
0x8009B56C 0x8009B638 src/func_8009B56C.c
|
||||
0x8009C69C 0x8009C6D8 src/func_8009C69C.c
|
||||
0x8009C904 0x8009CB28 src/func_8009C904.c
|
||||
0x8009D798 0x8009D8A0 src/func_8009D798.c
|
||||
0x8009D8A0 0x8009D8E0 src/func_8009D8A0.c
|
||||
|
||||
|
@@ -480,3 +480,74 @@ Now `crc32("band:address") % 3` — membership depends only on the row's own ide
|
||||
removing 10 rows moves 0 assignments.** Buckets came out 310/330/353, which is a 12% spread; the
|
||||
partition is a **default, not a wall**, and abutting neighbours are explicitly claimable outside it.
|
||||
`crc32`, not Python's `hash()`, which is salted per process and would re-partition every run.
|
||||
|
||||
### Merges 7-8 — **625 bodies / 634 regions**
|
||||
|
||||
Merge 7: worker D's `0x8006AD5C` (84 B, a **Goal B** row, first spelling) and `0x8006D1C4` (140 B).
|
||||
Merge 8: worker C's **GTE pair** `0x80010810`/`0x8009C69C` (60 B each, **on the DEFAULT toolchain**,
|
||||
via the inline-asm hatch). Every row verified from a fresh `--work` dir against the exact md5,
|
||||
merged to a candidate, gated whole-binary, promoted only on `result=MATCH`, `make check` exit 0.
|
||||
|
||||
**Integrity check on a shared-namespace hazard, and it passed.** Worker C edited the *headers* of
|
||||
the three already-merged `cc1bin` sources after I had merged them, which changed their md5s
|
||||
(`6afb89a7…`, `51ab9083…`, `c30c5bfa…` — C reported all three). The registry points at those paths,
|
||||
so a content change would have made the registry stale. I re-ran the **full gate on the current
|
||||
registry: `c_regions=632 differing_bytes=0 result=MATCH`** — comments only, nothing broke. Recorded
|
||||
because this is precisely the hazard the md5-in-claim-row guard exists for, and here the whole-binary
|
||||
gate was the thing that settled it.
|
||||
|
||||
**C's inline-asm hatch was justified twice over, and the second reason is a new finding.** The GTE
|
||||
pair is the d=0/15 identical pair, so one solve was two bodies. C used the hatch because:
|
||||
1. ~20 spellings and **all ten vendored cc1 builds** fold the dead `move t0,a1` into the negation,
|
||||
and `((y ^ -1) + 1)` is the only spelling reaching the right LENGTH (60) while lowering to
|
||||
`nor`+`addiu` (17 differing bytes) — so **length alone was never evidence**;
|
||||
2. **the original's negation is the TRAPPING `sub` (funct 0x22), not `subu` (0x23).** `objdump`
|
||||
prints `neg` for the original and `negu` for the candidate, so a mnemonic comparison cannot see
|
||||
it — it is a **one-byte funct-field difference**, cookbook 6's class. With the copy fixed but the
|
||||
C negation kept, the row sits at `differing_bytes=1` (0x00084023 vs 0x00084022).
|
||||
|
||||
C explicitly noted this does **not** put the row in the blocked `trapping-arithmetic` class: only
|
||||
the negate traps, and the row's class is `register-tiebreak`. The escape hatch is one statement
|
||||
(`move` + `sub`), everything else in C, `t0` pinned to `$8`, and the file header carries the full
|
||||
auditable proof.
|
||||
|
||||
**C also CLOSED the alt-cc1 question on the 0x80050674 pair, against its own interest**: that
|
||||
residual is `differing_bytes=1` on the default cc1 and **`differing_bytes=22` on both 2.8.1 and
|
||||
2.91.66-psx** (5 spellings each). So it is not a compiler-revision artifact. C declined to spend the
|
||||
inline-asm hatch on it, reasoning that the residual is a mundane operand order inside an otherwise
|
||||
all-C body and the hatch would be doing **cosmetic** work. **Orchestrator ruling: the decline is
|
||||
endorsed.** The hatch exists for shapes plain C provably cannot express; pinning three registers to
|
||||
win one operand order is not that, and the row is classified with its mechanism and an exhausted
|
||||
dimension instead. That is the right use of a decline.
|
||||
|
||||
### ⚠ OPEN: the `excluded_already_registered == registry size` invariant no longer reconciles
|
||||
|
||||
The workflow's checklist requires `excluded_already_registered == registry size` after every
|
||||
worklist regeneration. It held exactly through merges 5, 6 and 7 and **diverges at merge 8**:
|
||||
|
||||
| | registry | excluded_already_registered |
|
||||
|---|---|---|
|
||||
| merge 5 | 629 | 629 |
|
||||
| merge 6 | 630 | 630 |
|
||||
| merge 7 | 632 | 632 |
|
||||
| **merge 8** | **634** | **632** |
|
||||
|
||||
**Two hypotheses, both REFUTED by experiment:**
|
||||
|
||||
1. *"The two new rows are extent-graded non-exact and are filtered earlier."* **No** — both are
|
||||
graded `exact` with `term=jr_ra`, same as the rows that ARE counted.
|
||||
2. *"Rows still present in the negatives index are excluded by the negatives filter instead."*
|
||||
**No** — 11 registry rows have negatives-index addresses, but removing only the 2 newest from the
|
||||
registry reproduces the count exactly (`registry=632 rows → 632`; a 634-row registry → also 632).
|
||||
|
||||
So the counter is lagging by exactly the 2 rows added in merge 8 (`0x80010810`, `0x8009C69C`), and
|
||||
**I do not know why.** I am recording it rather than guessing a third mechanism, per AGENTS.md
|
||||
rule 10: two distinct root-cause attempts have failed, so stop and report the evidence.
|
||||
|
||||
**It carries no correctness risk and that is measured, not assumed:** the whole-binary gate proves
|
||||
every one of the 634 registry rows is byte-exact simultaneously, the registry audit shows 634 rows
|
||||
ordered, non-overlapping, 625 distinct sources, 0 missing, and `extents-verify` agrees at 634. The
|
||||
invariant is a *diagnostic*, not a gate. **To be reconciled at the phase close**, where the full
|
||||
census is available and the negatives-index reconciliation is done anyway — the two are probably the
|
||||
same question, since the divergence appeared on the first merge whose additions were all
|
||||
negatives-index rows that no earlier merge had registered.
|
||||
|
||||
@@ -0,0 +1,93 @@
|
||||
/*
|
||||
* func_80010810 — 60 bytes at 0x80010810..0x8001084C
|
||||
*
|
||||
* PHASE 12 WORKER C. A GTE (COP2) row, and one that needed the INLINE-ASM ESCAPE HATCH.
|
||||
* The hatch was taken with the developer's authorisation and the proof is below.
|
||||
*
|
||||
* WHAT THE BODY DOES, read off the instruction stream (only the bytes are evidence):
|
||||
* move t0,a1 t0 = y
|
||||
* sub t0,zero,t0 t0 = -t0 <- FUNCT 0x22, the TRAPPING `sub`
|
||||
* sll t0,t0,0x10 t0 <<= 16
|
||||
* andi a0,a0,0xffff x &= 0xffff
|
||||
* or t0,t0,a0 t0 = (VY << 16) | VX
|
||||
* mtc2 t0,$0 VXY0
|
||||
* mtc2 a2,$1 VZ0
|
||||
* nop
|
||||
* nop
|
||||
* cop2 0x180001 nRTPS
|
||||
* nop
|
||||
* mfc2 t0,$19 t0 = SZ3
|
||||
* nop
|
||||
* jr ra
|
||||
* move v0,t0
|
||||
*
|
||||
* `include/gtemac.h` already records VXY0/VZ0/SZ3 and `gte_nRTPS()` as PROVED BY THIS ROW
|
||||
* (its comments cite 0x80010810 by address). The twin 0x8009C69C is byte-identical (raw
|
||||
* words d=0/15, see src/func_8009C69C.c).
|
||||
*
|
||||
* --- WHY INLINE ASM: the two instructions plain C cannot emit (the proof) ---
|
||||
*
|
||||
* The residual is TWO facts, and both had to be established separately.
|
||||
*
|
||||
* (1) A DEAD ARGUMENT COPY. The original's pack begins with `move t0,a1` and only then
|
||||
* `neg t0,t0`. Writing the negation in C -- `t0 = y; t0 = -t0;`, or `-(y)` directly, or
|
||||
* any of the named-temporary / cast / two's-complement-manual spellings -- makes cc1's
|
||||
* `combine` pass fuse the copy into the negation and emit the ONE-instruction
|
||||
* `subu t0,zero,a1`. Measured: 20 spellings, ALL giving candidate_bytes=56 where the
|
||||
* original has 60. The spellings tried:
|
||||
* `-(y)` inline; `int ny; ny = y; ny = -ny;`; `ny = y; ny = 0 - ny;`;
|
||||
* `short vy = -y;`; `unsigned short vx = x, vy = -y;`; `int ny = y;` then `-ny`;
|
||||
* `register int t0 asm("$8")` with `t0 = y; t0 = -t0;`;
|
||||
* `register int t0 asm("$8")` with `t0 = y;` then `-t0` inside the pack expression;
|
||||
* `register int t0 asm("$8")` with an explicit `pack` temporary;
|
||||
* `register int y` / `register`-qualified parameters; `(int)(short)y`;
|
||||
* `y * -65536`; `(char *)`/`(int)` casts;
|
||||
* and `((y ^ -1) + 1)`, the ONLY spelling that reaches the correct LENGTH (60) --
|
||||
* it lowers to `nor`+`addiu` (17 differing bytes), which proves the length alone is
|
||||
* not evidence of the right sequence.
|
||||
* And it is not a compiler-revision artifact either: ALL TEN cc1 builds vendored in
|
||||
* `tools/old-gcc/` (2.5.7-psx, 2.6.0-psx, 2.6.3-psx, 2.7.2, 2.7.2-cdk, 2.7.2-psx,
|
||||
* 2.8.0-psx, 2.8.1-psx, 2.91.66-psx, 2.95.2-psx) give 56 on this source. There is no C
|
||||
* spelling and no vendored cc1 that emits a dead copy; the copy is required by the
|
||||
* original's bytes, so it is emitted explicitly.
|
||||
*
|
||||
* (2) THE NEGATION IS THE TRAPPING `sub`, NOT `subu`. The original's word 1 is 0x00084022:
|
||||
* funct 0x22 = `sub`, which traps on overflow. cc1's `-t0` always lowers to `negu`
|
||||
* (0x00084023, funct 0x23 = `subu`). A mnemonic comparison calls both of them `neg`
|
||||
* -- objdump prints `neg` for the original and `negu` for the candidate -- and the
|
||||
* difference is ONE byte in the funct field. This is cookbook finding 6's class
|
||||
* ("count right + residual 1-8 bytes: diff the ENCODING, not the mnemonic") with a
|
||||
* 1-byte residual that no C negation can reach. NOTE this is a local fact about this
|
||||
* row, not an argument that the row is in the blocked `trapping-arithmetic` class:
|
||||
* 0x80010810 is classed `register-tiebreak`, and only the negate here traps.
|
||||
* With (1) alone the row sits at 1 differing byte (0x00084023 vs 0x00084022); with
|
||||
* both, it is byte-exact.
|
||||
*
|
||||
* The two instructions therefore go through one asm statement that emits exactly them:
|
||||
* __asm__ volatile ("move %0,%1\n\tsub %0,$zero,%0" : "=r"(t0) : "r"(y));
|
||||
* Everything else in the body stays in C, and `t0` is a register variable pinned to $8 so
|
||||
* the COP2 operand registers are the original's (the idiom established for 0x800F3E18).
|
||||
* The `nop`s are written at the call site, per the gtemac.h convention for this file.
|
||||
*
|
||||
* LIMITS. The parameter names and the meaning of the packed coordinates are hypotheses
|
||||
* read off the instruction shape; the function's purpose (an nRTPS projection returning
|
||||
* SZ3) is inferred from the macro set that this row itself proved, not independently.
|
||||
* The command field is named by its VALUE, per the gtemac.h convention.
|
||||
*/
|
||||
#include "../include/gtemac.h"
|
||||
|
||||
int func_80010810(int x, int y, int z)
|
||||
{
|
||||
register int t0 __asm__("$8");
|
||||
|
||||
__asm__ volatile ("move %0,%1\n\tsub %0,$zero,%0" : "=r"(t0) : "r"(y));
|
||||
gte_ldVXY0(((unsigned int)t0 << 16) | ((unsigned int)x & 0xFFFF));
|
||||
gte_ldVZ0(z);
|
||||
__asm__ volatile ("nop");
|
||||
__asm__ volatile ("nop");
|
||||
gte_nRTPS();
|
||||
__asm__ volatile ("nop");
|
||||
gte_stSZ3(t0);
|
||||
__asm__ volatile ("nop");
|
||||
return t0;
|
||||
}
|
||||
@@ -0,0 +1,21 @@
|
||||
/*
|
||||
* func_8006AD5C — 84 bytes at 0x8006AD5C..0x8006ADB0 [DRAFT header]
|
||||
*/
|
||||
extern char *D_80121E88;
|
||||
extern unsigned char D_80122374;
|
||||
|
||||
void func_800C2D88(char *);
|
||||
|
||||
void func_8006AD5C(void)
|
||||
{
|
||||
char *p = D_80121E88;
|
||||
|
||||
while (p != 0) {
|
||||
char *n = *(char **)(p + 12);
|
||||
char *next = *(char **)(n + 396);
|
||||
|
||||
if (D_80122374 == 0)
|
||||
func_800C2D88(p);
|
||||
p = next;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,94 @@
|
||||
/*
|
||||
* func_8006D1C4 — 140 bytes at 0x8006D1C4..0x8006D250
|
||||
*
|
||||
* Byte-identical reconstruction. Initialises a 64-byte stack record (two 16-byte
|
||||
* vectors copied from the third argument when it is non-null, plus a flag word
|
||||
* whose bit 0 records whether that argument was non-null) and then calls the
|
||||
* FIRST ARGUMENT as a function pointer, passing the second argument and the
|
||||
* record's address.
|
||||
*
|
||||
* The observed instructions are:
|
||||
* addiu sp,sp,-88
|
||||
* move t0,a0 ; fn — must survive the call's a0 setup
|
||||
* move a3,a1 ; the value passed on as the call's first argument
|
||||
* sw ra,80(sp)
|
||||
* beqz a2,common ; if (src == 0) skip the copies
|
||||
* sw zero,20(sp) ; (delay slot) s.f1 = 0
|
||||
* lw v0,0(a2) / lw v1,4(a2) / lw a0,8(a2) / lw a1,12(a2)
|
||||
* sw v0,24(sp) / sw v1,28(sp) / sw a0,32(sp) / sw a1,36(sp) ; s.a
|
||||
* lw v0,16(a2) / lw v1,20(a2) / lw a0,24(a2) / lw a1,28(a2)
|
||||
* sw v0,40(sp) / sw v1,44(sp) / sw a0,48(sp) / sw a1,52(sp) ; s.b
|
||||
* common:
|
||||
* li v1,-2 ; mask for bit 0
|
||||
* move a0,a3 ; the call's first argument
|
||||
* lw v0,56(sp) ; s.flag
|
||||
* addiu a1,sp,16 ; the record's address
|
||||
* and v0,v0,v1
|
||||
* sltu v1,zero,a2 ; (src != 0)
|
||||
* or v0,v0,v1
|
||||
* jalr t0 ; fn(second argument, &s)
|
||||
* sw v0,56(sp) ; (delay slot) s.flag = ...
|
||||
* epilogue:
|
||||
* lw ra,80(sp) / addiu sp,sp,88 / jr ra / nop
|
||||
* (the return value is whatever the indirect call left in v0)
|
||||
*
|
||||
* FOUR things the bytes fix:
|
||||
*
|
||||
* 1. `s.f1 = 0` IS AN UNCONDITIONAL INITIALISATION BEFORE THE `if`, NOT A
|
||||
* STATEMENT INSIDE IT. Written inside the `if (src != 0)` block the row comes
|
||||
* out at the correct 140 bytes with **13 differing bytes**: the store loses
|
||||
* the `beqz` delay slot to the `move a3,a1` parameter copy, which then sits in
|
||||
* the fall-through block, and the whole entry-block order changes
|
||||
* (`sw ra` moves first). Moving the initialisation before the branch gives it
|
||||
* the delay slot and keeps BOTH parameter copies in the entry block — the
|
||||
* original exactly. The delay slot is what makes an otherwise-invisible
|
||||
* source-order choice load-bearing (cookbook 149/163/184).
|
||||
* 2. THE TWO VECTOR COPIES ARE 16-BYTE STRUCT ASSIGNMENTS. Each is 4 batched
|
||||
* loads then 4 batched stores (cookbook 96/102), so the source copies two
|
||||
* `int[4]` structs, not eight element stores; element stores would serialise
|
||||
* behind load-delay nops.
|
||||
* 3. THE FIRST ARGUMENT IS CALLED INDIRECTLY AND MUST BE COPIED OUT OF `a0`
|
||||
* FIRST, because the call's own argument setup overwrites `a0` with the
|
||||
* second argument (`move a0,a3`). Likewise the second argument is copied out
|
||||
* of `a1` because `a1` receives `&s`. Both copies land in caller-saved
|
||||
* registers (`t0`, `a3`) — no saved register is needed, only `ra` is stacked.
|
||||
* 4. THE RECORD IS 64 BYTES AND ONLY +4, +8..+39 AND +40 ARE TOUCHED. The frame
|
||||
* arithmetic fixes the size: 88 = 16 (outgoing argument area) + 64 (the
|
||||
* record at sp+16..sp+79) + 4 (`ra` at 80) + 4 padding. The bytes at +0
|
||||
* (sp+16) and +44..+63 are never written here, so the trailing `unknown[5]` is
|
||||
* a declaration sized to reproduce the frame, not a recovered field: the
|
||||
* address of the record escapes into the callee, which is what makes the
|
||||
* whole size load-bearing (cookbook 59/131).
|
||||
*
|
||||
* LIMITS: every displacement, the frame and the mask are read from the bytes.
|
||||
* Field names and types are hypotheses; in particular `f0` at +0 and the
|
||||
* trailing fields are unnamed because nothing writes them, and the flag word's
|
||||
* meaning is only "bit 0 = (src != 0), and the low bit is preserved otherwise".
|
||||
* The indirect call's signature is inferred from the argument registers that are
|
||||
* set, and its return value is whatever the callee leaves in v0.
|
||||
*/
|
||||
|
||||
struct func_8006D1C4_V {
|
||||
int v[4];
|
||||
};
|
||||
|
||||
struct func_8006D1C4_S {
|
||||
int f0; /* +0 — never written here */
|
||||
int f1; /* +4 — zeroed unconditionally */
|
||||
struct func_8006D1C4_V a; /* +8 — copied from src[0..3] */
|
||||
struct func_8006D1C4_V b; /* +24 — copied from src[4..7] */
|
||||
int flag; /* +40 — bit 0 = (src != 0), low bit cleared */
|
||||
int unknown[5]; /* +44..+63 — sized to reproduce the frame */
|
||||
};
|
||||
|
||||
int func_8006D1C4(int (*fn)(), int a1, int *src) {
|
||||
struct func_8006D1C4_S s;
|
||||
|
||||
s.f1 = 0;
|
||||
if (src != 0) {
|
||||
s.a = *(struct func_8006D1C4_V *)src;
|
||||
s.b = *(struct func_8006D1C4_V *)(src + 4);
|
||||
}
|
||||
s.flag = (s.flag & -2) | (src != 0);
|
||||
return fn(a1, &s);
|
||||
}
|
||||
@@ -0,0 +1,95 @@
|
||||
/*
|
||||
* func_8009C69C — 60 bytes at 0x8009C69C..0x8009C6D8
|
||||
*
|
||||
* PHASE 12 WORKER C. A GTE (COP2) row, and one that needed the INLINE-ASM ESCAPE HATCH.
|
||||
* The hatch was taken with the developer's authorisation and the proof is below.
|
||||
*
|
||||
* WHAT THE BODY DOES, read off the instruction stream (only the bytes are evidence):
|
||||
* move t0,a1 t0 = y
|
||||
* sub t0,zero,t0 t0 = -t0 <- FUNCT 0x22, the TRAPPING `sub`
|
||||
* sll t0,t0,0x10 t0 <<= 16
|
||||
* andi a0,a0,0xffff x &= 0xffff
|
||||
* or t0,t0,a0 t0 = (VY << 16) | VX
|
||||
* mtc2 t0,$0 VXY0
|
||||
* mtc2 a2,$1 VZ0
|
||||
* nop
|
||||
* nop
|
||||
* cop2 0x180001 nRTPS
|
||||
* nop
|
||||
* mfc2 t0,$19 t0 = SZ3
|
||||
* nop
|
||||
* jr ra
|
||||
* move v0,t0
|
||||
*
|
||||
* `include/gtemac.h` already records VXY0/VZ0/SZ3 and `gte_nRTPS()` as PROVED BY THIS ROW
|
||||
* (its comments cite 0x80010810 by address). The twin 0x80010810 is byte-identical (raw
|
||||
* words d=0/15, see src/func_80010810.c): the two 15-word bodies differ in ZERO words, so this
|
||||
* file is the twin's body with the symbol renamed, and the whole proof below applies
|
||||
* unchanged. Verified independently at its own address.
|
||||
*
|
||||
* --- WHY INLINE ASM: the two instructions plain C cannot emit (the proof) ---
|
||||
*
|
||||
* The residual is TWO facts, and both had to be established separately.
|
||||
*
|
||||
* (1) A DEAD ARGUMENT COPY. The original's pack begins with `move t0,a1` and only then
|
||||
* `neg t0,t0`. Writing the negation in C -- `t0 = y; t0 = -t0;`, or `-(y)` directly, or
|
||||
* any of the named-temporary / cast / two's-complement-manual spellings -- makes cc1's
|
||||
* `combine` pass fuse the copy into the negation and emit the ONE-instruction
|
||||
* `subu t0,zero,a1`. Measured: 20 spellings, ALL giving candidate_bytes=56 where the
|
||||
* original has 60. The spellings tried:
|
||||
* `-(y)` inline; `int ny; ny = y; ny = -ny;`; `ny = y; ny = 0 - ny;`;
|
||||
* `short vy = -y;`; `unsigned short vx = x, vy = -y;`; `int ny = y;` then `-ny`;
|
||||
* `register int t0 asm("$8")` with `t0 = y; t0 = -t0;`;
|
||||
* `register int t0 asm("$8")` with `t0 = y;` then `-t0` inside the pack expression;
|
||||
* `register int t0 asm("$8")` with an explicit `pack` temporary;
|
||||
* `register int y` / `register`-qualified parameters; `(int)(short)y`;
|
||||
* `y * -65536`; `(char *)`/`(int)` casts;
|
||||
* and `((y ^ -1) + 1)`, the ONLY spelling that reaches the correct LENGTH (60) --
|
||||
* it lowers to `nor`+`addiu` (17 differing bytes), which proves the length alone is
|
||||
* not evidence of the right sequence.
|
||||
* And it is not a compiler-revision artifact either: ALL TEN cc1 builds vendored in
|
||||
* `tools/old-gcc/` (2.5.7-psx, 2.6.0-psx, 2.6.3-psx, 2.7.2, 2.7.2-cdk, 2.7.2-psx,
|
||||
* 2.8.0-psx, 2.8.1-psx, 2.91.66-psx, 2.95.2-psx) give 56 on this source. There is no C
|
||||
* spelling and no vendored cc1 that emits a dead copy; the copy is required by the
|
||||
* original's bytes, so it is emitted explicitly.
|
||||
*
|
||||
* (2) THE NEGATION IS THE TRAPPING `sub`, NOT `subu`. The original's word 1 is 0x00084022:
|
||||
* funct 0x22 = `sub`, which traps on overflow. cc1's `-t0` always lowers to `negu`
|
||||
* (0x00084023, funct 0x23 = `subu`). A mnemonic comparison calls both of them `neg`
|
||||
* -- objdump prints `neg` for the original and `negu` for the candidate -- and the
|
||||
* difference is ONE byte in the funct field. This is cookbook finding 6's class
|
||||
* ("count right + residual 1-8 bytes: diff the ENCODING, not the mnemonic") with a
|
||||
* 1-byte residual that no C negation can reach. NOTE this is a local fact about this
|
||||
* row, not an argument that the row is in the blocked `trapping-arithmetic` class:
|
||||
* 0x80010810 is classed `register-tiebreak`, and only the negate here traps.
|
||||
* With (1) alone the row sits at 1 differing byte (0x00084023 vs 0x00084022); with
|
||||
* both, it is byte-exact.
|
||||
*
|
||||
* The two instructions therefore go through one asm statement that emits exactly them:
|
||||
* __asm__ volatile ("move %0,%1\n\tsub %0,$zero,%0" : "=r"(t0) : "r"(y));
|
||||
* Everything else in the body stays in C, and `t0` is a register variable pinned to $8 so
|
||||
* the COP2 operand registers are the original's (the idiom established for 0x800F3E18).
|
||||
* The `nop`s are written at the call site, per the gtemac.h convention for this file.
|
||||
*
|
||||
* LIMITS. The parameter names and the meaning of the packed coordinates are hypotheses
|
||||
* read off the instruction shape; the function's purpose (an nRTPS projection returning
|
||||
* SZ3) is inferred from the macro set that this row itself proved, not independently.
|
||||
* The command field is named by its VALUE, per the gtemac.h convention.
|
||||
*/
|
||||
#include "../include/gtemac.h"
|
||||
|
||||
int func_8009C69C(int x, int y, int z)
|
||||
{
|
||||
register int t0 __asm__("$8");
|
||||
|
||||
__asm__ volatile ("move %0,%1\n\tsub %0,$zero,%0" : "=r"(t0) : "r"(y));
|
||||
gte_ldVXY0(((unsigned int)t0 << 16) | ((unsigned int)x & 0xFFFF));
|
||||
gte_ldVZ0(z);
|
||||
__asm__ volatile ("nop");
|
||||
__asm__ volatile ("nop");
|
||||
gte_nRTPS();
|
||||
__asm__ volatile ("nop");
|
||||
gte_stSZ3(t0);
|
||||
__asm__ volatile ("nop");
|
||||
return t0;
|
||||
}
|
||||
Reference in New Issue
Block a user