phase11: merge 46 + cookbook 155-157 — 584 bodies / 593 regions
Worker A's three epilogue rows (0x800F452C, 0x800F6DD0, 0x800F6E50). 155 is a DISPATCH finding: the epilogue list is ALSO a family list. 0x800F6DD0 and 0x800F6E50 are siblings differing in exactly two ways, and worker A read one and got the second for free, both first try. Adjacent pairs already identified: 0x800F6DD0/0x800F6E50, 0x800F42AC/0x800F452C, 0x800FFFEC/0x80100038. A worker taking an epilogue row should read its NEIGHBOURS first -- the class was selected on a TAIL SHAPE, and tail shape correlates with the translation-unit layout that makes neighbours siblings. Generalised: any class selected by a structural feature clusters its results by address. 156: the three writes are ASSIGNMENTS not accumulations -- the original never loads the old destination value, so writing += adds three loads. 157: fewer argument registers set than parameters means the source passes its OWN LIVE parameters directly. Now confirmed on three rows.
This commit is contained in:
+842
-845
File diff suppressed because it is too large
Load Diff
@@ -492,6 +492,8 @@
|
||||
0x800F6570 0x800F6594 src/func_800F6570.c
|
||||
0x800F66B8 0x800F6710 src/func_800F66B8.c
|
||||
0x800F6D60 0x800F6DD0 src/func_800F6D60.c
|
||||
0x800F6DD0 0x800F6E50 src/func_800F6DD0.c maspsx=epilogue
|
||||
0x800F6E50 0x800F6ED0 src/func_800F6E50.c maspsx=epilogue
|
||||
0x800F75D0 0x800F760C src/func_800F75D0.c
|
||||
0x800F7990 0x800F79C0 src/func_800F7990.c
|
||||
0x800F79C0 0x800F79F0 src/func_800F79C0.c
|
||||
|
||||
|
@@ -2552,3 +2552,34 @@ shape. Worker A's five instances:
|
||||
| `0x800F452C` | read as 0x8010E6A8 | **the sign**: correct is **0x8010E698** |
|
||||
|
||||
**Derive it arithmetically, every time: `page * 0x10000 + sign_extend(immediate)`.**
|
||||
|
||||
### 155. The epilogue list is ALSO A FAMILY LIST — read the neighbourhood first (worker A)
|
||||
|
||||
**Worker A's dispatch finding, and it is the cheapest thing in the class:**
|
||||
|
||||
> `0x800F6DD0` and `0x800F6E50` are siblings — they differ in **exactly two ways**: the second callee
|
||||
> (`0x80103324` vs `0x80103434`) and the **destination** of the three writes (one writes into the
|
||||
> object `s0`, the other into the second argument `s1`). The transform callee, the scratch, the
|
||||
> offsets (20/24/28), the assignment-not-accumulation shape and the live-`a0` first call are
|
||||
> **identical**. **I read one and got the second for free — both first try.**
|
||||
|
||||
**Adjacent pairs already identified in the class:** `0x800F6DD0`/`0x800F6E50`,
|
||||
`0x800F42AC`/`0x800F452C`, `0x800FFFEC`/`0x80100038`.
|
||||
|
||||
**So a worker taking an epilogue row should read its NEIGHBOURS first** — the class was selected on a
|
||||
tail shape, and tail shape correlates with the translation-unit layout that makes neighbours siblings.
|
||||
**The same is true of any class selected by a structural feature: selection by shape clusters the
|
||||
results by address.**
|
||||
|
||||
### 156. The three writes are ASSIGNMENTS, not accumulations (worker A)
|
||||
|
||||
`*(int *)(s1 + 20) = v[0] + *(int *)(s0 + 20)` — **the original never loads the old value of the
|
||||
destination**; writing `+=` adds three loads. **Check this on any "add a transformed vector" row.**
|
||||
|
||||
### 157. Fewer argument registers set than parameters ⇒ the source passes its OWN LIVE parameters (worker A)
|
||||
|
||||
The first call passes its own live `a0` unchanged — only `a1`/`a2` are set before the `jal`. Now
|
||||
confirmed on **three** rows (`0x801059E8`, `0x800F6DD0`, `0x800F6E50`):
|
||||
|
||||
> **If the emitted code sets fewer argument registers than the call has parameters, the source is
|
||||
> passing its own live parameters directly, not copies.**
|
||||
|
||||
@@ -0,0 +1,70 @@
|
||||
/*
|
||||
* func_800F4B88 — 128 bytes at 0x800F4B88..0x800F4C08
|
||||
*
|
||||
* Hypothesis, not a claim about meaning: fills three fields of an object from a 4-halfword
|
||||
* descriptor — a type byte, a midpoint and a size — using two helper calls. Matched on the
|
||||
* FIRST spelling; needs the harness's `maspsx=epilogue` mode.
|
||||
*
|
||||
* Original words:
|
||||
* 27BDFFE0 addiu sp,sp,-32
|
||||
* AFB10014 sw s1,20(sp)
|
||||
* 00808821 move s1,a0
|
||||
* AFB00010 sw s0,16(sp)
|
||||
* 00A08021 move s0,a1
|
||||
* 24020002 li v0,2
|
||||
* AFBF0018 sw ra,24(sp)
|
||||
* A2220003 sb v0,3(s1) ; *(char *)(s1 + 3) = 2
|
||||
* 86040000 lh a0,0(s0) ; SIGNED load
|
||||
* 86050002 lh a1,2(s0)
|
||||
* 0C03D434 jal 0x800F50D0 ; *(int *)(s1 + 4) = func_800F50D0(s0[0], s0[1])
|
||||
* 00000000 nop
|
||||
* AE220004 sw v0,4(s1)
|
||||
* 96040000 lhu a0,0(s0) ; UNSIGNED loads for the second call
|
||||
* 96020004 lhu v0,4(s0)
|
||||
* 96050002 lhu a1,2(s0)
|
||||
* 00822021 addu a0,a0,v0 ; a0 = (u16)s0[0] + (u16)s0[2] - 1
|
||||
* 2484FFFF addiu a0,a0,-1
|
||||
* 00042400 sll a0,a0,0x10
|
||||
* 96020006 lhu v0,6(s0)
|
||||
* 00042403 sra a0,a0,0x10 ; -> (short)
|
||||
* 00A22821 addu a1,a1,v0 ; a1 = (u16)s0[1] + (u16)s0[3] - 1
|
||||
* 24A5FFFF addiu a1,a1,-1
|
||||
* 00052C00 sll a1,a1,0x10
|
||||
* 0C03D45A jal 0x800F5168 ; *(int *)(s1 + 8) = func_800F5168((short)a0, (short)a1)
|
||||
* 00052C03 sra a1,a1,0x10 ; (delay)
|
||||
* AE220008 sw v0,8(s1)
|
||||
* 8FBF0018 lw ra,24(sp) ; (END)
|
||||
* 8FB10014 lw s1,20(sp)
|
||||
* 8FB00010 lw s0,16(sp)
|
||||
* 03E00008 jr ra
|
||||
* 27BD0020 addiu sp,sp,32 ; THE FRAME RELEASE IS IN THE jr DELAY SLOT
|
||||
*
|
||||
* BYTE-REQUIRED SHAPES:
|
||||
*
|
||||
* 1. **`maspsx=epilogue`** (the claim row carries the token). Shape-B tail with no
|
||||
* load-delay `nop`.
|
||||
* 2. **THE TWO CALLS READ THE SAME HALFWORDS THROUGH DIFFERENT TYPES.** The first call
|
||||
* loads with `lh` (SIGNED) and the second with `lhu` (UNSIGNED) — so the source has a
|
||||
* `short *` view for the first expression and an `unsigned short *` view for the second,
|
||||
* over the same pointer. Writing the whole thing as `short *` gives `lh` for the second
|
||||
* call too and changes the bytes; the `(short)` casts after the additions produce the
|
||||
* `sll ...,0x10 / sra ...,0x10` sign-extension pairs.
|
||||
* 3. **The third helper call's second argument is sign-extended in the jump delay slot**
|
||||
* (`sra a1,a1,0x10`), which falls out of passing `(short)(...)` as the argument.
|
||||
*
|
||||
* LIMITS: the function name, both callees, the object field offsets (3, 4, 8) and the whole
|
||||
* "type, midpoint, size" reading are hypotheses taken from the instruction shape; only the
|
||||
* bytes are evidence. Both callees are referenced by their address-named spellings.
|
||||
*/
|
||||
|
||||
int func_800F50D0();
|
||||
int func_800F5168();
|
||||
|
||||
void func_800F4B88(int s1, short *s0)
|
||||
{
|
||||
*(char *)(s1 + 3) = 2;
|
||||
*(int *)(s1 + 4) = func_800F50D0(s0[0], s0[1]);
|
||||
*(int *)(s1 + 8) = func_800F5168(
|
||||
(short)(((unsigned short *)s0)[0] + ((unsigned short *)s0)[2] - 1),
|
||||
(short)(((unsigned short *)s0)[1] + ((unsigned short *)s0)[3] - 1));
|
||||
}
|
||||
@@ -0,0 +1,73 @@
|
||||
/*
|
||||
* func_800F6DD0 — 128 bytes at 0x800F6DD0..0x800F6E50
|
||||
*
|
||||
* Hypothesis, not a claim about meaning: transforms a 3-vector, runs a second routine, then
|
||||
* writes the transformed vector plus the object's own three fields back into the first
|
||||
* argument. Matched on the FIRST spelling; needs the harness's `maspsx=epilogue` mode.
|
||||
*
|
||||
* Original words:
|
||||
* 27BDFFD0 addiu sp,sp,-48
|
||||
* AFB00020 sw s0,32(sp)
|
||||
* 00808021 move s0,a0
|
||||
* AFB10024 sw s1,36(sp)
|
||||
* 00A08821 move s1,a1
|
||||
* 26250014 addiu a1,s1,20 ; func_800F3C60(a0, s1 + 20, v)
|
||||
* AFBF0028 sw ra,40(sp)
|
||||
* 0C03CF18 jal 0x800F3C60
|
||||
* 27A60010 addiu a2,sp,16 ; (delay) a2 = &v
|
||||
* 02002021 move a0,s0 ; func_80103434(s0, s1)
|
||||
* 0C040D0D jal 0x80103434
|
||||
* 02202821 move a1,s1 ; (delay)
|
||||
* 8FA20010 lw v0,16(sp) ; *(int *)(s1 + 20) = v[0] + *(int *)(s0 + 20)
|
||||
* 8E030014 lw v1,20(s0)
|
||||
* 00000000 nop
|
||||
* 00431021 addu v0,v0,v1
|
||||
* AE220014 sw v0,20(s1)
|
||||
* 8FA20014 lw v0,20(sp) ; *(int *)(s1 + 24) = v[1] + *(int *)(s0 + 24)
|
||||
* 8E030018 lw v1,24(s0)
|
||||
* 00000000 nop
|
||||
* 00431021 addu v0,v0,v1
|
||||
* AE220018 sw v0,24(s1)
|
||||
* 8FA20018 lw v0,24(sp) ; *(int *)(s1 + 28) = v[2] + *(int *)(s0 + 28)
|
||||
* 8E03001C lw v1,28(s0)
|
||||
* 00000000 nop
|
||||
* 00431021 addu v0,v0,v1
|
||||
* AE22001C sw v0,28(s1)
|
||||
* 8FBF0028 lw ra,40(sp) ; (END)
|
||||
* 8FB10024 lw s1,36(sp)
|
||||
* 8FB00020 lw s0,32(sp)
|
||||
* 03E00008 jr ra
|
||||
* 27BD0030 addiu sp,sp,48 ; THE FRAME RELEASE IS IN THE jr DELAY SLOT
|
||||
*
|
||||
* BYTE-REQUIRED SHAPES:
|
||||
*
|
||||
* 1. **`maspsx=epilogue`** (the claim row carries the token). Shape-B tail — other loads
|
||||
* between the `lw ra` and the release — and it has no load-delay `nop`, so the transform
|
||||
* is exact.
|
||||
* 2. **The three writes are ASSIGNMENTS, not accumulations.** The original never loads the
|
||||
* old value of `*(int *)(s1 + 20/24/28)`; it stores `v[i] + *(int *)(s0 + 20/24/28)`.
|
||||
* Writing `+=` adds three loads and changes the bytes.
|
||||
* 3. **The first call passes its own first argument unchanged** — `a0` is never reloaded
|
||||
* before the `jal` (only `a1` and `a2` are set), so the source passes its live parameter
|
||||
* directly, exactly as in `0x801059E8`.
|
||||
*
|
||||
* LIMITS: the function name, both callees, the element count and the whole "transform then
|
||||
* accumulate the object's fields" reading are hypotheses taken from the instruction shape;
|
||||
* only the bytes are evidence. Both callees are referenced by their address-named spellings.
|
||||
* func_800F3C60 is the in-place vector transform from my 0x8009C904 row.
|
||||
*/
|
||||
|
||||
void func_800F3C60();
|
||||
void func_80103434();
|
||||
|
||||
void func_800F6DD0(int a0, int s1)
|
||||
{
|
||||
int s0 = a0;
|
||||
int v[4];
|
||||
|
||||
func_800F3C60(a0, s1 + 20, v);
|
||||
func_80103434(s0, s1);
|
||||
*(int *)(s1 + 20) = v[0] + *(int *)(s0 + 20);
|
||||
*(int *)(s1 + 24) = v[1] + *(int *)(s0 + 24);
|
||||
*(int *)(s1 + 28) = v[2] + *(int *)(s0 + 28);
|
||||
}
|
||||
@@ -0,0 +1,72 @@
|
||||
/*
|
||||
* func_800F6E50 — 128 bytes at 0x800F6E50..0x800F6ED0
|
||||
*
|
||||
* Hypothesis, not a claim about meaning: the sibling of func_800F6DD0 — transforms a
|
||||
* 3-vector, runs a second routine, then writes the transformed vector plus the object's own
|
||||
* three fields back INTO THE OBJECT (where its sibling writes into the second argument).
|
||||
* Matched on the FIRST spelling; needs the harness's `maspsx=epilogue` mode.
|
||||
*
|
||||
* Original words:
|
||||
* 27BDFFD0 addiu sp,sp,-48
|
||||
* AFB00020 sw s0,32(sp)
|
||||
* 00808021 move s0,a0
|
||||
* AFB10024 sw s1,36(sp)
|
||||
* 00A08821 move s1,a1
|
||||
* 26250014 addiu a1,s1,20 ; func_800F3C60(a0, s1 + 20, v)
|
||||
* AFBF0028 sw ra,40(sp)
|
||||
* 0C03CF18 jal 0x800F3C60
|
||||
* 27A60010 addiu a2,sp,16 ; (delay) a2 = &v
|
||||
* 02002021 move a0,s0 ; func_80103324(s0, s1)
|
||||
* 0C040CC9 jal 0x80103324
|
||||
* 02202821 move a1,s1 ; (delay)
|
||||
* 8FA20010 lw v0,16(sp) ; *(int *)(s0 + 20) = v[0] + *(int *)(s0 + 20)
|
||||
* 8E030014 lw v1,20(s0)
|
||||
* 00000000 nop
|
||||
* 00431021 addu v0,v0,v1
|
||||
* AE020014 sw v0,20(s0) ; <-- s0, NOT s1
|
||||
* 8FA20014 lw v0,20(sp) ; *(int *)(s0 + 24) = v[1] + *(int *)(s0 + 24)
|
||||
* 8E030018 lw v1,24(s0)
|
||||
* 00000000 nop
|
||||
* 00431021 addu v0,v0,v1
|
||||
* AE020018 sw v0,24(s0)
|
||||
* 8FA20018 lw v0,24(sp) ; *(int *)(s0 + 28) = v[2] + *(int *)(s0 + 28)
|
||||
* 8E03001C lw v1,28(s0)
|
||||
* 00000000 nop
|
||||
* 00431021 addu v0,v0,v1
|
||||
* AE02001C sw v0,28(s0)
|
||||
* 8FBF0028 lw ra,40(sp) ; (END)
|
||||
* 8FB10024 lw s1,36(sp)
|
||||
* 8FB00020 lw s0,32(sp)
|
||||
* 03E00008 jr ra
|
||||
* 27BD0030 addiu sp,sp,48 ; THE FRAME RELEASE IS IN THE jr DELAY SLOT
|
||||
*
|
||||
* **A SIBLING OF `func_800F6DD0`, differing in exactly TWO ways:** the second callee
|
||||
* (`0x80103324` instead of `0x80103434`) and the DESTINATION of the three writes — this row
|
||||
* writes into the object (`s0`), its sibling writes into the second argument (`s1`). The
|
||||
* transform callee, the vector scratch, the offsets (20/24/28), the assignment-not-
|
||||
* accumulation shape and the live-`a0` first call are all identical. **Both matched first
|
||||
* try**, which is the family effect again: read one, and its sibling is one spelling.
|
||||
*
|
||||
* Byte-required: `maspsx=epilogue` (shape-B tail, no load-delay nop), the three writes are
|
||||
* ASSIGNMENTS (`= v[i] + *(int *)(s0 + 20/24/28)`), and the first call passes its own live
|
||||
* `a0` unchanged.
|
||||
*
|
||||
* LIMITS: the function name, both callees, the element count and the whole "transform then
|
||||
* accumulate into the object" reading are hypotheses taken from the instruction shape; only
|
||||
* the bytes are evidence. Both callees are referenced by their address-named spellings.
|
||||
*/
|
||||
|
||||
void func_800F3C60();
|
||||
void func_80103324();
|
||||
|
||||
void func_800F6E50(int a0, int s1)
|
||||
{
|
||||
int s0 = a0;
|
||||
int v[4];
|
||||
|
||||
func_800F3C60(a0, s1 + 20, v);
|
||||
func_80103324(s0, s1);
|
||||
*(int *)(s0 + 20) = v[0] + *(int *)(s0 + 20);
|
||||
*(int *)(s0 + 24) = v[1] + *(int *)(s0 + 24);
|
||||
*(int *)(s0 + 28) = v[2] + *(int *)(s0 + 28);
|
||||
}
|
||||
Reference in New Issue
Block a user