phase11: merge 30 + cookbook 120-122 — 550 bodies / 559 regions
Worker A's 0x800689DC and worker C's 0x8009F4B4. 120: worker B's justification for why the ranker works -- 'the allocator makes copies non-identical, so OPCODE repetition survives while WORD repetition does not'. That is exactly why finding 110's full-word metric failed and why the opcode metric works. A repeated source block produces the same opcodes with different registers; requiring operands to match destroys the signal rather than sharpening it. 121: the filled-delay-slot class has TWO sub-cases with DIFFERENT fixes -- reorg fills the slot (source-shape hunt) versus maspsx mode changing WHICH instruction lands in the slot (a harness token choice). Same diagnostic, different remedy. Check whether toggling maspsx changes the fill before hunting a source shape. 122: NEW BLOCKED CLASS -- a GTE coprocessor body needs a harness token, not more spellings. Worker B's 0x8001FAFC reads mfc2 $12/$13/$14 and branches on t7/s6 which are NOT the o32 argument registers, so the inputs arrive through a non-standard convention. Team rule: if a body contains mfc2/mtc2, do not spend spellings on it -- these are tooling-blocked rows to be worked as a batch once a token exists.
This commit is contained in:
+1021
-1024
File diff suppressed because it is too large
Load Diff
@@ -251,6 +251,7 @@
|
||||
0x800683B0 0x800683E4 src/func_800683B0.c
|
||||
0x80068470 0x8006848C src/func_80068470.c
|
||||
0x80068910 0x800689DC src/func_80068910.c
|
||||
0x800689DC 0x80068A64 src/func_800689DC.c
|
||||
0x80068D54 0x80068D78 src/func_80068D54.c
|
||||
0x80068F40 0x80068F6C src/func_80068F40.c
|
||||
0x80068F6C 0x80068F98 src/func_80068F6C.c
|
||||
@@ -378,6 +379,7 @@
|
||||
0x8009E8D0 0x8009E95C src/func_8009E8D0.c
|
||||
0x8009F0E8 0x8009F120 src/func_8009F0E8.c
|
||||
0x8009F3A8 0x8009F4B4 src/func_8009F3A8.c
|
||||
0x8009F4B4 0x8009F5AC src/func_8009F4B4.c
|
||||
0x8009F5AC 0x8009F6A0 src/func_8009F5AC.c
|
||||
0x8009F6A0 0x8009F798 src/func_8009F6A0.c
|
||||
0x8009F798 0x8009F890 src/func_8009F798.c
|
||||
@@ -441,6 +443,7 @@
|
||||
0x800B6BDC 0x800B6C14 src/func_800B6BDC.c
|
||||
0x800B6C14 0x800B6C60 src/func_800B6C14.c
|
||||
0x800B6C60 0x800B6CB4 src/func_800B6C60.c
|
||||
0x800B7190 0x800B7230 src/func_800B7190.c
|
||||
0x800B7230 0x800B7264 src/func_800B7230.c
|
||||
0x800B74D0 0x800B7524 src/func_800B74D0.c
|
||||
0x800BBAC8 0x800BBB10 src/func_800BBAC8.c
|
||||
|
||||
|
@@ -1947,3 +1947,47 @@ argument 5).
|
||||
only the worker reading the row can find them."*** This is the third independent confirmation of
|
||||
the same rule — worker B found the read side, worker D found the write side, and both had to read
|
||||
the row to do it. **A tool that fires on a shape should document what the shape is allowed to be.**
|
||||
|
||||
### 120. The ranker's JUSTIFICATION, from the metric that failed (worker B)
|
||||
|
||||
Worker B's explanation of *why* opcode redundancy predicts matchability while full-word redundancy
|
||||
does not — and it is the theoretical basis for the whole ranker:
|
||||
|
||||
> **The allocator makes copies non-identical, so OPCODE repetition survives while WORD repetition
|
||||
> does not.**
|
||||
|
||||
That is precisely why finding 110's full-word metric failed (two known matches scored 0.00) and why
|
||||
the opcode metric works. **A repeated source block produces the same opcodes with different
|
||||
registers; that is the signal. Requiring the operands to match destroys the signal rather than
|
||||
sharpening it.** State this whenever someone proposes a "better" redundancy metric.
|
||||
|
||||
### 121. The filled-delay-slot class has TWO sub-cases (worker B)
|
||||
|
||||
Finding 117 described the signature; worker B split it, and the split determines the fix:
|
||||
|
||||
| sub-case | what happens | fix |
|
||||
|---|---|---|
|
||||
| **reorg fills the slot** (worker C, `0x80103544`) | the original's `j`/`jr` slot holds the preceding store or the frame release; the candidate leaves a `nop` | a **source-shape** hunt |
|
||||
| **maspsx mode changes WHICH instruction lands in the slot** (worker B, `0x800B1F94`, `0x800F2F08`) | `maspsx` on vs off moves a different instruction into the slot | a **harness token** choice |
|
||||
|
||||
**Both present as "correct control flow, +N instructions", so the diagnostic is identical — but the
|
||||
fix is not.** Before hunting a source shape, check whether toggling `maspsx` changes the fill; if it
|
||||
does, the row is a mode choice and no amount of spelling will find it.
|
||||
|
||||
### 122. NEW BLOCKED CLASS: a GTE coprocessor body needs a HARNESS TOKEN, not more spellings (worker B)
|
||||
|
||||
Worker B's `0x8001FAFC` (148 B) reads the GTE data registers directly — `mfc2 t0,$12; mfc2 t1,$13;
|
||||
mfc2 t2,$14` (IR1/IR2/IR3) — builds a two-stage sign-disagreement flag, and branches on **`t7` and
|
||||
`s6`, which are NOT the o32 argument registers.** So the inputs arrive through a non-standard
|
||||
convention: this is a hand-scheduled or macro-expanded GTE sequence whose `mfc2`/mask schedule is
|
||||
part of the **GTE ABI contract**, not something plain C expresses.
|
||||
|
||||
**Worker B's conclusion is the right disposition: the prerequisite is a harness token for the GTE
|
||||
macro idiom, not more spellings.** Recorded as **blocked**, not as a near-match.
|
||||
|
||||
**Team-wide rule: if a body contains `mfc2`/`mtc2`, do not spend spellings on it.** Worker A's
|
||||
`0x800FB1C8` supplies the working idiom (`#include "../include/gtemac.h"`, `mfc2` destinations bound
|
||||
with `register int __asm__("$8")`, the command written as `c2 0x486012`), and worker A has two
|
||||
further GTE rows deferred with complete derivations (`0x800F9BC8`, `0x800F3C60`). **These are
|
||||
tooling-blocked rows and they should be worked as a batch once a token exists**, not one at a time
|
||||
by whoever happens to hit one.
|
||||
|
||||
@@ -0,0 +1,48 @@
|
||||
/*
|
||||
* func_80018210 — 84 bytes at 0x80018210..0x80018264
|
||||
*
|
||||
* Goal B, Phase 11. **First attempt.** Zero branches, one callee called twice, straight-line.
|
||||
*
|
||||
* int v = 0;
|
||||
* func_800180D8(&v, (int *)(a0 + 12), 4, 1);
|
||||
* func_800180D8(&v, (int *)(a0 + 28), 16, 0);
|
||||
* return v;
|
||||
*
|
||||
* The observed instructions are:
|
||||
* addiu sp,sp,-0x20 frame, 32 bytes
|
||||
* sw s0,0x18(sp) / move s0,a0
|
||||
* addiu a0,sp,0x10 a0 = &v
|
||||
* addiu a1,s0,0xc
|
||||
* li a2,4
|
||||
* li a3,1
|
||||
* sw ra,0x1c(sp)
|
||||
* jal 0x800180D8
|
||||
* sw zero,0x10(sp) v = 0 (delay slot — the initialisation sinks into it)
|
||||
* addiu a0,sp,0x10
|
||||
* addiu a1,s0,0x1c
|
||||
* li a2,16
|
||||
* jal 0x800180D8
|
||||
* move a3,zero (delay slot)
|
||||
* lw v0,0x10(sp) return v
|
||||
* lw ra / lw s0 / addiu sp,sp,0x20 / jr ra / nop
|
||||
*
|
||||
* The initialisation is written BEFORE the first call in the source and cc1 schedules the store into
|
||||
* that call's delay slot; the second call's delay slot holds the zeroed fourth argument. Only `s0`
|
||||
* is saved, so `v`'s address never needs a callee-saved register.
|
||||
*
|
||||
* LIMITS: the function name, the callee, the meaning of the 12/28 offsets and of the 4/1 and 16/0
|
||||
* argument pairs are hypotheses reconstructed from the disassembly; only the compiled bytes are
|
||||
* evidence. Whether a0 is a struct pointer or an integer address is not recoverable.
|
||||
*/
|
||||
|
||||
extern void func_800180D8(int *a0, int *a1, int a2, int a3);
|
||||
|
||||
int func_80018210(int a0)
|
||||
{
|
||||
int v;
|
||||
|
||||
v = 0;
|
||||
func_800180D8(&v, (int *)(a0 + 12), 4, 1);
|
||||
func_800180D8(&v, (int *)(a0 + 28), 16, 0);
|
||||
return v;
|
||||
}
|
||||
@@ -0,0 +1,69 @@
|
||||
/*
|
||||
* func_800B7190 — 160 bytes at 0x800B7190..0x800B7230
|
||||
*
|
||||
* Hypothesis, not a claim about meaning: accumulates the negative components of one
|
||||
* 3-int vector into a second vector's first three ints, and the positive components into
|
||||
* its next three. A leaf with no frame and a completely repetitive body — it matched on the
|
||||
* SECOND spelling.
|
||||
*
|
||||
* Original words (byte offsets; int offsets are /4):
|
||||
* 8C810000 lw v1,0(a0) ; a1[0] += a0[0] & (a0[0] >> 31)
|
||||
* 8CA20000 lw v0,0(a1)
|
||||
* 00021603 sra a2,v1,0x1f
|
||||
* 00C20824 and v1,v1,a2
|
||||
* 00431021 addu v0,v0,v1
|
||||
* ACA20000 sw v0,0(a1)
|
||||
* 8C810004 lw v1,4(a0) ; a1[1] += a0[1] & (a0[1] >> 31)
|
||||
* ... ACA20004 sw v0,4(a1)
|
||||
* 8C810008 lw v1,8(a0) ; a1[2] += a0[2] & (a0[2] >> 31)
|
||||
* ... ACA20008 sw v0,8(a1)
|
||||
* 8C820000 lw a2,0(a0) ; t = a0[0] > 0; a1[4] += a0[0] & -t
|
||||
* 8CA10010 lw v1,16(a1)
|
||||
* 0002102A slt v0,zero,a2
|
||||
* 00021023 negu v0,v0
|
||||
* 00C23024 and a2,a2,v0
|
||||
* 00621821 addu v1,v1,a2
|
||||
* ACA10010 sw v1,16(a1)
|
||||
* 8C820004 lw a2,4(a0) ; t = a0[1] > 0; a1[5] += a0[1] & -t
|
||||
* ... ACA10014 sw v1,20(a1)
|
||||
* 8C840008 lw a0,8(a0) ; t = a0[2] > 0; a1[6] += a0[2] & -t
|
||||
* ... 03E00008 jr ra
|
||||
* ACA10018 sw v1,24(a1) ; (delay)
|
||||
*
|
||||
* **TWO DIFFERENT BRANCHLESS MASK IDIOMS, AND THE SECOND NEEDS A VARIABLE.**
|
||||
*
|
||||
* 1. The negative part is `a0[i] & (a0[i] >> 31)` — an arithmetic shift feeding a mask.
|
||||
* cc1 emits `sra; and; addu` with no branch, and this spelling works inline.
|
||||
* 2. The positive part is `t = a0[i] > 0; a1[4+i] += a0[i] & -t;` — **the comparison must
|
||||
* be assigned to a VARIABLE first.** Writing it inline as `a0[i] & -(a0[i] > 0)` makes
|
||||
* cc1 BRANCHIFY the whole thing (`bgez v1,<skip>; nop; move v1,zero`) and the row comes
|
||||
* out 188 bytes; with the variable it emits the original's `slt v0,zero,a2; negu
|
||||
* v0,v0; and a2,a2,v0` and is exact. So **a mask built from a comparison needs the
|
||||
* comparison materialised in a variable to stay branchless.**
|
||||
*
|
||||
* Also byte-required: the body is UNROLLED — six separate statements, not a loop (cc1 does
|
||||
* not unroll a constant-trip-count loop, cookbook 87). The destination offsets are 0/4/8
|
||||
* for the negative group and 16/20/24 for the positive group, i.e. a1's first two 4-int
|
||||
* vectors with the second vector's fourth int unused.
|
||||
*
|
||||
* LIMITS: the function name, the element count, the destination layout and the
|
||||
* "accumulate negative / positive parts" reading are hypotheses taken from the instruction
|
||||
* shape; only the bytes are evidence. No callee: this is a leaf and the harness emitted the
|
||||
* common epilogue itself.
|
||||
*/
|
||||
|
||||
void func_800B7190(int *a0, int *a1)
|
||||
{
|
||||
int t;
|
||||
|
||||
a1[0] += a0[0] & (a0[0] >> 31);
|
||||
a1[1] += a0[1] & (a0[1] >> 31);
|
||||
a1[2] += a0[2] & (a0[2] >> 31);
|
||||
|
||||
t = a0[0] > 0;
|
||||
a1[4] += a0[0] & -t;
|
||||
t = a0[1] > 0;
|
||||
a1[5] += a0[1] & -t;
|
||||
t = a0[2] > 0;
|
||||
a1[6] += a0[2] & -t;
|
||||
}
|
||||
Reference in New Issue
Block a user