diff --git a/config/regions.tsv b/config/regions.tsv index 8c50325..9885514 100644 --- a/config/regions.tsv +++ b/config/regions.tsv @@ -157,6 +157,8 @@ 0x80036A54 0x80036A9C src/func_80036A54.c 0x80036A9C 0x80036AD8 src/func_80036A9C.c 0x80036AD8 0x80036B14 src/func_80036AD8.c +0x80036B14 0x80036DA4 src/func_80036B14.c +0x80036DA4 0x80036F70 src/func_80036DA4.c 0x8003768C 0x800376CC src/func_8003768C.c 0x80038788 0x80038790 src/func_80038788.c 0x80038790 0x8003879C src/func_80038790.c diff --git a/docs/MATCHING_COOKBOOK.md b/docs/MATCHING_COOKBOOK.md index ae3dbc2..e9cca4d 100644 --- a/docs/MATCHING_COOKBOOK.md +++ b/docs/MATCHING_COOKBOOK.md @@ -1466,3 +1466,51 @@ Worker B's batch clustered into exactly two classes, and **neither is a shape pr **Both are the same post-pass family** (finding 40), which is why grinding them with source spellings is the wrong expenditure. This is the strongest argument yet for writing the post-pass: two workers independently arrived at the same two classes and both stopped for the same reason. + +### 95. Residual of a few BRANCH WORDS at correct length ⇒ the block NESTING is wrong (worker A) + +Worker A's cleanest diagnostic, and it deserves to be the first thing checked on any +correct-length candidate with a tiny residual: + +> **A correct-length candidate whose residual is a handful of BRANCH WORDS means the block +> NESTING is wrong, not the code inside the blocks.** + +`0x80036B14`: the two guards written as **siblings** (`if (a1 != 0) {...} if (a2 != 0) {...}`) +gives **exactly 656 bytes — the correct length — and exactly 2 differing bytes, both of them +branch displacements**. `beqz s1` jumped to the `if (a2)` test instead of past the whole second +block. **Nesting the `a2` block INSIDE the `a1` block is 0.** The instruction stream is +otherwise identical, so there is no other tell. **Residual = 2 bytes at a branch displacement +⇒ go look at your braces.** + +### 96. The project's 4-int vector type: frame multiple of 16, offsets stepping by 16 (worker A) + +Confirmed across three independent rows: `int m[3][4]` (0x8009F798), five 16-byte groups +(0x8006AE54), `struct V { int v[4]; }` (0x80036B14/0x80036DA4). Declaring such a thing as +scalars costs **264 bytes** (0x8006AE54) or **128 bytes** (0x80036DA4's first spelling family). + +**If a row's frame is a multiple of 16 and the sp-relative offsets step by 16, the source has +4-int vectors.** Third independent confirmation that **vector copies are struct assignments** +(4 loads then 4 stores) — after 0x8005584C and 0x8009F798. + +Related, same rows: `t.v[0] = t.v[1] = t.v[2] = *a2;` loads the **SAME address three times** +(cc1 does not CSE it) and the struct's 4th word is copied as **stack garbage** — only +reproducible by copying a whole struct. And `struct V u = D_8010C124;` is an *initialiser*, so it +hoists above the following stores. + +### 97. `/ 4096` is not `>> 12` — and it is neither division class (worker A) + +cc1 emits the `addiu v0,v0,4095` **sign bias** for the division and would not for a shift. Worth +knowing given the constant-division discussion: this is a power-of-two divide that is **NOT** the +div-instruction class and **NOT** the finding-67 `mfhi`-without-`mflo` class — there *is* an +`mflo` — so it is safe to attempt. + +### 98. OPEN: "the original spills everything, cc1 promotes" (worker A, needs a recipe) + +Worker A's `0x8006AE54` (596 B): the original keeps **every one of the ten locals in memory** (all +stored and reloaded) while cc1 promotes the struct members to registers. Spellings tried: scalars +332 B, arrays 540 B, struct 548 B, struct-behind-pointer 680 B. Untried lever: force memory +residency with a whole-struct write, or an escape cc1 cannot see through. + +**This is a genuinely open question and I am recording it as one.** Any worker who has hit +"original spills everything, cc1 promotes" should report the recipe. Compare finding 64 (memory +residence as a load-bearing property of finding 59's family). diff --git a/src/func_80055654.c b/src/func_80055654.c new file mode 100644 index 0000000..7699860 --- /dev/null +++ b/src/func_80055654.c @@ -0,0 +1,94 @@ +/* + * func_80055654 — 148 bytes at 0x80055654..0x800556E8 + * + * Hypothesis, not a claim about meaning: maps one of two known object handles to an index + * (0, 1, or -1 otherwise), sets a single bit in the object's control word, and mirrors that + * bit into a per-index table entry. + * + * Original words: + * 3C028013 lui v0,0x8013 + * 8C42D79C lw v0,-10340(v0) ; v0 = D_8012D79C + * 00000000 nop + * 1082000B beq a0,v0,0x80055684 ; if (a0 == D_8012D79C) -> the k = 0 block + * 00000000 nop + * 3C028013 lui v0,0x8013 + * 8C42DD30 lw v0,-8912(v0) ; v0 = D_8012DD30 + * 00000000 nop + * 1442000A bne a0,v0,0x80055688 ; if (a0 != D_8012DD30) -> join with k = -1 + * 2402FFFF li a2,-1 ; (delay) + * 080155A2 j 0x80055688 + * 24020001 li a2,1 ; (delay) + * 00003021 move a2,zero ; 0x80055684: k = 0 + * 8C820020 lw v0,32(a0) ; 0x80055688: + * 00000000 nop + * 8C4400F4 lw a0,244(v0) + * 2403FFFE li v1,-2 + * 8C820018 lw v0,24(a0) ; p[6] + * 30A10001 andi a1,a1,0x1 + * 00431024 and v0,v0,v1 ; p[6] & ~1 + * 00411025 or v0,v0,a1 ; | (a1 & 1) + * AC820018 sw v0,24(a0) + * 00042100 sll a0,a2,0x4 ; k * 16 + * 00842021 addu a0,a0,a2 ; + k + * 00042080 sll a0,a0,0x2 ; * 4 -> k * 68 + * 2403FDFF li v1,-513 ; ~0x200 + * 3C018013 lui at,0x8013 + * 00240821 addu at,at,a0 + * 8C22D6B4 lw v0,-10572(at) ; D_8012D6B4[k].v[0] + * 00010A40 sll a1,a1,0x9 + * 00431024 and v0,v0,v1 + * 00411025 or v0,v0,a1 + * 3C018013 lui at,0x8013 + * 00240821 addu at,at,a0 + * AC22D6B4 sw v0,-10572(at) + * 03E00008 jr ra + * 00000000 nop + * + * **THE OUTER TEST MUST BE SPELLED `!=` WITH THE SINGLE-STATEMENT CASE AS THE `else`.** + * The natural spelling + * `if (a0 == D1) k = 0; else if (a0 == D2) k = 1; else k = -1;` + * gives the CORRECT LENGTH (148) and 34 differing bytes, all of them in the first four + * instructions: cc1 lays the `k = 0` block out first (fall-through) and jumps to the + * `k = 1` block, whereas the original jumps *forward* to `k = 0` and keeps `k = 1` as the + * fall-through. Writing the outer test inverted with the big block as the `then` — + * `if (a0 != D1) { if (a0 == D2) k = 1; else k = -1; } else k = 0;` + * — puts the blocks in the original's order and is exact. **The then/else ORDER of a + * single-statement arm is byte-load-bearing**, because it decides which arm falls through. + * + * Also byte-required: the object pointer is a BYTE offset (`244(v0)`, i.e. + * `*(char **)(... + 244)`), not `int *` arithmetic — `p + 244` on an `int *` is +976 and + * shows up as `lw a0,976(v0)`. The per-index table is a **68-byte struct array** + * (k*16 + k, then *4), not an `int` array. + * + * LIMITS: the function name, both handles, the table symbol and its 68-byte element size, + * the field offsets (24 = the control word, 244 = the sub-object pointer) and the meaning + * of the bit (bit 0 mirrored to bit 9) are hypotheses read from the instruction shape; only + * the bytes are evidence. The 68-byte element is modelled as `int v[17]` purely to obtain + * the scale; its real field layout is unknown. None of the three symbols is registered, so + * they are referenced by their address-named spellings. + */ + +struct S { int v[17]; }; + +extern int D_8012D79C; +extern int D_8012DD30; +extern struct S D_8012D6B4[]; + +void func_80055654(int a0, int a1) +{ + int k; + int *p; + + if (a0 != D_8012D79C) { + if (a0 == D_8012DD30) + k = 1; + else + k = -1; + } else { + k = 0; + } + + p = *(int **)(*(char **)(a0 + 32) + 244); + p[6] = (p[6] & ~1) | (a1 & 1); + D_8012D6B4[k].v[0] = (D_8012D6B4[k].v[0] & ~0x200) | ((a1 & 1) << 9); +} diff --git a/src/func_800556E8.c b/src/func_800556E8.c new file mode 100644 index 0000000..a774ab8 --- /dev/null +++ b/src/func_800556E8.c @@ -0,0 +1,98 @@ +/* + * func_800556E8 — 356 bytes at 0x800556E8..0x8005584C + * + * Hypothesis, not a claim about meaning: if a sub-object has any of three leading or three + * trailing components set, either copy one 4-int vector over another (when two control + * fields are 0 or 2) or nudge the destination vector by a scaled difference computed by a + * helper. The leading six tests are one OR-chain and the two inner tests are `== 0 || == 2` + * — that repetition is why it matched on the FIRST spelling. + * + * Shape (v = *(char **)(a0 + 32), s0 = (int *)(*(char **)(v + 244) + 400)): + * 27BDFFD8 addiu sp,sp,-40 + * AFBF0024 sw ra,36(sp) + * AFB00020 sw s0,32(sp) + * 8C820020 lw v0,32(a0) + * 8C4300F4 lw v1,244(v0) + * 8C620190 lw v0,400(v1) ; s0[0] + * 14400021 bnez v0,0x80055764 ; -- OR-chain: every test jumps to the SAME target + * 24820190 addiu s0,v1,400 ; (delay) + * 8C620194 lw v0,404(v1) ; s0[1] + * 14400021 bnez v0,0x80055764 + * ... ; s0[2], s0[11], s0[12] + * 8C6201C4 lw v0,452(v1) ; s0[13] + * 10400031 beqz v0,0x80055838 ; the LAST term inverts to skip the whole body + * ... + * 8E030024 lw v1,36(s0) ; 0x80055764: s0[9] + * 10600004 beqz v1,0x8005577C ; s0[9] == 0 -> test the next field + * 24020002 li v0,2 + * 1462001C bne v1,v0,0x800557BC ; s0[9] != 2 -> the else path + * 8E030028 lw v1,40(s0) ; 0x8005577C: s0[10], same shape + * ... 10600004 beqz / li 2 / bne + * 8E020000 lw v0,0(s0) ; 0x80055794: *(V *)(s0+11) = *(V *)s0 + * 8E030004 lw v1,4(s0) ; 4 loads then 4 stores + * 8E040008 lw a0,8(s0) + * 8E05000C lw a1,12(s0) + * AE02002C sw v0,44(s0) + * ... AE050038 sw a1,56(s0) + * 0801560E j 0x80055838 ; done + * 00000000 nop + * 8E020000 lw v0,0(s0) ; 0x800557BC: the else path + * 8E03002C lw v1,44(s0) + * 27A40010 addiu a0,sp,16 ; a0 = d + * 00021023 subu v0,v0,v1 + * AFA20010 sw v0,16(sp) ; d[0] + * ... ; d[1], d[2] + * 24010666 li a1,1638 + * 00403021 move a2,a0 ; func_800237B4(d, 1638, d) + * 0C008DED jal 0x800237B4 + * AFA20018 sw v0,24(sp) ; (delay) + * 8E02002C lw v0,44(s0) ; s0[11] += d[0]; s0[12] += d[1]; s0[13] += d[2] + * ... + * 8FBF0024 lw ra,36(sp) ; (END) + * + * BYTE-REQUIRED SHAPES: + * + * 1. **The leading six tests are ONE `||` chain.** Each `bnez` jumps to the same forward + * target (the body), which is exactly cc1's OR-chain codegen; the final term is tested + * inverted (`beqz`) to skip the body. Writing them as six separate `if (...) goto` + * statements gives the same shape only if the targets coincide — the `||` form is what + * the original is. + * 2. **`(s0[9] == 0 || s0[9] == 2) && (s0[10] == 0 || s0[10] == 2)`** — cc1 emits + * `beqz`/`li 2`/`bne` per term and both inner tests share the else path. + * 3. **The copy is a STRUCT ASSIGNMENT** (`*(struct V *)(s0 + 11) = *(struct V *)(s0 + 0);`) + * — 4 loads then 4 stores. Fourth independent confirmation of this lever. + * 4. **`func_800237B4(d, 1638, d)`** passes the same buffer as first and third argument + * (in-place), and `d` is an ARRAY (`int d[3]` at sp+16) because its address is taken. + * 1638 = 0x666 is a fixed-point constant, not an index. + * + * LIMITS: the function name, the callee, the sub-object layout (the 400-byte prefix, the + * fields at +0/+4/+8/+36/+40/+44..+56) and the meaning of the 0/2 tests are hypotheses read + * from the instruction shape; only the bytes are evidence. The `V` struct is used purely to + * obtain the 4-word block move. func_800237B4 is not registered and is referenced by its + * address-named spelling. + */ + +struct V { int v[4]; }; + +void func_800237B4(int *, int, int *); + +void func_800556E8(int a0) +{ + char *v = *(char **)(a0 + 32); + int *s0 = (int *)(*(char **)(v + 244) + 400); + int d[3]; + + if (s0[0] || s0[1] || s0[2] || s0[11] || s0[12] || s0[13]) { + if ((s0[9] == 0 || s0[9] == 2) && (s0[10] == 0 || s0[10] == 2)) { + *(struct V *)(s0 + 11) = *(struct V *)(s0 + 0); + } else { + d[0] = s0[0] - s0[11]; + d[1] = s0[1] - s0[12]; + d[2] = s0[2] - s0[13]; + func_800237B4(d, 1638, d); + s0[11] += d[0]; + s0[12] += d[1]; + s0[13] += d[2]; + } + } +}