phase10: merge 23 — 474 distinct bodies / 483 regions (ONE from the milestone)

+3 bodies (worker B2 claims 5-7). Candidate gate MATCH before promotion; md5 drift
check clean on all three (second use of the guard). make check green: regions=483
AGREE, differing_bytes=0 MATCH.

Region option granted: gp=-D_80121BFC on 0x80048128. This one carries PER-ACCESS
evidence in a single merge: D_80121BFC is read gp-RELATIVELY in worker B2's claim 7
row (lw v1,708(gp)) and ABSOLUTELY in claim 5's row (lui v1,0x8012 + lw v1,7164(v1)).
So the same symbol needs the override in one region and not in the other -- direct
confirmation of worker B's lever 12 that the access form is per-SITE, and an argument
that the override is a region property rather than a symbol property.

THREE NEW LEVERS (worker B2), each with a control:
1. A LOCAL SHARED BY TWO GUARD BLOCKS GETS COALESCED; TWO BRACE-SCOPED LOCALS DO NOT.
   0x8008BA80: the original's first guard loads the state byte into a0 (the
   parameter's own dead register) while the second guard loads *a2 into a FRESH v1.
   One function-scope local makes cc1 coalesce the live ranges into a0 and DIFFs;
   brace-scoping each reproduces it. A SCOPING lever, distinct from the
   named-locals family.
2. THE EVALUATION ORDER OF TWO SCALED TERMS IS BYTE-REQUIRED. 0x800504E4:
   base + a0*384 + a1*3072 scales a0 first; base + a1*3072 + a0*384 scales a1 first,
   which is the original. Worker B's claim-1 lever extended from the operands of one
   '+' to the ORDER OF TWO INDEX COMPUTATIONS.
3. AN UNSIGNED LOOP COUNTER SHOWS AS sltiu vs slti -- ONE BYTE (opcode 0x0b vs 0x0a).
   0x800504E4's i < 16 is sltiu, so the counter is unsigned int. The loop-test form of
   the signedness trap.

Also independently reproduced: worker C's lever 6, (unsigned)(c - 58) < 2 giving ONE
addiu+sltiu pair where c >= 58 && c <= 59 gives two tests (on 0x80048128).
This commit is contained in:
Christopher Williams
2026-09-24 08:11:03 -04:00
parent 475afcd1e7
commit ef29df9b58
5 changed files with 1458 additions and 1235 deletions
+1232 -1235
View File
File diff suppressed because it is too large Load Diff
+3
View File
@@ -162,6 +162,7 @@
0x800454C8 0x80045540 src/func_800454C8.c
0x800474B0 0x80047508 src/func_800474B0.c
0x80047984 0x80047A14 src/func_80047984.c
0x80048128 0x80048180 src/func_80048128.c gp=-D_80121BFC
0x80048350 0x80048394 src/func_80048350.c
0x80048E20 0x80048E70 src/func_80048E20.c
0x800491FC 0x80049298 src/func_800491FC.c
@@ -172,6 +173,7 @@
0x8004C0F0 0x8004C110 src/func_8004C0F0.c
0x8004CEEC 0x8004CF0C src/func_8004CEEC.c
0x8004E3FC 0x8004E440 src/func_8004E3FC.c
0x800504E4 0x80050548 src/func_800504E4.c
0x800516E0 0x800516FC src/func_800516E0.c
0x8005182C 0x80051864 src/func_8005182C.c
0x80052C98 0x80052CAC src/func_80052C98.c
@@ -274,6 +276,7 @@
0x8008B8E4 0x8008B8F4 src/func_8008B8E4.c
0x8008B8F4 0x8008B910 src/func_8008B8F4.c
0x8008B910 0x8008B960 src/func_8008B910.c
0x8008BA80 0x8008BAE0 src/func_8008BA80.c
0x8008D9AC 0x8008DA0C src/func_8008D9AC.c
0x8008DB44 0x8008DC04 src/func_8008DB44.c
0x8008F2E0 0x8008F338 src/func_8008F2E0.c
1 # Code-region registry: one C region per matched function.
162 0x800454C8
163 0x800474B0
164 0x80047984
165 0x80048128
166 0x80048350
167 0x80048E20
168 0x800491FC
173 0x8004C0F0
174 0x8004CEEC
175 0x8004E3FC
176 0x800504E4
177 0x800516E0
178 0x8005182C
179 0x80052C98
276 0x8008B8E4
277 0x8008B8F4
278 0x8008B910
279 0x8008BA80
280 0x8008D9AC
281 0x8008DB44
282 0x8008F2E0
+72
View File
@@ -0,0 +1,72 @@
/* func_80048128 — 0x80048128..0x80048180 (88 bytes).
*
* Original words (objdump of the validated payload, little-endian):
* sll v0,a0,0x2 \ a0 * 76, in cc1's strength-reduced form
* addu v0,v0,a0 | ((a0*5)*4 - a0)*4 = a0*19*4
* sll v0,v0,0x2 |
* subu v0,v0,a0 |
* lui v1,0x8012 |
* lw v1,7164(v1) | base = D_80121BFC (ABSOLUTE load, see point 2)
* sll v0,v0,0x2 |
* addu v0,v0,v1 / element = base + a0*76
* lbu v1,42(v0) c = *(unsigned char *)(element + 42)
* li v0,65
* beq v1,v0,0x80048174 if (c == 65) -> r = 1
* _move a1,zero (delay slot) r = 0
* addiu v0,v1,-58 \ (unsigned)(c - 58) < 2 ONE range test
* sltiu v0,v0,2 /
* bnez v0,0x80048174 if (c == 58 || c == 59) -> r = 1
* _nop
* li v0,36
* bne v1,v0,0x80048178 if (c != 36) -> return r
* _nop
* 0x80048174:
* li a1,1 r = 1
* 0x80048178:
* jr ra
* _move v0,a1 (delay slot) return r
*
* WHAT THE BYTES PIN DOWN:
*
* 1. THE THREE TESTS ARE ONE `||` CHAIN WITH A SINGLE RESULT LOCAL. Every
* condition branches to the SAME `li a1,1` block at 0x80048174, and the
* result lives in `a1` (not v0): `move a1,zero` sits in the FIRST branch's
* delay slot — i.e. `r = 0` is materialised once, unconditionally, before any
* test — and the epilogue's delay slot is `move v0,a1`. So the source is a
* single `int r = 0;` plus one combined condition, not three early returns.
* 2. THE RANGE TEST IS THE `(unsigned)(c - 58) < 2` SPELLING, NOT
* `c >= 58 && c <= 59`. The original emits ONE `addiu v0,v1,-58` plus
* `sltiu v0,v0,2`; the two-comparison spelling would emit two tests. This is
* worker C's lever 6 ("a range test has a single-test spelling") reproduced
* independently.
* 3. THE STRIDE IS 76 AND cc1 DECOMPOSES IT AS `((a0*5)*4 - a0)*4`. Writing the
* stride as a plain multiply is what produces that exact sequence, so the
* element type is 76 bytes wide (not a power of two, hence the shift/subtract
* chain rather than a shift-only scaling).
* 4. `D_80121BFC` IS READ ABSOLUTELY: `lui v1,0x8012` + `lw v1,7164(v1)` uses the
* absolute displacement 7164 = 0x1BFC from the lui base, NOT the gp-relative
* offset (which would be 708, since 0x80121BFC - 0x80121938 = 708). The
* registry row is gp-marked, so the region needs the per-site option
* `gp=-D_80121BFC`; verified with `--no-gp D_80121BFC`.
*
* LIMITS: the record layout (only the +42 byte is observed), the meaning of the
* three compared characters (65, 58/59, 36) and the stride 76 are read from the
* instruction shapes; the function is a classifier over a table of 76-byte
* records but what the characters and the record mean is not observable. Only
* the compiled bytes are evidence.
*
* REGION OPTION REQUIRED: `gp=-D_80121BFC` (per-site absolute access). Verified
* with `--no-gp D_80121BFC`.
*/
extern int D_80121BFC;
int func_80048128(int a0)
{
int r = 0;
unsigned char c = *(unsigned char *)(D_80121BFC + a0 * 76 + 42);
if (c == 65 || (unsigned)(c - 58) < 2 || c == 36)
r = 1;
return r;
}
+79
View File
@@ -0,0 +1,79 @@
/* func_800504E4 — 0x800504E4..0x80050548 (100 bytes).
*
* Original words (objdump of the validated payload, little-endian):
* li a2,9 i = 9
* li a3,-1 fill = -1
* sll v0,a1,0x1 \ a1 * 3072, cc1's strength-reduced form
* addu v0,v0,a1 | (a1*3) << 10
* sll v0,v0,0xa /
* sll v1,a0,0x1 \ a0 * 384, cc1's strength-reduced form
* addu v1,v1,a0 | (a0*3) << 7
* sll v1,v1,0x7 /
* lui a0,0x8013 \ base = D_8012FDD4 (SYMBOL form: %hi 0x8013,
* addiu a0,a0,-556 / %lo -556, i.e. the two-piece `la`)
* addu v1,v1,a0 v1 = base + a0*384
* addu v0,v0,v1 v0 = a1*3072 + (base + a0*384)
* addiu v1,v0,216 v1 = v0 + 9*24 (the first element)
* 0x80050518:
* sw zero,0(v1) \
* sw zero,4(v1) |
* sb zero,8(v1) | clear one 24-byte record
* sw zero,12(v1) |
* sh zero,20(v1) |
* sh a3,22(v1) / record[22] = -1
* addiu a2,a2,1 i++
* sltiu v0,a2,16 i < 16 (UNSIGNED compare, see point 3)
* bnez v0,0x80050518 loop
* _addiu v1,v1,24 (delay slot) v1 += 24
* jr ra
* _nop
*
* Clears a run of 24-byte records, nine of which are then given a -1 halfword.
*
* WHAT THE BYTES PIN DOWN:
*
* 1. THE TWO SCALED INDICES ARE COMPUTED IN A FIXED ORDER: a1's term FIRST,
* then a0's. Written `base + a0 * 384 + a1 * 3072` cc1 emits a0's scaling
* first and the region DIFFs in 6 bytes of pure register numbering; written
* `base + a1 * 3072 + a0 * 384` the two scales come out in the original's
* order and the row is byte-identical. Same commutative-operand family as
* worker B's claim 1, here on the *evaluation order* of two scaled terms
* rather than on the operands of one `+`.
* 2. THE BASE IS A SYMBOL, NOT A LITERAL: `lui a0,0x8013` + `addiu a0,a0,-556`
* is the two-piece `la` form (%hi carries because 0xFDD4 has the high bit
* set). The registry has no row for 0x8012FDD4, so the address-named symbol
* `D_8012FDD4` is what produces this encoding.
* 3. THE LOOP TEST IS UNSIGNED: the original emits `sltiu`, and a signed loop
* counter gives `slti` — ONE differing byte (opcode 0x0b vs 0x0a). So the
* counter is declared `unsigned int`.
* 4. THE FRAME IS EMPTY — no prologue at all — because the row makes no calls
* and uses only t-registers, so all four arguments stay in their incoming
* registers (a0/a1 are consumed by the scaling, a2 is the counter, a3 is the
* fill constant).
* 5. THE RECORD LAYOUT IS 24 BYTES with fields at +0, +4, +8, +12, +20, +22;
* the loop pointer is strength-reduced (`addiu v1,v1,24` in the branch's
* delay slot) and starts at +216 = 9*24, i.e. the loop begins at index 9.
*
* LIMITS: the record layout, the meaning of the run [9,16) and of the -1
* halfword are hypotheses read from the instruction shapes. Only the compiled
* bytes are evidence.
*/
extern char D_8012FDD4[];
void func_800504E4(int a0, int a1)
{
char *base = D_8012FDD4 + a1 * 3072 + a0 * 384;
unsigned int i;
for (i = 9; i < 16; i++) {
char *p = base + i * 24;
*(int *)(p + 0) = 0;
*(int *)(p + 4) = 0;
*(char *)(p + 8) = 0;
*(int *)(p + 12) = 0;
*(short *)(p + 20) = 0;
*(short *)(p + 22) = -1;
}
}
+72
View File
@@ -0,0 +1,72 @@
/* func_8008BA80 — 0x8008BA80..0x8008BAE0 (96 bytes).
*
* Original words (objdump of the validated payload, little-endian):
* beqz a0,0x8008BAA0 if (a0 == 0) skip the first block
* _li v0,4 (delay slot)
* lbu a0,38(a0) v = *(unsigned char *)(a0 + 38)
* nop
* beq a0,v0,0x8008BAD8 if (v == 4) return
* _li v0,14 (delay slot)
* beq a0,v0,0x8008BAD8 if (v == 14) return
* _nop
* 0x8008BAA0:
* li v0,28
* bne a1,v0,0x8008BAD8 if (a1 != 28) return
* _li v0,12 (delay slot)
* lw v1,0(a2) w = *a2
* nop
* beq v1,v0,0x8008BAD8 if (w == 12) return
* _li v0,16 (delay slot)
* beq v1,v0,0x8008BAD8 if (w == 16) return
* _li v0,3 (delay slot)
* beq v1,v0,0x8008BAD8 if (w == 3) return
* _li v0,19 (delay slot)
* beq v1,v0,0x8008BAD8 if (w == 19) return
* _li v0,15 (delay slot)
* sw v0,0(a2) *a2 = 15
* 0x8008BAD8: jr ra; _nop
*
* A guard chain: the object's state byte and the caller's mode decide whether a
* single field is overwritten with 15.
*
* WHAT THE BYTES PIN DOWN:
*
* 1. THE TWO BLOCKS USE DIFFERENT REGISTERS BECAUSE THEY ARE DIFFERENT LOCALS IN
* DIFFERENT SCOPES. The first block's byte lands in `a0` (the parameter's own
* register is reused, since the parameter is dead after the `lbu`), while the
* second block's `*a2` lands in `v1` — a FRESH register, not `a0`. Written
* with one function-scope local shared by both blocks, cc1 coalesces the two
* live ranges into `a0` and the region DIFFs in 5 bytes
* (`lw a0,0(a2)` / `beq a0,v0` instead of `lw v1,0(a2)` / `beq v1,v0`).
* Declaring each in its own brace-scoped block is what keeps them apart.
* 2. EACH CONDITION IS A `||` CHAIN BRANCHING TO THE SHARED EPILOGUE, so the
* guards are early `return`s and the only store is the last statement.
* 3. THE COMPARED CONSTANTS ARE 4/14 for the state byte and 12/16/3/19 for the
* field, in that order, and the mode constant is 28. The order is observable
* (each `li` sits in its own branch's delay slot) and the source order matches
* it.
*
* LIMITS: the record layout (only the +38 byte is observed), the meaning of the
* constants 4, 14, 12, 16, 3, 19, 28 and 15, and the type of `a2`'s pointee are
* hypotheses read from the instruction shapes. Only the compiled bytes are
* evidence.
*/
void func_8008BA80(int a0, int a1, int *a2)
{
if (a0 != 0) {
int v = *(unsigned char *)(a0 + 38);
if (v == 4 || v == 14)
return;
}
if (a1 != 28)
return;
{
int w = *a2;
if (w == 12 || w == 16 || w == 3 || w == 19)
return;
*a2 = 15;
}
}