phase-36: S104 s104_e26 — func_8017E764 banked at 0 through the whole-object gate + propagated — $3 pin → 0: the early temp renamed to val (sched.c:2489-2490) + i = 0 moved up; the byte-needed volatile kept unmarked (decision 3)

This commit is contained in:
Drew T
2026-09-11 04:02:40 -06:00
parent 32f641aca0
commit 834dd2f1bd
11 changed files with 489 additions and 191 deletions
@@ -0,0 +1,21 @@
void func_801823BC(a0, a1)
s32 a0;
s32 a1;
{
extern void func_80049CAC(s32 a0, s32 a1);
extern void func_8017EF68(s32 a0, s32 a1, s32 a2);
extern s16 D_801AEB52[];
extern s32 D_801B20E0;
extern s32 D_801A158C;
extern s32 D_801AEB1C;
s16 *q;
s32 *r;
q = D_801AEB52;
*q = a0;
func_80049CAC((s32)q - 2, (s32)q - 78);
D_801AEB1C = (s16)a1;
r = (s32 *)((s32)q - 82);
*r = 0;
func_8017EF68((s32)&D_801B20E0, (s32)&D_801A158C, (s32)r);
}
@@ -0,0 +1,41 @@
# func_801823BC (ov_SC06_000_jr_8017AE2C.c) — e28, P36 T7 S104
**Score: 4 (sweep best 2) -> 0, zero levers** (the `$17` pin AND the launder asm both deleted; plain C). First `--try`.
## (a) Residual
COUNT +1 (29 vs 28): the store `*(s32 *)((s32)q - 82) = 0` came out absolute (`lui at,%hi(D_801AEB52-82); sw zero,..(at)`)
where the target stores through the callee-saved register holding `&D_801AEB52` (`sw zero,-82(s1)`, in the `jal` delay
slot, after `addiu a2,s1,-82`). The pin was only there because the launder needed q in a known register; the register
itself (`$s1`) is what global alloc gives q anyway once the call argument keeps it live across the first call.
## (b) Pass and decision (the e12 mechanism, same TU; e12 proved it on .cse dumps for func_80182560)
cse1 `find_best_addr` (gcc-2.7.2 cse.c:2622): the free text's address `(plus (reg q) -82)` is not a REG, so `fold_rtx`
runs first (cse.c:2662-2664), substitutes q's constant equivalent `(symbol_ref D_801AEB52)` and folds the address to
`(const (plus sym -82))` — a CONSTANT address, never replaced again (cse.c:2656) -> `lui at; sw`. A REG address
`(mem (reg r))`, `r = q - 41`, is not folded; the lookup of r's class finds `(plus (reg q) -82)` at the same ADDRESS_COST
as a REG (1: mips.h:2895, mips.c:1652-1653) and a higher rtx cost, which the tie-break prefers (cse.c:2711-2721); the
symbolic constant costs 2 (mips.c:1631) and loses. r stays live as the third call argument, so cse2 does not re-fold
(e12's func_80184A68 (b)3: a derived pointer with no other live use is re-folded in cse2).
Proven on bytes here (score 0); the dump reading is e12's (not re-dumped for this function).
## (c) Move
Name the derived pointer and pass it: `r = (s32 *)((s32)q - 82); *r = 0; func_8017EF68(&D_801B20E0, &D_801A158C, (s32)r);`
(`r = (s32 *)(q - 41);` also scores 0 — scratch/v2.c.)
## (d) Generator proposal
When a launder/`la` pointer's constant-offset access comes out as `lui at; sw/lw %lo(SYM+k)(at)` where the target has
`sw k(reg)`, and the same `p + k` is also passed to a later call, name `r = p + k`, access `*r` and pass `r`.
## (e) What did not work
The sweep's R4/R7/R9/R12 moves stay at 2 (ORDER): none of them turns the address into a REG. Nothing else was needed.
## (f) Method
Counting first + e12's sibling text closed it on the first try; the method worked as written (brief + neighbours).
## (g) Structs
A struct on the global would give CONSTANT field addresses (cse.c:2656), the opposite of the target; a struct pointer
based at D_801AEB52 with fields at -2/-78/-82 is not expressible (negative offsets) — the base would move to
D_801AEB52-82 and the offsets change. The deciding fact is REG vs `reg+const` address in find_best_addr, not aliasing.
## Copies
func_80182464 (same TU) closes with the identical text (its own pack). No other copy of the `q - 2, q - 78` call in src/.
@@ -0,0 +1,19 @@
void func_80182464(s32 a0, s32 a1)
{
extern void func_80049CAC(s32 a0, s32 a1);
extern void func_8017EF68(s32 a0, s32 a1, s32 a2);
extern s16 D_801AEBAA[];
extern s32 D_801B20E0;
extern s32 D_801A3A0C;
extern s32 D_801AEB74;
s16 *q;
s32 *r;
q = D_801AEBAA;
*q = a0;
func_80049CAC((s32)q - 2, (s32)q - 78);
D_801AEB74 = (s16)a1;
r = (s32 *)((s32)q - 82);
*r = 0;
func_8017EF68((s32)&D_801B20E0, (s32)&D_801A3A0C, (s32)r);
}
@@ -0,0 +1,41 @@
# func_80182464 (ov_SC06_000_jr_8017AE2C.c) — e28, P36 T7 S104
**Score: 4 (sweep best 2) -> 0, zero levers** (the `$17` pin AND the launder asm both deleted; plain C). First `--try`.
## (a) Residual
COUNT +1 (29 vs 28): the store `*(s32 *)((s32)q - 82) = 0` came out absolute (`lui at,%hi(D_801AEBAA-82); sw zero,..(at)`)
where the target stores through the callee-saved register holding `&D_801AEBAA` (`sw zero,-82(s1)`, in the `jal` delay
slot, after `addiu a2,s1,-82`). The pin was only there because the launder needed q in a known register; the register
itself (`$s1`) is what global alloc gives q anyway once the call argument keeps it live across the first call.
## (b) Pass and decision (the e12 mechanism, same TU; e12 proved it on .cse dumps for func_80182560)
cse1 `find_best_addr` (gcc-2.7.2 cse.c:2622): the free text's address `(plus (reg q) -82)` is not a REG, so `fold_rtx`
runs first (cse.c:2662-2664), substitutes q's constant equivalent `(symbol_ref D_801AEBAA)` and folds the address to
`(const (plus sym -82))` — a CONSTANT address, never replaced again (cse.c:2656) -> `lui at; sw`. A REG address
`(mem (reg r))`, `r = q - 41`, is not folded; the lookup of r's class finds `(plus (reg q) -82)` at the same ADDRESS_COST
as a REG (1: mips.h:2895, mips.c:1652-1653) and a higher rtx cost, which the tie-break prefers (cse.c:2711-2721); the
symbolic constant costs 2 (mips.c:1631) and loses. r stays live as the third call argument, so cse2 does not re-fold
(e12's func_80184A68 (b)3: a derived pointer with no other live use is re-folded in cse2).
Proven on bytes here (score 0); the dump reading is e12's (not re-dumped for this function).
## (c) Move
Name the derived pointer and pass it: `r = (s32 *)((s32)q - 82); *r = 0; func_8017EF68(&D_801B20E0, &D_801A3A0C, (s32)r);`
(`r = (s32 *)(q - 41);` scored 0 on func_801823BC.)
## (d) Generator proposal
When a launder/`la` pointer's constant-offset access comes out as `lui at; sw/lw %lo(SYM+k)(at)` where the target has
`sw k(reg)`, and the same `p + k` is also passed to a later call, name `r = p + k`, access `*r` and pass `r`.
## (e) What did not work
The sweep's R4/R7/R9/R12 moves stay at 2 (ORDER): none of them turns the address into a REG. Nothing else was needed.
## (f) Method
Counting first + e12's sibling text closed it on the first try; the method worked as written (brief + neighbours).
## (g) Structs
A struct on the global would give CONSTANT field addresses (cse.c:2656), the opposite of the target; a struct pointer
based at D_801AEBAA with fields at -2/-78/-82 is not expressible (negative offsets) — the base would move to
D_801AEBAA-82 and the offsets change. The deciding fact is REG vs `reg+const` address in find_best_addr, not aliasing.
## Copies
func_801823BC (same TU) closes with the identical text (its own pack). No other copy of the `q - 2, q - 78` call in src/.
@@ -0,0 +1,35 @@
void func_801842CC(s32 arg0)
{
s32 i;
for (i = 0; i < 0x49; i += 0x18) {
s32 e = func_80132EF4(arg0, 0x22);
if (e != 0) {
s32 t;
s32 v;
s32 d;
s32 z;
t = *(u16 *)(arg0 + 0x84) & 4;
z = *(u16 *)(e + 0xE);
d = i + t * 3 - 0x2A;
do { // !FAKE: do-while — its LOOP notes are sched1 barriers (sched.c:2053-2080): the +6 load stays after d, the +E store before the sign extension (P36 S104 e28)
v = *(u16 *)(e + 0x6);
z += 0x28;
*(u16 *)(e + 0xE) = z;
} while (0);
v -= d;
*(u16 *)(e + 0x6) = v;
v = *(u16 *)(e + 0xA);
v -= 0x20;
*(u16 *)(e + 0xA) = v;
v = (s16)d;
if (v >= 0) {
v = v << 7;
} else {
v = -(v << 7);
}
*(u16 *)(e + 0x34) = v + 0x1000;
*(s32 *)(e + 0x14) = 0xFFFE0000;
*(s32 *)(e + 0x18) = 0x30000;
}
}
}
@@ -0,0 +1,78 @@
# func_801842CC (ov_SC06_000_jr_8017AE2C.c) — e28, P36 T7 S104
**Score: 24 free (sweep best 5) -> 0 with ZERO pins/asm and ONE marked `do { … } while (0)`** (step 8's allowed-but-marked
construct). Tree: 3 levers (pin v1 + pin a0 + launder) -> 0 levers + 1 marked do-while. NOT a strict plain-C close: every
do-while-free spelling I found stays at >= 21 (see (e)).
The same text closes the class's two other copies (`--try` 0 each; rename + the copy's own constants):
ov_SC03_024 `func_80182108` (identical) and ov_SC03_014 `func_801889B4` (`z -= 0x28`, `0xFFFD0000`) —
scratch/copy_80182108_dw.c, scratch/copy_801889B4_dw.c.
Alternative at 0 (kept for the coordinator's choice): the tree's shape with ONE pin, `register s32 v0 __asm__("v0")` on
the reused variable, launder and the v1/a0 pins deleted — scratch/body_onepin.c (also 0 on both copies:
scratch/copy_80182108.c, copy_801889B4.c). And two pins (v1, a0) + no launder: scratch/t4.c.
## (a) Residual
COUNT-equal register rotation in the loop body: free text x/(x&4) in `a0`, d in `v0`, the +E value in `v1`; the target has
x in `v0`, d in `v1`, +E in `a0`, and (consequence) `sh v0,52(a1)` sinks below the `lui v0` constants. Plus
`addu v0,v0,s0` vs the target's `addu v1,s0,v1`.
## (b) Passes and decisions (PROVEN on dumps: scratch/dumps_free, dumps_c1, dumps_c2, dumps_e2, dumps_h1; localalloc_sim)
1. **local-alloc** (eligibility local-alloc.c:472 = one block AND one death; qty order :1486-1507; find_free_reg lowest
free). In the free text ONE variable `v0` holds x, x&4, the +6 value, the +A value and the sign-extended result: it
dies in 4 places -> global. Block 2 then has two local qtys, d (33846) and +E (15000): d takes `$2`, +E `$3`, global
gives v0 `$4`. The target needs a `$2` holder that local-alloc SEES.
2. Giving x its own variable `t` makes THREE local qtys, t=q0, +E=q1, d=q2; the three-qty switch (it compares qty
NUMBERS 0/1/2 while swapping POSITIONS, :1489-1500) orders them [t, d, +E] -> `$2/$3/$4`, the target's
(localalloc_sim on dumps_h1: q0 37500 v0, q2 33846 v1, q1 20625 a0, 0 mismatches).
3. **sched1** then breaks it. `priority()` (sched.c:1425) = max over LOG_LINKS of pred priority + cost - 1. With x in the
same pseudo as the +6 load, that load carried REG_DEP_ANTI links to the x uses (priority 2) and stayed after the d
chain; with `t` split it has no predecessor, priority 1, and is hoisted above the d chain (dumps_c2 trace T-11:
"58 (2) 62 (1)" -> 58), overlapping t, so global cannot give the reused `v` `$2`; the subu then gets 2, the `sll d,16`
ties the +E store at 2 and the store wins the potential-hazard tie (schedule_select sched.c:2616-2686, "insn 68 has a
greater potential hazard"). Needs 1-2 (x its own single-death local) and 3 (x sharing the +6 load's pseudo) contradict
in plain C.
4. **The do-while**: its NOTE_INSN_LOOP_BEG/END are sched1 barriers (sched.c:2053-2080: the insn after a loop note gets a
dependence on every earlier set/use, `reg_pending_sets_all`, `flush_pending_lists`). LOOP_BEG makes the +6 load depend
on the d chain (stays after it); LOOP_END makes the +E store the barrier, so the sign extension and everything after
it must follow the store. dumps_h1's block-2 trace is serialised to single-element ready lists. One note is not enough:
a lone barrier before the +6 load (scratch/e1-e5.c) scores 2 — the barrier insn is then the LOAD, every later insn
inherits its latency (+40 gets priority 3) and the +E store again ties the `sll` and wins by potential hazard.
5. `addu v1,s0,v1`: expand_binop swaps a commutative op whose target equals op1 (optabs.c:409-418), so `v1 = s0 + v1`
became `(plus v1 s0)`; one expression `d = i + t * 3 - 0x2A` keeps the counter first. The launder's only job was to keep
combine from turning `(v0<<1) + v0` (disjoint bits of `x & 4`) into `or` (score 1 without it, scratch/t1.c); `t * 3`
in one expression keeps `addu` — proven on bytes; which combine test declines the ior was NOT dumped.
## (c) Moves (body.c)
- x split into its own `t = *(u16 *)(arg0 + 0x84) & 4;` (the rest of the reused variable is `v`);
- `v1 = v0 * 2; launder; v1 = v1 + v0; v1 = s0 + v1; v1 = v1 - 0x2A;` -> `d = i + t * 3 - 0x2A;`;
- the +6 load and the +E update wrapped in a marked `do { v = *(u16 *)(e + 6); z += 0x28; *(u16 *)(e + 0xE) = z; } while (0);`
- readability: `if (e != 0) { … }` instead of `continue`, `v = (s16)d;` instead of `<< 16`/`>> 16`, `v = -(v << 7);`,
`*(u16 *)(e + 0x34) = v + 0x1000;` — each proven neutral on bytes (scratch/g3, g4, g8, g9, g10 -> h1 = body.c).
- Not neutral (keep as written): `*(u16 *)(e + 6) = v - d;` (21), `*(u16 *)(e + 0xE) = z + 0x28;` inside the do (21),
`*(u16 *)(e + 0xA) -= 0x20;` (30).
## (d) Generator proposal
When a register residual needs a variable SPLIT for local-alloc (a multi-death local, :472) but the split lets sched1
hoist a load that the shared variable used to pin with anti-dependences, wrap [that load .. the next statement group the
hazard tie reorders] in a marked `do { … } while (0)` — and try the two pins -> one pin on the reused variable as the
lever-count fallback.
## (e) What did not work (bytes)
- 2160-body plain-C sweep (scratch/gen.py -> gen_results.txt): x split / fused, three d spellings, three +E spellings,
+6 and +A updates inline / via temp, both sign-extension spellings, all 6 field-update orders, +E load before/after
the x load: best 21. s16 `d` (d1-d5.c): 21-27. Single pin on v1 or a0 (t7/t8.c): 21.
- do-while around the d chain only, empty do-while, do-while around the +6 load only (e1-e5.c): 2.
- do-while around only the +40 with the +6 load before it (f3.c): 21.
## (f) Method
Counting first said "no missing instruction"; alloc_table.py was the wrong tool (the decider is local-alloc) —
`tools/localalloc_sim.py` + the `.sched` trace settled it. The method has no row for "a variable reuse that the
SCHEDULER needs (anti-dependences) but LOCAL-ALLOC forbids (multi-death)"; that contradiction is what the tree's pins,
the one-pin fallback and the do-while each resolve.
## (g) Structs
Plausibly NO for this decision: typing `e` as an entity struct (`e->x -= d; e->y -= 0x20; e->z += 0x28; …`) produces the
fresh-temp form (the sweep's inline `*(u16 *)(e + K) -= …` spellings: 21+). The aggregate channel (expr.c:4568-4577) only
changes MEMORY dependences; here every load/store pair is already disambiguated by base+offset (memrefs_conflict_p) and
the deciding links are a REGISTER anti-dependence and sched1's loop-note barrier. Not tested with a body-local struct.
@@ -0,0 +1,17 @@
void func_80185354(void *param_1)
{
s32 i;
u8 *e;
u32 vA, vB;
for (i = 0; i < 200; i++) {
e = D_801AED08 + i * 0x2C;
((void (*)(void *, s32))func_80185520)(e, i);
vA = rand() & 0x1F;
*(u32 *)(e + 0x18) = (vA << 2) | ((vA << 0x12) | (vA << 10));
vB = rand() & 7;
*(u32 *)(e + 0x20) = vB;
}
func_8012BF4C((s32 *)param_1, 4);
func_8012AD50(param_1);
}
@@ -0,0 +1,43 @@
# func_80185354 (ov_SC06_000_jr_8017AE2C.c) — e28, P36 T7 S104
**Score: 22 free (sweep best 18) -> 0, zero levers** (the `$16` pin deleted; plain C). Second `--try`.
## (a) Residual
COUNT +4 (43 vs 39) and one extra callee-saved register: the free text walks TWO pointers — `s2` = `e` (the call
argument, stepped `+44`) and `s0` = `e + 32` (`addiu s0,s2,32`, stores at `-8(s0)`/`0(s0)`) — plus an extra `s3` for the
parameter. The target walks ONE pointer `s0` (initialised `lui/addiu` to `D_801AED08`, stepped `+44` in the branch delay
slot) and addresses both stores as `24(s0)`/`32(s0)`.
## (b) Pass and decision
loop.c strength reduction (`strength_reduce`, gcc-2.7.2 loop.c; givs recorded by `record_giv` :4341, merged by
`combine_givs` :5494). In the free text `e` is itself a BIV (`e = e + 0x2C`); its field addresses `e+0x18`/`e+0x20` are
DEST_ADDR givs of that biv, and the pass reduces one of them into a new register (`s0 = e + 32`) while `e` must survive
for the call argument — two stepped registers. The tree's `$16` pin hid `e` from loop.c (a hard register is never a biv),
which is why the pin was NEEDED. In the target, `e` is a GIV of the counter biv `i` (`D_801AED08 + i*0x2C`): the call
argument and both field addresses are givs with the same multiplier, `combine_givs` folds them into ONE reduced register
with constant offsets, and `i` stays as the call's second argument and the exit test (`slti v0,s1,200`).
Proven on bytes (score 0); the `-dL` dump was NOT taken (the prediction from S103 c2/d16 closed on the first spelling).
## (c) Move
Index the array by the loop counter instead of walking a pointer:
`for (i = 0; i < 200; i++) { e = D_801AED08 + i * 0x2C; func_80185520(e, i); … }` (the do-while and `e += 0x2C` deleted).
A do-while with the same indexed `e` (scratch/a3.c) scores 23 — the `for` (rotated loop, loop.c sees the canonical
biv/giv shape) is part of the close; writing each access as `D_801AED08 + (i-1)*0x2C + K` (a2.c) scores 32.
## (d) Generator proposal
When the residual shows two callee-saved registers stepped by the same stride (one `addiu sK,sJ,K` before the loop) and
the tree pins the walked pointer, rewrite `p = BASE; do { … p += S; } while (++i < N)` as
`for (i = 0; i < N; i++) { p = BASE + i * S; … }` — the pointer becomes a giv of the counter and combines into one.
## (e) What did not work
The sweep's R4/R6/R7/R8/R9/R10/R12 moves reach 18-19: none turns the walked pointer into a counter giv.
## (f) Method
S103 c2 / S104 d16-d18 in METHOD_S103.md named this exact class ("a walked destination pointer -> index the array by the
loop counter"); counting first (+4, one extra `$s`) pointed straight at it. Worked as written.
## (g) Structs
Yes, plausibly the ORIGINAL shape: `D_801AED08` is an array of 0x2C-byte records (`struct { u8 pad[0x18]; u32 f18;
u32 pad1c; u32 f20; … } D_801AED08[200];`, see func_801853F0's `+0x1C/+0x20/+0x24` fields) and `&D_801AED08[i]` is
exactly the indexed giv the close needs. A struct type would not change the pass decision beyond what the indexed form
already does (the decisive fact is biv vs giv in loop.c, not aliasing); not tested with a body-local struct.
+185 -185
View File
@@ -1,7 +1,7 @@
{
"head": "6b2e585dd",
"head": "32f641aca",
"stamp": "15956e4a96c4",
"generated": "2026-09-11 04:01",
"generated": "2026-09-11 04:02",
"aliases": [
"main",
"ov_SC03_014",
@@ -21,207 +21,207 @@
"main": {
"objects": 85,
"identical": 85,
"seconds": 8.579999999999998,
"mean_s": 0.101
"seconds": 7.531999999999999,
"mean_s": 0.089
},
"ov_SC03_014": {
"objects": 32,
"identical": 32,
"seconds": 5.74,
"mean_s": 0.179
"seconds": 4.885999999999999,
"mean_s": 0.153
},
"ov_SC03_015": {
"objects": 32,
"identical": 32,
"seconds": 5.2860000000000005,
"mean_s": 0.165
"seconds": 4.711999999999999,
"mean_s": 0.147
},
"ov_SC04_011": {
"objects": 28,
"identical": 28,
"seconds": 4.6899999999999995,
"mean_s": 0.167
"seconds": 4.067,
"mean_s": 0.145
}
},
"per_object_seconds": {
"build/src/800.o": 0.845,
"build/src/800_b.o": 0.107,
"build/src/800_b_2.o": 0.33,
"build/src/800_b_o0a.o": 0.106,
"build/src/800_c.o": 0.235,
"build/src/800b2.o": 0.074,
"build/src/apicard1.o": 0.088,
"build/src/apicard2.o": 0.11,
"build/src/apicard3.o": 0.113,
"build/src/apicard4.o": 0.088,
"build/src/apicard5.o": 0.096,
"build/src/apicard6.o": 0.061,
"build/src/apicard7.o": 0.084,
"build/src/boot.o": 0.142,
"build/src/gap.o": 0.062,
"build/src/libapi1.o": 0.067,
"build/src/libapi2.o": 0.07,
"build/src/libc2_1.o": 0.095,
"build/src/libc2_2.o": 0.065,
"build/src/libcd1.o": 0.093,
"build/src/libcd2.o": 0.066,
"build/src/libetc.o": 0.069,
"build/src/libgpu.o": 0.075,
"build/src/libgpu2.o": 0.079,
"build/src/libgs1.o": 0.09,
"build/src/libgs2.o": 0.067,
"build/src/libgs3.o": 0.077,
"build/src/libgs4.o": 0.07,
"build/src/libgs5.o": 0.073,
"build/src/libgs6.o": 0.085,
"build/src/libgs7.o": 0.109,
"build/src/libgs8.o": 0.061,
"build/src/libgte1.o": 0.07,
"build/src/libgte10.o": 0.107,
"build/src/libgte11.o": 0.084,
"build/src/libgte12.o": 0.074,
"build/src/libgte13.o": 0.113,
"build/src/libgte14.o": 0.068,
"build/src/libgte15.o": 0.074,
"build/src/libgte16.o": 0.075,
"build/src/libgte17.o": 0.085,
"build/src/libgte18.o": 0.083,
"build/src/libgte19.o": 0.079,
"build/src/libgte2.o": 0.064,
"build/src/libgte20.o": 0.115,
"build/src/libgte21.o": 0.099,
"build/src/libgte22.o": 0.07,
"build/src/800.o": 0.815,
"build/src/800_b.o": 0.077,
"build/src/800_b_2.o": 0.301,
"build/src/800_b_o0a.o": 0.081,
"build/src/800_c.o": 0.208,
"build/src/800b2.o": 0.091,
"build/src/apicard1.o": 0.1,
"build/src/apicard2.o": 0.081,
"build/src/apicard3.o": 0.092,
"build/src/apicard4.o": 0.076,
"build/src/apicard5.o": 0.084,
"build/src/apicard6.o": 0.087,
"build/src/apicard7.o": 0.086,
"build/src/boot.o": 0.101,
"build/src/gap.o": 0.076,
"build/src/libapi1.o": 0.087,
"build/src/libapi2.o": 0.059,
"build/src/libc2_1.o": 0.067,
"build/src/libc2_2.o": 0.061,
"build/src/libcd1.o": 0.107,
"build/src/libcd2.o": 0.091,
"build/src/libetc.o": 0.072,
"build/src/libgpu.o": 0.077,
"build/src/libgpu2.o": 0.08,
"build/src/libgs1.o": 0.059,
"build/src/libgs2.o": 0.079,
"build/src/libgs3.o": 0.069,
"build/src/libgs4.o": 0.072,
"build/src/libgs5.o": 0.062,
"build/src/libgs6.o": 0.084,
"build/src/libgs7.o": 0.068,
"build/src/libgs8.o": 0.058,
"build/src/libgte1.o": 0.065,
"build/src/libgte10.o": 0.061,
"build/src/libgte11.o": 0.061,
"build/src/libgte12.o": 0.07,
"build/src/libgte13.o": 0.064,
"build/src/libgte14.o": 0.079,
"build/src/libgte15.o": 0.068,
"build/src/libgte16.o": 0.067,
"build/src/libgte17.o": 0.069,
"build/src/libgte18.o": 0.076,
"build/src/libgte19.o": 0.101,
"build/src/libgte2.o": 0.068,
"build/src/libgte20.o": 0.069,
"build/src/libgte21.o": 0.098,
"build/src/libgte22.o": 0.065,
"build/src/libgte23.o": 0.071,
"build/src/libgte24.o": 0.07,
"build/src/libgte25.o": 0.084,
"build/src/libgte26.o": 0.072,
"build/src/libgte27.o": 0.087,
"build/src/libgte28.o": 0.075,
"build/src/libgte29.o": 0.075,
"build/src/libgte3.o": 0.081,
"build/src/libgte30.o": 0.085,
"build/src/libgte4.o": 0.082,
"build/src/libgte5.o": 0.075,
"build/src/libgte6.o": 0.109,
"build/src/libgte7.o": 0.084,
"build/src/libgte8.o": 0.087,
"build/src/libgte9.o": 0.094,
"build/src/libmcrd1.o": 0.119,
"build/src/libmcrd2.o": 0.09,
"build/src/libpad1.o": 0.084,
"build/src/libpad2.o": 0.129,
"build/src/sgap.o": 0.085,
"build/src/sgap_2.o": 0.098,
"build/src/sgap_3.o": 0.087,
"build/src/sgap_4.o": 0.137,
"build/src/sgap_5.o": 0.065,
"build/src/sgap_6.o": 0.111,
"build/src/sgap_8.o": 0.09,
"build/src/libgte24.o": 0.066,
"build/src/libgte25.o": 0.071,
"build/src/libgte26.o": 0.065,
"build/src/libgte27.o": 0.066,
"build/src/libgte28.o": 0.068,
"build/src/libgte29.o": 0.071,
"build/src/libgte3.o": 0.058,
"build/src/libgte30.o": 0.068,
"build/src/libgte4.o": 0.07,
"build/src/libgte5.o": 0.073,
"build/src/libgte6.o": 0.074,
"build/src/libgte7.o": 0.069,
"build/src/libgte8.o": 0.084,
"build/src/libgte9.o": 0.065,
"build/src/libmcrd1.o": 0.095,
"build/src/libmcrd2.o": 0.092,
"build/src/libpad1.o": 0.066,
"build/src/libpad2.o": 0.081,
"build/src/sgap.o": 0.094,
"build/src/sgap_2.o": 0.129,
"build/src/sgap_3.o": 0.071,
"build/src/sgap_4.o": 0.082,
"build/src/sgap_5.o": 0.062,
"build/src/sgap_6.o": 0.061,
"build/src/sgap_8.o": 0.073,
"build/src/snd1.o": 0.083,
"build/src/snd10.o": 0.074,
"build/src/snd11.o": 0.095,
"build/src/snd12.o": 0.111,
"build/src/snd2.o": 0.097,
"build/src/snd3.o": 0.084,
"build/src/snd4.o": 0.077,
"build/src/snd5.o": 0.104,
"build/src/snd6.o": 0.113,
"build/src/snd7.o": 0.1,
"build/src/snd8.o": 0.125,
"build/src/snd9.o": 0.079,
"build/src/ov_SC03_014/ov_SC03_014.o": 0.156,
"build/src/ov_SC03_014/ov_SC03_014_after.o": 0.573,
"build/src/ov_SC03_014/ov_SC03_014_jr_8012ACE0.o": 0.454,
"build/src/ov_SC03_014/ov_SC03_014_jr_80135888.o": 0.08,
"build/src/ov_SC03_014/ov_SC03_014_jr_80135A4C.o": 0.103,
"build/src/ov_SC03_014/ov_SC03_014_jr_80135D20.o": 0.162,
"build/src/ov_SC03_014/ov_SC03_014_jr_801380E0.o": 0.213,
"build/src/ov_SC03_014/ov_SC03_014_jr_8013C98C.o": 0.131,
"build/src/ov_SC03_014/ov_SC03_014_jr_8013F350.o": 0.113,
"build/src/ov_SC03_014/ov_SC03_014_jr_8013FFD8.o": 0.091,
"build/src/ov_SC03_014/ov_SC03_014_jr_80140608.o": 0.239,
"build/src/ov_SC03_014/ov_SC03_014_jr_8015444C.o": 0.095,
"build/src/ov_SC03_014/ov_SC03_014_jr_80154C24.o": 0.198,
"build/src/ov_SC03_014/ov_SC03_014_jr_801588CC.o": 0.103,
"build/src/ov_SC03_014/ov_SC03_014_jr_80159C84.o": 0.098,
"build/src/ov_SC03_014/ov_SC03_014_jr_8015A3C8.o": 0.109,
"build/src/ov_SC03_014/ov_SC03_014_jr_8015AE2C.o": 0.106,
"build/src/ov_SC03_014/ov_SC03_014_jr_8015C32C.o": 0.547,
"build/src/ov_SC03_014/ov_SC03_014_jr_8016AB6C.o": 0.309,
"build/src/ov_SC03_014/ov_SC03_014_jr_80171B4C.o": 0.119,
"build/src/ov_SC03_014/ov_SC03_014_jr_801734BC.o": 0.231,
"build/src/ov_SC03_014/ov_SC03_014_jr_801789AC.o": 0.08,
"build/src/ov_SC03_014/ov_SC03_014_jr_80178D40.o": 0.143,
"build/src/ov_SC03_014/ov_SC03_014_jr_8017A4AC.o": 0.08,
"build/src/ov_SC03_014/ov_SC03_014_jr_8017AE2C.o": 0.207,
"build/src/ov_SC03_014/ov_SC03_014_jr_8017EB7C.o": 0.242,
"build/src/ov_SC03_014/ov_SC03_014_jr_80184440.o": 0.058,
"build/src/ov_SC03_014/ov_SC03_014_jr_801848E4.o": 0.329,
"build/src/ov_SC03_014/ov_SC03_014_o0b.o": 0.086,
"build/src/ov_SC03_014/ov_SC03_014_o0c.o": 0.091,
"build/src/ov_SC03_014/ov_SC03_014_o0d.o": 0.096,
"build/src/ov_SC03_014/ov_SC03_014_o0e.o": 0.098,
"build/src/ov_SC03_015/ov_SC03_015.o": 0.148,
"build/src/ov_SC03_015/ov_SC03_015_after.o": 0.574,
"build/src/ov_SC03_015/ov_SC03_015_jr_8012ACE0.o": 0.432,
"build/src/ov_SC03_015/ov_SC03_015_jr_80135888.o": 0.055,
"build/src/snd10.o": 0.063,
"build/src/snd11.o": 0.073,
"build/src/snd12.o": 0.072,
"build/src/snd2.o": 0.07,
"build/src/snd3.o": 0.077,
"build/src/snd4.o": 0.069,
"build/src/snd5.o": 0.097,
"build/src/snd6.o": 0.076,
"build/src/snd7.o": 0.076,
"build/src/snd8.o": 0.072,
"build/src/snd9.o": 0.07,
"build/src/ov_SC03_014/ov_SC03_014.o": 0.138,
"build/src/ov_SC03_014/ov_SC03_014_after.o": 0.536,
"build/src/ov_SC03_014/ov_SC03_014_jr_8012ACE0.o": 0.404,
"build/src/ov_SC03_014/ov_SC03_014_jr_80135888.o": 0.057,
"build/src/ov_SC03_014/ov_SC03_014_jr_80135A4C.o": 0.057,
"build/src/ov_SC03_014/ov_SC03_014_jr_80135D20.o": 0.152,
"build/src/ov_SC03_014/ov_SC03_014_jr_801380E0.o": 0.16,
"build/src/ov_SC03_014/ov_SC03_014_jr_8013C98C.o": 0.143,
"build/src/ov_SC03_014/ov_SC03_014_jr_8013F350.o": 0.072,
"build/src/ov_SC03_014/ov_SC03_014_jr_8013FFD8.o": 0.065,
"build/src/ov_SC03_014/ov_SC03_014_jr_80140608.o": 0.178,
"build/src/ov_SC03_014/ov_SC03_014_jr_8015444C.o": 0.101,
"build/src/ov_SC03_014/ov_SC03_014_jr_80154C24.o": 0.174,
"build/src/ov_SC03_014/ov_SC03_014_jr_801588CC.o": 0.099,
"build/src/ov_SC03_014/ov_SC03_014_jr_80159C84.o": 0.07,
"build/src/ov_SC03_014/ov_SC03_014_jr_8015A3C8.o": 0.073,
"build/src/ov_SC03_014/ov_SC03_014_jr_8015AE2C.o": 0.097,
"build/src/ov_SC03_014/ov_SC03_014_jr_8015C32C.o": 0.489,
"build/src/ov_SC03_014/ov_SC03_014_jr_8016AB6C.o": 0.281,
"build/src/ov_SC03_014/ov_SC03_014_jr_80171B4C.o": 0.097,
"build/src/ov_SC03_014/ov_SC03_014_jr_801734BC.o": 0.211,
"build/src/ov_SC03_014/ov_SC03_014_jr_801789AC.o": 0.058,
"build/src/ov_SC03_014/ov_SC03_014_jr_80178D40.o": 0.113,
"build/src/ov_SC03_014/ov_SC03_014_jr_8017A4AC.o": 0.077,
"build/src/ov_SC03_014/ov_SC03_014_jr_8017AE2C.o": 0.181,
"build/src/ov_SC03_014/ov_SC03_014_jr_8017EB7C.o": 0.217,
"build/src/ov_SC03_014/ov_SC03_014_jr_80184440.o": 0.053,
"build/src/ov_SC03_014/ov_SC03_014_jr_801848E4.o": 0.282,
"build/src/ov_SC03_014/ov_SC03_014_o0b.o": 0.064,
"build/src/ov_SC03_014/ov_SC03_014_o0c.o": 0.062,
"build/src/ov_SC03_014/ov_SC03_014_o0d.o": 0.057,
"build/src/ov_SC03_014/ov_SC03_014_o0e.o": 0.068,
"build/src/ov_SC03_015/ov_SC03_015.o": 0.129,
"build/src/ov_SC03_015/ov_SC03_015_after.o": 0.521,
"build/src/ov_SC03_015/ov_SC03_015_jr_8012ACE0.o": 0.394,
"build/src/ov_SC03_015/ov_SC03_015_jr_80135888.o": 0.05,
"build/src/ov_SC03_015/ov_SC03_015_jr_80135A4C.o": 0.053,
"build/src/ov_SC03_015/ov_SC03_015_jr_80135D20.o": 0.129,
"build/src/ov_SC03_015/ov_SC03_015_jr_801380E0.o": 0.179,
"build/src/ov_SC03_015/ov_SC03_015_jr_8013C98C.o": 0.14,
"build/src/ov_SC03_015/ov_SC03_015_jr_8013F350.o": 0.089,
"build/src/ov_SC03_015/ov_SC03_015_jr_8013FFD8.o": 0.077,
"build/src/ov_SC03_015/ov_SC03_015_jr_80140608.o": 0.195,
"build/src/ov_SC03_015/ov_SC03_015_jr_8015444C.o": 0.094,
"build/src/ov_SC03_015/ov_SC03_015_jr_80154C24.o": 0.187,
"build/src/ov_SC03_015/ov_SC03_015_jr_801588CC.o": 0.104,
"build/src/ov_SC03_015/ov_SC03_015_jr_80159C84.o": 0.072,
"build/src/ov_SC03_015/ov_SC03_015_jr_8015A3C8.o": 0.084,
"build/src/ov_SC03_015/ov_SC03_015_jr_8015AE2C.o": 0.121,
"build/src/ov_SC03_015/ov_SC03_015_jr_8015C32C.o": 0.521,
"build/src/ov_SC03_015/ov_SC03_015_jr_8016AB6C.o": 0.301,
"build/src/ov_SC03_015/ov_SC03_015_jr_80171B4C.o": 0.114,
"build/src/ov_SC03_015/ov_SC03_015_jr_801734BC.o": 0.23,
"build/src/ov_SC03_015/ov_SC03_015_jr_801789AC.o": 0.07,
"build/src/ov_SC03_015/ov_SC03_015_jr_80178D40.o": 0.11,
"build/src/ov_SC03_015/ov_SC03_015_jr_8017A4AC.o": 0.097,
"build/src/ov_SC03_015/ov_SC03_015_jr_8017AE2C.o": 0.192,
"build/src/ov_SC03_015/ov_SC03_015_jr_8017EB7C.o": 0.221,
"build/src/ov_SC03_015/ov_SC03_015_jr_80184440.o": 0.075,
"build/src/ov_SC03_015/ov_SC03_015_jr_801848E4.o": 0.313,
"build/src/ov_SC03_015/ov_SC03_015_o0b.o": 0.074,
"build/src/ov_SC03_015/ov_SC03_015_o0c.o": 0.074,
"build/src/ov_SC03_015/ov_SC03_015_o0d.o": 0.072,
"build/src/ov_SC03_015/ov_SC03_015_o0e.o": 0.089,
"build/src/ov_SC04_011/ov_SC04_011.o": 0.151,
"build/src/ov_SC04_011/ov_SC04_011_after.o": 0.481,
"build/src/ov_SC04_011/ov_SC04_011_jr_8012ACE0.o": 0.416,
"build/src/ov_SC04_011/ov_SC04_011_jr_80135888.o": 0.084,
"build/src/ov_SC04_011/ov_SC04_011_jr_80135A4C.o": 0.058,
"build/src/ov_SC04_011/ov_SC04_011_jr_80135D20.o": 0.141,
"build/src/ov_SC04_011/ov_SC04_011_jr_801380E0.o": 0.176,
"build/src/ov_SC04_011/ov_SC04_011_jr_8013C98C.o": 0.127,
"build/src/ov_SC04_011/ov_SC04_011_jr_8013F350.o": 0.098,
"build/src/ov_SC04_011/ov_SC04_011_jr_8013FFD8.o": 0.076,
"build/src/ov_SC04_011/ov_SC04_011_jr_80140608.o": 0.217,
"build/src/ov_SC04_011/ov_SC04_011_jr_8015444C.o": 0.091,
"build/src/ov_SC04_011/ov_SC04_011_jr_80154C24.o": 0.182,
"build/src/ov_SC04_011/ov_SC04_011_jr_801588CC.o": 0.105,
"build/src/ov_SC04_011/ov_SC04_011_jr_80159C84.o": 0.094,
"build/src/ov_SC04_011/ov_SC04_011_jr_8015A3C8.o": 0.084,
"build/src/ov_SC04_011/ov_SC04_011_jr_8015AE2C.o": 0.105,
"build/src/ov_SC04_011/ov_SC04_011_jr_8015C32C.o": 0.405,
"build/src/ov_SC04_011/ov_SC04_011_jr_8016AB6C.o": 0.237,
"build/src/ov_SC04_011/ov_SC04_011_jr_80171B4C.o": 0.113,
"build/src/ov_SC04_011/ov_SC04_011_jr_801734BC.o": 0.198,
"build/src/ov_SC04_011/ov_SC04_011_jr_801789AC.o": 0.075,
"build/src/ov_SC04_011/ov_SC04_011_jr_80178D40.o": 0.114,
"build/src/ov_SC04_011/ov_SC04_011_jr_8017A4AC.o": 0.077,
"build/src/ov_SC04_011/ov_SC04_011_jr_8017AE2C.o": 0.119,
"build/src/ov_SC04_011/ov_SC04_011_jr_8017D494.o": 0.503,
"build/src/ov_SC04_011/ov_SC04_011_o0b.o": 0.087,
"build/src/ov_SC04_011/ov_SC04_011_o0c.o": 0.076
"build/src/ov_SC03_015/ov_SC03_015_jr_80135D20.o": 0.122,
"build/src/ov_SC03_015/ov_SC03_015_jr_801380E0.o": 0.155,
"build/src/ov_SC03_015/ov_SC03_015_jr_8013C98C.o": 0.115,
"build/src/ov_SC03_015/ov_SC03_015_jr_8013F350.o": 0.086,
"build/src/ov_SC03_015/ov_SC03_015_jr_8013FFD8.o": 0.071,
"build/src/ov_SC03_015/ov_SC03_015_jr_80140608.o": 0.181,
"build/src/ov_SC03_015/ov_SC03_015_jr_8015444C.o": 0.079,
"build/src/ov_SC03_015/ov_SC03_015_jr_80154C24.o": 0.177,
"build/src/ov_SC03_015/ov_SC03_015_jr_801588CC.o": 0.095,
"build/src/ov_SC03_015/ov_SC03_015_jr_80159C84.o": 0.07,
"build/src/ov_SC03_015/ov_SC03_015_jr_8015A3C8.o": 0.073,
"build/src/ov_SC03_015/ov_SC03_015_jr_8015AE2C.o": 0.097,
"build/src/ov_SC03_015/ov_SC03_015_jr_8015C32C.o": 0.461,
"build/src/ov_SC03_015/ov_SC03_015_jr_8016AB6C.o": 0.287,
"build/src/ov_SC03_015/ov_SC03_015_jr_80171B4C.o": 0.108,
"build/src/ov_SC03_015/ov_SC03_015_jr_801734BC.o": 0.196,
"build/src/ov_SC03_015/ov_SC03_015_jr_801789AC.o": 0.058,
"build/src/ov_SC03_015/ov_SC03_015_jr_80178D40.o": 0.087,
"build/src/ov_SC03_015/ov_SC03_015_jr_8017A4AC.o": 0.071,
"build/src/ov_SC03_015/ov_SC03_015_jr_8017AE2C.o": 0.18,
"build/src/ov_SC03_015/ov_SC03_015_jr_8017EB7C.o": 0.207,
"build/src/ov_SC03_015/ov_SC03_015_jr_80184440.o": 0.053,
"build/src/ov_SC03_015/ov_SC03_015_jr_801848E4.o": 0.276,
"build/src/ov_SC03_015/ov_SC03_015_o0b.o": 0.063,
"build/src/ov_SC03_015/ov_SC03_015_o0c.o": 0.066,
"build/src/ov_SC03_015/ov_SC03_015_o0d.o": 0.06,
"build/src/ov_SC03_015/ov_SC03_015_o0e.o": 0.071,
"build/src/ov_SC04_011/ov_SC04_011.o": 0.123,
"build/src/ov_SC04_011/ov_SC04_011_after.o": 0.429,
"build/src/ov_SC04_011/ov_SC04_011_jr_8012ACE0.o": 0.364,
"build/src/ov_SC04_011/ov_SC04_011_jr_80135888.o": 0.05,
"build/src/ov_SC04_011/ov_SC04_011_jr_80135A4C.o": 0.055,
"build/src/ov_SC04_011/ov_SC04_011_jr_80135D20.o": 0.122,
"build/src/ov_SC04_011/ov_SC04_011_jr_801380E0.o": 0.163,
"build/src/ov_SC04_011/ov_SC04_011_jr_8013C98C.o": 0.115,
"build/src/ov_SC04_011/ov_SC04_011_jr_8013F350.o": 0.073,
"build/src/ov_SC04_011/ov_SC04_011_jr_8013FFD8.o": 0.064,
"build/src/ov_SC04_011/ov_SC04_011_jr_80140608.o": 0.188,
"build/src/ov_SC04_011/ov_SC04_011_jr_8015444C.o": 0.085,
"build/src/ov_SC04_011/ov_SC04_011_jr_80154C24.o": 0.163,
"build/src/ov_SC04_011/ov_SC04_011_jr_801588CC.o": 0.093,
"build/src/ov_SC04_011/ov_SC04_011_jr_80159C84.o": 0.071,
"build/src/ov_SC04_011/ov_SC04_011_jr_8015A3C8.o": 0.071,
"build/src/ov_SC04_011/ov_SC04_011_jr_8015AE2C.o": 0.094,
"build/src/ov_SC04_011/ov_SC04_011_jr_8015C32C.o": 0.353,
"build/src/ov_SC04_011/ov_SC04_011_jr_8016AB6C.o": 0.233,
"build/src/ov_SC04_011/ov_SC04_011_jr_80171B4C.o": 0.097,
"build/src/ov_SC04_011/ov_SC04_011_jr_801734BC.o": 0.172,
"build/src/ov_SC04_011/ov_SC04_011_jr_801789AC.o": 0.058,
"build/src/ov_SC04_011/ov_SC04_011_jr_80178D40.o": 0.088,
"build/src/ov_SC04_011/ov_SC04_011_jr_8017A4AC.o": 0.07,
"build/src/ov_SC04_011/ov_SC04_011_jr_8017AE2C.o": 0.118,
"build/src/ov_SC04_011/ov_SC04_011_jr_8017D494.o": 0.452,
"build/src/ov_SC04_011/ov_SC04_011_o0b.o": 0.047,
"build/src/ov_SC04_011/ov_SC04_011_o0c.o": 0.056
},
"ok": true,
"seconds": 2.7
"seconds": 2.4
}
File diff suppressed because one or more lines are too long
+8 -6
View File
@@ -5117,16 +5117,18 @@ void func_8017E764(s32 param_1)
}
ent = D_801202A0;
i = 0;
pos[0] = *(u16 *)(param_1 + 6);
{
register u16 *vp __asm__("$3"); // !FAKE: pin $3 — NEEDED DIFFERS (P36 rung B tus7)
u16 t;
vp = (u16 *)&D_80126B66;
t = *(volatile u16 *)vp;
u16 tt;
/* `val` is the loop's table-value temp reused here (both live in $v1 in the target): its second set keeps
* this address load from sched1's birthing priority (sched.c:2469-2545), so it is scheduled above the pos[0]
* store and global-alloc gives it $v1 (P36 S104 e26). */
val = (s32)&D_80126B66;
tt = *(volatile u16 *)val; /* a byte-needed volatile, kept as ordinary C (gate-1 decision 3): a volatile MEM fails combine's recog (combine.c:491 init_recog_no_volatile, recog.c:807), so the address stays in a register (la + lhu 0(reg)) (P36 S104 e26 minimum-lever) */
pos[1] = *(u16 *)(param_1 + 0xA);
pos[2] = t + 8;
pos[2] = tt + 8;
}
i = 0;
np = (u16 *)D_801152A8;
t = (ratan2(*(s16 *)&D_80126B62 - *(s16 *)(param_1 + 0xA),
*(s16 *)&D_80126B5E - *(s16 *)(param_1 + 6)) - 0x400) & 0xFFF;