From 6c2a3bb43b7fd8093ce781f73cb2e0c6c22811cf Mon Sep 17 00:00:00 2001 From: Drew T <50529377+Druthulu@users.noreply.github.com> Date: Wed, 2 Sep 2026 15:24:29 -0600 Subject: [PATCH] =?UTF-8?q?docs(cookbook):=20=C2=A7432=20=E2=80=94=20defea?= =?UTF-8?q?t=20cse's=20mask=20merge=20with=20a=20shift=20pair=20that=20com?= =?UTF-8?q?bine=20folds=20back?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit From main/func_8002DC68 (MATCH 198/198). The target masks one value twice (a compare, plus a second andi that reorg steals for a beqz delay slot). Written as param_2 & 0x7F on both sides, cse merges the two (and:SI) and the delay slot comes out EMPTY. Spelling ONE as (param_2 << 25) >> 25 hides it from cse — different RTX — and combine's simplify_shift_const folds it back to andi. Two masks in the RTL, one instruction each out. Byte-verified on either side. The inverse of the usual advice: normally you make two expressions identical so cse merges them; here you make them different to cse and identical to combine, exploiting pass order. Any x & ((1< **Generated by `tools/cookbook_index.py` — do not hand-edit** (R33). Regenerate after adding a cookbook section. > -> `docs/matching-cookbook.md` is ~716 KB / 1099 sections. Grepping it blind is how three P30 wave-1 agents each "discovered" an idiom that was already written down. **Start here, then read the section.** A section appears under every symptom it addresses. +> `docs/matching-cookbook.md` is ~716 KB / 1100 sections. Grepping it blind is how three P30 wave-1 agents each "discovered" an idiom that was already written down. **Start here, then read the section.** A section appears under every symptom it addresses. **How to use:** name what you SEE in the diff (a stolen delay slot, an extra `la`, a swapped register pair, a `conflicting types` error), find that symptom below, read those sections first. If nothing fits, THEN grind — and add a section when you win. @@ -329,7 +329,7 @@ - **§417** — ★★★ — A REGISTER PIN CAN BLOCK `jump.c`'s SELECT COLLAPSE, AND UNPINNING THEN EXPOSES A `cse` SKIP-BLOCKS MERGE (P31 S71; byte-proven `ov_SC03_013/func_8017E6F4`, 182 ins) L33605 - **§419** — ★★★ — WHEN A PIN IS IMPOSSIBLE, WIN THE local-alloc DENSITY CONTEST INSTEAD (P31 S71; byte-proven `ov_SC01_000/func_8017DD04`, 297 ins) L33658 -### CSE / redundancy / rematerialization (45) +### CSE / redundancy / rematerialization (46) - **§46** — The `func_80178D40` crack (890 ins ×134, the heaviest core in the game): four LOOP-STRUCTURE levers cheap-Opus found by reading loop.c/jump.c/cse.c (Phase 26 session 8, 2026-07-13) L3347 - **§83d** — CSE's quantity budget is WHOLE-FUNCTION, so a local rewrite cannot fix a local symptom L6486 @@ -376,6 +376,7 @@ - **§3-1.** — DEAD-RESET CSE-BREAKER — the zero-footprint replacement for a §195-I asm re-tie L31979 - **§385** — ★★★ — THE **SCHED2 PRIORITY-DONOR ASM**: closing the "hoisted-invariant vs IV-init preheader swap" class (P31 S69; byte-proven main/func_80038A58, 347 ins, fable escalation 2 → 0) L32309 - **§417** — ★★★ — A REGISTER PIN CAN BLOCK `jump.c`'s SELECT COLLAPSE, AND UNPINNING THEN EXPOSES A `cse` SKIP-BLOCKS MERGE (P31 S71; byte-proven `ov_SC03_013/func_8017E6F4`, 182 ins) L33605 +- **§432** — ★★★ — DEFEAT cse's MERGE OF TWO IDENTICAL MASKS BY SPELLING ONE AS A SHIFT PAIR (P31 S72/S73; `main/func_8002DC68`, MATCH 198/198) L34128 ### loops & induction variables (48) @@ -2793,6 +2794,7 @@ - **§428a** — ★★★ — TWO RESIDUALS THAT MOVE IN OPPOSITE DIRECTIONS UNDER EVERY LEVER USUALLY SHARE ONE CAUSE (P31 S72; `main/func_8001B0D4`, NEAR/53 -> MATCH; **my first answer here was WRONG and is kept below as the refutation**) L33992 - **§430** — ★★★ — A SHARED TAIL IS A LATE CROSS-JUMP MERGE, NOT A SOURCE `goto` — AND SPELLING IT AS ONE CAN INVALIDATE A LOOP (P31 S72; `main/CdReadSectorReadyCB`, 422/424 ins, verified with `cc1 -dL`) L34027 - **§431** — ★★★ — SPLITTING A 27,000-LINE TU AT ITS ORIGINAL BOUNDARIES: THE JTBL SPANS TELL YOU WHERE, AND THE COMPILER TELLS YOU WHAT CROSSES (P31 S72; `src/800.c` -> `800.c`/`800_b.c`/`800_c.c`, byte-identical with nothing banked) L34065 +- **§432** — ★★★ — DEFEAT cse's MERGE OF TWO IDENTICAL MASKS BY SPELLING ONE AS A SHIFT PAIR (P31 S72/S73; `main/func_8002DC68`, MATCH 198/198) L34128 --- @@ -3904,3 +3906,4 @@ Notes routinely quote that as a section id. This table resolves it. Grep bait: ` | L33992 | §428a | ★★★ — TWO RESIDUALS THAT MOVE IN OPPOSITE DIRECTIONS UNDER EVERY LEVER USUALLY SHARE ONE C | | L34027 | §430 | ★★★ — A SHARED TAIL IS A LATE CROSS-JUMP MERGE, NOT A SOURCE `goto` — AND SPELLING IT AS O | | L34065 | §431 | ★★★ — SPLITTING A 27,000-LINE TU AT ITS ORIGINAL BOUNDARIES: THE JTBL SPANS TELL YOU WHERE | +| L34128 | §432 | ★★★ — DEFEAT cse's MERGE OF TWO IDENTICAL MASKS BY SPELLING ONE AS A SHIFT PAIR (P31 S72/S | diff --git a/docs/matching-cookbook.md b/docs/matching-cookbook.md index 11fbbfed9..fadfede08 100644 --- a/docs/matching-cookbook.md +++ b/docs/matching-cookbook.md @@ -34124,3 +34124,45 @@ caught this, and the project already mandates it for exactly this reason. **The payoff, measured:** the carve + split banked **14 main functions** in one session, including **10 of the 11** the previous session had recorded as "PROVEN gate-rejects". + +## §432 ★★★ — DEFEAT cse's MERGE OF TWO IDENTICAL MASKS BY SPELLING ONE AS A SHIFT PAIR (P31 S72/S73; `main/func_8002DC68`, MATCH 198/198) + +**The shape.** The target masks the same value TWICE — a compare into `$v1` and a second `andi` into +`$v0` that `reorg` then steals for a `beqz` delay slot: + +``` +andi $v1, $s2, 0x7F # the compare +andi $v0, $s2, 0x7F # a SECOND, identical mask -> reorg fills the branch delay slot with it +``` + +Written the obvious way — `param_2 & 0x7F` on both sides — **cse merges the two `(and:SI)` RTXs**, +there is only one mask left, and the delay slot comes out EMPTY. No amount of statement reordering +recovers it, because the two expressions really are identical and cse is right. + +**The lever.** Spell ONE of them as a shift pair: + +```c +(param_2 << 25) >> 25 /* instead of param_2 & 0x7F */ +``` + +cse sees `(lshiftrt (ashift ...))`, not `(and ...)`, so it does not merge — and then **combine's +`simplify_shift_const` folds the pair straight back to `andi 0x7F`**. You get two masks in the RTL +and one instruction each in the output. Byte-verified with the shift on EITHER side, so use whichever +reads more naturally. + +**Why it generalises.** This is the inverse of the usual advice. Normally you fight gcc by making two +expressions *identical* so cse merges them; here you must make them *textually different in RTL but +identical after combine*. Any `x & ((1<