docs(cookbook): §432 — defeat cse's mask merge with a shift pair that combine folds back

From main/func_8002DC68 (MATCH 198/198). The target masks one value twice (a compare,
plus a second andi that reorg steals for a beqz delay slot). Written as param_2 & 0x7F on
both sides, cse merges the two (and:SI) and the delay slot comes out EMPTY. Spelling ONE
as (param_2 << 25) >> 25 hides it from cse — different RTX — and combine's
simplify_shift_const folds it back to andi. Two masks in the RTL, one instruction each
out. Byte-verified on either side.

The inverse of the usual advice: normally you make two expressions identical so cse
merges them; here you make them different to cse and identical to combine, exploiting
pass order. Any x & ((1<<n)-1) has a shift-pair twin with this property.
This commit is contained in:
Drew T
2026-09-02 15:24:29 -06:00
parent 5b6a2e86dd
commit 6c2a3bb43b
2 changed files with 47 additions and 2 deletions
+5 -2
View File
@@ -2,7 +2,7 @@
> **Generated by `tools/cookbook_index.py` — do not hand-edit** (R33). Regenerate after adding a cookbook section.
>
> `docs/matching-cookbook.md` is ~716 KB / 1099 sections. Grepping it blind is how three P30 wave-1 agents each "discovered" an idiom that was already written down. **Start here, then read the section.** A section appears under every symptom it addresses.
> `docs/matching-cookbook.md` is ~716 KB / 1100 sections. Grepping it blind is how three P30 wave-1 agents each "discovered" an idiom that was already written down. **Start here, then read the section.** A section appears under every symptom it addresses.
**How to use:** name what you SEE in the diff (a stolen delay slot, an extra `la`, a swapped register pair, a `conflicting types` error), find that symptom below, read those sections first. If nothing fits, THEN grind — and add a section when you win.
@@ -329,7 +329,7 @@
- **§417** — ★★★ — A REGISTER PIN CAN BLOCK `jump.c`'s SELECT COLLAPSE, AND UNPINNING THEN EXPOSES A `cse` SKIP-BLOCKS MERGE (P31 S71; byte-proven `ov_SC03_013/func_8017E6F4`, 182 ins) <sub>L33605</sub>
- **§419** — ★★★ — WHEN A PIN IS IMPOSSIBLE, WIN THE local-alloc DENSITY CONTEST INSTEAD (P31 S71; byte-proven `ov_SC01_000/func_8017DD04`, 297 ins) <sub>L33658</sub>
### CSE / redundancy / rematerialization (45)
### CSE / redundancy / rematerialization (46)
- **§46** — The `func_80178D40` crack (890 ins ×134, the heaviest core in the game): four LOOP-STRUCTURE levers cheap-Opus found by reading loop.c/jump.c/cse.c (Phase 26 session 8, 2026-07-13) <sub>L3347</sub>
- **§83d** — CSE's quantity budget is WHOLE-FUNCTION, so a local rewrite cannot fix a local symptom <sub>L6486</sub>
@@ -376,6 +376,7 @@
- **§3-1.** — DEAD-RESET CSE-BREAKER — the zero-footprint replacement for a §195-I asm re-tie <sub>L31979</sub>
- **§385** — ★★★ — THE **SCHED2 PRIORITY-DONOR ASM**: closing the "hoisted-invariant vs IV-init preheader swap" class (P31 S69; byte-proven main/func_80038A58, 347 ins, fable escalation 2 → 0) <sub>L32309</sub>
- **§417** — ★★★ — A REGISTER PIN CAN BLOCK `jump.c`'s SELECT COLLAPSE, AND UNPINNING THEN EXPOSES A `cse` SKIP-BLOCKS MERGE (P31 S71; byte-proven `ov_SC03_013/func_8017E6F4`, 182 ins) <sub>L33605</sub>
- **§432** — ★★★ — DEFEAT cse's MERGE OF TWO IDENTICAL MASKS BY SPELLING ONE AS A SHIFT PAIR (P31 S72/S73; `main/func_8002DC68`, MATCH 198/198) <sub>L34128</sub>
### loops & induction variables (48)
@@ -2793,6 +2794,7 @@
- **§428a** — ★★★ — TWO RESIDUALS THAT MOVE IN OPPOSITE DIRECTIONS UNDER EVERY LEVER USUALLY SHARE ONE CAUSE (P31 S72; `main/func_8001B0D4`, NEAR/53 -> MATCH; **my first answer here was WRONG and is kept below as the refutation**) <sub>L33992</sub>
- **§430** — ★★★ — A SHARED TAIL IS A LATE CROSS-JUMP MERGE, NOT A SOURCE `goto` — AND SPELLING IT AS ONE CAN INVALIDATE A LOOP (P31 S72; `main/CdReadSectorReadyCB`, 422/424 ins, verified with `cc1 -dL`) <sub>L34027</sub>
- **§431** — ★★★ — SPLITTING A 27,000-LINE TU AT ITS ORIGINAL BOUNDARIES: THE JTBL SPANS TELL YOU WHERE, AND THE COMPILER TELLS YOU WHAT CROSSES (P31 S72; `src/800.c` -> `800.c`/`800_b.c`/`800_c.c`, byte-identical with nothing banked) <sub>L34065</sub>
- **§432** — ★★★ — DEFEAT cse's MERGE OF TWO IDENTICAL MASKS BY SPELLING ONE AS A SHIFT PAIR (P31 S72/S73; `main/func_8002DC68`, MATCH 198/198) <sub>L34128</sub>
---
@@ -3904,3 +3906,4 @@ Notes routinely quote that as a section id. This table resolves it. Grep bait: `
| L33992 | §428a | ★★★ — TWO RESIDUALS THAT MOVE IN OPPOSITE DIRECTIONS UNDER EVERY LEVER USUALLY SHARE ONE C |
| L34027 | §430 | ★★★ — A SHARED TAIL IS A LATE CROSS-JUMP MERGE, NOT A SOURCE `goto` — AND SPELLING IT AS O |
| L34065 | §431 | ★★★ — SPLITTING A 27,000-LINE TU AT ITS ORIGINAL BOUNDARIES: THE JTBL SPANS TELL YOU WHERE |
| L34128 | §432 | ★★★ — DEFEAT cse's MERGE OF TWO IDENTICAL MASKS BY SPELLING ONE AS A SHIFT PAIR (P31 S72/S |
+42
View File
@@ -34124,3 +34124,45 @@ caught this, and the project already mandates it for exactly this reason.
**The payoff, measured:** the carve + split banked **14 main functions** in one session, including
**10 of the 11** the previous session had recorded as "PROVEN gate-rejects".
## §432 ★★★ — DEFEAT cse's MERGE OF TWO IDENTICAL MASKS BY SPELLING ONE AS A SHIFT PAIR (P31 S72/S73; `main/func_8002DC68`, MATCH 198/198)
**The shape.** The target masks the same value TWICE — a compare into `$v1` and a second `andi` into
`$v0` that `reorg` then steals for a `beqz` delay slot:
```
andi $v1, $s2, 0x7F # the compare
andi $v0, $s2, 0x7F # a SECOND, identical mask -> reorg fills the branch delay slot with it
```
Written the obvious way — `param_2 & 0x7F` on both sides — **cse merges the two `(and:SI)` RTXs**,
there is only one mask left, and the delay slot comes out EMPTY. No amount of statement reordering
recovers it, because the two expressions really are identical and cse is right.
**The lever.** Spell ONE of them as a shift pair:
```c
(param_2 << 25) >> 25 /* instead of param_2 & 0x7F */
```
cse sees `(lshiftrt (ashift ...))`, not `(and ...)`, so it does not merge — and then **combine's
`simplify_shift_const` folds the pair straight back to `andi 0x7F`**. You get two masks in the RTL
and one instruction each in the output. Byte-verified with the shift on EITHER side, so use whichever
reads more naturally.
**Why it generalises.** This is the inverse of the usual advice. Normally you fight gcc by making two
expressions *identical* so cse merges them; here you must make them *textually different in RTL but
identical after combine*. Any `x & ((1<<n)-1)` has a shift-pair twin that survives cse and folds back:
the mask and the shift pair are the same value to `combine` and different values to `cse`, and the
passes run in that order. Reach for it whenever a duplicated narrowing operation is missing from your
output and the second use is a delay-slot filler.
**Verification standard this one met (§405-A + §409):** the agent checked the **18-entry
`jtbl_80072F3C` order** and **all 51 reloc symbols** by hand before reporting, because `match_one`
sees neither.
**Companion finding — a §378c declaration blocker, and it is the norm on main now.** The draft is
`s32 func_8002DC68(u32, u32)` while `src/800_b.c` carries **five** `extern void
func_8002DC68(s32, s32);` declarations. That is a §376 gate blocker, not a body problem: align the
declarations to the definition, prove the alignment byte-neutral with NO draft substituted, COMMIT
it, and only then gate (the gate reverts `src/*.c` as its first action — §431).