phase11: merge 27 + cookbook 109-110 — 547 bodies / 556 regions

Worker D's 0x800307FC (92 B) and 0x80031EBC (112 B), both first spelling by adjacency.

109 records worker D's calibration claim -- 'adjacency finds the ROW, redundancy predicts the
PRICE' -- and the harder discipline behind it: D read 0x8006C044, identified it as a
tie-break-dense 3D-maths routine, and RELEASED it in favour of two small adjacent rows that
together cost less context than a first draft and returned two bodies instead of zero-to-one.

110 records a NEGATIVE RESULT from the coordinator. Worker C found a real false positive (the
ranker's top row is tie-break-dense, its score inflated by a repeated multu/mflo/sra idiom)
and proposed comparing full instruction words instead of opcodes. I implemented that and
measured it: it scores two KNOWN matches at ZERO and the known false positive HIGHEST. The
reason is fatal -- a genuine repeated source block does not produce identical instruction
words across copies because the allocator assigns different registers, so 'same opcodes,
different operands' describes a repeated block and a repeated idiom equally well. They are
indistinguishable at the instruction level. The opcode metric stays.
This commit is contained in:
Christopher Williams
2026-09-24 10:16:20 -04:00
parent 17e3b4496f
commit f14744e76c
2 changed files with 53 additions and 0 deletions
+1
View File
@@ -141,6 +141,7 @@
0x800308C4 0x800309CC src/func_800308C4.c
0x800319F0 0x80031A48 src/func_800319F0.c
0x80031BBC 0x80031CC0 src/func_80031BBC.c
0x80031EBC 0x80031F2C src/func_80031EBC.c
0x80031F2C 0x80031F78 src/func_80031F2C.c
0x80031F78 0x80031FC4 src/func_80031F78.c
0x80031FC4 0x800320D8 src/func_80031FC4.c
1 # Code-region registry: one C region per matched function.
141 0x800308C4
142 0x800319F0
143 0x80031BBC
144 0x80031EBC
145 0x80031F2C
146 0x80031F78
147 0x80031FC4
+52
View File
@@ -1680,3 +1680,55 @@ Worker B's caution, now pinned in `tools/sf3_rank`'s docstring: the score is the
*"repetitive at every scale"*. **Any implementation using a mean or a different normalisation
produces non-comparable numbers** — compare scores only against the reference tool's output. This
matters because workers were sharing rankings across partitions.
### 109. "Adjacency finds the row; redundancy predicts the price" (worker D)
Worker D's calibration, and it is a sharper claim than finding 99: **adjacency and redundancy do
different jobs.**
> **Adjacency finds the ROW. Redundancy predicts the PRICE.**
Measured across worker D's session: every redundant body cost **1–4 spellings** (11 matches); both
non-redundant above-ceiling attempts (`0x800FF5A8`, `0x80065980`) are **still open after 6–9
spellings**. The filter has never been wrong about *cost* on a row where it made a prediction.
**And worker D used it to DECLINE work, which is the harder discipline.** It read `0x8006C044`
(736 B, adjacent to its 700 B match — apparently the ideal adjacency pick) and identified it as a
**3D-maths culling routine**: 4 loops, 6 `mult`/`mflo` pairs, pointer walks — **the tie-break-dense
class, not the redundant class.** So it released the row and took two small adjacent rows instead,
which **together cost less context than a first draft of the 736 B row** and returned two bodies
rather than zero-to-one.
**A released row with a stated reason is worth more than a started row with a hope.**
### 110. The redundancy metric CANNOT be fixed by comparing operands — MEASURED (coordinator)
Worker C reported a false positive: the ranker's #1 row (`0x8001C8AC`, 652 B, 0.87) is
tie-break-dense, its score inflated by a repeated **instruction idiom** — `multu`/`mflo`/`sra 0xc`
fourteen times over nine *distinct* expressions. C proposed ranking on repeated **maximal
instruction runs** instead, on the theory that comparing full words (opcode **and** operands)
would separate a genuine repeated source block from a repeated idiom.
**I implemented exactly that and measured it. It is WORSE.** Coverage by repeated runs of 3
consecutive full words, against known rows:
| row | known | opcode metric | full-word metric |
|---|---|---|---|
| A match `0x8009C904` | high | 0.81 | 0.49 |
| A match `0x80036DA4` | high | 0.77 | 0.43 |
| A match `0x800556E8` | high | 0.67 | **0.00** |
| A match `0x80068910` | high | 0.82 | **0.00** |
| A fail `0x80017C6C` | low | 0.33 | 0.14 |
| C false positive `0x8001C8AC` | low | 0.57 | **0.58** |
**The full-word metric scores two known matches at ZERO and scores the known false positive
HIGHEST.** The reason is fatal to the approach: a genuine repeated source block does **not**
produce identical instruction words in its copies, because the register allocator assigns
different registers. So "same opcodes, different operands" describes a repeated source block
**and** a repeated idiom equally well — they are **indistinguishable at the instruction level**.
**Conclusion: the false positive is real, but it is not fixable by operand comparison.** The
discriminator has to come from source structure, not from the binary. Until something better
exists, **the opcode metric stays** — its record (13 of 23 first-spelling for one worker, 11
matches for another) is far stronger than the alternative's, and a single false positive at the
top of a 1000-row list is a tolerable error rate.