I banked the crack agent's story that a wrong-TU citation CAUSED
func_8017F2D4's seven gate refusals, and relayed it to Drew, without checking
it. corpus.stubs() derives each stub's TU from the actual INCLUDE_ASM site and
gate_stage splices via corpus — the harness was always editing the right file.
Only the PROSE was wrong.
Measured: func_8017F2D4 is still a stub, still classifies DIFF, and is a
has_mid_jr function referencing jtbl_801CC504 — so it carries a jump table the
standalone gate cannot see. The real residual is CAUSE NOT DETERMINED.
The ORACLE stands on its own evidence (the asm subdir's third component IS the
TU stem, by construction from the split config). The causal story does not, and
is now marked as such. This entry was written to stop a tool printing an
unmeasured cause and its first draft printed one.
Wave 6 added 110 (24 cracks + 23/24 families propagated). Fleet 95.1% instr /
89.7% distinct / 96.51% fn-count. R22 clean-fleet run 9x this session, 213/213
every time. Bank rate across six waves: 67/79/69/68/73/60%.
Records §166a (the destination-TU oracle) and the four-instance pattern it
completes: a tool asserting a conclusion it never reached. A confident wrong
label costs more than a missing one.
gate_stage labelled every "standalone MATCH / whole-binary DIFF" with
"(declaration/TU plumbing)". The tool never checked for a declaration conflict —
that was a GUESS printed as a diagnosis, and func_8017F2D4 carried it through
SEVEN attempts across five waves while every agent hunted codegen. The body was
byte-correct from the first attempt; the notes had simply named the wrong
destination TU (a file holding only a caller + prototype), and splicing there is
a no-op that leaves the INCLUDE_ASM bytes in place.
- gate_stage now says only what is true (the two oracles disagree) and hands
over the check that resolves it, instead of naming a cause it did not measure.
- §166a banks the oracle: asm/<ov>/nonmatchings/<TU_stem>/<fn>.s => the
INCLUDE_ASM is in src/<ov>/<TU_stem>.c. The third path component IS the TU
stem, derived from the split config, and it beats any prose citation — a grep
for the function name also hits callers and prototypes in OTHER TUs and reads
exactly like a destination hit.
- Plus the two probe gotchas that cost wave-5/6 agents real time: the wrong
--aspsx-version fakes ~32 ori-vs-addiu mismatches, and a collateral-drift
check must filter to sized symbols (nm -S) or the zero-size .NON_MATCHING
aliases all report false drift.
cookbook_index.py: 506 sections.
Wave 6 (wf_9729fd89-c16, 76 agents, 6.7M tok): 40 targets -> 35 agent-MATCH,
1 refuted, 0 UNVERIFIED (the new field works), 4 NEAR -> 24 BANKED.
*** THE FINDING OF THE CAMPAIGN, and it is not a compiler idiom ***
func_8017F2D4 had been "MATCH standalone / DIFF at gate" SEVEN times across
five waves. Every attempt hunted codegen. The body was byte-correct the whole
time. The fault was that the prior notes named the WRONG DESTINATION TU:
..._jr_8017C340.c holds only a CALLER and a prototype; the INCLUDE_ASM lives in
..._jr_8017ED5C.c. Splicing into the wrong file is a no-op, the binary differs,
and harvest recorded its canned string "match_one MATCH but gate rejected
(declaration/TU plumbing)" — A GUESS, NOT A MEASUREMENT. No declaration
conflict ever existed.
NEW LAW (zero-cost oracle): THE SPLAT ASM SUBDIR NAMES THE DESTINATION TU.
asm/<overlay>/nonmatchings/<TU_basename>/<fn>.s => the INCLUDE_ASM is in
src/<overlay>/<TU_basename>.c, always — the third path component IS the TU
stem. Never accept a prose TU citation that disagrees with the --asm-subdir you
were handed: a grep for the function name also hits callers and prototypes in
OTHER TUs and reads exactly like a destination hit. Corollary: re-derive the TU
path from the asm subdir BEFORE hunting codegen on any gate-refused backlog
entry.
The agent proved it properly: in-situ splice into BOTH TUs, full pinned triple,
279/279 with 0 masked diffs, plus a collateral-drift check (71/71 other sized
symbols byte-identical). It also corrected the reach to x2 (only two .s files
for that function exist, byte-identical modulo the overlay name).
Bank rate across six waves: 67% -> 79% -> 69% -> 68% -> 73% -> 60%.
A markdown backtick inside the prompt's template literal terminates the string
and the workflow dies at parse time. Cheap (0 agents, 0 tokens) but it has now
cost two launches, so the file carries a warning at the top: use double quotes
for inline code in prompt prose.
Wave 5 wrote its drafts into .run/wave4/ because the script was built by
sed-ing the wave-3 copy. Harmless (per-function dirs) but wrong. Pass
{wave:'waveN'} on any target; defaults to '.run/wave/' so a missing field can
never silently reuse a prior wave's directory. Meta name/description are now
generic too — this is THE crack-wave script, not wave 4's copy.
A prior attempt that was GATE-REFUSED (standalone MATCH, whole-binary DIFF) has
an in-TU residual, so the standalone gate cannot see it. A wave-5 agent closed
exactly this by splicing into the real destination TU, running the pinned
triple end-to-end, and masked-diffing the function out of the WHOLE-TU object
(0 mismatched, 279/279). That is now instruction, not luck.
Carries both gotchas it paid for: the wrong --aspsx-version produces ~32
spurious mismatches all of the ori-vs-addiu li-form shape (the fingerprint of a
version mismatch, not a codegen residual), and engine_core.h resolves relative
to the including file so the spliced TU needs a directory + shared symlink.
Also warns that two wave-5 agents found their predecessor's file/line citations
wrong while its idioms were right — re-derive the TU, don't trust quoted lines.
Wave 5 added 152 (29 cracks + 28/29 families propagated) — the session's
largest. Fleet 95.0% instr / 89.6% distinct / 96.48% fn-count; stubs
13,345 -> 12,771. R22 clean-fleet run 8x this session, 213/213 every time.
P30's milestone is '>=95% instr fleet, or every remaining overlay stub on a
named ledger'. THE FIRST HALF IS NOW MET — T5 (phase close) is a live option.
Bank rate across five waves: 67% -> 79% -> 69% -> 68% -> 73%.
crack_wave.js classified anything without verdict_check.confirmed as 'refuted'.
When the usage limit killed 22 verifiers mid-wave, the result read
'refuted: 22' — 22 good drafts reported as rejected, with evidence 'verifier
died' as the only tell. Acting on that would have discarded the wave.
Same disease as the no-diagnostic classifier (S47) and this session's
poisoned-tree 0/17: a tool stating a conclusion it never reached. UNVERIFIED is
now its own outcome, returned with the draft path and sha1 and the instruction
'VERIFIER NEVER RAN — re-verify, do not discard'.