tools/journal_notes.py mines the agent journals per (binary, fn) and appends a
PAST ATTEMPTS section to the pack; claude_wave_packs.py calls it automatically, so
it is the default rather than a step to remember. Idempotent, and R48-safe (a note
stamped with a different binary is never served — §238 homonyms).
Measured before adopting (S71 wave 1, 50 one-agent workflows over the 210-function
real frontier where every target had already refused an earlier wave):
* 38/39 MATCH at closeness 0 (97.4%) vs S70's 124/131 (94.7%) on an EASIER pool
* 29/39 agents cite a prior attempt as what they used
* 4/39 banked by RECOVERING a body that already matched, from a path a note named
* 11/39 matched on the first compile
The two costs it removes are re-testing a measured-inert lever (§406 lists twelve,
§407 fifteen, §410 four — each paid for by an agent and never seen again) and
re-deriving a body that already exists on disk.
Also: jr_isolate_all places file-local `static` definitions with the region that uses
them instead of refusing the whole file. A `static inline` helper (§82.1) has no
address by construction, which is not a defect; the R32 guard was refusing these and
blocking the isolate on 4 of the 6 overlays whose CARVE-REFUSED functions it is the
named remedy for. Two regions using one static is still a hard refusal (duplicating a
used static is a byte change, R43).
docs: cookbook §411, wave-playbook step 3b, accelerators entry.
* Every pack carried PAST ATTEMPTS ON THIS EXACT FUNCTION, mined per-function from the
historical agent journals (52 of 60 targets, 131 notes). Every landed agent returned
MATCH at closeness 0 on the hardest frontier we have.
* §409 — the wave and the nine laws it produced. Law 1: a relocation-stream
TRANSPOSITION is invisible to match_one, the permuter scorer and every similarity
tier (HI16/LO16 masking; the §195-D blind spot for a different reloc class), and it
retroactively explains "MATCH but the gate rejected it" verdicts.
* §410 — COPY THEN ACCUMULATE ON THE COPY: satisfies the $s2 in-place destination and
the sched1 birthing boost at once, with the agent's measured refutation list.
* gate 1 (all 64 across 33 binaries): 12 banked — main 11 + ov_SC07_006 1.
* gate 2 tested "a bad draft kills its binary's good ones" by re-staging only the 25
that recover_integration --probe-only called MATCH in their real TU: 0 banked.
An honest null — that probe compiles and diffs bytes but never LINKS or CARVES,
so it is a third oracle with its own blind spot.
* triage (25/25 accounted): CARVE-REFUSED 10, undefined-reference 4, DIFF 3,
CC1-FAIL-no-diagnostic 2, PARSE 1; gate 1 adds 7 func-decl / 4 data-decl /
6 type-decl conflicts.
* R37 probe of the carve class: 6 of 8 are one refusal — a subseg would host
NON-CONTIGUOUS .rodata carves — whose named remedy is jr_isolate_all (§8b).
tools/restage_matching.py — rebuild a gate plan from probe verdicts.
tools/gate_triage.py — route a gate's verdicts to the lane each one names (R47).
The S70 patch was refused by its own adversarial review for sorting rows by recency:
a pair's ledger rows are several PROBES about one draft, alternating between
`closeness 4` and `won't compile standalone`, so max(ts) serves whichever probe ran
last — often the least informative. This form keeps both.
* the ts-newest verdict is still selected (file order made the per-binary bulk ledger
always win regardless of age: 25 pairs mis-selected),
* AND the best measurement ever taken on the pair rides alongside it, so a later
uninformative probe can no longer erase an earlier residual: 981 of 2,605 pairs
gain a line they were previously denied.
* BASELINE-RED is a fact about a binary at a moment (R51), frozen into an append-only
ledger and replayed forever — 2,676 rows all stamped 2026-08-26. gate_feedback now
reads the same live red union gate_stage consults, so a pack and the next gate run
cannot disagree: 173 expired claims retired, 0 binaries currently red.
R39 control 3/3 (expired-when-green, harness-line-when-red, measurement-survives).
* `git status --porcelain -- src/<binary>/` finds nothing for main, whose TUs are
src/800.c, src/boot.c, ... — so a main worker returned `files: {}` while the bank
oracle (the stub disappeared) still counted the banks. parallel_gate printed
"12 banked across 2 binaries" and committed one of them.
* src_scope() takes the scope from the binary's own stub rows (each names its TU),
captured BEFORE the gate because a bank deletes the stub that names it, and keeps
the directory prefix for overlays that have one.
Negative control: main 0 -> 54 TUs, ov_SC07_006 1 -> 3 (superset, no regression).
* A reused worktree kept the previous job's .run/harvest_failed*.classified.txt, so
verdicts surfaced under the wrong binary; the worker clears them first.
* tools/gate_triage.py — routes a gate's verdicts to the repair lane each names (R47),
with the staged-draft denominator asserted (R32/R41).
Re-gated main: 11 banked (commit:3586), main real frontier 64 -> 53.
* `binof = {c["fn"]: c["binary"]}` was last-writer-wins, and `status`, `det` and `subof`
had the same shape — a draft of a name carried by two binaries was stamped with
whichever card came last and then reloc-checked against the OTHER binary's symbols.
* Resolve per draft instead: the shard's own target list first
(`.run/wave_<tag>_targets.<i>.json` = `targets[i::workers]`, each row carrying its
binary), a unique-name card second, a counted refusal when neither can answer (R43).
* R39 negative control over every historical wave: 42,655 drafts, 0 regressions,
2,317 (5.4%) previously mis-stamped; 2,107 homonym card names fleet-wide.
Intra-shard ambiguity: 0 of 50,684 (shard, name) pairs over 302,370 shard files.
docs: §408 — §406 refuted as a sweep (0 MATCH / 14 applied, 0 / 210). The 134-member
census counted main's 960 LINKED library stubs and matched a symmetric SHAPE; derived
from the mine-vs-target residual the addressable set is 15 / 210. Decision-log entry
records the pivot: 64 of 210 (30.5%) already match standalone, so the frontier's
largest lane is §376 integration, not codegen.
tools/weave_sweep.py — the derived-selector sweep (R32 coverage, R41 denominators,
--lever-all ablation control).
gate_feedback selected the newest reloc_rejects row with shape=='MATCH' and printed
its mismatches under "your instruction stream already matched; ONLY these names were
wrong". But reloc_identity's binding condition is `aligned` (shape=='MATCH' AND equal
relocation-stream lengths); when that fails it downgrades status to "MISMATCH?" and
stamps the row ADVISORY. Gating on `shape` alone therefore republished ADVISORY rows
as binding per-index instructions — and when the streams are not index-aligned, draft
index i is compared to target index i of a DIFFERENT stream, so every "the target
references 0x..." line is arithmetic on the wrong word.
MEASURED (agent-run, not predicted):
* 15 of 130 S70 targets were served this block; 15 of 15 were aligned=False, i.e.
100% carried reloc_identity's own "verdicts are ADVISORY" caveat while the pack
text told the agent the opposite.
* 55 of the 66 printed lines (83%) name a value that is not an address at all
(0x82020084, 0x880801C0, ...).
* Of the 4 whose .s is on disk, 4 of 4 named symbols the target never relocates.
* Whole index: 182 servable (binary, fn) rows, 149 aligned=False; 140 of those 149
print >=1 non-address vs 1 of the 33 aligned=True.
This reproduces both S70 agent reports verbatim (ov_SC02_035:func_8017D3F4 "cross-
overlay contamination"; ov_SC06_020:func_8017D918 "those symbols are absent from
this .s").
Root cause has a second half, still OPEN upstream: ox_campaign.reloc_filter stamps the
row's binary from `binof = {c["fn"]: c["binary"]}` — a BARE-NAME dict (R48). Wave `el`
carried 44 names in >=2 binaries, so func_8017D918's row was stamped ov_SC06_020 while
the draft it checked belonged to ov_SC01_074. A correct read key cannot repair a wrong
write-side stamp, which is why the fix validates against the TARGET'S OWN bytes.
Adversarially reviewed (sound=True) and controlled here: the known-true aligned=True
case ov_SC07_011:func_8016AB6C is STILL SERVED; ov_SC02_035:func_8017D3F4 is withheld
with a loud reason. The reviewer's own first attempt validated draft_symbol against the
.s and rejected that good block — a false positive caught only by a known-true case.
§405 — the generalisable residue of three waves, grouped by lever family:
A. match_one compares .text ONLY, so a switch's jump table is invisible to it —
a draft can score 110/110 with a PERMUTED table emitted as identity
(resident/func_800D02D0, byte-witnessed). Some historical 'MATCH but gate
rejected' verdicts were the ORACLE being wrong (R34 in our most-trusted tool).
B. the scheduler dials, incl. the birthing-boost re-tie's PLACEMENT rule (must be
a LATER basic block) and reorg's stop_search_p halting at any asm.
C. regalloc from C without pins: variable identity picks global- vs local-alloc;
a cross-arm join value loses first-fit and should be duplicated per arm for
cross_jump to refund; pass-through params reserve arg regs at zero cost.
D. integration is still the bottleneck — ~1/3 had byte-correct bodies blocked
only by declarations; the TU is the authority.
E. what the agents REFUTED: §137 invariance is false for CONFLICT-driven ties;
§153's 'cse2 puts it back' fails for the dead-def case; §257-8 volatility
polarity is per-site, not portable.
§406 — the prologue-weave class: 134 of 1,237 open stubs (11%) share the
sw / move , / sw shape, cause traced in cc1's .i.sched2 dump
(memrefs_conflict_p finds no dependence, potential_hazard picks sw $ra early), 12
variants measured inert, and ONE working lever (non-volatile memory-clobber asm
after the param copy). One lever x 134 known targets = a sweep, not an idiom.
MEASURED over the 130 S70 targets: 110 cards printed "This card has NO banked twin
— derive the structure from the .s", and **75 of them (68%) had that function
already BANKED at the same address in a sibling overlay.**
seed_ref joins on signature hashes and is blind to indexed-global relocs (§389), so
a reloc-only twin of an already-banked body hashes differently and reads as a
singleton. Overlays share code at the same VRAM, so "is this address banked
elsewhere?" is a one-line question the card never asked. An S70 agent found its
answer at src/ov_SC02_000/ov_SC02_000_jr_8018173C.c:4827 and reported the card was
simply wrong: "a cross-overlay same-address grep as step 0 would have returned this
for ~0 tokens" — which is exactly what the wave playbook prescribes and what nothing
was supplying.
_same_addr_banked() derives it from the corpus invariant (R33: banked == in sig and
not an INCLUDE_ASM stub), memoized once per process. The card now names the binaries
and tells the agent to READ IT FIRST, while warning that a same-address function in
another overlay is usually — not always — the same function (verify per law 1c).
Controls: positive ov_SC02_003:func_80185840 -> ['ov_SC02_000', 'ov_SC03_091'] (the
first is the very binary the agent found by hand); bogus address -> []; a named
symbol -> []. Failure returns [] so this only ever ADDS fuel.
claude_wave_packs writes out_dir/SYS.md + out_dir/packs/<fn>.md, so out_dir is the
WAVE dir. The playbook documented plus an mv to undo the
resulting packs/packs nesting — which also put SYS.md at <wave>/packs/SYS.md while
claude_wave_draft.js tells every agent to read <wave>/SYS.md.
Net effect: the laws file did not exist where any agent looked, in every wave run
this way, and the brief silently degraded to the pack alone. Two S70 agents said so
verbatim; the rest never noticed. Passing the wave dir fixes it and removes the mv.
api_agent.prior_draft's law-1c guard had two holes, both measured live in the S70
wave where FOUR independent agents reported discarding the warm-start as "a
different function entirely":
* `len(syms) >= 2` exempted every body referencing 0 or 1 symbols — exactly the
small-function case. func_80182438 (21 ins, ONE symbol) sailed through carrying
ov_SC02_028's body for the SAME ADDRESS, and its agent reported that as the
reason its PRIOR attempt failed outright.
* requiring a strict majority foreign let a body sharing half its symbols pass.
A correct draft can only reference what the target's .s actually relocates, so ANY
foreign symbol disqualifies.
NEGATIVE CONTROL over all 50 S70 targets: 46 admitted -> 42, and the 4 rejected are
exactly the bodies the agents flagged (func_80182438 foreign func_801330E0,
func_800D0664, func_801831D0 foreign func_80182570, func_80185F4C). No collateral.
Cost of the hole: every agent reading a poisoned warm-start burns compiles
discarding it, and a weaker model follows it instead. R48 again — never key by bare
function name.
`migrated_tables()` can flag a function whose table is actually in the DATA TAIL,
and the §154-A island branch then refused the whole batch with "a tail carve cannot
help ... the carve model covers jump tables only, not an island of mixed included
data". That reads as a permanent toolchain wall. It is a ROUTING error: island_probe
classifies the same function 'tail', and its own detail says "standard §8a carve at
gate time" -- i.e. it names the ordinary lane as the owner (R43: each probe kind
names the lane that owns it). The refusal was about the branch we entered, not the
function.
Consult the probe first and let a 'tail' function fall through to build_carve.
Byte-proven immediately: ov_SC02_000/func_8017F950 -- three full parallel_gate
passes had booked it CARVE-REFUSED -- now reports `[jtbl] carved func_8017F950`,
`verified 1 / failed 0`, BYTE-IDENTICAL, and corpus.stubs confirms it banked. No
config change was needed: its carve was already committed and merely PENDING an
owner (the class identified while fixing jr_inventory), so banking the function
completed the 1:1 ownership the assertion wanted.
This was blocker 3 of 3 on the §322b route; 1 and 2 were cleared earlier in S70.
jr_inventory R32-aborted on 36 committed .rodata carves across 8 binaries with
"ownership is not 1:1 — a stranded/duplicated carve", blocking the whole §322b
carve route. Every one of those binaries is BYTE-GREEN (R22 213/213), so the config
was right and the MODEL was blind (R34). Measured, the two blind spots:
1. THE SUBSEG NAME IS THE OWNERSHIP RECORD -- 32 of 36 (89%). The isolate
convention writes the owner into the name: a carve in `<ov>_jr_<ADDR>` belongs
to func_<ADDR>. Several owners are RESIDENT-range (0x80135D20, 0x8015C32C)
instantiated through a shared macro, so they are not overlay-local definitions
and parse_overlay_c cannot see them at all. Reading the name is R33.
2. A CARVE FOR A STILL-STUBBED FUNCTION IS PENDING, NOT STRANDED -- the other 4.
ov_SC07_010's func_8016AB6C references its carve at 0x801A6460 from an
INCLUDE_ASM stub.
A carve with none of the three still aborts loudly -- that is the real corruption
the assertion exists to catch (§8b func_801734BC class).
jr_isolate_all --dry-run over the carve set: 7 PASS / 10 FAIL -> 15 PASS / 2 FAIL.
The 2 remaining are the §323 file-local-type class §322b already predicted
(ov_SC02_017 typedef, ov_SC03_029 "carry the naming type").
The carve dispatch has three branches keyed on jtbl_carve's FIRST refusal message.
A function whose first refusal mentions a leading .rodata island is sent down the
§260 island-split branch — but island-split can then refuse with "... is 'tail',
not 'island-end' — table(s) in the data tail — standard §8a carve at gate time",
i.e. it NAMES the branch that should have handled it. That was booked CARVE-REFUSED:
a verdict about the ROUTE WE CHOSE, not about the function, and no branch ever ran
the carve the tool actually asked for.
The isolate has already run at that point, so the standard route is just
re-extract + re-carve (the tail of the _ISO_WALLS branch). If that also refuses, it
now prints the TERMINAL reason instead of the routing one.
Measured on ov_SC02_000/func_8017F950: the fallback fires and reaches the real
answer — "jump tables only, not an island of mixed included data" — a genuine
structural refusal. So this fixes the DIAGNOSIS for that function rather than
unlocking it, and should unlock any tail case whose table is a pure jump table.
Those two objects were originally assembled in REORDER mode (the assembler filled
the delay slots). maspsx force-emits `.set noreorder`, making that unreachable, so
a whole class there read as a permanent compiler wall (§332) when the property
belongs to the OBJECT, not the toolchain.
For REORDER_TUS only, swap maspsx for tools/reorder_passthrough.py + `as -O2` --
the pipeline tools/oracle_reorder.py already proved byte-exact (0 diffs on
func_80061FA8 where the pinned path gives 57). Everything else is untouched.
Verified:
* branch selection BOTH ways: 800c3 -> reorder_passthrough, 800.c -> maspsx
* tools/reorder_passthrough.py --selftest, incl. a negative control (a line
merely CONTAINING "move", e.g. `jal remove_thing`, must not be rewritten)
* BYTE-INERT: main rebuilds BYTE-IDENTICAL via verify_binary (§384, re-extracts)
Note the first patch used `ifeq ($(filter $*,...))`, which make evaluates at PARSE
time when $* is empty -- it would have silently always taken the maspsx branch.
`$(if ...)` expands per-target, which is why the rule already uses that form for
JTBL_PADS.
Drew, correctly: "this is a tools health test, not a full regression test."
audit-cdecl re-parsed every declaration in all 4,168 TUs and handed each to real
gcc — ~787s of pure-Python collection before the first cc1 call. It made
`make tools-health` unrunnable: >15 min, killed twice, never completed once.
`--limit` already existed and its own help calls it "a fast smoke run"; nothing
was using it. Sampled by default (CDECL_AUDIT_TUS ?= 60); the exhaustive form
stays as `audit-cdecl-full` for when cdecl.py itself changes.
audit-cdecl : >9 min -> 61s (4,777 declarations adjudicated, 0 rejected)
tools-health: never completed -> 333s, rc=0, all green
Known limit, recorded not hidden: --limit takes the FIRST N TUs, not a random
sample, so the smoke run always exercises the same files. Randomising the sample
(or rotating by seed) is the follow-up.
Drew asked why `make tools-health` runs 15+ min. Measured per step rather than
guessed (I guessed wrong twice first, and both are recorded in the comments):
sig-overlays ~52s serial -> 3.9s wall / 51.8s user (xargs -P$(JOBS), 32 cores)
sig-modules 0s sig-resident 0s audit-corpus 17s
audit-cdecl >9 MINUTES <-- the actual bottleneck, and NOT the gcc probes:
the `[gcc] N distinct declarations` line never printed inside a
10-minute run, so not one cc1 call had happened. `tu_statements`
over 4,168 TUs is ~787s single-core, all of it before the probes.
SHIPPED
* sig-overlays / sig-modules: xargs -P$(JOBS), same pattern extract-all and
check-all already use in this file. sig_image has exactly one write path
(its own per-alias .jsonl), verified before fanning out. NEGATIVE CONTROL:
141/141 sig files BYTE-IDENTICAL to the serial output. Also adds the failure
detection the serial loops never had -- a sig_image crash used to vanish (R32).
* cdecl._gcc_probe: `probe_{tag}.c` was ONE FIXED FILENAME PER TAG, correct only
while _sift is serial. Now unique per call, so concurrent probes cannot
overwrite each other's source between write and compile and return a verdict
about another chunk's declarations.
* cdecl._sift: threads over chunks + over the bisection probes (gcc is a
subprocess, so the GIL is released), results written back BY INDEX so the
output stays deterministic. A/B on --limit 6: IDENTICAL verdicts (829/829).
NOT SHIPPED, and the measurement is left in the code
A ProcessPoolExecutor over the collection phase was tried and REVERTED: 12 TUs
yield 32,352 statements, so the full pass ships ~11M strings through IPC and the
pickling costs more than the parse it saves. The fix is to dedupe/filter INSIDE
the worker or memoise per-TU by content hash -- left measured, not guessed.
make tools-health audits DATA integrity (corpus/cdecl/binaries/digest/text) and
nothing audited TOOL BEHAVIOUR -- the gap all four S70 defects fell through. Each
reported success while doing nothing or doing harm, and none would have been found
by reading the source: a wrong instrument returns a plausible NUMBER, not an error.
tools/work_evidence.py, three assertions on OBSERVABLE CONSEQUENCE:
assert_inputs zero readable inputs is a DEFECT, not a zero-yield result. "0 of 0"
is a fact about the harness; "0 of 57" is a fact about the subject.
assert_floor work claiming a compile/gate cannot beat physics -- the ONLY tell on
the pgate defect was a 1-2s wall clock (§402).
assert_effect N claimed successes must show a persistent effect; verification is
not banking (§404).
Self-test is a negative control both directions (11/11): each assertion PASSES the
already-succeeded case and FAILS the known-bad one, and non-strict warns instead of
raising. Wired into `make tools-health` so it cannot rot (R54).
Wiring on the wave critical path:
* parallel_gate: per-worker wall-clock floor; a sub-floor worker is flagged
"BLIND SUSPECT" in the summary line instead of passing as a clean zero.
* gate_stage: the silent `if not draft_fns: return {...}` -- the exact point the
pgate defect flowed through -- is now loud and marks the result `refused`.
* harvest_verify: says at the point of confusion that "verified" is not "banked"
and names gate_stage as the entrypoint that persists.
Negative control: empty drafts dir -> loud + refused. Positive control: a real
2-draft dir still gates normally (drafts:2, no false refusal).
The ledger write recorded every entry in `ready` as `gated:rc<N>` on ANY rc. When
the gate REFUSES to start (parallel_gate on a dirty tree, a worker missing its link
inputs) it examines nothing -- yet S69's Gate37 refused with rc=1, gated nothing,
and both of its functions were recorded as gated and silently skipped on the retry.
The phantom entries had to be cleared by hand.
"Attempted" and "never looked at" are different facts and only the first justifies
suppressing a re-gate. A binary now counts as EXAMINED when its worker banked
something, wrote per-function verdict rows, or reported a draft count -- i.e. got
far enough to have an opinion (R32). Everything else stays eligible and is named
loudly rather than dropped silently (R55).