api_agent.prior_draft's law-1c guard had two holes, both measured live in the S70
wave where FOUR independent agents reported discarding the warm-start as "a
different function entirely":
* `len(syms) >= 2` exempted every body referencing 0 or 1 symbols — exactly the
small-function case. func_80182438 (21 ins, ONE symbol) sailed through carrying
ov_SC02_028's body for the SAME ADDRESS, and its agent reported that as the
reason its PRIOR attempt failed outright.
* requiring a strict majority foreign let a body sharing half its symbols pass.
A correct draft can only reference what the target's .s actually relocates, so ANY
foreign symbol disqualifies.
NEGATIVE CONTROL over all 50 S70 targets: 46 admitted -> 42, and the 4 rejected are
exactly the bodies the agents flagged (func_80182438 foreign func_801330E0,
func_800D0664, func_801831D0 foreign func_80182570, func_80185F4C). No collateral.
Cost of the hole: every agent reading a poisoned warm-start burns compiles
discarding it, and a weaker model follows it instead. R48 again — never key by bare
function name.
`migrated_tables()` can flag a function whose table is actually in the DATA TAIL,
and the §154-A island branch then refused the whole batch with "a tail carve cannot
help ... the carve model covers jump tables only, not an island of mixed included
data". That reads as a permanent toolchain wall. It is a ROUTING error: island_probe
classifies the same function 'tail', and its own detail says "standard §8a carve at
gate time" -- i.e. it names the ordinary lane as the owner (R43: each probe kind
names the lane that owns it). The refusal was about the branch we entered, not the
function.
Consult the probe first and let a 'tail' function fall through to build_carve.
Byte-proven immediately: ov_SC02_000/func_8017F950 -- three full parallel_gate
passes had booked it CARVE-REFUSED -- now reports `[jtbl] carved func_8017F950`,
`verified 1 / failed 0`, BYTE-IDENTICAL, and corpus.stubs confirms it banked. No
config change was needed: its carve was already committed and merely PENDING an
owner (the class identified while fixing jr_inventory), so banking the function
completed the 1:1 ownership the assertion wanted.
This was blocker 3 of 3 on the §322b route; 1 and 2 were cleared earlier in S70.
jr_inventory R32-aborted on 36 committed .rodata carves across 8 binaries with
"ownership is not 1:1 — a stranded/duplicated carve", blocking the whole §322b
carve route. Every one of those binaries is BYTE-GREEN (R22 213/213), so the config
was right and the MODEL was blind (R34). Measured, the two blind spots:
1. THE SUBSEG NAME IS THE OWNERSHIP RECORD -- 32 of 36 (89%). The isolate
convention writes the owner into the name: a carve in `<ov>_jr_<ADDR>` belongs
to func_<ADDR>. Several owners are RESIDENT-range (0x80135D20, 0x8015C32C)
instantiated through a shared macro, so they are not overlay-local definitions
and parse_overlay_c cannot see them at all. Reading the name is R33.
2. A CARVE FOR A STILL-STUBBED FUNCTION IS PENDING, NOT STRANDED -- the other 4.
ov_SC07_010's func_8016AB6C references its carve at 0x801A6460 from an
INCLUDE_ASM stub.
A carve with none of the three still aborts loudly -- that is the real corruption
the assertion exists to catch (§8b func_801734BC class).
jr_isolate_all --dry-run over the carve set: 7 PASS / 10 FAIL -> 15 PASS / 2 FAIL.
The 2 remaining are the §323 file-local-type class §322b already predicted
(ov_SC02_017 typedef, ov_SC03_029 "carry the naming type").
The carve dispatch has three branches keyed on jtbl_carve's FIRST refusal message.
A function whose first refusal mentions a leading .rodata island is sent down the
§260 island-split branch — but island-split can then refuse with "... is 'tail',
not 'island-end' — table(s) in the data tail — standard §8a carve at gate time",
i.e. it NAMES the branch that should have handled it. That was booked CARVE-REFUSED:
a verdict about the ROUTE WE CHOSE, not about the function, and no branch ever ran
the carve the tool actually asked for.
The isolate has already run at that point, so the standard route is just
re-extract + re-carve (the tail of the _ISO_WALLS branch). If that also refuses, it
now prints the TERMINAL reason instead of the routing one.
Measured on ov_SC02_000/func_8017F950: the fallback fires and reaches the real
answer — "jump tables only, not an island of mixed included data" — a genuine
structural refusal. So this fixes the DIAGNOSIS for that function rather than
unlocking it, and should unlock any tail case whose table is a pure jump table.
Those two objects were originally assembled in REORDER mode (the assembler filled
the delay slots). maspsx force-emits `.set noreorder`, making that unreachable, so
a whole class there read as a permanent compiler wall (§332) when the property
belongs to the OBJECT, not the toolchain.
For REORDER_TUS only, swap maspsx for tools/reorder_passthrough.py + `as -O2` --
the pipeline tools/oracle_reorder.py already proved byte-exact (0 diffs on
func_80061FA8 where the pinned path gives 57). Everything else is untouched.
Verified:
* branch selection BOTH ways: 800c3 -> reorder_passthrough, 800.c -> maspsx
* tools/reorder_passthrough.py --selftest, incl. a negative control (a line
merely CONTAINING "move", e.g. `jal remove_thing`, must not be rewritten)
* BYTE-INERT: main rebuilds BYTE-IDENTICAL via verify_binary (§384, re-extracts)
Note the first patch used `ifeq ($(filter $*,...))`, which make evaluates at PARSE
time when $* is empty -- it would have silently always taken the maspsx branch.
`$(if ...)` expands per-target, which is why the rule already uses that form for
JTBL_PADS.
Drew, correctly: "this is a tools health test, not a full regression test."
audit-cdecl re-parsed every declaration in all 4,168 TUs and handed each to real
gcc — ~787s of pure-Python collection before the first cc1 call. It made
`make tools-health` unrunnable: >15 min, killed twice, never completed once.
`--limit` already existed and its own help calls it "a fast smoke run"; nothing
was using it. Sampled by default (CDECL_AUDIT_TUS ?= 60); the exhaustive form
stays as `audit-cdecl-full` for when cdecl.py itself changes.
audit-cdecl : >9 min -> 61s (4,777 declarations adjudicated, 0 rejected)
tools-health: never completed -> 333s, rc=0, all green
Known limit, recorded not hidden: --limit takes the FIRST N TUs, not a random
sample, so the smoke run always exercises the same files. Randomising the sample
(or rotating by seed) is the follow-up.
Drew asked why `make tools-health` runs 15+ min. Measured per step rather than
guessed (I guessed wrong twice first, and both are recorded in the comments):
sig-overlays ~52s serial -> 3.9s wall / 51.8s user (xargs -P$(JOBS), 32 cores)
sig-modules 0s sig-resident 0s audit-corpus 17s
audit-cdecl >9 MINUTES <-- the actual bottleneck, and NOT the gcc probes:
the `[gcc] N distinct declarations` line never printed inside a
10-minute run, so not one cc1 call had happened. `tu_statements`
over 4,168 TUs is ~787s single-core, all of it before the probes.
SHIPPED
* sig-overlays / sig-modules: xargs -P$(JOBS), same pattern extract-all and
check-all already use in this file. sig_image has exactly one write path
(its own per-alias .jsonl), verified before fanning out. NEGATIVE CONTROL:
141/141 sig files BYTE-IDENTICAL to the serial output. Also adds the failure
detection the serial loops never had -- a sig_image crash used to vanish (R32).
* cdecl._gcc_probe: `probe_{tag}.c` was ONE FIXED FILENAME PER TAG, correct only
while _sift is serial. Now unique per call, so concurrent probes cannot
overwrite each other's source between write and compile and return a verdict
about another chunk's declarations.
* cdecl._sift: threads over chunks + over the bisection probes (gcc is a
subprocess, so the GIL is released), results written back BY INDEX so the
output stays deterministic. A/B on --limit 6: IDENTICAL verdicts (829/829).
NOT SHIPPED, and the measurement is left in the code
A ProcessPoolExecutor over the collection phase was tried and REVERTED: 12 TUs
yield 32,352 statements, so the full pass ships ~11M strings through IPC and the
pickling costs more than the parse it saves. The fix is to dedupe/filter INSIDE
the worker or memoise per-TU by content hash -- left measured, not guessed.
make tools-health audits DATA integrity (corpus/cdecl/binaries/digest/text) and
nothing audited TOOL BEHAVIOUR -- the gap all four S70 defects fell through. Each
reported success while doing nothing or doing harm, and none would have been found
by reading the source: a wrong instrument returns a plausible NUMBER, not an error.
tools/work_evidence.py, three assertions on OBSERVABLE CONSEQUENCE:
assert_inputs zero readable inputs is a DEFECT, not a zero-yield result. "0 of 0"
is a fact about the harness; "0 of 57" is a fact about the subject.
assert_floor work claiming a compile/gate cannot beat physics -- the ONLY tell on
the pgate defect was a 1-2s wall clock (§402).
assert_effect N claimed successes must show a persistent effect; verification is
not banking (§404).
Self-test is a negative control both directions (11/11): each assertion PASSES the
already-succeeded case and FAILS the known-bad one, and non-strict warns instead of
raising. Wired into `make tools-health` so it cannot rot (R54).
Wiring on the wave critical path:
* parallel_gate: per-worker wall-clock floor; a sub-floor worker is flagged
"BLIND SUSPECT" in the summary line instead of passing as a clean zero.
* gate_stage: the silent `if not draft_fns: return {...}` -- the exact point the
pgate defect flowed through -- is now loud and marks the result `refused`.
* harvest_verify: says at the point of confusion that "verified" is not "banked"
and names gate_stage as the entrypoint that persists.
Negative control: empty drafts dir -> loud + refused. Positive control: a real
2-draft dir still gates normally (drafts:2, no false refusal).
The ledger write recorded every entry in `ready` as `gated:rc<N>` on ANY rc. When
the gate REFUSES to start (parallel_gate on a dirty tree, a worker missing its link
inputs) it examines nothing -- yet S69's Gate37 refused with rc=1, gated nothing,
and both of its functions were recorded as gated and silently skipped on the retry.
The phantom entries had to be cleared by hand.
"Attempted" and "never looked at" are different facts and only the first justifies
suppressing a re-gate. A binary now counts as EXAMINED when its worker banked
something, wrote per-function verdict rows, or reported a draft count -- i.e. got
far enough to have an opinion (R32). Everything else stays eligible and is named
loudly rather than dropped silently (R55).
Both tools restored with `text.replace(after, before, 1)` -- the FIRST occurrence.
--any-proto (and sync-decls) collapse DISTINCT declarations of one function to the
SAME `after` text, so occurrence N received entry N's `before` in JOURNAL order,
not file order: the originals land on the wrong occurrences and the file is
corrupted while the tool prints full success.
Byte-witnessed twice in S70:
* fix_arity_callers: "restored 382, kept 0, missing 0" left the FLEET-SHARED
src/shared/engine_core.h with 97 insertions / 97 deletions (func_8012A828
rotated between three declaration sites).
* cast_self_callers: "reverted 10 edit(s)" left src/800.c with the two decls of
func_80031988 swapped.
Both were caught only by `git diff` AFTER the success line (R40: the tool's own
report is not evidence).
The occurrence->original mapping is NOT recoverable from either journal format, so
the undo now REFUSES (rc=2) when one (file, after) group maps back to differing
`before` texts, naming the file and telling the caller to git checkout it (R43:
refuse, never mishandle). Journals additionally record per-file sha_before, and a
clean undo hash-verifies its own result and reports HASH-MISMATCH loudly. Old
list-form journals are still read.
gather_externs carries file-scope externs out of the EXEMPLAR's TU and prepends
them. When the destination TU already declares the same symbol with a DIFFERENT
spelling that is a `conflicting types` error -- the documented cap on this lane.
Build the rename table BEFORE gathering so each carried extern is judged under its
DESTINATION name (R48), then drop only a GENUINE conflict; a duplicate-identical
extern is legal C and is kept, so nothing the body needs is ever removed.
HONEST SCOPE: negative-controlled A/B over all 53 d<=1 twin candidates -- 52/52
generated drafts BYTE-IDENTICAL to the pre-fix output, 0 changed. 37 of the 52 do
carry externs (152 total), so the filter had inputs and found no conflict: the decl
environment is NOT the binding constraint for this population. Kept as a correct
defensive guard, not as an unlock. Verified the guard actually runs (dest TU
resolves, tu_decls returns 2,712 symbols) rather than silently no-opping.
gate_stage runs with cwd=<worktree>, so a RELATIVE --drafts path resolved inside
the worktree. .run/ is deliberately not linked into a worktree, so every plan
pointing at the project's own scratch convention (R12: scratch lives under .run/)
landed on a nonexistent path: gate_stage found 0 drafts, banked 0, exited rc=0.
A clean success reporting a TRUE number about an EMPTY world -- the dominant
defect class in this codebase (silently-narrowed-tool-scope).
Measured: 35 binaries / 57 drafts all "banked 0" in 1-2s each, while the SAME
drafts gated IN-TREE banked 15/16 (ov_SC06_011) and 3/6 (ov_SC06_029). After the
fix the same worktree job takes 100s instead of 1s -- it is actually building.
Also refuse a job whose drafts are unreadable (R32/R43) rather than let it report
"banked 0" as though the drafts had failed -- the same shape as the existing
missing-generated-inputs refusal directly below it.
- ran residual_rules_b over the WHOLE open frontier (1,312 cases, 0 errors, ~2min, $0)
instead of a 50-row sample; artifacts in .run/S70_*
- DENOMINATOR (Drew's correction, R41): main's 960 PsyQ LINKED stubs are not
matching targets; true frontier = 355 (67 main REAL + 288 non-main), partitioned
with progress.linked_subsegs() rather than a hand-rolled filter (R33)
- discriminating test settles population-vs-coverage: fire rate DOES rise as
residuals get clean (35.3% at <=8 vs 6.1% at >64) but 57% of the cleanest band
is still UNKNOWN -> coverage binds where rules are worth writing
- hand-label 4/4 labelable to existing cookbook buckets; WIDTH/lhu!=lh has its
discriminating sig already computed and still returns top=None
- 86 REAL standalone MATCHES (closeness 0) = 24% of the frontier, blocked on TU
plumbing only -- outranks the rule work (standalone-match-is-not-bankable)
- logs 4 instrument defects in my own probe, incl. one wrong answer reported to
Drew before checking: 4 of S68's 10 autodecl MATCH drafts are STILL OPEN
- tools/r22_verify.sh from a clean tree: clean rc=0, extract-all 212+main rc=0,
check-all 213 passed / 0 failed of 213 (2m49s). Clears the S69 --no-r22 debt.
- R38 read of the recorded measurement behind the "1-2% ceiling" (S68 eval set +
.run/rules_b/eval_results.jsonl) before designing the queued probe:
* citation fix: the design is Fable-1 (.run/S69_fable/report.md:93), not Fable-2 §7.7
* denominator fix (R41): shape rules can only fire on the 39 near rows, not 113;
real fire rate 3/39 = 7.7% (5/39 with REDRAFT), and 14 are UNKNOWN
* the probe as written is unrunnable: backlog has 125 rows / 20 with residual text
and the UNKNOWN pile is 14 -- sampling 50 would report a narrower world (R41/R32)
§400 — a baseline check that conflates "absent everywhere" with "changed under
us" silently drops new files. The general law: when a comparison uses two
different sentinels for "nothing" ("" from a failed command, None from a missing
file), it reports a difference that does not exist — and in a GUARD, a phantom
difference becomes a refusal, which looks exactly like the guard working.
Corollary recorded in both §400 and the carve-state memory: "never blanket-add"
covers SHARED carve state (overlays.mk, splat yamls). It does NOT cover a carve's
own new per-binary source file, which is named by a committed yaml and whose 31
siblings are tracked — that one must be adopted with the bank that created it.
Docstring correction: parallel_gate does NOT use `git add -u src/` (that is
gate_stage's form); it adds exactly the adopted paths. My first diagnosis of this
bug blamed `-u` on the strength of that stale line and was WRONG — the cause was
the baseline comparison. Noted in the docstring so the next reader is not
misdirected the same way.
Root cause of the 8 untracked src/ files. The merge-safety check compared:
base = sh(["git","show", pin:path]).stdout -> "" when the path is NOT at the pin
cur = open(path).read() if exists else None -> None when absent from the main tree
if cur != base: REFUSE
For a file that exists in NEITHER — exactly what a jtbl carve creates when it
splits a TU into src/<bin>/<bin>_jr_<addr>.c — that is `None != ""`, so every
carve-created file was refused as "main tree moved under them" and never added.
Nothing failed locally: the file is on disk and R22 passes. But config/splat.<bin>.yaml
names the subseg and IS committed, and 31 sibling _jr_ files in the same binary are
tracked — so a fresh clone (or a push) got the config without the source. Eight
accumulated in one session and only surfaced because the dirty-tree guard refused a
later run.
Fix: distinguish "not at the pin" from "empty at the pin" via git show's RETURN
CODE, so absent-in-both compares equal and the file is adopted. New adoptions are
reported explicitly ("N NEW file(s) created by a carve, now tracked") rather than
merged silently — adopting a brand-new source file should never be invisible (R32).
The `git add -- <adopted>` step was always correct; it simply never received these
paths.
`parallel_gate` commits with `git add -u src/`, which updates TRACKED files and
cannot add NEW ones. A jtbl carve SPLITS a TU, creating `src/<bin>/<bin>_jr_<addr>.c`
— so every carve landed its yaml change (tracked) while leaving the new source
file UNTRACKED.
Why this mattered: `config/splat.<bin>.yaml` is committed and names the subseg
(`- [0x577f8, c, ov_SC02_000_jr_8017F950]`), and 31 sibling `_jr_` files in that
same binary are tracked — so these are source by convention, not build artifacts.
R22 passed locally only because they exist on disk. A fresh clone, or Drew's
push, would have the yaml without the file.
Found because parallel_gate REFUSED to run with an unclean tree (rc=1) and listed
them — the guard did its job; the earlier `REFUSED 8 (main tree moved under them)`
line in the carve gate was the same eight files.
TODO for the tool: parallel_gate's commit step must add NEW files under
src/<binary>/ that its own carve produced (narrowly, per-binary — never a blanket
`git add src/`, per the carve-state discipline).
Caught by Drew asking whether the last waves were harvested. They were not: I
banked 1 of 5 (§398b) and left four lever sets in the notifications. Also found
two paid-for MATCHes that were never staged or gated.
(a) a fence BETWEEN two prologue loads, where source reorder does nothing —
the order is fixed before statement order matters (md_MAIN_013/func_800CB56C)
(b) SINK a call into BOTH arms and let cross_jump keep only the jal suffix;
88ins/close86 -> 92/13, then §3-T2 field order let each sh $zero fill an lhu
load-delay. Duplicate in source so the compiler merges, rather than writing
the merged form yourself (ov_SC07_001/func_8017EDC0)
(c) a $v0->$a0->$s3 DOUBLE COPY is a two-pseudo tell: SImode temp for the compare
+ separate HImode var for the tail (70->37); plus §195-N precondition 5 —
nesting `return 1` with ONE trailing `return 0` blocks jump.c's store-flag
transform so reorg fills both delay slots (18->0) (ov_SC02_017/func_8018347C)
(d) the re-tie as a BIV KILLER: a second set makes n_times_set>1 so loop.c
refuses the pseudo as a biv, killing the combined address giv. volatile was
worse, a dead read did nothing (ov_SC07_001/func_8017E4DC)
(d) makes THREE distinct uses of the zero-byte re-tie in one session — §380
un-hoists a move_movables invariant, §393 kills the scheduler's birthing boost,
§399d denies a biv. One line, three passes: when a single-set pseudo is being
treated specially, give it a second set.