THE FIX (harvest_verify._jtbl_prep_one): isolate WITH THE BODY STILL SPLICED.
jr_isolate_all accumulates each object's file-scope decls as the new region's
`ambient` set, so partitioning around an INCLUDE_ASM stub hands the region a
DIFFERENT decl context than the draft's body needs — and the gate then produced a
byte-DIFF rather than a compile error, which is why a batch of these read as
"9 compile / 0 bank" and looked like a codegen wall. §61b's law extends one step:
THE CARVE MUST FOLLOW THE SPLICE — AND SO MUST THE ISOLATION.
The un-splice existed only because the tool is stub-centric (a spliced fn is no
longer in corpus.stubs). New _unsplice_body() handles that properly: it finds the
file that NOW holds the body (isolation may have MOVED it into a fresh region file)
and writes back the stub line for THAT subseg, so the gate re-splices identical text.
BANKED via the corrected path (each whole-binary byte-gated, one per invocation):
func_80135D20 (100 ins, reach 138)
func_801749C8 (105 ins, reach 136)
func_8019059C (673 ins, reach 3 — a giant)
Together with func_80135888 and func_8018F694 earlier: 5 of the 11 preserved
t5wave cracks are now banked.
THE REMAINING 6, honestly classed: 2 genuine DIFF (func_80135260, func_80191C50),
2 CC1-FAIL (func_801365B8, func_80165CA0), 2 §57 self-decl plumbing
(func_801299C8 `prototype declaration`, func_8012AAAC own-name conflict).
⚠️ RETRACTION: the earlier "the ladder converts 0/10, which prices Task 14 stages
2-3" is WRONG and is withdrawn. Three of those ten bank through harvest_verify
alone with the SAME transformed drafts that gate_stage rejected — so that 0/10 was
measuring gate_stage's own interference, not the residuals. Prime suspect is the
arity pre-pass (fix_arity_callers --apply --any-proto edits engine_core.h before
the gate; the §19 regression mode). GATE_NO_ARITY=1 is the ready A/B. Task 14
stages 2-3 remain UNPRICED until that runs.
- R22 clean-fleet 140/140 BYTE-IDENTICAL; tree clean before and after
THE DIAGNOSIS (the recommended next task) did not find an image-level difference:
performed manually, there is no difference. The recipe
splice draft -> jtbl_carve (NON-CONTIGUOUS) -> jr_isolate_all --only <fn>
WITH THE BODY STILL SPLICED -> make extract -> jtbl_carve -> make extract -> build
yields BYTE-IDENTICAL (R22 clean-fleet 140/140). The draft was never wrong.
THE DIFFERENCE from the ladder path is one line: harvest_verify._jtbl_prep_one
UN-SPLICES before isolating, so jr_isolate_all partitions a TU in which the
function is still INCLUDE_ASM. jr_isolate_all accumulates each object's file-scope
decls as the new region's `ambient` set, so partitioning around a stub gives the
region a DIFFERENT decl context than the one the draft's body needs — and the gate
then builds a byte-DIFF rather than a compile error, which is why it read as
"9/10 compile, 0 bank" and looked like a codegen wall.
So §61b's law — THE CARVE MUST FOLLOW THE SPLICE — extends one step further:
THE ISOLATION MUST FOLLOW THE SPLICE TOO. The un-splice exists because after
isolating with the body in, the function is a real C def and corpus.stubs no longer
lists it (the tool is stub-centric). That is a fixable plumbing problem, not a
reason to isolate around a stub.
- func_80135888: 113 ins x reach 138 = 15,594 instruction-instances (~0.12pp) once
swept; banked x1 here, sweep is its own batch (§55b)
- R22 clean-fleet 140/140 BYTE-IDENTICAL; tree verified clean before and after
- the other 9 residuals are now expected to be the same class — to be re-run through
the corrected sequence, one per invocation (§61c fault 2 stands)
The session-7 checkpoint gated the entire jtbl track behind one finding: the
carve+isolation path yields a bank that is incrementally valid and clean-invalid
(139/140, [FAIL] ov_SC06_018, "twice, identically"). The prescribed diagnosis
(diff the incremental vs clean object set) never ran, because the failure does
not reproduce.
MEASURED, with the bank applied through the single-function automated path
(harvest_verify --chunk 1 -> [jtbl] carved -> + chunk(1) -> BYTE-IDENTICAL):
per-binary clean (rm asm+build; extract; build) -> BYTE-IDENTICAL cbbc4f44
make clean && extract-all && check-all (run 1) -> 140 passed, 0 failed of 140
make clean && extract-all && check-all (run 2) -> 140 passed, 0 failed of 140
ATTRIBUTION (best-supported; the failing tree is gone): the 139/140 runs were
taken on the tree left by the BATCH _jtbl_prep (6 table-bearing -> 1 carved,
4 isolate-FAILED, 1 stale-asm carve fail) — five failed preps' residue of
stranded carves + half-applied isolations. The per-function snapshot-restore
that removes exactly that residue landed AFTER those runs, in commit:0803, the
same commit that named the blocker.
THE LESSON (R35 on ourselves, -> decision-log): "twice, identically" was not a
replication — two reads of the SAME contaminated state is one observation. A
replication must RE-CREATE the state, not re-run the check. Standing guard:
re-apply a fault from a known-clean tree before writing it down as a property
of the mechanism. Sixth "structural wall" to resolve to our own tree/tooling.
- BANKED: func_80135A4C (181 ins) x1 in ov_SC06_018 — isolated into its own
code subseg + .rodata carve (single-table, no JTBL_PADS; tail3..tail18 renumber)
- §61c faults 1-2 STAND: a stranded carve poisons the overlay; per-function undo
is unsound in a batch -> ONE jtbl draft per harvest_verify invocation.
jr_inventory's 1:1 ownership assertion was right and is unchanged.
- UNFROZEN: this family = 138 members / PURE / 24,978 ins ~ +0.19pp (jtbl_family_bank,
§53 carve law); the 9 preserved t5wave cracks (Task 14 stages 2-3, §57 plumbing)
- R22 clean-fleet 140/140 x2; tools-health OK (dedup 1848/0, C1 234481/234481,
cdecl 53189/53189, audit-binaries 140); 0 NON_MATCHING (G4)
- fleet 78.0% instr / 66.5% distinct / 87.95% fn-count
- also: preserve the 4 untracked wave-4 .o0 drafts (R20); killed an orphaned cc1
from the Jul-21 session burning a full core for 13h23m
PROPAGATION (§55b, its own targeted batch): dedup_propagate --addr 0x80141B90 --recover
-> "138 overlays byte-identical after propagation"; 117 remaining stubs -> 0; 1 new
dedup group. This was the ONLY one of the 21 directed-run banks worth propagating.
THE REPRICING (R14 — measure a bucket's VALUE, not just its conversion rate):
the directed run converted 27% (21/77) but moved the fleet ~0.03pp, because h_exact
reach of the 21 is: func_80141B90=138, TEN at reach-1 (nothing to propagate), rest 2-10.
Instruction-weighted, the ENTIRE permuter bucket is worth ~0.36pp at 100% conversion.
The mechanism is validated; the fuel was small. Priced frontier (ins-weighted / 13.08M):
LENGTH-DRIFT |d|<=2 472,178 ~3.6pp (339 fns) <- the real permuter-adjacent lever
integration 419,162 ~3.2pp (305 fns) <- Task 14's ladder
WIDTH 71,593 ~0.55pp (45)
permuter (current) 46,571 ~0.36pp (74)
BRANCH-POLARITY 9,462 ~0.07pp (22)
So WIDTH/BRANCH-POLARITY are NOT worth prioritizing; my earlier "~200 candidates"
framing undersold LENGTH-DRIFT 10x and oversold WIDTH.
NEW: permuter_weights._LENGTH profile (perm_temp_for_expr/perm_expand_expr are the only
passes that change instruction COUNT; the reorder/decl-order levers that dominate the
regalloc+schedule profiles cannot, so they are down-weighted here) + residual_class
._drift_route (|d|<=2 -> permuter/`length`, larger stays structural — same class,
opposite tool) + classify() accepts a PROFILE NAME directly (the measured profile beats
re-parsing a free-text label). 17 unit tests green.
grinder: --profile filter (probe ONE residual class's conversion) + a PERSISTENT attempt
ledger. `tried` was in-process only, so every fresh --once run re-permuted the previous
run's losers — the permuter is deterministic given (base.c, target.o), so that CPU can
never produce a new win. Measured: a 20-target probe drew 19 already-tried targets.
Keyed by draft_sig so an improved draft legitimately re-opens the function.
Propagated the func_8015C32C jump-table exemplar to 110 more overlays via
jtbl_family_bank (carve + ld_interleave + remap, per-sibling whole-binary gated).
Run halted by an uncaught exception at member 111 (ov_SC06_018); 110/119 attempted
banked clean. Remaining ~19 members to follow.
First jtbl core of crack-wave3: gcc emits the switch jump table into .rodata, so
each overlay needs jtbl_carve + ld_interleave. Path validated 9/9 BANKED.
-O0 core, single jump table jtbl_801D82FC carved into the o0 object's .rodata via
jtbl_carve (fit contiguously with the existing o0 carve). ov_SC01_077 byte-identical (R22 d19c9580).
These reference a per-overlay tail work-buffer whose base address DIFFERS per
overlay (h_seq relocated data). remap_hseq keyed the byte-OFFSET addresses
(D_801D9C21..) not the base symbol D_801D9C20/D_801D9C60 the exemplar C uses, so
it left them unresolved -> 0/137. Per overlay: base = symbol_map[D_801D9C21]-1;
declare dlabel D_<base> in config/symbols.<ov>.txt (byte-neutral, re-extract emits
the linker def), remap D_801D9C20->D_<base1> / D_801D9C60->D_<base2>. 3 SC07 stragglers
needed the carried ApplyMatrixSV/RotMatrixYXZ externs dropped (TU already declares
them with a different sig -> conflicting types). Each gated whole-overlay byte-identical.
The hexR=138 dedup core (banked ×1 in commit:0716) -> dedup_propagate --addr
0x80150170 --source-overlay ov_SC01_077 --recover: 138 overlays rebuilt
byte-identical, 1 new group registered in config/dedup.us.yaml (0 stubs left).
~+13k ins (95 ins × 137 new members). (First attempt SIGTERM'd mid-gate at the
2-min timeout -> reverted the half-gated state, re-ran clean fail-closed.)
Chunk stopped at 8/50 (3 BANKED / 5 gate-fail) to read the real per-sibling error instead of
churning the ladder (§55b). Mid-flight ov_SC01_080 reverted clean.
Single-table carve jtbl_801D836C (27e, 4-mod-8 first-table = placement-only) into the _o0
subseg; draft spliced clean, no reconciles needed (its only blocker was the missing carve).
Whole-binary gate [ OK ] sha1 d19c9580 == check.
The full 138-overlay family (424 ins) is now banked: exemplar + 137 siblings, 0 failures across
all 3 chunks — the first jtbl giant family completed through the §8e pad-spec mechanism.
Logs .run/sweep_80131340_c{1,2,3}.log.
The 2 Task-3 cores banked x1 but skipped by dedup_propagate ("not self-contained: local types").
Lifted EntSC01077 (func_8014E284) + P_TAG_80137DD4 (func_80137DD4) into src/shared/engine_types.h
(fleet-included via engine_core.h), and inlined func_80137DD4's file-local `#define OTE` into the body
(byte-neutral macro expansion, re-evaluated per use to preserve codegen). Both now self-contained ->
dedup_propagate --recover = 138 overlays byte-identical, 0 stragglers, 2 new dedup groups.
~+32.7k ins (108+129 x138). R22 clean-fleet 140/140 byte-identical; tools-health OK; dedup 1846->1852;
C1 coverage 234205. §55c local-type propagation cap lifted for these 2.
SESSION-2 close: this session banked ~94k ins across 4 fns x~137 overlays (2 non-jtbl giants fully
propagated + this 2-core type-lift); fleet 71.4->72.1% instr (+0.7pp), 140/140 throughout. The 4 jtbl
giants remain deferred on the byte-proven 8-align jtbl-carve gap (root cause half-pinned: cc1+maspsx
both emit .align 2, so the +4B pad is a downstream as/ld_interleave artifact) -> teed up as the next task.
Task-4 giant-bank #2. func_8014F4C0 (141 ins) banked x1 in ov_SC01_077_after.c (its earlier
"gate reject" was pure §55b propagate-damage — gated clean on the healthy tree, no fleet change).
h_exact family -> dedup_propagate --addr 0x8014F4C0 --recover: ov_SC01_000 was a cross-overlay
straggler (all-or-nothing h_exact), --recover reconciled the conflicting caller externs and kept
it -> 134 overlays byte-identical after propagation, +1 dedup group in config/dedup.us.yaml.
R22 clean-fleet 140/140 byte-identical; tools-health OK; ~+19k ins.
GIANT TAXONOMY (session finding, cookbook §56 + CURRENT_PHASE): the 12 preserved giants split into
- NON-jtbl (func_8013FAF8, func_8014F4C0): bank clean on a healthy tree, propagate x137 via macro/h_seq.
- jtbl (func_80131340/func_80159C84/func_8013C414/func_8013F350): each needs a per-overlay jtbl carve
x137 AND hits an 8-align gap -> func_80131340 DEFERRED (byte-proven: gcc emits a non-first jump
table .align 3 while the original packs it 4-aligned -> +4B padding shifts the whole data island,
+5B/3077-diff image-wide %lo breakage). A jtbl_carve 8-align/isolation fix unlocks ~4 giants x137.
- targeted dedup_propagate --addr per core (NOT --auto-from), --recover for stragglers:
0x8014ADE0 -> 138 overlays byte-identical
0x801325B8 -> 134 (ov_SC07_011 byte-diverges -> auto-excluded, kept x1 — what --recover is for)
0x801387B8 -> 138 overlays byte-identical
= ~410 member-instances; 3 new dedup groups (1840 -> 1843), C1 coverage 233795/233795.
- 2 of the 5 banked cores (func_8014E284, func_80137DD4) stay ×1: "not self-contained (local types)"
-> blocked on the build_engine_types type-lift (the §19/§20 propagation cap). Carried.
- R22 clean-fleet 140/140 BYTE-IDENTICAL; audit-binaries OK; dedup 1843/0; 0 NON_MATCHING (G4).
Fleet instr 71.0 -> 71.4% / fn-count 86.30 -> 86.42% / distinct-code 53.3%.
- SELF-CORRECTION (R14/R35), now fixed in cookbook §55c + CURRENT_PHASE: my earlier claim that this
propagate "needs ~2h+" was WRONG. That timing was taken while the tree still carried the partial
damage of a killed --auto-from (90/140 overlays broken), so every member-gate was failing/retrying.
On a HEALTHY tree a targeted --addr propagate is ~233s/core (all 3 = ~27 min) — ~20x faster. Only
--auto-from is genuinely fleet-slow. A timing taken on a broken tree measures the breakage, not the
tool — recover the tree FIRST, then measure.
- cookbook §55: the wave's new byte-proven levers (§49-variant birthing-boost suppression via
reg_n_sets 1->2; sched1 birthing/LUID + "cc1 -dL" movable introspection; switch-tree vs jtbl
CASE_VALUES_THRESHOLD=5; block-scope-extern beats *(T*)&sym) + the GATE-ORCHESTRATION law
(--no-propagate per group then ONE targeted --addr; commit banks BEFORE propagating; a reverted src
needs a re-extract; gate_stage's default harvest_verified.txt accumulates -> phantom banks).
Completes T4 and corrects two defects I introduced, both landed in commit:0649.
- WIRED: 006 1543/1614 · 007 1544/1615 · 010 1544/1614 · 011 1543/1614 = 6174/6457 = 95.6%,
~0 agent tokens. Stubs/overlay ~2400 -> 831/984/898/825. Fleet instr 67.0 -> 68.9%,
fn-count 82.16 -> 83.94%. dedup-check 1840 validated / 0 failed; groups now read
"138 members [138 binaries]" (was 134); C1 coverage 227211 -> 233385 = exactly +6174.
R22 make clean && extract-all && check-all -> 140 passed, 0 failed of 140 at every stage.
- FIX#1 — I DESTROYED THE REGISTRY'S DOCUMENTATION, AND EVERY GATE CALLED IT GREEN (H5).
The first cut wrote config/dedup.us.yaml with yaml.safe_dump, round-tripping the whole file:
47 comment lines -> 0 (including the curated Phase-11 header explaining WHY the share is
source-level) and 1832 `vram: 0x80162FF4` -> `vram: 2148937716` (PyYAML parses YAML-1.1 hex to
int; dumps int as decimal). 25,948 lines rewritten. It passed dedup-check 1840/0 AND check-all
140/140 because _addr() accepts both forms: THE DATA WAS CORRECT AND THE DOCUMENT WAS RUINED.
Fixed forward (R6, no history rewrite): restored from commit:0649~1 and re-applied the 6174
memberships via a surgical text edit (add_members_surgical). Verified: 1545 insertions / 1545
deletions, 0 non-`binaries:` lines changed, 47 comments + 1908 hex fields intact, and the
rebuilt fleet is byte-identical to the destructive version (140/140).
THE LESSON: every oracle this project owns measures BYTES, so a formatting-destructive write is
invisible to all of them by construction. R34 says the byte-gate is a null COVERAGE oracle; this
is the same hole one layer out — it is a null DOCUMENT oracle too.
- FIX#2 — I MIS-REPORTED THE DIFFs, TWICE (R14).
(a) commit:0649 claims ov_SC07_006's 71 non-banks were "ALL PLUMBING, ZERO DIFF". FALSE — I read
head -6 of the classified file and generalized. It has the same 4 DIFFs as the others.
(b) I then built the jr guard assuming those 4 were the §53 jr class BECAUSE ov_SC01_077 hosts
them in _jr_8017A4AC.c / _jr_80182268.c. has_mid_jr is FALSE for all four (33-52 ins, no
jump table): they merely live in a carved jr-REGION split, which sweeps in every function in
its address range. HOSTING FILE != FUNCTION CLASS.
The guard is KEPT (preventive, §53-correct, currently skips 0 — no jr fn is in the extendable
set) with its docstring corrected to record what it is NOT. The 12 DIFFs (0.19%) are UNDIAGNOSED
and logged, correctly left as stubs by the gate — not dressed in a story.
- The 283 non-banks: 271 PLUMBING (the loose-typing conflict class + the whale, whose body lives
in src/shared/func_80144B9C.h so no DEFINE macro exists to expand) + 12 DIFF. Existing tools
cover the plumbing (cast_call_sites / canon_sig_reconcile / reconcile_tu).
The 4 SC07 overlays P27 onboarded were byte-clean but NOT citizens: their .c included only
common.h (never ../shared/engine_core.h), so no shared body could reach them, and they
appeared in ZERO dedup groups (1689 groups read "134 binaries", never 138). Each sat at ~80
matched / ~2400 stubs while its siblings were ~2150 matched.
- NEW tools/dedup_extend.py — the missing mode. dedup_propagate is built for CRACK -> AUTHOR
MACRO -> INSTANTIATE: --auto-from scans INLINE DEFS (planned only 11 here; the ~1600 shared
bodies are ALREADY DEFINE_func_* macros in engine_core.h) and --addr dies "no source overlay
has it matched" because no overlay holds an inline def. Extending an existing MACRO-BACKED
group to a newly-onboarded binary is a different operation and nothing implemented it.
- SAFETY (explicit — this feeds the byte-gate): h_exact is the SHA1 of RAW INSTRUCTION BYTES, so
two instances sharing one are identical INCLUDING their jal/lui/%lo reloc immediates — same
callees, same data addresses, same symbols. The body that compiles byte-identically at one
member does so at the other with NO remap. (Exactly why dup_report calls h_exact "guaranteed
byte-match" and h_norm "candidate-only".) A bug here can only FAIL TO BANK, never falsely bank.
- REUSE, DON'T REBUILD (R33): owns only the set computation + the registry edit. The splice and
the gate are harvest_verify verbatim (it already derives each stub's home TU from the corpus
oracle, chunks + bisects, reverts on failure). h_exact members are byte-identical by
construction -> the happy path is ~1 build per binary, not one per function.
- RESULT ov_SC07_006: 1543 / 1614 banked = 95.6%, ~0 agent tokens. Stubs 2374 -> 831.
The 71 non-banks are ALL PLUMBING, ZERO DIFF, in two named classes with existing tools:
* func_80144B9C "undefined reference" — the whale's body lives in src/shared/func_80144B9C.h
(the -O0 shared header), not engine_core.h, so no DEFINE macro exists to expand.
* "conflicting types for D_800A5E60 / func_8012C750 / func_8012C0EC" — the loose-typing
conflict class (cast_call_sites / canon_sig_reconcile / reconcile_tu already exist for it).
- GATES: R22 make clean && extract-all && check-all -> 140 passed, 0 failed of 140, 0 FAIL lines.
dedup-check 1840 validated / 0 failed; groups now read "135 members [135 binaries]" (was 134);
C1 coverage 227211 -> 228754 = exactly +1543. The second oracle accepts the extension.
- Mechanism had been proven by hand first (probe-before-investing): +include + ONE stub ->
DEFINE_func_80128158() -> ov_SC07_006 built 7ca772be BYTE-IDENTICAL, then reverted.