mirror of
https://github.com/Druthulu/BFM-decomp
synced 2026-09-27 22:45:39 -04:00
85fb289db582d842fc41dc059fa187bb992e76ea
118 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bcc3130eb4 |
feat(phase-30 S49): the A-prop word-diff card + aprop_wave — 56 banked from the >=16 head (§170)
- NEW family_cousins.py --aprop-cards + tools/wave/aprop_wave.js: lane A (1,700 open fns / 76,419 ins) had NO card type — cousin diffs are empty for h_seq-identical members, so the card is a positional WORD diff vs the matched sibling, grouped BY FAMILY (one agent, N drafts). Head cards: 13 families / 433 members, median TWO differing words each. - calibration 9 batches / 108 members: 98 agent-MATCH (91%, best of any wave) -> 56 BANKED (57%), ~80k tok/banked fn vs 157k (cousin card) vs 400k+ (crack wave). R22 213/213 BYTE-IDENTICAL. - HONEST GAP (R14): 91% agent -> 57% gate is the worst conversion measured; 14 groups banked 0. Hypothesis TESTABLE not proven — family batching concentrates members per destination TU, the §169 collision. Re-gate unbanked ONE PER TU before scaling the remaining 320. - >=16 head diagnosed: 3 of 4 blockers are plumbing — the --band substantial default hid 5 of 13 families from every prior sweep; one missing file-scope extern (D_801ED98C) gates 56 PURE members; dedup_extend is macro-only. Only func_8017C294 is a genuine crack. - fleet 96.56% fn / 95.2% instr / 89.9% distinct; stubs 12,535 -> 12,468; dedup 2,043/0. - cookbook §170. |
||
|
|
79b7ff2cbf |
chore(phase-30 S49): wave 7b — adapt lane scaled, 44 banked (92% MATCH->bank); the TU-spread law
- thresholds relaxed to <=6 blocks/<=16 tokens UNION edit-fraction <=0.20: cards 518 -> 721, MIXED 310 -> 50 skeletons; the 753-ins func_8017BEBC (0.987 sim) became reachable. - 59 cards -> 48 agent-MATCH (81%) -> 44 BANKED (92% MATCH->bank, 75% end-to-end), 6.9M tok. - FINDING (the actionable one): 7b's bank rate crushed 7a's because it SPREAD 48 drafts over 35 destination TUs; 7a's failures were per-TU declaration collisions between sibling drafts. Cookbook §169 updated with the spread law. - R22 213/213 BYTE-IDENTICAL from clean; fleet 96.55% fn / 95.2% instr / 89.9% distinct; stubs 12,584 -> 12,535; dedup 2,035/0. - incidents 3 & 4 recorded: an agent wrote a TRACKED header (guard caught it, prose is not enforcement); my own gate_lane filtered on the wrong key and printed 'gating 0 drafts' as a result (R32 silent skip) — fixed with a coverage assertion that refuses to report 0. |
||
|
|
dcba5d0f4f |
feat(phase-30 S49): the micro-adapt lane — adapt cards + adapt_wave.js (wave 7a)
- family_cousins.py --adapt-cards: per seeded-unit member, drift classified vs the seed (LI-ONLY 27 / SMALL-EDIT 491 / MIXED 310 excluded); cards carry the seed C location + the aligned diff blocks with the member's raw words + disasm (the new constant is readable in the card). 518 cards / 1,101 instances / 23,820 ins; 514 haiku-band. - tools/wave/adapt_wave.js: the EDIT-contract wave (crack_wave contracts preserved: per-agent dirs, sha1-last, UNVERIFIED != refuted); symbol surface from the TARGET .s; haiku<=60/sonnet. - regen chain absorbed the 48 lane-A banks (A-prop open ins -6,475 == the report's instr delta exactly — two independent derivations agree); pilot slate .run/wave7a_pilot.json (30). - R37: pilot before scaling to the 518-card pool. |
||
|
|
b3713cc3ca |
feat(phase-30 S49): the cousin tier — family_cousins.py similarity map + seeded wave-7 slate (§168)
- FINDING (Drew's smell, byte-verified): the '4,513 unique singletons' picture is substantially an h_seq exact-hash artifact — 86/120 near-pairs in the 0.85-0.99 band differ by PURE insertion/deletion (li-expansion tell in 25). Specimen: ov_SC06_010:0x8017bebc (753 ins, 'singleton') is 0.987-similar to a MATCHED fn in the same binary. - NEW tools/family_cousins.py: distinct open skeletons -> shingle index -> >=0.85 union-find -> matched-seed attachment -> .run/family_cousins.json + docs/family-cousins.md. R32 BOTH ways (independent stub recount fails loud on a stale map — negative-control-proven; partition assert). Reproduced the probe within +-1%; totals EXACT (11,627 inst / 584,448 ins). - Unit table: A-prop 197u/68,729ins · seeded 418u/50,422 · cousin-multi 1,552u/249,799 · cold 3,240u/215,498 — the genuinely-unique tail is 37% of the remainder, not 90%. Main's 'structurally barren' HOLDS at the similarity tier (94% mass <0.70). - --targets wave slate: .run/wave7_targets.json = 40 targets / 33,304 unit ins (+33% vs family-ranked), 9 resolved seed C paths, size-routed 2 haiku/20 sonnet/18 opus. - LAWS (§168): a cousin is a SEEDED CRACK never a remap; rank waves by UNIT weight; discount short-fn similarity. Byte-gate stays the sole arbiter (G3/P9). - docs/family-hseq.md: this session's frontier regen (post-S48 propagations) rides along. - cookbook §168 + SETUP inventory row (R16/R21/R30); CURRENT_PHASE S49 entry. |
||
|
|
4144eabc74 |
chore(phase-30 S48): checkpoint — 684 banked (13,345 -> 12,661), R22 213/213
Wave 6 added 110 (24 cracks + 23/24 families propagated). Fleet 95.1% instr / 89.7% distinct / 96.51% fn-count. R22 clean-fleet run 9x this session, 213/213 every time. Bank rate across six waves: 67/79/69/68/73/60%. Records §166a (the destination-TU oracle) and the four-instance pattern it completes: a tool asserting a conclusion it never reached. A confident wrong label costs more than a missing one. |
||
|
|
4181b865f2 |
chore(phase-30 S48): checkpoint — 574 banked, FLEET CROSSED 95% instr, R22 213/213
Wave 5 added 152 (29 cracks + 28/29 families propagated) — the session's largest. Fleet 95.0% instr / 89.6% distinct / 96.48% fn-count; stubs 13,345 -> 12,771. R22 clean-fleet run 8x this session, 213/213 every time. P30's milestone is '>=95% instr fleet, or every remaining overlay stub on a named ledger'. THE FIRST HALF IS NOW MET — T5 (phase close) is a live option. Bank rate across five waves: 67% -> 79% -> 69% -> 68% -> 73%. |
||
|
|
19a138c318 |
chore(phase-30 S48): checkpoint — 422 banked (13,345 -> 12,923), R22 213/213
Wave 4 added 105 (27 cracks + 27/27 families propagated). Fleet 94.9% instr / 89.4% distinct / 96.44% fn-count. R22 clean-fleet run 7x this session, 213/213 every time. Bank rate now measured four times: 67% -> 79% -> 69% -> 68%. Prior-notes seeding 10/12 (was 7/9). func_8017C294 — the x16 family, largest item on the board — is NEAR at 2 ins after three seeded attempts (18 -> 11 -> 2). Also records the 4th comment-blindness defect and its blast radius (one draft comment refused a binary's stub oracle, failed 5 later binaries, and left drafts spliced in src/ so 17 re-gates read a poisoned tree as 0/17), and that the wave harness now lives in tools/wave/ with its contracts written down. |
||
|
|
72932899d5 |
chore(phase-30 S48): checkpoint — 317 banked (13,345 -> 13,028), R22 213/213
Wave 3 added 84 (27 cracks + 21 propagated families). Fleet 94.8% instr / 89.2% distinct / 96.41% fn-count. R22 clean-fleet run 6x this session, 213/213 every time. 96 commits. Bank rate measured three times: 67% -> 79% -> 69%. The dip is the cost curve (wave 3's tier was 29 Opus-band / 14 jr vs wave 2's 8 / 5, median reach x6 -> x3-4), not a regression. Two levers proved out and belong in every future wave: the hardened harness contract (0 drafts lost vs 21) and prior-notes seeding (7 of 9 previously failed targets converted, incl. both long-standing NEARs and all three wave-2 gate misses). NEAR is a resumable state, not a write-off. |
||
|
|
35d7b11d03 |
chore(phase-30 S48): checkpoint — 233 banked (13,345 -> 13,112), R22 213/213
Session close state. Three parts: stage 0b (91, zero decompilation), wave 1 (26), wave 2 (116). Fleet 94.4% -> 94.7% instr, 88.3% -> 88.9% distinct, 13,345 -> 13,112 stubs. R22 clean-fleet run 5x, 213/213 every time. The campaign now has a MEASURED rate, twice: 67% (wave 1, all-Opus) then 79% (wave 2, 20 of 28 Sonnet) of cracks survive the whole-binary gate. The Sonnet band beating the all-Opus wave is the session's most useful economic finding and sets wave 3's routing. Resume order changed on evidence, twice over: - harden the wave harness FIRST (per-agent dirs, sha1-last verifier, and a tools/recover_drafts.py built from the transcript-replay method that recovered 21/21 today); - then wave 3, sized on 79%, not on the reach-15 prior. Error ledger grew to 6. The two that matter: I wrote off 21 verified cracks as lost when the run transcripts held every one of them, and my first two recovery passes both failed by reading a single tool record instead of replaying the file's mutation history. |
||
|
|
551239bd3c |
chore(phase-30 S48): checkpoint — 117 banked (13,345 -> 13,228), R22 213/213
Stage 0b closed (91, zero decompilation) + Stage-1 wave 1 (8 cracks -> 26 instances). Fleet 94.6% instr / 88.8% distinct / 96.36% fn-count. Resume order changed on measured evidence: FIX THE md_ MODULE LANE FIRST. 16 of the wave's 42 member slots were unreachable for tooling reasons, not matching reasons — 12 on a carve that assumes raw data lives in <binary>/data/*.data.s (modules do not), 4 on an uncarried extern (`D_8011511A' undeclared). Both are named with verbatim errors; probe one of each before pricing (R37). Precedent: 0b's three repairs banked 91 for ~0 agent tokens; the wave spent 3.36M for 26. Also recorded: the frontier re-derivation (1,955 zero-crack families / 330,622 templatable ins), the tier-ordering correction (ins-per-crack is flat across x5-x8, so rank by templatable weight, not by tier), and the §162 harvest with its two in-place cookbook corrections. |
||
|
|
d5fbd2630f |
fix(phase-30 S47): family_hseq derives its own SCOPE, not just its own count; + a zero-crack glossary
The targeting oracle stamped its scope as "the N OVERLAYS only (no main, no resident)" while
load() has scanned the md_* modules and the resident since S44. Measured at this HEAD: 141 location
overlays + 70 md_* modules + the resident = 212 binaries. That is the §159 coverage law broken by
the file that documents coverage, on the repo's most load-bearing targeting instrument — and it is
how "main is structurally barren" survived two phases unexamined.
The COUNT beside it was already derived, with a comment saying "report the scope we ACTUALLY
scanned, never a hardcoded count". The PROSE describing what the count meant was hardcoded and
rotted. Both are derived now.
Caught while fixing it: my first cut read glob(".run/sig.main.jsonl") and stamped "main INCLUDED"
the moment that file existed — while load() still did not glob it. Same defect one layer down: a
stamp describing the filesystem instead of the run. Now derived from the loaded instances.
Also added a glossary line: "zero-crack" means n_matched == 0 (needs its FIRST crack) in this map,
and the OPPOSITE (a matched exemplar awaiting propagation) in roadmap §3 T3 — a ~30x mis-scope risk
for any session reading one against the other.
NOT DONE — main inclusion (0c) is still blocked on settling the attribution. Confirmed the
mechanism: main has 49 LINKED PsyQ subsegs, corpus.stubs('main') returns 2,002 INCLUDING them,
progress.py correctly excludes them and reports 1,034 game-code stubs. progress.linked_subsegs'
own docstring records this exact trap ("an importer then classifies ~1,300 already-byte-identical
LINKED library stubs as outstanding game-code work") — and my sig-main seeded from corpus.stubs,
so it inherited the LINKED rows, which is why G2's 207-family finding was inflated.
My partition probe is NOT trustworthy: 954 of 2,002 stubs returned no asm path from
corpus.asm_path, so 199 LINKED / 849 game / 954 unresolved does not reconcile with 1,034. Fix the
probe before trusting any main-scope number.
|
||
|
|
d3f3d8ba22 |
feat(phase-30 S47-W1b/G2): 3 retries banked, 2 new rules; main signed for the first time
W1b — the 3 targets whose agents died on API rate limiting, retried with cookbook §160 in the prompt: func_801EFBF4 (reach 12), func_801EFDC8 (12), func_8018CC40 (10, jr). 3/3 confirmed by an independent verifier, all banked, R22 clean-fleet 213 passed / 0 failed of 213. func_8018CC40 failed the first gate with `too many arguments to function func_80178970` — which its own crack agent had PREDICTED in its report, naming the §17a-1 remedy. Dropped the draft's empty-paren externs and cast 6 call sites instead; banked. Read the agent's integration notes before diagnosing a gate failure — it has already seen the TU. Cookbook §161a-c (index 469 sections): §161a case 0: break; is LOAD-BEARING when a jump table is indexed from zero. The natural case 1..5 makes gcc-2.7.2 pick minval=1, emit `addiu $v1,-1`, and shift every table index — 58 of 77 mismatched on a byte-perfect body. Tell: the table's FIRST entry points at the function's own end address. Family-wide (10 members). §161b aliasing a parameter into a local can force a SECOND callee-saved register (+8 frame, +3 ins) even when uses are mutually exclusive. Suspect it before reaching for register pins. §161c loose-prototype engine helpers: don't fight the TU's (void) decl, cast at the call site. G2 — THE MAIN EXPERIMENT. family_hseq excludes main as "structurally barren — zero h_exact overlap". True and irrelevant: an h_exact claim guarding an h_seq tool. There is not even a sig-main target — main had never been signed for this pipeline. Signed it (2,002 fns, seeded from splat boundaries via corpus.stubs rather than --bootstrap, which glues functions around jtbl dispatch and would have corrupted the hashes under test). Result: main is ~85% singleton work, not 100%. internal h_seq families (>=2): 207 families / 748 fns / 11,537 ins (13.7%) shapes shared with the fleet: 161 fns / 1,346 ins (1.6%) genuine x1 remainder: ~71,034 ins (84.6%) IMMEDIATELY ACTIONABLE: 44 classes / 151 main functions / 1,239 ins already have a matched exemplar in the fleet — free propagation, invisible only because main is not in the map. Long-term: 748 of main's 2,002 functions (37%) are templatable once one exemplar per family is cracked, which refutes "2,002 independent cracks" as the planning assumption for the 79k-ins tail. OPEN, deliberately not done unilaterally: adding a sig-main target and dropping main's exclusion from family_hseq.load() changes a fleet-shared oracle every targeting tool reads. Needs Drew's call. |
||
|
|
ff11fc556c |
feat(phase-30 S47-W1s): the reach-15 wave templates to 140 members (81% conversion)
The 10 exemplars from W1 flipped modal -> matched in the regenerated map, so family_sweep could
template them. 9 non-jr families swept: BANKED 140 member-matches / 32 failed across 50 overlays.
Derived net = report = 140 (no untracked carve files this time, so the two counts agree).
R22 clean-fleet: check-all 213 passed / 0 failed of 213.
Fleet 94.3 -> 94.4% instr / 88.2 -> 88.3% distinct / 96.22 -> 96.27% fn-count; stubs 13,713 -> 13,563.
WAVE ONE, FULLY ACCOUNTED: 10 agent cracks + 140 templated members = 150 functions for 1.36M
tokens (~9k tokens/function). Still owed from this wave: 73 member-slots in 2 NEAR families,
34 in 3 rate-limited targets, 9 in the jr family (routes to jtbl_family_bank, §53).
TWO MEASUREMENTS THAT CORRECT MY OWN FORECASTS (R14):
1. Conversion was 81%, not the 58% I projected from this morning's propagation run. Today's
plumbing fixes (alias-drop, cpp-derived TU type map, group-level draft-vs-draft aliasing) are
paying off in a population they were not tuned for.
2. The effective multiplier was 15x, not the 2-3.5x I predicted. That estimate used the MEAN
family size across the whole zero-crack pool (3.55); this wave deliberately targeted the TOP of
the reach distribution, where families run 10-28 members. Ordering waves by reach is what
produced the difference — the mean was the wrong statistic for a wave that selects on the tail.
The regen step is load-bearing and now byte-proven twice: a fresh crack reads as `modal` until sigs
+ family_hseq are rebuilt, and family_sweep templates only from `matched`. Skipping it sweeps a
stale map and the multiplier evaporates (the Phase-26 finding, whose surviving qualifier is that
remap works BEHIND a fresh crack).
|
||
|
|
f212ebcc28 |
chore(phase-30 S47): refresh frontier docs at HEAD commit:1565
Fleet 94.3% instr / 88.2% distinct / 96.22% fn-count; INCLUDE_ASM stubs 13,713. Frontier (overlays): 6,701 families / 12,679 instances / 685,757 ins. siblings + matched exemplar (propagate): 166 fams / 1,122 members / 61,466 ins siblings + zero-crack: 1,969 fams / 6,991 members / 342,004 ins singleton + matched exemplar: 53 / 53 / 4,451 singleton + zero-crack: 4,513 / 4,513 / 277,836 Zero-crack by size band: <30 ins 2,184 fams/89,785 ins - 30-49 1,738/118,104 - 50-199 2,353 fams/3,800 members/325,223 ins - 200-399 179/67,115 - 400+ 28/19,613. |
||
|
|
efec1b9b71 |
fix(phase-30 S47-A1): asm-label aliases must never be dropped by §8d; +148 members
scope_data_externs §8d drops the draft's decl of any symbol the TU already declares at file scope.
It keys on the SYMBOL, but a §37 asm-label ALIAS binds a DIFFERENT C identifier to that symbol:
the TU declares `D_801851BC`, it does NOT declare `tbl_D_80187044`. Dropping the alias left the
body referencing an undeclared name, which cc1 reports with no `error:` prefix — so the sweep
classified all 132 siblings as CC1-FAIL(no-diagnostic), i.e. as a codegen wall.
The bitter part: the alias exists PRECISELY BECAUSE the TU declares that symbol with a conflicting
type (a `void (*[])(void)` dispatch table vs this function's 20-byte-stride view). The drop rule
fired on exactly the declarations written to survive it. Why 1 of 2 died was fully determined:
tbl_D_80187048's symbol is not in the TU, so it demoted normally.
Fix: is_asm_alias() — an alias is demoted into the body, never dropped (the identifiers differ, so
it cannot collide with the TU's decl). Control-tested 6 ways incl. self-labels and plain externs.
Measured: func_80132018 3/135 -> 135/135; full re-sweep +16 more. Total +148 members.
R22 clean-fleet 213 passed / 0 failed of 213. tools-health OK, dedup-check 1949/0.
Fleet 96.11 -> 96.15% fn-count, 87.8 -> 87.9% distinct; stubs 14,120 -> 13,972 = -148 (2nd oracle).
CORRECTION TO MY OWN CLAIM (R14): after the probe I said the 58% aggregate was concealing a broad
problem. The re-sweep refuted it — only 16 more banks fleet-wide. The alias class really was one
family; the first read ("outlier") was right and the correction was wrong.
875 sweep failures classified: 231 PLUMBING-other, 141 DIFF (real divergence, only 16%),
136 `conflicting types for cdFileLocTable` (ONE symbol — biggest single class left),
77 CC1-FAIL(no-diagnostic), 26 memcpy, 12 D_80114F24, 11 D_800AE620, 9 D_800183E0.
STILL UNFIXED, and the most dangerous instrument left: the sweep's failure classifier greps for
`error:`, which gcc-2.7.2 never emits on hard errors. Every hard error therefore reads
CC1-FAIL(no-diagnostic). That is how a missing declaration looked like a codegen wall across 132
functions. rtu_match was fixed for this at T0(b); this classifier was not.
|
||
|
|
5f001a9392 |
chore(phase-30 S47): refresh derived frontier docs at HEAD commit:1543
Regenerated after the S47-B/C banks (family_hseq.py + report): docs/family-hseq.md,
docs/progress.fleet.md, docs/backlog.md. Numbers only — no analysis change.
Frontier at this HEAD (overlays only): 7,085 families / 14,508 instances / 752,073 ins.
with siblings (>=2): 2,429 fams / 9,852 members / 467,634 ins (62.2%)
- matched exemplar (propagate, ~0 tok): 460 fams / 2,861 members / 125,630 ins
- zero-crack (crack 1 -> templates to N): 1,969 fams / 6,991 members / 342,004 ins
singletons: 4,656 fams / 4,656 members / 284,439 ins (37.8%)
- matched exemplar: 143 / 6,603 ins - zero-crack (pays x1): 4,513 / 277,836 ins
Structural: the x138 era is over — 3 fleet-wide families remain and ALL 3 already have matched
exemplars, so no fleet-wide CRACK is left, only propagation. 82% of remaining code now sits in
the two worst cost profiles (x2-9 zero-crack 45.5%, singleton zero-crack 36.9%).
|
||
|
|
788f33d523 |
feat(phase-30 S45 L3-p3): SC02/9 = the Steam Knight boss module — decoded, captured, retro-verified, onboarded; parked = 5
- the gate DECODED from matched C (func_8012832C case 0x300E -> func_80128998 -> streaming API with &cdFileLocTable[144]) -> scene arithmetic named the 1ST-BOSS arena -> ONE targeted load captured it at 0x801E4C60 - RETRO-VERIFIED: Phase-3's dumps/ram_castle.bin (2026-06-14) holds it at the SAME address, same 6,764-B exact prefix — R10 two independent datapoints two months apart; bossHp_SteamKnight (0x801E4398) lives inside this module's image - onboarded md_SC02_009 (id 0x3E, TLO 0x4): BYTE-IDENTICAL first build; fleet 213; R22 213/213; tools-health OK; audit-disc UNCLAIMED 6 -> 5, residue 0 - the last 5 (MAIN/7, MAIN/9, SC03/53/54/56) reclassified emulator->STATIC-RE targets with decoded leads (memory-map §S45 p3); loc-id map appended to docs/debug-menu-list.txt - negatives banked: pause menu, memory-box prompt, new-game intro, high/low game, Minku spawn (slot-A actor 0x15 = md_MAIN_015 candidate naming) |
||
|
|
fa7b9d4c71 |
feat(phase-30 S45 L3): the emulator tour — all 28 script modules + MAIN/3 onboarded; fleet 212, R22 212/212
- THE TOUR (Drew driving the retail debug menu; mode-7 hammer over the Redux web API): all 28 script modules captured live at four byte-verified per-chapter slots (SC03/73-79 @0x801EF468 ch2-period, SC03/132-138 @0x801E25E8 ch3, SC04/24-30 @0x801E7B28, SC05/23-29 @0x801ED988); the routing law: debug-menu AREA selects the chapter, each CITY interior streams its own module (member k <-> interior k). md_MAIN_011/DISELECT byte-proven 24,236/24,240 in RAM; slots A/B/boot R34-verified live. - MAIN/3 DISCOVERED: the main-menu module (id 0x39, 121,884 B), mis-bucketed as data by BOTH audit oracles; live byte-proven @0x800CEDF8 (42,632-B exact prefix); onboarded. - 29 onboardings BYTE-IDENTICAL on first build -> fleet 212; R22 212/212 after three md_MAIN_003 catches: the A4 DsMix leak; an extract-order-sensitive splat boundary (bytes: a 1-word data sentinel in .text + fn at +4 -> pinned in symbols file); corpus.stubs now treats D_*/jtbl_* INCLUDE_ASM as blob includes (mirrors progress.py) - module-id census (offline, disc-wide): 77 id-law code payloads, 0 further misses; SC03/55 = confirmed DATA. audit-disc: UNCLAIMED 34 -> 6, residue 0 — the 6 carry byte-checked negative evidence; next tier = the CD-read tracer - docs: memory-map §S45 (slots + routing + debug-menu ops), disc-completeness S45 addendum, decision-log R31 entry, docs/debug-menu-list.txt (Drew's transcription) - .run/s45 evidence allowlisted (tour logs/scripts/rosters); 104 ram dumps LOCAL-ONLY - new baseline: 93.8% instr / 95.68% fn / 87.2% distinct over 212 |
||
|
|
4cadac4e11 |
feat(phase-30 S45 II.1c): module batch dedup-banked + verified — 408 banks, R22 183/183, audit-disc 75->34 (parked-only)
- dedup measure (R37 probe): 69/1,113 module fns h_exact-match matched corpus (~6%, LOW as
planned — modules are novel frontier); dedup_extend inapplicable (same-vram group model) ->
family_sweep --hseq --band all over the 57 matched-exemplar families: 408 member-matches
banked (182 into modules, 226 into the big 3 — families Part I's --only scoping missed),
169 failed + 77 STRUCT = genuine per-member frontier
- R22 clean-fleet 183/183 BYTE-IDENTICAL; audit-disc UNCLAIMED 75->34 residue 0 (34 = 31
parked-for-L3 + SC03/53,54,56 — 3 rows Discovery-3 never tiered, now parked with evidence)
- three instrument fixes, each negative-control-proven:
- family_sweep --hseq stub map derives ov_*+md_*+resident (was sig.ov_* glob -> module
members silently 'not-stub', R32 class) [committed earlier as commit:1506]
- sig-modules seeds from the built ELF's func_* symbols (bootstrap GLUES adjacent fns
around jtbl dispatch -> 24 false TRUNCATED; perturbed-sig control still bites)
- corpus.audit counts CODE lines only (module .s carries its header jtbl as .word lines);
progress.py buckets INCLUDE_RODATA symbols as blobs (unbucketed R32 hole)
- NEW HONEST BASELINE (183 binaries): 94.0% instr / 95.96% fn-count / 87.6% distinct;
tools-health OK, audit-digest OK
|
||
|
|
c1d5670f36 |
feat(phase-30 S44 I.2c2): the h_norm/template tier — +290 members banked into the new binaries
- family_sweep --hseq scoped by --only to the 637 families with a matched exemplar AND a member in the new 3 (2,232 stageable; avoids re-gating the swept-dry fleet). BANKED 290 / 81 failed / 6 skipped (unresolved immediates), every one whole-binary byte-gated; all three SHAs green. - Session total into the big 3: 4,836 h_exact + 290 template = 5,126 member-functions. - Remaining stubs 624+717+710 = 2,051 = the ~802 novel functions x instances + the genuinely failed/unstageable tier (the new frontier). |
||
|
|
f6bbe7272a |
feat(phase-30 S44 I.2a): the big 3 onboarded BYTE-IDENTICAL — ov_MAIN_012, ov_SC02_037, ov_SC03_107
- Three uncompressed (PAC type-1) overlays at the standard 0x80128158 slot, onboarded via the new
tools/new_binary.sh, each byte-identical at 100% INCLUDE_ASM on the FIRST build:
ov_MAIN_012 d6b3e8b9 (383,783 B, 2,324 fns)
ov_SC02_037 b0c5394a (661,903 B, 2,434 fns)
ov_SC03_107 87d02b57 (474,087 B, 2,414 fns)
This also BYTE-PROVES the statically derived base (the §S44 loader table + the 500:1 h_exact
vote): a wrong vram could not have produced byte-identical images once symbols resolve.
- Fleet: 140 -> 143 binaries. audit-binaries currently FAILS on all three by design (no
engine_core.h include yet — the SC07-blindness check working as built); dedup_extend is the fix
and the next commit.
- Registered by the script: overlays.mk blocks, check.sha, symbols seeds, the 3 BINARIES dicts.
family map regenerated (3,577 target families / 279 with a matched sib — the new binaries'
members now visible).
|
||
|
|
369dd14f4f |
fix(phase-30 S44 I.1d): the module class reaches every enumerating consumer
- family_hseq: widened from src/ov_*+sig.ov_* to every non-main binary (resident + md_*); the map now carries 139 binaries incl. resident (was overlays-only — which is exactly why the R36 gate's CHECK 4 could never see them). Self-count uses the SAME widened globs (cannot drift). - progress --weighted :647 + audit_frontier :57: + sig.md_* globs. - corpus.sig_is_independent: md_* sigs are sig_image-signed => independent (R34 trust). - backlog alias regex + prefetch_fleet (md_* derived from splat configs) + dedup_propagate (reads modules.mk alongside overlays.mk — excluding modules would re-create the SC07 invisible-work bug one class over). - VERIFIED: family map regenerated with resident (139 binaries); audit-binaries OK over 140; all six tools parse. |
||
|
|
9f61cd33c5 |
feat(phase-30 S6): BOTH GIANT WALLS CRACKED ×138 (+50,094 ins) — the verdicts were stale, not wrong
The two functions the roadmap has carried as PERMANENT WALLS since Phase 24 are matched in all 138
overlays. Neither needed a siege. Both matched from drafts ALREADY ON DISK.
func_80178004 165 ins x 138 = 22,770 Phase 26: Fable5, ~477k tokens, "intrinsic 3-integer
regalloc wall". THREE stored drafts report match_one
MATCH today; one banked first try, no new work.
func_801412A8 198 ins x 138 = 27,324 close=29/110 since Phase 24. Matched from 1 of 31 stored
drafts + the §37/§124 alias.
WHY func_801412A8 LOOKED INTRINSIC (worth understanding — match_one is structurally blind to it):
the TU declares `extern int func_801412A8(int,int,int,int,int,int)` and its callers USE the return
(`param_1 = func_801412A8(...)`), while the byte-true definition is
`Prim_1412A8 *(Prim_1412A8 *, int, int, int, u16, u16)`. Narrow params cannot agree with an `int`
prototype and the no-prototype escape is illegal once a param promotes, so NEITHER side can move --
and the resulting byte difference is in the CALLERS, which match_one never compiles. The §37/§124
def-side asm-label alias decouples them: the TU decl keeps governing the call sites (codegen
untouched), the definition keeps its byte-true signature.
THEN PROPAGATION RETURNED 0/137 TWICE, both times a missing TYPE, not codegen:
family_remap's `_carry_macros` carries file-scope #defines but (a) NOT typedefs, and (b) is NOT
TRANSITIVE -- it brought addPrim_1412A8 and stopped, though that macro calls setaddr/getaddr and
getaddr casts to PTag_1412A8. Lifted Env_1412A8 / PTag_1412A8 / Prim_1412A8 + OT/getaddr/setaddr
into src/shared/engine_types.h (inside the include guard) -> 137/137, 0 failed.
MY ERROR, CAUGHT BY THE GATE: I lifted the typedefs but did not STRIP them from ov_SC01_077.c, so
they were declared twice and gcc-2.7.2 rejects a repeated typedef even when identical -- the lesson
already recorded at the foot of engine_types.h. R22 came back 139/140 with [FAIL] ov_SC01_077 (the
exemplar's own overlay). Stripped, re-verified, 140/140. A proper lift strips the source;
build_engine_types --strip does both and I did it by hand.
Also a measurement error worth recording: I checked whether the draft defined Prim_1412A8 with a
plain `grep -c` -- which matches inside `addPrim_1412A8` -- and briefly concluded the carry worked.
Substring false positive; the same shape as reading a `return` as a declaration.
VERIFIED: make clean && make extract-all && make check-all -> 140 passed, 0 failed of 140.
Fleet 12432941 -> 12483035 instr (+50,094 -- EXACTLY the two giants x138); fn-count +276;
instr-weighted 94.5% -> 94.8%. audit-digest OK. 0 NON_MATCHING (G4).
THE RULE THIS BUYS: re-measure a wall before respecting it, and SCAN every stored draft rather than
sampling (my first pass checked 8 of 31 and reported "closeness 40" for a function whose MATCH was
in the 9th). Four minutes of re-measurement was worth 50,094 instructions.
|
||
|
|
669367dab0 |
feat(phase-30 S40): propagate the 19 wave exemplars — 61/87 members banked (+7,087 ins), R22 140/140
Propagation behind every crack, same session (the multiplier the waves exist for). 19 newly-banked exemplars from waves 1+2, all in the family_sweep lane (0 has_mid_jr): 87 candidate members / 10,212 ins -> 61 BANKED / 26 failed across 39 overlays The 26 that did not bank are the known plumbing shapes, not codegen: 20 CC1-FAIL + 5 callee `conflicting types` (func_8017EFA0 x3, func_8012B23C x2) -- the same classes the S40 recovery ladder already has levers for (§17a-1 no-proto + call-site cast; recover_giant block-scoping). Left open deliberately rather than force-banked (P9); they are the cheapest fuel on the board next session. TOOLING GAP RECORDED: the sweep's classifier writes "CC1-FAIL: make: *** Error 33" WITHOUT the actual cc1 message, so 20 of 26 failures carry no actionable reason. Diagnosing one currently requires manually splicing the draft into its TU and rebuilding (done twice this session). The classifier should capture cc1 stderr the way harvest_verify already does -- worth fixing before the next big sweep, or every CC1-FAIL costs a manual reproduction. VERIFIED: make clean && make extract-all && make check-all -> 140 passed, 0 failed of 140. Fleet 12425854 -> 12432941 instr (+7,087); distinct +6,145 / +51 uniq; fn-count +61. instr-weighted back to 94.5% ON THE HONEST (post-main-regen) denominator of 13,160,961. audit-digest OK. 0 NON_MATCHING (G4). |
||
|
|
443a3e3afe |
feat(phase-30 S40): waves 1+2 bank 24/24 after recovery — ZERO codegen walls; +5,479 ins
Two ultracode waves over the open-only h_norm clusters (the pool nobody had ever aimed a wave at),
pool VERIFIED from the sigs first (R14).
wave 1 8 targets 8/8 match_one 5/8 gate first pass -> 8/8 after recovery
wave 2 16 targets 16/16 match_one 14/16 gate first pass -> 16/16 after recovery
THE HEADLINE IS NOT 24/24 -- IT IS THAT NOT ONE FAILURE WAS CODEGEN. All six first-pass gate
failures were TU-integration plumbing, each with an already-documented lever:
func_801802EC redefinition of morph_lerp strip the §77 PROBE LAYER (the draft carries types +
a static inline so match_one can compile standalone;
the real TU already defines them -- scaffolding is
not part of the bank)
func_8018B238 conflicting types D_80115158 recover_giant: draft declared it file-scope as a
struct array, TU declares u8[] BLOCK-scope inside
other functions -> block-scope the draft's externs
func_8017EF54 conflicting types (SELF) §37/§124 def-side asm-label alias (TU declares
void f(void) for no-arg callers; byte-true def takes
s32 in $a0; no-proto escape illegal once a param
promotes)
func_80183D78 conflicting types (callee) recover_giant
func_8017F278 conflicting types func_80146C3C §17a-1: the fleet canonical is the NO-PROTOTYPE
form + the intended signature applied AT THE CALL
SITE; a concrete prototype collides with it
(wave-2's 14 first-pass banks needed nothing -- the wave-1 lessons were folded into the prompt)
=> the gate number measures INTEGRATION, not matching. Run the recovery ladder before recording a
wave's yield or the metrics under-report the drafters and send the next wave hunting walls that are
not there. docs/wave-metrics.md S40-1.
POOL VERIFICATION (R14, and it cut both ways): the frontier report's cluster pool MEASURED
1,677 clusters / 5,795 fns / 319,755 ins at a 3.68x multiplier vs its claimed 1,689 / 5,956 /
326,261 at 2.7x -- within 2-4%, and the multiplier is BETTER than claimed. The SAME document's whale
claim was 3/4 wrong. Verify each claim separately; do not accept or reject a source wholesale.
ALSO: 24/24 members propagated from wave 1's 5 banked exemplars (0 failed) -- the same machinery
that returned 0/39 before this session's cast_call_sites fix.
NEW IDIOMS, distilled in-session (R16/R30):
§144 the LITERAL'S SPELLING picks the immediate encoding (`cnt + 0xff` vs `cnt - 1`: mod-256
identical, both one addiu, but gcc emits 0x00FF vs 0xFFFF from the source text)
§145a combine_givs ANCHOR RULE -- the address-giv group anchors on the LAST address-giv in SOURCE
order (record_giv prepends, combine_givs takes the head); store order decides the base and a
wrong choice spawns a third induction register
§145b a bare `p = r;` is a COMBINE BARRIER (can_combine_p/use_crosses_set_p) -- it preserves a
pointer-bump addiu that combine would otherwise fold into every MEM offset
§145c chained assignment `a=b=c=0` emits stores RIGHT-TO-LEFT
VERIFIED: make clean && make extract-all && make check-all -> 140 passed, 0 failed of 140.
Fleet 12420375 -> 12425854 instr (+5,479); distinct +5,479 / +43 uniq; fn-count +43.
audit-digest OK. 0 NON_MATCHING (G4). Cost: 3.73M subagent tokens across 24 agents, 0 errors.
|
||
|
|
6e0b1605c6 |
fix(phase-30 S40): cast_call_sites read a RETURN as a prototype and deleted it — 0/39 sweep becomes 18/39
THE BUG. tools/cast_call_sites.py classifies a declaration line with
^([ \t]*)(extern\s+)?([A-Za-z_][\w \t\*]*?)\b([A-Za-z_]\w*)\s*\(([^;{]*)\)\s*;
Feed it a return statement and `return` is a perfectly good identifier where a type is expected:
return func_8012CB64((s32)out, -0xC0, 0x40, -0x60, 0);
^^^^^^ captured as the return TYPE, func_8012CB64 as the DECLARED NAME
so the "rewrite this decl to canonical" path REPLACED the statement with
`extern s32 func_8012CB64(s32,s32,s32,s32,s32);`, DELETING the return. In C89 a declaration after a
statement is a parse error, so the damage surfaced as a bare syntax error in the DRAFT -- reading as
the draft's fault, not the tool's. 9 of 9 staged members of family 0x801848dc lost their return.
fix: a keyword guard (a declaration's type-specifier can never begin with a statement keyword)
family_sweep --hseq --band all over 5 families: 0/39 -> 18/39 banked (only the guard changed)
⚠️ AND THE TRAP INSIDE THE FIX: the obvious R33 move is "route it through cdecl". CHECKED, and it is
WRONG -- cdecl.parse() is a DECLARATOR-GRAMMAR parser that assumes it was handed a declaration; it
reports `return func_X(...);` as declaring func_X and `if (f(a));` as declaring `if`.
Statement-vs-declaration is a question cdecl does not answer. Routing there would have been a silent
non-fix that looked principled. §134's law still holds for line-SHAPE masking; this is a different
question.
BLAST RADIUS (measured, not assumed -- R14): cast_call_sites is in gate_stage's DEFAULT pipeline
(canon_resident_calls -> cast_call_sites -> sig_unify -> harvest_verify) and has been since Phase 20.
Of 44,833 stored drafts, 318 (0.7%) carry a `return f(...);` line this mis-reads, across 67 callees
(func_8014F468 x41, func_8014F6F4 x37, func_8014F74C x32, ratan2 x25). Every one, every time it
passed the gate pipeline, lost its return and failed as PLUMBING. Part of the historical plumbing
tail is this bug.
ALSO IN THIS COMMIT
- S5 CALIBRATION WAVE (8 agents, ultracode, 1.31M tokens). Pool VERIFIED FIRST (R14 -- Fable's whale
claim was 3/4 wrong): measured 1,677 clusters / 5,795 fns / 319,755 ins at a 3.68x multiplier vs
its claimed 1,689 / 5,956 / 326,261 at 2.7x -- its numbers hold, and the multiplier is BETTER.
Result: 8/8 match_one MATCH (close=0), and 5/8 banked whole-binary -- the §52b/§61 gap is
integration, not codegen. Banked: func_801822E0 func_8017EC98 func_801851A8 func_80189A34
func_80188E10 (693 ins x1 before propagation). Not banked: func_8018B238 (FAILED),
func_8017EF54 + func_801802EC (NEAR) -- drafts kept in .run/wave-s40/ for recovery.
- 18 member-banks from the re-run sweep (the cross-address free-h_exact pool: h_exact-identical at
DIFFERENT addresses, which dedup_propagate correctly refuses since it assumes position-locking --
family_sweep is the right lane).
- cookbook §143 (this bug + the cdecl trap + the blast radius); index regenerated.
VERIFIED: make clean && make extract-all && make check-all -> 140 passed, 0 failed of 140.
Fleet 12419169 -> 12420375 instr; distinct +1,526 / +5 uniq; fn-count +23. audit-digest OK.
0 NON_MATCHING (G4).
NEW IDIOM FROM THE WAVE, not yet folded into §31 (agent was told to write only its draft): a byte
counter must be spelled `cnt + 0xff`, NOT `cnt - 1`. Both are mod-256 identical and both compile to
one addiu, but gcc-2.7.2 picks the immediate encoding from the SOURCE SPELLING (0xFFFF vs 0x00FF).
Also flagged: .run/ghidra_c/func_8017EF54.c is a stale decompile of the WRONG function.
|
||
|
|
f2696653ef |
feat(phase-30 S39/S4): 4 wave-6 drafts bank UNCHANGED — the block was our tooling, not the code (+1,905 ins)
The 6 still-open wave-6 drafts were triaged against S38's own diagnosis table; 4 banked,
R22 clean-fleet 140/140.
func_801919A0 ov_SC06_032 710 ins (was: undefined ref func_8018B878 -- "alias class")
func_80189030 ov_SC03_001 557 ins (was: undefined ref func_80186F88 -- "alias class")
func_801878E8 ov_SC04_018 513 ins (was: undefined ref func_801848DC -- "alias class")
func_8018A564 ov_SC02_027 125 ins (was: CC1-FAIL Error 33)
THE FINDING: all four banked with NO change to the drafts. S38 recorded them blocked on a
class that needed cracking ("cracking this one class frees 6 drafts at once"); they had
ALREADY been freed by S38's own tool repairs -- the jr_isolate_all/overlay_src_split
alias-DEFINITION-deletion blindness and harvest_verify._reload_corpus. The drafts were
correct all along; the instruments were failing them. That is the FIFTH recorded "wall"
this phase to resolve to our own tooling.
=> RE-GATE STORED DRAFTS AFTER ANY TOOL REPAIR before treating a stored verdict as a
fact about the code. A verdict is only as current as the instrument that produced it
(R35 applied to the backlog, not just to metrics).
Each bank also performed a jtbl carve, so config/ changed => fleet blast radius => full R22
(clean + extract-all + check-all) = 140 passed, 0 failed of 140.
Metrics move exactly as the model predicts: instr 12406172 -> 12408077 = +1,905, the exact
sum of the four (710+557+513+125); distinct +1,905 / +4 unique fns; fn-count +4.
audit-digest OK. 0 NON_MATCHING (G4).
LEFT ON THE BACKLOG as genuine codegen residuals, not forced (P9):
func_8017C974 (ov_SC01_077, 947 ins, close=47, REGALLOC-PERM, 12 permuter variants inert)
func_80188C68 (ov_SC03_124, 551 ins, close=370, the only target with no twin anywhere)
NEXT: the func_801878E8 family (4 open siblings x 513 ~= +2,052). family_sweep --hseq
correctly REFUSED it via the §53 interlock (has_mid_jr => jtbl carve route; "a 0% from this
path would be a TOOL artifact, not a wall"), and jtbl_family_bank.py requires a clean tree
because it reverts from HEAD per sibling -- which is why this commit lands first.
|
||
|
|
65f0184f7d |
chore: regenerate docs/family-hseq.md after the wave-6 banks
Derived digest; committed so the next family map regen has a clean base. |
||
|
|
e5380ff6fb |
fix(tools): family_hseq's stdout claimed "fleet" for an OVERLAYS-ONLY number
load() globs .run/sig.ov_*.jsonl, so main and the resident are absent from every denominator — the numbers run ~0.3-1.0pp above the authoritative `make report` fleet (measured today: 96.6/94.6/ 89.7 vs the digest's 96.28/94.1/88.7). docs/family-hseq.md has always carried the "(overlays)" qualifier; the stdout print did not. That asymmetry matters because stdout is the channel a session actually reads and transcribes into a checkpoint — a right number under a wrong label is how a wrong number propagates (R35: the measurement was fine, the instrument's LABEL was the defect; R14: a checkpoint that disagrees with the digest must lose, and it can only lose if the disagreement is visible). Print-only; the metric itself is unchanged and correct for its scope. |
||
|
|
af16c38a62 |
feat(phase-30 S37): wave 5 banks 16/16 with ZERO reconcile + 26 members; wave metrics logged
Fleet 96.27 -> 96.28% fn-count / 94.1% instr / 88.6 -> 88.7% distinct (77,765
uniq). R22 clean-fleet: 140 passed, 0 failed of 140. dedup 1910/0.
16 targets / 16,884 templatable ins. 19 agents, 2.71M tokens, 82 min wall.
The gate banked 16/16 — the FIRST perfect gate of the session, and the first
needing NO reconcile at all. Sweep: +26 members / 3 failed across 17 overlays.
NEW: docs/wave-metrics.md — the wave-by-wave performance log, with the derivation
commands so future rows are COMPUTED, not hand-transcribed (R33). Four findings,
each recorded with its caveat rather than as a bare number:
1. THE PROMPT IS THE LEVER, AND THE AGENTS WRITE IT. Bank rate 76 -> 77 -> 100
-> 100 -> 100% with models and gate held constant. The jump was STEP 0 (a
magic-literal grep of src/, ahead of engine_core.h) — which came from a
wave-2 agent's index_gap report. Caveat recorded: waves 3-5 targets also
trended easier, so the mechanism is the durable claim, not the exact %.
2. pipeline() vs batched parallel(): 136 min/14 targets -> 82 min/16 targets,
parallelism 2.5x -> 3.8x. The two-batch design was a hard barrier with 37-50
min dead gaps; the harness already caps at 16 so it bought nothing. Floor
recorded honestly: the slowest agent is still ~50 min of real match_one
iteration, so the lever there is target SELECTION, not concurrency.
3. ECONOMICS: ~170-300k tokens per banked head in the stable regime — but a head
is not the unit of value. Head + propagated members is, and sweep yield is
BIMODAL not average (21/21 vs 18/165), because it is a property of the FAMILY.
Averaging those two predicts nothing.
4. A perfect gate is a signal the prompt rules landed. Waves 1-4 each needed 1-2
post-gate reconciles; wave 5 needed zero. The reconcile lane is the fallback,
not the plan. Lifetime 21/22.
|
||
|
|
9064840757 |
feat(phase-30 S36): wave 4 banks 14/14 + 26 members — step 0 is now the agents' default move
Fleet: R22 clean-fleet 140 passed, 0 failed of 140. dedup 1910/0. 14 targets / 16,844 templatable ins. 17 agents, 3.0M tokens. Claimed 14/14; the whole-binary gate banked 13, the 14th on reconcile. Sweep: +26 members / 0 failed across 18 overlays. Reconcile lane 21/22 lifetime. STEP 0 HAS BECOME THE AGENTS' DEFAULT MOVE. Nearly every wave-4 verdict cites the cross-overlay magic-literal grep BY NAME, several reporting `index_gap: none` because it resolved the target in one pass with no cookbook derivation needed: - func_8018BED0: grep 0xE100000A -> func_80188C04 (ov_SC03_089), verbatim, MATCH first try - func_8018BAB4: grep D_800A6610/D_800B9A02 -> func_801887E8, verbatim + callee swap - func_8017ED54: grep named all 5 family members -> reused func_8017D9F0's body, 14 data remaps - func_8017C290: grep found byte-identical twins ALREADY banked in two other overlays Bank rate by wave, same models + same gate, prompt the only variable: 76% -> 77% -> 100% -> 100%. THE ONE FAILURE IS THE §138 RECONCILE-DIRECTION RULE, in its purest form: `redefinition of struct B16_8018A758` — the agent copied its sibling's struct tag verbatim, and that sibling had banked into the SAME TU earlier in THIS wave. Decl ABOVE the splice => DELETE the duplicate (do not rename it). Worth noting the mechanism: a wave can create its own reconcile work when two targets share a TU. func_8018A808's own family swept 0/14 — its members are the per-location kind that do not template (the settled h_seq ceiling), not a plumbing failure. |
||
|
|
cbf0bce26c |
feat(phase-30 S35): wave 3 banks 13/13 — the flywheel paid off one wave later
Fleet 96.24 -> 96.25% fn-count / 94.0% instr / 88.4 -> 88.5% distinct.
R22 clean-fleet: 140 passed, 0 failed of 140. dedup 1910/0.
13 targets / 17,644 templatable ins. 14 agents, 2.7M tokens. Claimed 13/13;
the whole-binary gate banked 12, the 13th on reconcile (another §37/§124
SELF-axis alias — the TU declares `(void)`, the def takes an s32). Reconcile
lane 20/21 lifetime. Sweep: +21 members / 0 failed across 16 overlays.
THE FLYWHEEL, MEASURED ACROSS THREE WAVES (same models, same gate):
wave 1 baseline prompt 15/17 claimed -> 13 banked (76%)
wave 2 + the S33 rules 11/13 -> 10 (77%)
wave 3 + S34 magic-grep as STEP 0 13/13 -> 13 (100%)
Multiple wave-3 agents report the cross-overlay magic-literal grep landing the
answer on the FIRST search. One found a banked twin whose own header comment
already documented it as byte-identical to the new target, so the body
transferred verbatim with only file-local type suffixes renamed. That is the
wave-2 discovery paying off one wave later (R16).
Sweep quality also differed for a reason worth keeping: 21/21 here vs 18/165 in
wave 2. Wave 2's two big families are the per-location kind I then probed and
ruled out (BUILD OK + byte diff = genuine per-member codegen, not plumbing);
wave 3's are genuinely templatable. The sweep rate is a property of the FAMILY,
not of the wave.
TOOLING: an agent left 8 scratch files (test_licm*.c) in the drafts dir and the
gate driver died on `int('full', 16)`, taking the whole gate with it. Hardened to
treat a non-conforming filename as a NAMED, COUNTED skip rather than a crash
(R32) — a drafts dir is agent-writable by design, so it must not be trusted to
contain only deliverables.
|
||
|
|
7a4abddbfa |
feat(phase-30 S34): wave 2 — 10 heads + 18 members; the search order had a cross-overlay hole
Fleet 96.24% fn-count / 93.9 -> 94.0% instr / 88.4% distinct. R22 clean-fleet: 140 passed, 0 failed of 140. dedup 1910/0. WAVE 2: 13 targets / 37,943 templatable ins. 19 agents, 4.5M tokens. Claimed 11 MATCH; the whole-binary gate banked 9, +1 on reconcile (func_8017E5D0 via the §37/§124 DEFINITION-side alias — the TU declares it `(void)`, the byte-true def takes a pointer). Reconcile lane now 19/20 lifetime. 18 members swept. THE FINDING (an agent caught a hole in our own procedure). §136c's search order — engine_core.h near-twin -> same-TU banked sibling -> the .s — is entirely SAME-TU or SHARED-HEADER scoped, so no step can reach a banked twin in a DIFFERENT overlay's TU. But the large template classes live cross-overlay by construction. func_80188C04 (328 ins) turned out byte-identical to an already-banked func_801833F0 in ov_SC02_028, and ONE command found it: `grep -rn "E100000A" src/` — a magic word lifted from the target .s. The body was then reused verbatim, only file-local suffixes renamed. Promoted to STEP 0 of §136c, ahead of engine_core.h. That compounds with the manifest finding this session: the family map's `exemplar` is an IN-FAMILY pointer, so a family whose twin is banked elsewhere looks un-cracked — and the pointer can itself name an ALREADY-BANKED instance, hiding the family from any ranking built on it. Derive open sites from corpus.stubs over the member list instead. Measured on this wave: ranking off the map's exemplar gave 16,696 templatable ins; deriving from corpus.stubs gave 41,023, including a 55-ins family open in 138 overlays and a 46-ins one in 133. HONEST ON THE SWEEP: those two big families templated 18/165. That is the known h_seq refusal ceiling, not a new wall. One agent reported "all 10 members distance 0" — that is NORMALIZED distance, not h_exact, which is why dedup_propagate correctly answered reach<2. Do not read a normalized-distance claim as an h_exact guarantee. LEDGERED (real residual, not paperwork): func_8017F7B4 — needed its sibling's type names AND a data asm-label alias for a u8-shaped symbol, and still refuses. Plus func_8017C294 (DIFF close=12: 4 register/schedule permutations + a frame where I can get the 0x138 size OR pEnd's slot at 0x108, not both) and func_801898E4. |
||
|
|
138e21c7d4 |
feat(phase-30 S33e): both gate-refused drafts reconciled — the lane is 18/18 lifetime
Fleet 96.23 -> 96.24% fn-count / 93.9% instr / 88.3 -> 88.4% distinct. R22 clean-fleet: 140 passed, 0 failed of 140. dedup 1910/0. func_8017D318 (184 ins) + func_80181EE0 (198 ins) both banked, + 6 members swept (6 per-overlay variants failed — ledger material, not a lever). THE RECONCILE DIRECTION DEPENDS ON WHERE THE TU'S DECL IS, and picking wrong CREATES the next error (-> cookbook §138): - decl ABOVE the splice point -> DELETE the draft's duplicate (§100). func_8017D318: the TU defines MATRIX_/SVECTOR_8017C290, D_801EA8C0 AND a `struct PW8017C290` tag above it; I missed the tag on the first pass, so it took two rounds. - decl BELOW the splice point -> KEEP a decl in the TU's EXACT shape and cast at the use (§17a-1 D2). func_80181EE0: I removed its decl assuming the TU provided one; the TU's `extern int func_80143C74(short *, int);` is at L5082, ~180 lines BELOW the splice at 4901, so the identifier went undeclared. Grep the TU for the symbol and compare line numbers with the stub line first. MY OWN §136a VIOLATION, recorded: the blocker-capture filtered the build log for `error|conflicting|undefined reference` and reported "NO COMPILE ERROR" on a build that was failing with `redefinition of struct PW8017C290` and `'func_80143C74' undeclared` — neither phrase matched. A narrow keyword filter is exactly how a real error goes unseen, which is the thing §136a exists to say. Widened to keep any line naming a source position. |
||
|
|
0bdc7f44a0 |
feat(phase-30 S33d): Sonnet wave — 13 heads + 65 members banked (78 instances)
Fleet 96.21 -> 96.23% fn-count / 93.8 -> 93.9% instr / 88.0 -> 88.3% distinct
(+73 unique fns). R22 clean-fleet: 140 passed, 0 failed of 140. dedup 1910/0.
THE WAVE. Re-ran S10's 17 unbanked targets (26,227 templatable ins) at LOW
concurrency in two batches of ~9 — S10's finding was that 14 of 30 agents were
SERVER-throttled, i.e. the limiter is capacity, not capability. Targets
re-derived against corpus.stubs first (R35): all 17 still live, paths verified.
25 agents, 6.87M tokens, ~2.8h. Claimed 15 MATCH; the whole-binary gate — the
sole arbiter (G3/P9) — banked 13, then family_sweep propagated 65 members across
38 overlays.
HEADLINE: func_8017D174 (793 ins) — the largest single crack of this phase. Its
agent closed two compiler-internal residuals jointly: a §137 allocno-priority tie
between &g.sz0/&g.sz1 (R=7, L=607 vs 606 -> 230/231) that spilled the wrong one
and cost a load-delay nop in BOTH switch arms, and a sched2 rotation in the
outer-loop head block that survived 470+ statement orderings. Fix was four
zero-byte asms: two `"=r"/"0"` re-ties splitting wz's live range, plus two
volatile sliders placed in a DIFFERENT basic block so they lift the live-length
count without perturbing the head schedule.
THE 4 NON-BANKS SPLIT CLEANLY (§136b — none is a wall on one attempt):
- func_8017E2EC (close=20) and func_80186E24 (close=187): honest DIFF verdicts,
real codegen residuals, ledger material.
- func_8017D318 and func_80181EE0: claimed MATCH, gate refused -> the known
match_one->gate gap, which is DECLARATION plumbing (agents cannot run the
gate, so a TU-level conflict is invisible to them). Routed to the reconcile
lane, not retired.
AGENT-REPORTED INDEX GAPS worth acting on (the flywheel closing on itself):
- no symptom key for "schedule rotation at a loop-head block that NO statement
permutation reaches" — the index's nearest line points at §76 regalloc, and
the decisive doc was gcc-2.7.2-map/sched.md, which no scheduling symptom
cross-references.
- §137 is written as a two-compile arithmetic on ONE contender pair; the real
fix here was an N-zero-byte-insn budget that ties only for N in {1,3,4} and
splits the WRONG way for N=2, so a naive "add one slider, add another" walk
silently regresses.
- no key for "gcc hoists a loop-invariant SYMBOL_REF base out of a loop the
target keeps in the `sym(reg)` macro form" (~105 of func_80186E24's 187).
|
||
|
|
1859266d60 |
feat(phase-30 S10): Sonnet wave — 13 heads + 57 members; the §136i ~120 boundary is too LOW
- DREW'S CALL (2026-08-03): route the 30-target x2-9 wave to SONNET instead of Opus. The §136i >=120-ins Opus threshold was MY EXTRAPOLATION, never measured; this wave (125-793 ins) probes exactly the region where there was no data. - RESULT: 13 banked of 16 that ran = **81%**, vs Opus's 10/13 = 77% on the comparable S8-3 slice. At least 8 banked SONNET-DIRECT (65 agents spawned: 60 sonnet, 5 opus escalations). Propagated 57 member-matches / 4 failed across 42 overlays. 70 instances. R22 clean-fleet 140/140. FLEET 96.01% fn / 93.6% instr / 88.0% distinct (77,550 uniq). => **Sonnet is at least as capable as Opus on 125-793 ins. The ~120 boundary is too low.** NOT rewriting it to a specific number yet: 16 samples under a throttle confound cannot name a cliff. The controlled A/B (task #12) is how that number gets fixed properly. - THE REAL LIMITER IS CAPACITY, NOT CAPABILITY: 14 of 30 agents were killed by SERVER-side throttling ("Server is temporarily limiting requests (not your usage limit)") that 30 concurrent Opus agents did not trigger. Practical rule: run Sonnet waves at ~12-16 concurrency, not 30. The 14 unrun targets are listed in the checkpoint for a smaller-batch retry. - Sonnet's work quality was not shallow — three examples: func_8018797C read local-alloc.c and forced loads into an AGGREGATE to stop find_free_reg greedily taking 3 callee-saved regs; func_8018D870 used §136c sibling-first for ~70% of the body then blocked a coalesce with a pin; func_8017E3AC diagnosed an RC-3 callee-saved-order swap and noted the pin must be s32 or a stray `andi 0xffff` appears. - MY ERROR, RECORDED: `until [ -s <output> ]` fires at the FIRST LINE of output, not at completion. It fired mid-propagation and I ran `make clean` on top of a live family_sweep, deleting asm/ and aborting both the regen and the sweep (corpus's R32 assertion refused to answer rather than return a wrong stub set — working as designed). No bad bytes: R22 verified 140/140 immediately after, and the propagation simply re-ran clean. Correct waiter is `pgrep -x make` (exact process name), which also cannot self-match the way `pgrep -f <pattern>` did when it leaked 4 waiter shells earlier. Third instance today of ONE root cause: trusting a proxy instead of the thing itself (a weight column vs a probe §136h; an exit status vs build output §136a; file-existence vs process exit). |
||
|
|
cf10d28459 |
feat(phase-30 S9): x2-9 calibration wave — 22/25 banked + 67 members; the grind rate is MEASURED
- CALIBRATION (25 stratified targets: 12 head-by-weight + 13 sampled across the band, so the
measurement captures the DECAY, not just the head): 28 agents -> 22 claimed -> gate BANKED 22/25
(88%) -> propagated 67 member-matches / 13 failed across 38 overlays. 89 instances.
**22,937 templatable instructions banked in one wave.**
HEAD 9/12 -> 19,492 of 28,584 templ ins
BODY 13/13 -> 3,445 of 3,445 templ ins (the small ones are EASY; all 3 misses were 611-793 ins)
- FIRST SONNET DATA (§136i ladder's new middle rung): **Sonnet 6/6 · Haiku 6/6 · Opus 10/13.**
The two cheap tiers went 12/12 and Opus absorbed every hard failure — consistent with correct
size-routing rather than luck. Small n; the controlled A/B stays parked (task #12).
- THE PROJECTION for the >=95% instr bar (Drew's decision input): 411 of 1,872 families cover the
212,594-instruction gap = ~19 waves optimistic, 20-30 realistic. Mean templ ins/family decays
1,844 (top-25) -> 1,046 (top-100) -> 525 (top-400) -> 193 (band-wide), so early waves look like
this one and later ones bank MORE functions for FEWER instructions.
- DECISION (Drew): NO phase close — keep grinding. Campaign tracked as task #15.
- Carried failures -> next lanes: func_8017D174 (793 ins, closeness 5 after ~90 variants; diagnosed
a backward-scheduler priority race -> permuter, correctly NOT ledgered a wall), func_80186E24
(611 ins, 133 of 139 diffs pure register numbers -> a natural §137 test), func_8017E2EC.
- §136c sibling-first paid again: func_8017DF84 (766 ins) MATCHED because a banked byte-matched twin
existed in the same TU; its 697 index-diffs traced to ONE root cause (a bare 0xFFFFFF literal that
loop.c hoisted to the OUTER preheader, stealing $s3) — closed by binding it to a local declared as
the FIRST statement of the inner loop body. Verified via rtu_match (real-TU), not just match_one.
|
||
|
|
a4bc49a23d |
feat(phase-30 S8): the x10-99 band closes 23/23; §137 makes REGALLOC-PERM arithmetic, not a permuter job
- S8-3 (23 fresh x10-99 families, 121-328 ins — the hardest band this session): draft 16/23 ->
capture (1 PLUMBING / 6 DIFF) -> reconcile 1/1 -> redraft 6/6 => **23/23 (100%)**.
Propagated 206 + 81 = 287 member-matches across 80+54 overlays. R22 clean-fleet 140/140.
FLEET 95.97% fn / 93.4% instr / 87.5% distinct (77,404 uniq); dedup 1905/0; 0 NON_MATCHING.
- §136b CLOSES AT 15/15 — no function ledgered "genuine byte-DIFF" survived a redraft, all session.
- §137 (NEW, the session's most reusable result): REGALLOC-PERM — a clean 2-register swap — is a
TWO-COMPILE ARITHMETIC PROBLEM. global.c:allocno_compare ranks by floor_log2(R)*R/L*1e4*size;
read R and L out of `cc1 -dl -dg` for BOTH contenders AND their ranked neighbours to get the
admissible priority WINDOW, then place a zero-byte `__asm__ __volatile__("" ::"r"(v))` so L lands
inside it. func_801833F0: contenders ONE unit apart (1297 vs 1296), window (1228,1296), five
placements probed, only L=219 -> pri 1232 worked. R and L are FORCED BY THE EMITTED CODE (L is
recomputed post-sched1), which is exactly why source-reordering is a dead end for this class.
Converts a class the permuter banked 0 from all session into a deterministic calculation.
Companion: floor_log2 makes ref-count a STEP function (5/6/7 refs are worthless, you must reach 8)
— func_8017EFA8 closed 30 register-name mismatches by taking a pseudo 4 refs -> 8 with a dead read.
- §136j — the failure MIX FLIPS WITH SIZE: <=120 ins fails ~70% on declarations; 121-328 ins fails
86% on genuine codegen. Budget reconcile for the small band, redraft for the big one — and do NOT
read 70% on a big-function wave as a broken pipeline; that is the expected shape.
- §137a — a gate verdict has a TIMESTAMP. Two "DIFF" ledger entries were STALE (draft rewritten 28
min after the gate ran, never re-gated); both were already byte-perfect. Compare verdict time to
draft mtime before redrafting. Plus two offline oracles an agent built: a FULL RELOCATION RESOLVE
(catches wrong jal/%hi/%lo targets that match_one's mask hides) and a COLLATERAL CHECK (whole-TU
objdump with/without splice). Together they discriminate all three causes of "match_one says MATCH
but the overlay SHA differs" without running make.
- §136f addendum — the collider is often an already-banked SIBLING BELOW the splice; locate it by
arithmetic (draft grows the file N lines, so TU line L reports at L+N).
- cookbook-index 380 -> 382 sections.
|
||
|
|
18fc50d0f8 |
docs(phase-30): §136i — insert SONNET between Haiku and Opus in the drafter ladder (Drew 2026-08-03)
- MEASURED BASIS (P30 S7, 144-target campaign): the two-tier rule from the 2026-06-29 A/B left the ~50-120-ins band unassigned, and every wave since defaulted it to Haiku-with-Opus-escalation. Haiku-direct banked 3/8 on that band while Opus-escalation-after-a-Haiku-miss banked 10/11 — i.e. Haiku was acting as EXPENSIVE TRIAGE (a wasted draft + a full Opus redraft), not a cheap drafter. The original A/B only proved parity <=52 ins; everything above that was extrapolation. - LADDER: haiku <=~50 ins · SONNET ~50-120 · opus >=~120 or escalation · fable5 for a genuinely NEW wall class only. Never haiku->opus directly; never default a whole wave to opus because the band "looks hard" (the same extrapolation in the other direction). - WIRED, not just documented: s7_manifest.py routes by the new thresholds; s7_wave4b.js escalates haiku->sonnet->opus instead of haiku->opus, and its meta/prose say so. - Boundaries (~50/~120) are current best estimates — re-measure per-tier from the journal + the gate, never from the workflow's by_tier (it counts claims, not banks — §136). - Byte-gate remains the sole arbiter, so a weaker drafter is a throughput risk, never a correctness risk (G3/P9). cookbook-index 378 -> 379. |
||
|
|
1b8f26113c |
feat(phase-30 S7): close the B-shape queue — 144/144 drafts banked; §136b closes 9/9
- FINAL LANES: reconcile ×5 (5/5) + redraft ×1 (1/1) -> gate BANKED 6/6 -> propagated 50 members across 28 overlays. **ALL 144 DRAFTED TARGETS BANKED (100%); zero stubs remain in the queue.** R22 clean-fleet 140/140 (seventh time this session). FLEET 95.88% fn / 92.9% instr / 86.5% distinct (77,106 unique fns); dedup 1905/0; 0 NON_MATCHING. - LANE RECORDS: reconcile 15/15 lifetime · redraft 9/9 · §136b closes at 9 FOR 9 (every function ever ledgered "genuine byte-DIFF" banked on redraft). - THE CAPTURE CLASSIFIER, third and final defect (§136a): it decided PLUMBING by matching a regex against cc1's PROSE, and cc1's vocabulary is open-ended — `too many arguments to function` matched nothing, so a trivially reconcilable function sat UNKNOWN through two gate rounds. Now DERIVES the class from the closed invariant (did the compile produce an object: `make ... Error N` + `Deleting file`). Re-running it moved 5 PLUMBING / 1 UNKNOWN -> 5 PLUMBING / 1 DIFF, and BOTH reclassified functions then banked. Three defects in one small tool in one session — an unreachable exit-status branch, a missed phrasing, and the prose-matching design behind both — each SILENTLY MIS-ROUTING REAL WORK. R33 in one line: if an invariant answers it, never re-parse. - §136f — two declaration sub-cases: (1) a symbol you call may be DEFINED, not just declared, BELOW your splice point (func_8017D540 is defined 275 lines below as int(int); the draft guessed void(s32) from a bare jal); (2) an ARITY clash on the symbol you are DEFINING cannot be fixed by a cast — use the §37/§124 asm-label alias (func_801848DC; in-TU precedent at :8872). - §136g — TWO INDEX ROUTINGS BYTE-REFUTED (func_801863B4). The index sends BRANCH-POLARITY to §3-T4 (invert) and §34 (zero-byte fence); the agent tested BOTH at zero, read the gcc-2.7.2 source, and found jump.c:1806 `if (foo) bar; else break` range-swap — which runs long BEFORE reorg, so a fence CANNOT block it. Real lever: put a label between the if-join and the return label (wrap the loop in the guard). Also: same-address lh+lhu is MIPS LOAD_EXTEND_OP==ZERO_EXTEND (mips.h:1163), and combine collapses the pair unless the HImode pseudo has two reaching defs. REFUTED ROUTINGS ARE RECORDED NEXT TO THE CORRECT ONE — otherwise the next agent re-runs them. - cookbook-index 375 -> 377 sections (§136 .. §136g earned this session). |
||
|
|
b61d805b8b |
feat(phase-30 S7): wave 4b batch 3 — 35 heads + 305 members; the 144-family B-shape queue is worked
- BATCH 3 (37 targets, 41 agents, 2.75M tok -> 35 claimed): gate BANKED 35; family_sweep propagated
305 member-matches / 39 failed across 77 overlays (14 STRUCT skipped by design). 340 instances.
R22 clean-fleet 140/140 (sixth time this session).
FLEET 95.86% fn / 92.9% instr / 86.5% distinct (77,061 unique fns); dedup 1905/0; 0 NON_MATCHING.
- WAVE 4b COMPLETE: b1 32/37 + b2 35/37 + b3 35/37; with wave 4a (30/33) the whole 144-family
B-shape queue that opened this session is worked through — 138 of 144 drafts banked (96%).
- §136e — batch 3's two HONEST NEGATIVES, worth as much as the wins:
(1) §136c SIBLING-FIRST HAS A PRECONDITION. func_801899AC's family has all 13 members still
unmatched and no engine_core.h twin, so there IS no byte-verified sibling and the search is
pure cost. Check a banked sibling EXISTS before spending the greps.
(2) A loop increment in the loop-back DELAY SLOT + a compensating negative addiu is a SOURCE
SHAPE, not a reorg artefact — MIPS1 has no annulling, so reorg CANNOT invent the
compensation. Write `p += 2; if (t == cur) break; ... p -= 2;`. combine's reg_n_sets==1 guard
stops the addiu folding into the following lw. The index's delay-slot entries point at reorg,
which is a dead end for this class.
Plus a new §136-L1 application on the RETURN axis (an over-scoped temp became a global allocno and
swapped $v0/$v1 with the returned local, collapsing the target's `j` + `addu` return).
- COMPOSITION, demonstrated on func_8017D5F4 (46 ins): flat early-returns -> dead-local frame pad ->
s16 locals -> operand order -> 3 register pins -> 2 zero-byte re-ties -> permuter for the last 2.
THE PERMUTER IS THE LAST STEP ON AN ALREADY-PINNED BASE, not the first.
- cookbook-index 374 -> 375 sections. 6 stubs remain; per §136b none is a wall on one attempt.
|
||
|
|
7f70b6850a |
feat(phase-30 S7): batch 2 + reconcile + redraft — 41 heads + 431 members; §136b closes 8/8
- THREE LANES: wave 4b batch 2 (37 targets, 46 agents, 3.44M tok -> 35 claimed) + the reconcile lane
on 3 PLUMBING failures (3/3) + a REDRAFT lane on 4 DIFF-ledgered failures (4/4). Combined gate
BANKED 41; family_sweep propagated 431 member-matches / 29 failed across 79 overlays.
103 of 107 drafts banked (96%). R22 clean-fleet 140/140 (fifth time this session).
FLEET 95.77% fn / 92.9% instr / 86.4% distinct (76,824 unique fns); dedup 1905/0; 0 NON_MATCHING.
- §136b CLOSES AT 8/8: every function ledgered "genuine byte-DIFF" banked on redraft — wave 3's
four, the THREE I classified from wave 4a's capture, and one from batch 1. The classifier is
right about what it measures ("this draft compiles clean and differs in bytes"); reading that as
"this function resists matching" is the error. A DIFF verdict is a fact about ONE DRAFT.
- §136a CORRECTED (a reconcile agent refuted me against the bytes): I wrote "70% of gate refusals
are paperwork, not codegen". WRONG. A declaration conflict ABORTS THE COMPILE, so a PLUMBING
verdict says NOTHING about the body. Two of three second-round PLUMBING drafts had a real codegen
residual behind the conflict (func_80188694 DIFF/4 SCHEDULE-REORDER, closed with a §21 zero-byte
re-tie after six other variants failed; func_8018C638 DIFF/6 ADDRESSING/cse). Both agents ran
match_one on the untouched draft FIRST and rejected my premise — which is what §135 asks for.
- §136c SIBLING-FIRST IS A DERIVATION SHORTCUT, not just a conflict fix: grep engine_core.h's
DEFINE_func_* bodies for a byte-verified NEAR-TWIN before deriving from the .s. func_801859D8
found DEFINE_func_80185978 (identical offset chain, 3 differing constants), reused its expression
forms verbatim -> FIRST-DRAFT MATCH, and the twin generalizes to its whole 10-member family.
Search order: near-twin -> banked same-TU sibling -> the .s -> the Ghidra seed LAST (byte-proven
an entirely different body twice this session).
- §136d, four new gcc-2.7.2 levers from the redraft lane, each with its REFUTED axis recorded:
RC-12 $0-add opaque copy (cse.c canonical-copy promotion; do NOT pin the pair to real regs);
jump.c if-then-else -> conditional-overwrite collapse (defeat with TWO SEPARATE CALLS, not a
ternary); fix the STORE not the load for a load hoisted above a constant-address store (the
INDIRECT_REF reshape is the wrong half of the /s lattice, 2 -> 32 mismatched); a branchless flag
is -(a != b) & 0xFF, never a ternary.
- cookbook-index 372 -> 374 sections. Batch 3 staged with all of the above promoted into its prompt.
|
||
|
|
cc3c49d7af |
feat(phase-30 S7): wave 4b batch 1 — 32 heads + 365 members; §136b (a DIFF verdict is not evidence)
- WAVE 4b BATCH 1 (37 volume-lane targets, 10-19 members, <=60 ins; wave 4a's §136 idioms promoted
into the drafting prompt per the measured 83%->93% law): 50 agents / 4.35M tokens / 32 min ->
34 claimed MATCH -> gate BANKED 32 -> family_sweep propagated 365 member-matches / 4 failed
across 84 overlays (3 STRUCT skipped by design). 397 function-instances from 37 targets.
- §136b — THE FINDING THAT CHANGES THE BACKLOG: all FOUR functions wave 3 ledgered as "genuine
byte-DIFF" BANKED on redraft. The recorded causes were never codegen:
func_801845B0 a branch to the EPILOGUE misread as an inner early-exit -> the whole tail was
hoisted out of its enclosing if (control-flow misread)
func_80184A94 a declaration conflict on a symbol declared BELOW the splice point; closed by
copying an already-banked family sibling's decl forms verbatim (§71)
func_8017BEBC the cached Ghidra seed was an ENTIRELY DIFFERENT body and the prior draft
followed it; the .s was the only usable source
func_8018480C re-derived clean
=> a DIFF verdict describes THE DRAFT THAT WAS ATTEMPTED, never the function's matchability.
Never retire a target on one; route it to REDRAFT. And re-GATING an unchanged draft is not a
retry — which is exactly why wave 4a's 3 DIFFs stayed stubs through this gate (same bytes
resubmitted); they still owe an actual redraft and are now likely winnable.
Corollary: backlog entries carrying an old closeness/class are stale by construction (P29
measured 77% of stored drafts decayed) — re-verify before valuing one.
- The wave-4b prompt handed each retry its prior verdict EXPLICITLY LABELLED "a data point, not a
verdict — re-derive from the .s". Every retry agent did exactly that and refuted it.
- R22 clean-fleet 140/140 (fourth time this session). FLEET 95.63% fn / 92.8% instr / 86.3% distinct
(76,499 unique fns); dedup 1905/0; C1 240496/240496; 0 NON_MATCHING (G4).
- Orchestration: this batch's workflow script was GENERATED from the manifest files rather than
hand-pasted — transcription had already cost this session one dead launch (args-as-string) and
cost the prior session three agents' time (hand-typed _jr_* paths). Generate the artifact; do not
ask yourself to be careful. cookbook-index 371 -> 372 sections.
|
||
|
|
d00dfe363b |
feat(phase-30 S7): reconcile lane 7/7 — wave 4a closes at 30/33 (91%), +76 members
- RECONCILE LANE: all 7 PLUMBING failures FIXED and banked (329K tokens — ~13x cheaper than the drafting wave's 4.44M). Propagated +76 member-matches / 0 failed across 51 overlays. Wave 4a final: 30/33 heads (91%) + 327 members = 357 function-instances from 33 drafted targets. - THE CAPTURE CLASSIFICATION WAS EXACTLY PREDICTIVE: all 7 PLUMBING banked, all 3 DIFF stayed stubs (func_8017E978 / func_80184494 / func_80184960 -> redraft lane, their C is wrong). That is what makes the ~10-build capture step worth running before any reconcile fan-out. The lane is now 19/19 across three waves. - EVERY reconciled draft had a HIDDEN SECOND CONFLICT cc1 never reached (it reports only the first) -> "grep the whole TU in one pass" must be in the RECONCILE prompt, not just the drafting prompt. One agent additionally assembled the spliced TU and masked-compared its function IN TU CONTEXT (67/67) — proving the casts byte-neutral in situ, not merely standalone. - NEW HAZARD, agent-surfaced (§136a corollary): an agent chose a SHARED scratch path, a concurrent agent overwrote it, and its verification silently compiled ANOTHER agent's TU and returned a meaningless rc=0. It caught the swap only because the emitted .s lacked its own function. A shared scratch path yields a CONFIDENT WRONG VERDICT, and no tool fix reaches it — the choice happens inside the agent, so the PROMPT must mandate a process-unique path. This is the Phase-28 match_one fake-isolation defect recurring one level up. - R22 clean-fleet 140/140 (third time this session). FLEET 95.52% fn / 92.7% instr / 86.1% distinct (76,273 unique fns); dedup 1905/0; C1 240496/240496; 0 NON_MATCHING in any default build (G4). - .run/s7_extra.txt: wave 4a's idioms compiled into the wave-4b drafting prompt (the promotion that measured 83%->93% between waves 1 and 2). |
||
|
|
fdb5813fc1 |
docs(phase-30 S7): S6c banked ×12 (all in the P27 SC07 quartet); blockers classified 7/3; §136a
- S6c (deterministic, ~0 agent tokens): 12 sibling banks across 3 of 9 jr zero-crack families (func_80178D40 890ins 4/4, func_801734BC 4/4, func_8012ACE0 4/4). The other 6 are ledgered: 5 gate-fail (genuine byte DIFF) + 1 carve-fail (span table starts do not fit the span). R22 clean-fleet 140/140 over the whole S6c series. - FINDING: all 12 banks landed in ov_SC07_006/007/010/011 — the four overlays P27 discovered and P28 made citizens (R36). P28 drained their h_exact backlog via dedup_extend; the jr/h_seq propagation lane was still owed. R14 GUARD AGAINST OVER-READING IT: the quartet are the top four overlays by remaining zero-crack residue (2,190-2,355 ins each vs 500-870 typical) but hold only 7% of the 2,114 remaining slots — a per-overlay priority signal, NOT a bulk lever. - BLOCKER CAPTURE for the 10 wave-4a gate failures -> .run/s7_blockers.json: 7 PLUMBING (all `conflicting types for func_X`) / 3 genuine byte-DIFF. 70% of "the gate refused" is declaration paperwork. New tool .run/s7_capture.py (any overlay/any draft dir; reverts the TU in a finally:). - MY DEFECT, FIXED AND DISTILLED (§136a): the capture tool first classified on the EXIT STATUS, so its `rc == 0 => byte DIFF` branch was UNREACHABLE — `make build` runs `check`, so a draft that compiles perfectly and merely differs in bytes also exits non-zero, and all 3 real DIFFs were filed as "unknown". Now classifies on the OUTPUT ([FAIL]/got/want vs a non-warning error line); the warning-exclusion matters because `conflicting types` also appears benignly for builtins. - Also probe-discipline: my first S6c probe reported 1/9, which was 1 bank + 8 CORRECT REFUSALS — jtbl_family_bank refuses on a dirty config/+src/ (its per-sibling revert restores from HEAD). Driver now commits between families. A uniform failure across N functions is a statement about the mechanism, not the functions (§134). - cookbook-index 364 -> 371 sections, --check green. CURRENT_PHASE SESSION-31 checkpoint refreshed with the queue re-derived at HEAD (the S30 ROI-floor trigger stays REFUTED — do not close on it). |
||
|
|
09d96b1531 |
feat(phase-30 S7): wave 4a — 23 heads + 251 members banked ×N; cookbook §136 (the local-variable lever)
- WAVE 4a (T6, the 33 high-value B-shape families, 61-120 ins / >=10 members):
33 targets, 46 agents, 4.44M tokens, 29 min -> 29 claimed match_one MATCH.
Whole-binary gate BANKED 23/33 (70%); family_sweep --hseq --band all propagated
251 member-matches across 69 overlays (13 failed, 4 STRUCT skipped by design).
Total 274 function-instances from 33 drafted targets.
- R22 clean-fleet: make clean + extract-all + check-all -> 140 passed, 0 failed of 140.
make report: fn-count 95.49% / instr 92.6% / distinct 85.9% (76,180 unique fns);
dedup 1905 validated / 0 failed; 0 NON_MATCHING in any default build (G4).
- COOKBOOK §136 (R30, distilled in-session from 25 banked functions' index-gap reports):
the wave's finding is that in the 60-120-ins band most "regalloc residuals" are decided
by HOW MANY C LOCALS AND AT WHAT SCOPE, not by register pins (local-alloc.c:472 promotes
any pseudo with REG_N_DEATHS>1 to a global allocno). 19 byte-verified idioms: 6 splitting/
merging rules, 6 type-form rules, 5 scheduling rules refining §135-2/§135-4, 2 declaration-
surface rules. One case explicitly REFUTES the pin as the lever for a redundant copy.
cookbook-index regenerated 364 -> 370 sections, --check green.
- TWO SELF-CORRECTIONS (R37/R14), both caught before they could mislead sizing:
(1) I wrote the tier split from the workflow's by_tier, which counts CLAIMED matches (29)
not banks (23). Derived per-function: Opus-direct 10/14, Haiku-direct 3/8, Opus
escalation-after-Haiku-miss 10/11. The operative number is the 10-of-11 rescue rate;
on this band Haiku is triage, not a substitute (it is == Opus only at <=50 ins).
(2) The gate printed "1/1 banked FAILED: func_X" on single-draft groups (the known
double-list artifact) -> bank set DERIVED from corpus.stubs instead. Totals agreed.
- TOOLING: the wave scripts now parse args-as-string and assert Array.isArray, so the
roadmap's standing "args must be an array" gotcha cannot silently kill a future wave
(it killed wave 4a's first launch in 60ms with 0 agents).
- tools-health green + fail-closed before matching (corpus+resident 0 PHANTOM/0 TRUNCATED,
audit-binaries 140/140 citizens, cdecl, report/lint/dedup).
|
||
|
|
0ffb2c04d7 | docs: regenerate family-hseq digest at HEAD (the checkpoint's queue is derived from it) | ||
|
|
6fe9b66f2d |
feat(phase-30 S6h): wave 3 — 34/38 banked, +639 members, reconcile lane now 12/12 (R22 140/140)
- 38 targets / 44,297 templ ins, model-routed (Haiku <=89 + Opus escalation, Opus direct >=90):
52 agents, ~4.1M tokens -> gate 27/38 (71%). All 11 failures captured + classified: 8 declaration/
link plumbing, 3 genuine byte-DIFF. An 8-agent Opus reconcile wave fixed 8/8 (7 banked) ->
wave-3 total 34/38 = 89%. Propagation +639 members / 1 failed / 83 overlays.
R22 clean-fleet 140/140. Fleet 95.42% fn / 92.4% instr / 85.5% distinct.
- DESIGN (S27 law applied BEFORE it bit): six of eight reconcile targets share ONE TU, so this wave
FORBADE agents any build — six concurrent splice-builds would have clobbered a tracked file.
- THE AGENTS OUT-DIAGNOSED MY BLOCKERS:
* func_801848B0 — an agent REJECTED MY PREMISE: I said byte-correct + decl-blocked; it ran
match_one first, found a real 1-ins DIFF, fixed both. R14 aimed back at me, correctly.
* func_8017C5F0 — the "invented symbol" D_801DA0F0 is an INTERIOR ADDRESS: offset 0x6C into
D_801DA084 (0x801DA084..0x801DA103). The lui/addiu pair builds an interior pointer.
* func_8018A860 — the TU declares memcpy THREE times with incompatible signatures, with a latent
byte bug behind it. One symbol declared three ways is a defect awaiting the next draft.
- Carried (4): func_80184A94 (match_one MATCH, gate-refused) + 3 genuine byte-DIFFs
(func_801845B0, func_8017BEBC@ov_SC02_026, func_8018480C).
|
||
|
|
372dc62d35 |
feat(phase-30 S6g): wave 2 — 93% bank rate (was 83%), all 4 reconciles closed, +342 members (R22 140/140)
- 15 targets (11 fresh Haiku + 4 gate-failed reconciles on Opus), 15 agents, ~0.74M tokens.
Gate banked 14/15 (93%) vs wave 1's 20/24 (83%); ALL 4 RECONCILES BANKED.
Propagation +328 members / 0 failed / 76 overlays. R22 clean-fleet 140/140.
Fleet 95.23% fn-count / 92.1% instr / 85.0% distinct (phase opened 92.00 / 87.5 / 78.0).
- THE 83->93% CAME FROM THREE FIXES, ONE PER WAVE-1 FAILURE (the S27 finding reproducing):
(1) args pasted from the DERIVED manifest, never typed — all 30 paths verified on disk first;
(2) blocker-capture BEFORE the reconcile fan-out (S29 law: agents cannot run the gate, so a
match_one-MATCH draft dying on `conflicting types` reads to them as a codegen wall) —
each got the exact symbol+line plus the two byte-neutral levers;
(3) wave-1's Opus DISCOVERIES became wave-2's Haiku INSTRUCTIONS (ori-vs-addiu unsigned
destination; store-sinking scheduler order).
- THE RECONCILES OUT-DIAGNOSED MY CAPTURE: func_80189B78's error named ONE symbol; the agent found
SIX invented prototypes, two AFTER the splice point where cc1 had not yet reached — all fixed by
copying the TU's decls verbatim + casting at the call site, zero bytes changed. func_8018584C had
lever (A) blocked in BOTH directions (the draft must also compile standalone for match_one) and
closed with the DATA form of the asm-label alias. func_80180A4C was one character class (s32[] vs
the TU's u8[], declared 11 lines after the splice point).
- Carried: func_80189C4C (the one agent that returned no structured result; gate refused).
|
||
|
|
6e181db771 |
feat(phase-30 S6f): B-shaped wave — Haiku drafts, Opus closes, +544 members (R22 140/140)
- POOL (derived from the regenerated map): 36 families / 28,829 templatable ins, kind=modal (no member matched ANYWHERE so no sweep could reach them), >=20 members, <=60 ins, non-jr, and NOT ONE exemplar in ov_SC01_077. Hand-calibrated 3/3 one-shot before scaling (Phase-15/18 discipline). - WAVE (ultracode; Haiku drafters + Opus escalation, 24 targets): 31 agents, 0 errors, ~2.0M tokens, 12.5 min. Agents claimed 24/24 MATCH; the whole-binary gate banked 20/24 (83%); propagation +524 members / 0 failed / 91 overlays. 17 of 20 banks were HAIKU, 3 Opus — the cheap-tier-ab-validated call (Haiku == Opus at <=~50 ins, ~4.8x cheaper) held on real work. - WHAT OPUS BOUGHT: (1) a `sh` of a constant with the stored width's top bit set needs a u16 destination — via s16 gcc folds it sign-extended and li emits addiu, via u16 force_fit_type keeps it positive and li emits ori; (2) a schedule-reorder closed by STATEMENT ORDER not the permuter (gcc's list scheduler preserves relative order of disambiguable stores); (3) three loose-typing fn-ptr casts a cheap drafter had misread as delay-slot/permuter residuals. - MY ERROR (R37/R14): I generated the manifest to .run/s6f_wave_targets.json then HAND-TRANSCRIBED the args into the Workflow call, pattern-filling _jr_8017BEBC across overlays where no such split exists (corpus.stubs says _jr_8017AE2C). Three agents lost time rediscovering real paths. The gate driver written after (.run/s6f_gate.py) DERIVES every TU/split from corpus.stubs and asserts nothing. Assert nothing you can derive. - The 24->20 gap is the known match_one->gate gap (standalone compile cannot see a TU decl conflict; Phase 19 measured 88-92% -> 60-71%). 4 carried: func_8018584C, func_80180A4C, func_8017CC80, func_80189B78. - R22 clean-fleet 140/140. Fleet 95.13% fn-count / 92.0% instr / 84.9% distinct (phase opened 92.00 / 87.5 / 78.0). |