Commit Graph

438 Commits

Author SHA1 Message Date
Drew T 9c8bb51831 feat(phase-24): T4-flagship gate — +1 fns x0 propagated (fleet 64.86%) 2026-07-02 23:38:32 -06:00
Drew T 0abcd250ab feat(phase-24): T3 — permuter setup fixes (pins/GTE-asm, typedefs, -O0); flagship closed
- p16_permute.hide_asm: b64literal-pragma carrier for register pins + GTE __asm__ blocks.
  pycparser parses the pragma; decomp-permuter's process_pragmas decodes it back so cc1 sees
  the real pins/asm (regalloc steered, mvmva compiles). No submodule edit (reuses its own carrier).
- drop_preproc_and_scalar_typedefs: keep #define + custom struct/typedefs, drop only #include
  (fixes func_801412A8 'OTLINK undeclared'); +f32 typedef gap; make_base_c hides asm.
- compile.sh/compile_o0.sh: prepend .include "macro.inc" so GTE mvmva assembles (the real build
  gets it via include_asm.h, which base.c omits + -DPERMUTER disables). Byte-neutral for non-GTE.
  compile_o0.sh (-O0) auto-selected for _o0 targets.
- run_masked.py: expr_type->int fallback for hidden-pin vars -> 0 internal-permuter-failures
  (perm_split_assignment/perm_temp_for_expr no longer KeyError on pinned drafts).
- ALL 5 seeds parse+compile (base 4/110/36/52/77). FLAGSHIP func_80132784 (4/400) CLOSED to a
  masked-0 that match_one confirms MATCH (400 ins) -> .run/wave/func_80132784.win.c for T4 gate.
- no build-input changed (136/136 untouched)
2026-07-02 23:32:38 -06:00
Drew T 5984749421 feat(phase-24): T2 — floor-free relocation-masked permuter scorer (-drz)
- tools/masked_diff.py: shared objdump -drz masking oracle (jal/j 26-bit + HI16/LO16
  immediate mask; -z keeps nop runs -> no GTE under-count). object-vs-object (permuter,
  +reloc-operand equality) and object-vs-.s (match_one) modes.
- match_one.py refactored onto masked_diff (-dr -> -drz); regression-clean on count-exact
  seeds (35/52/72/110, leaf-MATCHes), output format unchanged (gate_stage/grinder compatible)
- tools/masked_scorer.py MaskedScorer: drop-in for decomp-permuter's Scorer; scores masked
  .text closeness (bottoms out at 0) not the stock mnemonic-diff floor
- tools/permuter/run_masked.py: rebinds src.main.Scorer -> MaskedScorer (NO submodule edit,
  R3/R20); p16_permute.run_permuter wired to it
- VALIDATED: masked_self=0; masked_cand~match_one; STOCK floor 1750-2100 vs masked 36-77
  (the wander cause); live permuter base score = masked 77 (not stock 1930), descends to 74
- R14 SELF-CORRECTION: -drz confirms func_80132784 is 4/400 count-exact (T1's -dr 204 was a
  16-nop-collapse artifact); backlog re-logged close 4. The flagship IS 4 ins away.
- no build-input changed (136/136 untouched); 2 compile FAILs surfaced -> T3 (GTE asm, typedefs)
2026-07-02 23:10:54 -06:00
Drew T 394e3a81b3 chore(phase-24): T1 — re-log map-wave seeds (byte-verified) + backlog/gate_stage hygiene
- re-logged 7 map-wave seeds in .run/backlog.jsonl with match_one-verified
  closeness (dropped 19 stale/unreproducible records, .bak kept):
  func_8014E048=35 (was mis-logged 'failed'), func_80176D94=52,
  func_80148094=72 (best draft vG2, not the named file), func_801412A8=110;
  T6 leaf-MATCHes func_8014F4C0/func_80155800=0; best_draft -> .run/backlog_drafts/
- R14 FINDING: the 'func_80132784 is 4/400' premise is NOT reproducible from any
  on-disk draft (best=204-off, wrong instr count 384!=400 -- match_one -dr collapses
  ~16 nops on this GTE seed); true closeness pends T2's -drz scorer -> T4 reframed
- gate_stage.py: fixed the stale 'sig_unify SKIPPED with --src-file' doc-drift
  (code runs it; sig_unify.py:151 has --src-file support)
- wave_targets --class REGALLOC + grinder now surface the 4 count-exact permuter
  seeds with true closeness; no build-input changed (136/136 invariant untouched)
- added T5b (Fable5 S11/RC-6 map-extension spike) per Drew
2026-07-02 22:55:05 -06:00
Drew T 80b89d8a2d feat(phase-23): LLM matching tier + the Fable5 wall-breaker + the gcc-2.7.2 codegen map (v1.22.0)
- LLM TIER (T1-T10.9): tools/serve_local + api_draft + lora_grind + bulk_harvest + the LoRA
  pipeline (v3 model). Local v3 saturated <=15 (~1300 banks, $0). Frontier GLM-5.2 (~$4)
  confirmed the def-side wall INTRINSIC + the idiom well DRY at the leaf (both angles).
- THE PIVOT (07-02): Fable5Max reading the gcc-2.7.2 SOURCE matched func_8014EE14, a §20/§10
  store-vs-load giant "CONFIRMED unsteerable" for 22 phases -> the "wall is intrinsic" verdict
  is MODEL-RELATIVE. Generalized to Opus agents applying §30 (2 more giants, 3-5x cheaper).
- §31 THE CODEGEN MAP: 4 Fable5 agents read the whole gcc-2.7.2 source -> docs/gcc-2.7.2-map/
  {sched,regalloc,loop,cse_expr}.md (935 lines, byte-proven residual->lever catalogs) +
  cookbook §31 (index+triage). Broke hoist-vs-remat/delay-slot/coalescing walls. Correction:
  gcc-papermario is 2.8.1 not 2.7.2 -> tools/reference/gcc-2.7.2/ (SETUP §5.6).
- THE WAVE (8 reach-134 near-misses, Opus + §31): 3 leaf-MATCH by lookup (47-137k tok vs
  200-250k walls) + 5 tight permuter seeds (func_80132784 240->4-off!), zero dead-ends.
- 5 giants/hard-fns banked x134 (func_8014EE14/F2E0/150528/149374/144090); fleet 63.66->64.86%,
  136/136 byte-identical, 0 NON_MATCHING. cookbook §30/§30a/§31. No new governance rules.
- FINDING: matching is SOLVED by the map; bottlenecks are now the permuter (S11 seeds) +
  whole-binary integration (auto-bank leaf-MATCHes) -> Phase 24. worklog -> logs/Phase23.md.
2026-07-02 21:51:25 -06:00
Drew T da4eb3ff6c feat(phase-23): map-wave gate — +1 fns x0 propagated (fleet 64.86%) 2026-07-02 21:19:07 -06:00
Drew T 3b6f2c5f94 docs(phase-23): §31 — the gcc-2.7.2 codegen map (4 Fable5 agents read the source)
- docs/gcc-2.7.2-map/{sched,regalloc,loop,cse_expr}.md: source-cited, byte-proven
  pass -> residual -> C-lever catalogs (935 lines). cookbook §31 = the index + triage.
- WALLS BROKEN (byte-proven steerable, were "CONFIRMED unsteerable"): §10/§20
  hoist-vs-remat (cross-call address-caching), store-vs-load (/s), dbr delay-slot
  (D1+S2 fresh-local: func_801770E0 53->49), birthing-boost both directions.
  Incidental bank: func_80149374 x134 (fleet 64.86%).
- GENUINE walls -> permuter: S3 chain-priority sink, S11 LUID(x)alloc coupling,
  RC-6 pressure-lock, cse 1000-insn table flush.
- CORRECTION (loop agent): tools/reference/gcc-papermario is gcc 2.8.1 NOT 2.7.2
  (behavioral biv-elim diff). Vanilla gcc-2.7.2 -> tools/reference/gcc-2.7.2/
  (gitignored); SETUP §5.6 + §31 flag it. All banked levers stand (match_one-validated
  vs the real cc1). Also corrected: spill-slot = declaration order; §25 4th rank rule.
- the payoff: the cheap tier (Opus agents + local model) can now apply compiler-internal
  levers by triage-table lookup, without reading 80k lines of source.
2026-07-02 20:12:19 -06:00
Drew T fcfe3afea6 feat(phase-23): gccmap-remat gate — +1 fns x1 propagated (fleet 64.86%) 2026-07-02 20:08:10 -06:00
Drew T 337a30a88f docs(phase-23): CURRENT_PHASE — the giant campaign (Fable5 wall-break + toolkit generalization, 3 giants ×134, reframe) 2026-07-02 17:44:39 -06:00
Drew T e74dd52b19 docs(phase-23): cookbook §30a — §30 generalizes via Opus agents + 2 new levers (IV-combine, inline-limit) + mechanical macro-widen integration 2026-07-02 17:34:53 -06:00
Drew T 396fd28398 feat(phase-23): toolkit-storeload gate — +2 fns x2 propagated (fleet 64.82%) 2026-07-02 17:33:17 -06:00
Drew T 6678c6868a feat(phase-23): Fable5Max cracks §20 "unsteerable" giant func_8014EE14 ×134
- Fable5Max agent (Agent model=fable) matched a 248-ins reach-134 GIANT on the
  §20/§10 store-vs-load wall (22 phases "CONFIRMED unsteerable") by reading the
  gcc-2.7.2 source (tools/reference/gcc-papermario) + RTL -da dumps. Leaf
  MATCH(248 ins) -> whole-binary banked:1 -> dedup_propagate ×134. Verified:
  check-all 136/136 byte-identical, dedup-check 1780 validated/0 failed.
- 3 byte-proven idioms -> cookbook §30 (corrects §29's "not a bigger model"):
  (1) store-vs-load is a deterministic MEM_IN_STRUCT_P /s aliasing flag, not a
      scheduler tie-break; steer via ((struct{s32 f;}*)p)->f (anon struct keeps
      /s AND propagates ×134) to grant, *p to deny
  (2) def-side return-type wall has a MACRO escape: widen a discarding caller
      macro's extern void->s32 (byte-neutral, check-all-verified) -- extends §29
  (3) birthing-boost prologue-order lever: __asm__("":"=r"(x):"0"(x)) re-tie in a
      later bb kills sched.c's REG_N_SETS==1 priority boost
- tools/glm_parallel.sh: K concurrent OpenRouter/GLM cloud drafters (parallel
  api_draft), key read from .env at runtime
- §10/§20 "store-vs-load unsteerable" backlog now re-test candidates:
  func_8014F2E0, func_80150528, func_8014EA4C
2026-07-02 16:56:56 -06:00
Drew T 13122771a0 feat(decomp): bulk_harvest — +45 fns across 13 binaries, 6 propagated (fleet 64.7%) 2026-07-02 15:29:38 -06:00
Drew T 0d1e8d2372 feat(decomp): bulk_harvest — +44 fns across 18 binaries, 3 propagated (fleet 64.69%) 2026-07-02 14:57:27 -06:00
Drew T 09e80573b7 feat(phase-23): glm-small40 gate — +2 fns x2 propagated (fleet 64.67%) 2026-07-02 11:08:49 -06:00
Drew T 83d3c762ee docs(phase-23): T10.9 verdict + clean stopping point — idiom well dry (both angles); LLM tier concluded; next = PhaseEnd + public flip 2026-07-02 02:02:56 -06:00
Drew T 00d9d9d491 feat(phase-23): T10.9 — 2 fresh hard-band GLM banks; fresh hand-solve confirms idiom well is DRY
- GLM deep-solved 15 fresh 25-118-ins hard fns: 4/15 match_one, 2 whole-binary banks (func_8017DE28,
  func_8015EEE0), $2.00. def-side wall caps the other 2 correct bodies.
- IDIOM VERDICT (Drew's fair test): GLM's correct bodies reason about KNOWN gcc mechanics (delay slots,
  callee-saved $s0, reload-after-call aliasing, switch jump tables) — cookbook §10/§17/jump-table.
  NO new idiom. The quirk space is largely mapped (22 phases of Opus-Max mining). Well DRY confirmed
  from BOTH angles: failed-residual (T10.8) AND fresh-hand-solve (T10.9).
2026-07-02 02:02:25 -06:00
Drew T ecd8741b94 docs(phase-23): CURRENT_PHASE checkpoint — T10.6-T10.9 (GLM frontier exploration) + resume brief 2026-07-01 23:48:57 -06:00
Drew T d71305fa5f feat(phase-23): T10.8 idiom-hunt harness + calibration — GLM confirms our idioms, no NEW ones
- tools/idiom_hunt.py: group backlog near-misses by residual class -> GLM names the reusable idiom +
  emits corrected C -> byte-gate to validate; captures reasoning; HARD --budget cap
- CALIBRATION ($0.51 total, 2 classes): struct + regalloc-order -> 0 banks. GLM re-derives our OWN
  idioms (register-pin §17, array-of-struct §18, type-width §25) and confirms walls, but banks nothing
  new — the backlog near-misses are the residual our idioms already failed on (irreducible/def-side wall).
- VERDICT: the new-idiom well is DRY (Fable5 review confirmed empirically for $0.51, not $300 overnight).
  GLM's value stays: direct drafter for FRESH def-conflict-free hard fns (~22%), not an idiom generator.
2026-07-01 23:13:29 -06:00
Drew T 6127367017 docs(phase-23): T10.7 conclusion — def-side wall intrinsic (GLM 1/7); GLM = drafter+teacher not wall-breaker
- cookbook §29: reasoning-model reconciliation idioms (match-pointer-type-to-TU-decl, call-site cast
  for value mismatch, cast-a-callee-definition, data-type match) + the narrow-param hard limit
- gen2-mips-matching-model.md + CURRENT_PHASE: Option-3 verdict (GLM reasons the wall expertly but
  banks 1/7; wall INTRINSIC, Fable5 §3c triple-confirmed); GLM role = $0.03-0.08/fn hard-band drafter
  + idiom teacher; real lever past the wall = public flip, not a bigger model
2026-07-01 21:05:18 -06:00
Drew T 8562aa89c1 feat(phase-23): T10.7 Option-3 — GLM reasons the def-side wall; +1 bank (func_80175184)
- tools/glm_reconcile.py (NEW): aim GLM's reasoning at the DEF-side loose-typing wall (body + conflicting
  TU decls + reconciliation toolkit -> consistent buildable byte-identical decls); captures reasoning
  (.run/glm_reason/, idiom source R16); relax-in-any-TU-file + crash-robust call
- api_draft: REASON=1 saves the reasoning trace per draft (idiom mining on any GLM run)
- fix_arity_callers: --any-proto (relax any prototype, not just (void))
- RESULT: GLM's reasoning is expert-level (store-width/sh-vs-sw awareness, K&R promotion, independently
  derives the cast idiom) but banks only 1/7 reconciliations; mechanical relaxation 0/7. The def-side
  wall is INTRINSIC (narrow-param + byte-level addressing defeat reconciliation) — Fable5 §3c re-test
  CONFIRMS the wall holds even vs a frontier reasoning model aimed directly at it. func_80175184 banked,
  check-all 136/136
2026-07-01 21:03:08 -06:00
Drew T bce95a13cd docs(phase-23): T10.7 GLM5.2 A/B result — 10x codegen edge, but def-side wall caps banks (3/18)
- GLM5.2 vs v3 on 18 hard-band fns: 10/18 vs 1/18 match_one; 3/18 vs ~1/18 whole-binary bank
- def-side loose-typing wall (Phase 16/20) caps banking for ANY drafter (Fable5 review §3c: HOLDS)
- 6 hardest beyond GLM too (0/6 at MAXTOK=16000); ~$1.02 of $25 spent
- 3 strategic options handed to Drew (direct-drafter / flywheel / wall-reconciliation)
2026-07-01 20:01:14 -06:00
Drew T b1e0391c52 feat(phase-23): T10.7a — recover 1 more GLM bank via fix_arity_callers (func_801577C8)
- def-side conflict (engine_core.h forward-declared func_801577C8(void) vs GLM's byte-correct
  (s32) def) resolved by relaxing the caller decl to no-proto; strip GLM externs + gate. 136/136.
- FINDING: only +1 of 8 stranded recovers mechanically; the other 7 are the intrinsic Phase-16/20
  DEF-side loose-typing wall (5 have non-(void) conflicting forward-decls, 1 narrow-param) — the
  wall caps ANY drafter, not a v3-tuning artifact (Fable5 review §3c re-test: wall HOLDS)
2026-07-01 19:27:45 -06:00
Drew T 42719160e7 feat(phase-23): T10.7 — first GLM5.2 cloud-model banks (2 hard-band fns, byte-gated)
- GLM5.2 via OpenRouter banked func_8013373C + func_8012F8C8 (ov_SC01_077_a.c), 16-22 ins
  hard-band fns v3 could not match; whole-binary byte-gate verified, check-all 136/136
2026-07-01 19:17:39 -06:00
Drew T 1caff3dc4c feat(phase-23): api_draft MAXTOK env + OpenRouter cost capture (T10.7 reasoning-model support)
- MAXTOK env (default 512 = local v3 unchanged); reasoning models (GLM5.2) need a high cap or
  they spend the budget on reasoning tokens and return empty content
- accumulate usage.cost from the response -> per-run $ + $/fn readout (OpenRouter reports it)
2026-07-01 18:29:27 -06:00
Drew T 2d0d63953e docs(phase-23): Fable5 strategy review — scope-vs-reality scorecard, wall re-tests, resource map
- scorecard: original 2026-06-10 scope vs 22 phases of byte-verified reality (what held,
  what emerged beyond scope, what deviated and should be revisited)
- adversarial pass: P21 no-shortcut + giant scheduler walls HOLD; P16 loose-typing wall has
  a TIMESTAMP GAP (declared 06-19, pre-dating cast_call_sites/block-scope-externs/v3) -> re-test
- July-2026 resources: frontier-on-hard-band (T10.7 re-aim), continuous architect-tier judgment,
  RE-ELEVATE THE PUBLIC FLIP (community labor = the only lever that scales into the proven tail)
- strategic fork: posture A/B/C on the byte-match goal; recommends dual-metric (B), Drew's call
- ranked recs 1-6 + explicit endorsements of what not to change
2026-07-01 17:25:47 -06:00
Drew T 14fbc0039d docs(phase-23): T10.6 campaign result — 1297 banked, fleet 64.6%, ~92% ≤15 saturated (crash-graceful) 2026-07-01 14:23:09 -06:00
Drew T 70aaad4efb chore(phase-23): regenerate backlog + fleet digests (≤15 campaign end, fleet 64.6%) 2026-07-01 14:20:12 -06:00
Drew T 6069eaa727 feat(decomp): bulk_harvest — +5 fns across 5 binaries, 0 propagated (fleet 64.6%) 2026-07-01 14:17:42 -06:00
Drew T 8f15e2bad9 feat(decomp): bulk_harvest — +38 fns across 23 binaries, 1 propagated (fleet 64.59%) 2026-07-01 14:10:47 -06:00
Drew T 34ef8ad960 feat(decomp): bulk_harvest — +35 fns across 21 binaries, 0 propagated (fleet 64.58%) 2026-07-01 13:28:40 -06:00
Drew T d53b6ae56a feat(decomp): bulk_harvest — +45 fns across 28 binaries, 2 propagated (fleet 64.57%) 2026-07-01 12:45:44 -06:00
Drew T 8ad248fa3d feat(decomp): bulk_harvest — +48 fns across 35 binaries, 1 propagated (fleet 64.56%) 2026-07-01 12:10:38 -06:00
Drew T f2b9c24735 feat(decomp): bulk_harvest — +56 fns across 38 binaries, 2 propagated (fleet 64.54%) 2026-07-01 11:32:03 -06:00
Drew T a95c19447d feat(decomp): bulk_harvest — +58 fns across 39 binaries, 1 propagated (fleet 64.53%) 2026-07-01 10:44:57 -06:00
Drew T ecf01e32a3 feat(decomp): bulk_harvest — +58 fns across 38 binaries, 0 propagated (fleet 64.51%) 2026-07-01 09:59:28 -06:00
Drew T ef27dedb83 feat(decomp): bulk_harvest — +51 fns across 36 binaries, 2 propagated (fleet 64.49%) 2026-07-01 09:21:07 -06:00
Drew T eb696bf35d feat(decomp): bulk_harvest — +49 fns across 41 binaries, 0 propagated (fleet 64.48%) 2026-07-01 08:45:39 -06:00
Drew T 416b3f37a6 feat(decomp): bulk_harvest — +56 fns across 46 binaries, 2 propagated (fleet 64.46%) 2026-07-01 08:05:11 -06:00
Drew T 67a0718148 feat(decomp): bulk_harvest — +54 fns across 48 binaries, 3 propagated (fleet 64.45%) 2026-07-01 07:23:56 -06:00
Drew T 00064af879 feat(decomp): bulk_harvest — +47 fns across 43 binaries, 1 propagated (fleet 64.43%) 2026-07-01 06:46:07 -06:00
Drew T 7e9dd8f03a feat(decomp): bulk_harvest — +53 fns across 48 binaries, 2 propagated (fleet 64.42%) 2026-07-01 06:05:00 -06:00
Drew T c1afac9f76 feat(decomp): bulk_harvest — +70 fns across 58 binaries, 2 propagated (fleet 64.4%) 2026-07-01 05:17:55 -06:00
Drew T f55052e2b5 feat(decomp): bulk_harvest — +64 fns across 55 binaries, 1 propagated (fleet 64.38%) 2026-07-01 04:39:21 -06:00
Drew T b4693430a0 feat(decomp): bulk_harvest — +66 fns across 55 binaries, 1 propagated (fleet 64.36%) 2026-07-01 04:00:29 -06:00
Drew T 3d90707d41 feat(decomp): bulk_harvest — +69 fns across 62 binaries, 1 propagated (fleet 64.34%) 2026-07-01 03:22:46 -06:00
Drew T 7fa2134326 feat(decomp): bulk_harvest — +75 fns across 64 binaries, 0 propagated (fleet 64.32%) 2026-07-01 02:41:58 -06:00
Drew T d2d15c9e41 feat(decomp): bulk_harvest — +78 fns across 69 binaries, 1 propagated (fleet 64.3%) 2026-07-01 02:03:59 -06:00
Drew T 7efe5ce10d feat(decomp): bulk_harvest — +72 fns across 63 binaries, 5 propagated (fleet 64.28%) 2026-07-01 01:29:20 -06:00
Drew T 95586a6751 feat(decomp): bulk_harvest — +69 fns across 67 binaries, 6 propagated (fleet 64.26%) 2026-07-01 00:54:27 -06:00