Files
BFM-decomp/phase-ends/logs/Phase23.md
T
Drew T 80b89d8a2d feat(phase-23): LLM matching tier + the Fable5 wall-breaker + the gcc-2.7.2 codegen map (v1.22.0)
- LLM TIER (T1-T10.9): tools/serve_local + api_draft + lora_grind + bulk_harvest + the LoRA
  pipeline (v3 model). Local v3 saturated <=15 (~1300 banks, $0). Frontier GLM-5.2 (~$4)
  confirmed the def-side wall INTRINSIC + the idiom well DRY at the leaf (both angles).
- THE PIVOT (07-02): Fable5Max reading the gcc-2.7.2 SOURCE matched func_8014EE14, a §20/§10
  store-vs-load giant "CONFIRMED unsteerable" for 22 phases -> the "wall is intrinsic" verdict
  is MODEL-RELATIVE. Generalized to Opus agents applying §30 (2 more giants, 3-5x cheaper).
- §31 THE CODEGEN MAP: 4 Fable5 agents read the whole gcc-2.7.2 source -> docs/gcc-2.7.2-map/
  {sched,regalloc,loop,cse_expr}.md (935 lines, byte-proven residual->lever catalogs) +
  cookbook §31 (index+triage). Broke hoist-vs-remat/delay-slot/coalescing walls. Correction:
  gcc-papermario is 2.8.1 not 2.7.2 -> tools/reference/gcc-2.7.2/ (SETUP §5.6).
- THE WAVE (8 reach-134 near-misses, Opus + §31): 3 leaf-MATCH by lookup (47-137k tok vs
  200-250k walls) + 5 tight permuter seeds (func_80132784 240->4-off!), zero dead-ends.
- 5 giants/hard-fns banked x134 (func_8014EE14/F2E0/150528/149374/144090); fleet 63.66->64.86%,
  136/136 byte-identical, 0 NON_MATCHING. cookbook §30/§30a/§31. No new governance rules.
- FINDING: matching is SOLVED by the map; bottlenecks are now the permuter (S11 seeds) +
  whole-binary integration (auto-bank leaf-MATCHes) -> Phase 24. worklog -> logs/Phase23.md.
2026-07-02 21:51:25 -06:00

102 lines
46 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CURRENT PHASE — Phase 23: The offline LLM matching tier (free local-model grinder + the data flywheel)
**Generation:** Gen2 (15th phase of the arc) · **Started:** 2026-06-29 · **Effort:** Max (synthesis/decisions/debug) ↔ xHigh (mechanical) · **Plan:** set up at PhaseEnd_Phase22 close (Drew); substance was already built across the 2026-06-29 session.
> Per-task crash-recovery log (P3). This file is the committable state a fresh session resumes from. Phase 22 (giants) closed → `PhaseEnd_Phase22.md` + `phase-ends/logs/Phase22.md`. The deep design + measured results of this phase live in `docs/gen2-mips-matching-model.md` (+ `docs/history/cheap-tier-ab-experiment.md`); the working-knowledge in the `cheap-tier-ab-validated` memory.
## Goal
Stand up a **free, local, fine-tuned matching model** that grinds the small/medium-function bulk and banks byte-verified matches for **$0** — with (a) a **data flywheel** (banks grow the corpus → retrain → the model improves *during* the run), (b) **escalation** (the permuter `grinder.py` closes the regalloc/schedule near-misses the model leaves; bigger models / your Opus take the tail), and (c) the **whole-binary byte-gate as the incorruptible sole arbiter** (a weak model is a *throughput* risk only, never correctness). **On-demand, NOT 24/7** (run only when Drew says go).
## Decisions (Drew)
- **On-demand runs**, bounded — never a standing 24/7 daemon. Queue a large batch, run a single shot, measure, scale the next.
- **Serve the model on the GPU via `tools/serve_local.py`** (Unsloth/torch, OpenAI endpoint at `http://127.0.0.1:1234/v1`) — **LM Studio was ejected 2026-06-30**; the prebuilt `llama-cpp-python` CUDA wheel SIGILLs on this no-AVX-512 CPU, so the Unsloth path is the reliable one (no build, loads in ~6 s).
- **Measure before investing** (the discipline that killed the v1 dead-end): every corpus/model change is **retested free on the 7B first**, and only scaled to a dense model if the cheap retest pays. *(Now satisfied for the 14B step: v3-on-7B paid.)*
- **Corpus-v3 DONE** (macro bodies + struct types; v3 banks the small/setter bulk, fleet→63.82%). The **14B gate is satisfied** → **corpus-v4 (struct-giant types) + a dense Qwen2.5-Coder-14B on cloud** for the >15-ins/giant band is the warranted next investment — Drew's call vs. path A (reach≥2 with v3).
## The operating loop (when running)
A. `lora_grind.py` (GPU) drafts open ≤N-ins stubs with the served fine-tuned model → banks via `gate_stage` → near-misses to the backlog.
B. `grinder.py` (CPU permuter) closes the backlog's regalloc/schedule near-misses → banks. **Run A+B concurrently** (GPU vs CPU; the histogram already shows 94 schedule/regalloc nears waiting).
C. Periodic propagate sweep (`dedup_propagate`) multiplies each bank fleet-wide.
D. Periodically: `export_pairs → format_finetune → train_lora → redeploy` (the model learns its own new wins); read the near-by-class histogram → pick the corpus-v3 / recovery-tooling target; raise `--max-nins` as the band lifts.
## Tasks
- [x] **T1 — Cheap-tier A/B** (Opus orchestrator + cheap agents) — Haiku 4.8× more matches/$ than Opus on the bulk; tools `tools/workflows/ab_match.js` + `tools/ab_score.py`. (`docs/history/cheap-tier-ab-experiment.md`.)
- [x] **T2 — Stock-local floor** — Qwen3.6-35B-A3B (LM Studio): structurally smart but **0 reliable byte-matches** (can't refine to byte-exact); format-robust; full-cookbook context made it *worse* (dilution). The floor to beat.
- [x] **T3 — The LoRA specialist pipeline** — `export_pairs.py` (mine banked asm↔C from build objects) → `format_finetune.py` (Qwen chat-template, compile-filtered) → `train_lora.py` (Unsloth QLoRA, Qwen2.5-Coder-7B, GPU on the 3080 Ti) → `eval_lora.py` (held-out gate-true). Corpus `datasets/match_pairs/` (gitignored).
- [x] **T4 — Corpus-v2 = the extern-block fix (THE unlock)** — capture the `extern <type> D_xxx;` block the src declares above each def (correct byte-verified types) → self-contained completions, compile 52%→92%, train 638→1111 (non-trivial 257→813). Same 7B retrained: held-out **6–15 ins 0%→85%**, non-trivial 0→26, meaningful(>15) 0→3. **Data was the bottleneck, confirmed.**
- [x] **T5 — First REAL banks on open stubs** — the v2 model drafted real open ov_SC01_077 stubs (LEAN, raw `.s` — model is format-robust, no bridge needed); whole-binary gate **banked 4** (func_80160B34, func_8015CC74, func_8016084C, func_801705C0; @commit:0320). Sample rate: 9/22 match_one proxy → **4/22 (18%) whole-binary** (the proxy→gate TU-plumbing gap).
- [x] **T6 — The mass-run driver** — `tools/lora_grind.py`: rotates every binary (config/check.*.sha), drafts the served model, banks via gate_stage (commit), defers/periodic-propagates, writes the **classified near-miss histogram** (the flywheel "missing idioms" signal). STOP/heartbeat/stats like grinder.py.
- [x] **T7 — Calibration run + the 0/222 puzzle — DEBUGGED + FIXED (2026-06-30).** Root cause = TWO independent harness bugs in `lora_grind`'s use of `gate_stage.run_gate` (R14, by reading the code + the run's backlog — which resolved a direct contradiction between two scout agents): **Bug A** — `good_sha()` passed the whole sha1sum line `"<sha> <name>"` vs harvest_verify's bare `sha1()` → **0 banks for EVERY binary incl. 077** (the "0/12" was a bug artifact, NOT an exhausted tail — the prior "`.sha` files carry the filename" note was the unfollowed thread); **Bug B** — the gate call left `src/asm/out` at the hardcoded ov_SC01_077 defaults → non-077 drafts dropped at the 077 stub-filter, **silently** (and the asm mis-resolution contaminated the backlog near-miss classes). Fix (`tools/gate_stage.py`): `run_gate` is binary-agnostic (resolve src/asm/out/good_sha from `binary`; good_sha bare-hash normalized) + a **loud negative-control guard**; `lora_grind.good_sha` fixed at source; byte-neutral (check-all 136/136). **Proof:** ov_SC01_000 spot-run banked **7/15 (47%) byte-identical** (was 0; @commit:0322). **ROI finding:** 6/7 banks are reach-1 (overlay-unique ×1) → broad rotation = high bank-RATE, low fleet-% ROI; the fleet lever is reach≥2 targeting (T9) + corpus-v3 (T8). Full write-up: `docs/gen2-mips-matching-model.md` → "T7 RESULT".
- [x] **T8 — Corpus-v3 — DONE (2026-06-30).** Mined the **1623 `engine_core.h` `DEFINE_func` macro bodies** (the shared setters/return-const the model was blind to — 96.6% of v2 was overlay-unique) + `format_finetune` inlines `engine_types.h` structs → corpus **1312→2891**, trainable **2534+291** (2.5× v2). **v3 trained** (loss 1.275→0.085), held-out eval **23/40 (57.5%)**; production batch banked ~352 fns + 45 shared groups → fleet **63.67→63.82%** ($0). The empty-leaf/setter class v2 couldn't draft is now banked. Struct-GIANT types deferred → **corpus-v4** (with a 14B). Details: `docs/gen2-mips-matching-model.md`.
- [x] **T9 — Operating loop — WIRED + RUN (2026-06-30/07-01).** Local GPU serving stood up (`tools/serve_local.py`, Unsloth); retrain cycle + reach≥2 targeting + propagate sweeps run end-to-end. Grinder concurrent = 0 banks (permuter tail exhausted). **2026-07-01:** measured v3-on-reach≥2 (banks 6/15), built `inject_capped_externs` (+0.34% fleet, 11 reach-134 fns ×134), and found the saturation campaign needs a throughput rebuild (≤15 band = ~4.6k unique funcs; current tooling ~48s/func + GPU-idle-during-gate) → **T10**.
- [x] **T10 — PHASE-SEPARATED + PARALLEL-GATE HARVESTER — BUILT + MEASURED (2026-07-01).** `tools/bulk_harvest.py` (bulk-draft GPU → `ProcessPool` parallel-gate over DISTINCT binaries → dedupe-once + ONE commit) + the back-compat `gate_stage`/`harvest_verify` parallel-safe params (T10.1: `lock_path`/`verified_out`/`failed_out`/`compute_fleet`; defaults = serial, so lora_grind/grinder unaffected). Validated (12→9 banked, 136/136, R23-clean commit) then MEASURED on 80 fresh SC03 ≤15-ins: **65% bank-rate** (52/80; 75% on the tiny band), gate **0.4s/fn** (8 workers, ~75× the serial gate) → **drafting (13.5s/fn serial) is now the SOLE bottleneck**. Commits: tooling `commit:0379`, validation `commit:0380`, measurement `commit:0381`; fleet 64.19→64.20%, dedup 1687→1715/0, 136/136 throughout. Details: `docs/gen2-mips-matching-model.md` "T10 RESULT". `lora_grind` kept as the serial fallback.
- [ ] **T10.5 — vLLM batched serving (the drafting lever) — DEFERRED, Drew-go.** The measurement JUSTIFIES it (65% bank-rate; the full ~4,087-unique ≤15 campaign is ~15h serial drafting @ 13.5s/fn; vLLM continuous-batching 5–20× → ~1–3h). Path: separate `.venv-vllm`; try bnb-direct + runtime LoRA first, fall back to merge→fp16→AWQ 4-bit (12 GB fit); serve the same `/v1/chat/completions` on :1234 so `bulk_harvest`/`api_draft` are unchanged. Quantization = throughput risk only, never correctness (byte-gate). **Do NOT install without Drew's go** (P5d toolchain add). Meanwhile the campaign can run on the serial path in resumable background chunks ($0).
- [ ] **T11 — Progress honesty + PhaseEnd** — track fleet % + bank-rate (function-count AND the decomp.dev-comparable byte-weighted ~30%); PhaseEnd_Phase23 at a clean checkpoint (Tier-1 Max).
## ▶ RESUME HERE (fresh session)
**State:** Phase 23 in progress (NOT a phase end). **v3 is the current model** — `bfm-match-7b-v3` (Qwen2.5-Coder-7B QLoRA on **corpus-v3**), adapter at `models/bfm-match-7b-v3` (v2 kept as fallback at `models/bfm-match-7b`). **LM Studio is EJECTED** — serve via **`tools/serve_local.py`** (Unsloth GPU, OpenAI endpoint), NOT LM Studio. The **8-hour autonomous run (2026-06-30)** built local serving + the prompt fix + corpus-v3 + v3 + a production batch → **fleet 63.82%** (+502 byte-identical, $0), 136/136 byte-clean, 27 commits this session (local — **Drew pushes**, R6). Pipeline validated end-to-end: a free local model banks the small/setter bulk, including the empty-leaf class v2 couldn't. Corpus `datasets/match_pairs/` + `.venv-train` gitignored. (Phase 22 close `PhaseEnd_Phase22.md` is committed `commit:0325`.) **UPDATE 2026-07-01 (this session):** measured v3-on-reach≥2 (banks 6/15, the band is reachable) + built **`tools/inject_capped_externs.py`** (the macro-extern-injection lever) and freed **11 reach-134 propagation-capped fns ×134** → **fleet 64.16%** (+0.34%, $0), 136/136 byte-clean, 1687 groups. v3 still served via `serve_local.py`. See the 2026-07-01 progress-log entry for the full arc + the deferred type-walled-14 + the strategic fork. **2026-07-01 close: fleet 64.19%** (injection lever +0.34% + a small campaign bank), 136/136 byte-clean. **The session ended at an ARCHITECTURE DECISION (Drew confirmed option c — build the factory fresh):** the saturation campaign needs a throughput rebuild before it's worth running. `serve_local.py` may still be UP holding the GPU (~8.7 GB) — stop it (`pkill -f serve_local`) if you want the GPU free.
---
## ★ NEXT TASK (fresh session) — BUILD THE PHASE-SEPARATED + BATCHED HARVESTER (Drew-confirmed, 2026-07-01)
> This supersedes the old "A then B / reach≥2 / 14B" plan. **Tier-1, plan-mode, Max.**
**Drew's directive (verbatim intent):** saturate **every small (≤15-ins) function across the whole project with v3** ($0) — target ~10,000 — so the heavier model (v4) only has to face the giants/struct band. *Then* v4-7B on giants as a test; if good, 14B+ in the **cloud** (money/compute no object). Exhaust ≤15 **fleet-wide** before scaling `--max-nins` to bigger funcs (Drew was explicit on this).
**WHY a rebuild (the blocker, byte-measured this session):** the current `lora_grind` is **~48s/function** (≈18s serial drafting + ≈30s recovery+gate on big overlays) and **idles the GPU during every gate** (the "low GPU usage" Drew noticed) — so the full ≤15 band (below) is a multi-day grind on tooling that wastes the hardware. **Drew's phase-separation idea is the fix.**
**THE BUILD (in priority order):**
1. **[BIGGEST LEVER] vLLM batched serving** to replace the serial Unsloth `serve_local.py`. Unsloth/transformers does ~18s/func *serial*; vLLM continuous-batching drafts *many at once* → 5–20×. Serve base+LoRA (`models/bfm-match-7b-v3`) as the same OpenAI endpoint `api_draft`/`lora_grind` already hit (`http://127.0.0.1:1234/v1`), so nothing downstream changes. (Install vLLM in a venv; mind the no-AVX-512 CPU + the 12 GB card. If the **Titan X** is plugged in: serve-batch on one card, gate on the other's CPU.)
2. **`tools/bulk_harvest.py`** — phase-separated driver: **(A) bulk-draft** K fresh fns → one drafts dir (GPU, batched); **(B) bulk-gate IN PARALLEL across cores** (`ProcessPool` over distinct binaries — different overlays build to different `build/<bin>` dirs → parallel-safe; gate with `commit=False`, collect banks, **commit once** at the end to avoid git-index races); **(C) dedupe once** (`dedup_propagate --auto-from` per banked binary). Reuse `lora_grind.draft` + `gate_binary` (refactor `gate_binary` to take `commit=False`). Optional v2: **pipeline** (draft N+1 while gating N).
3. **FIRST measurement (do before scaling):** v3's real **bank-rate on fresh location code** is STILL UNMEASURED (the SC03 run hung on the slow gate). Draft ~50–100 fresh SC03 fns through the new fast path → bank-rate decides **~8,000 vs ~2,000** yield → whether the multi-day campaign is worth it. (Use `match_one` as a fast proxy if the whole-overlay gate is still slow.)
**THE FUEL (measured 2026-07-01, `.run/unique_fuel.py` / `.run/fuel_by_group.py`):** **~32,636 open ≤15-ins stub-instances / ~4,597 unique-by-name / 4,087 untried.** Per group: SC01 2,133 (harvested-dregs, ~0 bank) · SC02 3,978 · **SC03 12,497** · SC04 4,580 · SC05 3,616 · SC06 4,668 · SC07 1,164. Almost all untried ≤15 are **reach-1 (location-unique, ×1)** — the shared band was already matched/propagated, so the fleet-% yield is direct-not-multiplied (banking ~25% of 4k uniques ≈ +1–2%, the "10,000" comes from counting per-overlay banks). `lora_grind` persists `tried` (`.run/auto/lora_grind_tried.json`, currently ~632) → drafts each unique once, RESUMABLE.
**FIXES ALREADY APPLIED THIS SESSION (don't redo):** `api_draft` max_tokens 4096→**512** (killed the no-stop-token ramble, ~80–130s→~10s on the bad cases); `lora_grind` runs `progress.py --fleet` **only on propagate sweeps** (was every batch = ~14s overhead); `lora_grind --min-nins` (target a fresh band w/o re-grinding small failures — **for the LATER >15 band, NOT now**). `tools/inject_capped_externs.py` (the macro-extern-injection lever) is DONE + tapped for 077 main+`_a` (11 freed ×134); the **14 type-walled reach-134 fns** are deferred (logged in the 2026-07-01 progress entry — `build_engine_types --strip` is collision-blocked on `typedef vec`; manual reconciliation, low-ROI).
**Serve + run (current serial path — works but slow; the rebuild replaces serving + the driver):**
```
# serve v3 (serial Unsloth — TO BE REPLACED by vLLM):
LD_LIBRARY_PATH=$(ls -d .venv-train/lib/python3.12/site-packages/nvidia/*/lib | tr '\n' :) \
nohup .venv-train/bin/python tools/serve_local.py --adapter models/bfm-match-7b-v3 --name bfm-match-7b-v3 --port 1234 >.run/serve_v3.log 2>&1 &
# the new driver (to build): bulk_harvest --count K -> draft-all -> parallel-gate -> dedupe-once
# stop a run: touch .run/auto/STOP (or pkill -f lora_grind / serve_local)
```
**PARKED (Drew, for later — captured, not now):** `docs/community-matching-model-plan.md` (publish the drafter; universal toolchain-conditioned model on decomp.me data) · `docs/sister-decomps.md` (contribute to Vagrant Story = CC0 same-compiler twin; the gitignored-fork-clone layout, pending Drew's go) · memory `cross-project-idiom-discovery` (when BFM's idiom well dries up, mine a same-compiler sibling). The **14B-for-giants** step is now AFTER the ≤15 saturation (corpus-v4 struct types → cloud A100/H100).
## Verification invariant (every bank)
The **whole-binary byte-gate** (`gate_stage`/`harvest_verify`, G3/P9) is the sole arbiter — a wrong/weak draft can NEVER bank (it reverts to the stub). `make check-all` 136/136 byte-identical from a clean tree (R22); `dedup-check` 0 failed; the `db.*.gbf` churn is R23 restart-noise (do NOT stage). The fine-tuned model only affects *throughput*, never correctness.
## Reuse (no rewrites)
*New (this phase):* **`tools/serve_local.py`** (Unsloth GPU serving, OpenAI endpoint — **replaces LM Studio**) · `api_draft.py` (LEAN drafter + the "translate every instruction, never empty" prompt fix) · `ab_match.js`/`ab_score.py` (the cost A/B) · `export_pairs.py` (now mines the `engine_core.h` macro bodies — corpus-v3) / `format_finetune.py` (inlines `engine_types.h` structs) / `train_lora.py` / `eval_lora.py` (the LoRA pipeline) · `lora_grind.py` (the mass-run; `--min-reach`/`--min-nins`; progress.py only on propagate-sweeps) · **`inject_capped_externs.py`** (2026-07-01, cookbook §28d — frees reach-134 inline matches for ×134 propagation by block-scope-injecting their file-scope externs; byte-neutral, fail-safe). **TO BUILD (T10):** `bulk_harvest.py` + vLLM serving.
*Existing (now Phase-23 binary-agnostic):* `gate_stage` (`GATE_PHASE`-tagged + resolves src/asm/out/good_sha from `binary`) / `grinder.py` (per-binary + grouped gating) + `auto_supervisor.sh` / `dedup_propagate` / `backlog` (fleet-aware `load_best`) / `harvest_verify` / `match_one`; cookbook §17–§28.
## Guardrails
- **On-demand only** — never leave a 24/7 daemon running; bounded `--max-batches`, STOP-sentinel safe-exit.
- **Measure before investing** — retest every corpus/model change free on the 7B before any cloud spend.
- **Disk** — GGUF conversion writes ~30 GB intermediates onto the WSL vhdx (C:); `train_lora` now auto-cleans them; keep only the q4 GGUF. (A disk-full crash cost a session on 2026-06-29.)
- **Serve on GPU via `tools/serve_local.py`** (Unsloth/torch, the in-repo replacement for the ejected LM Studio) — the prebuilt llama-cpp-python CUDA wheel SIGILLs on this no-AVX-512 CPU; don't try it.
- Model size matches data size — no "massive" models on ~1–2k examples (overfit); dense > MoE for a limited-data LoRA.
## Blockers
- **(current) Throughput — the saturation campaign can't scale on the current tooling.** The interleaved `lora_grind` is ~48s/func (≈18s serial draft + ≈30s recovery/gate on big overlays) and idles the GPU during gating → the ~4.6k-unique ≤15 band is multi-day. **Not a bug — an architecture gap.** Unblocked by **T10** (vLLM batched serving + the parallel-gate `bulk_harvest`). The current serial tooling still WORKS (banks correctly, byte-gated) — usable as a slow fallback if needed.
- **(open, low-ROI) The 14 type-walled reach-134 fns** (`build_engine_types --strip` collision-blocked on `typedef vec`) — manual type-collision reconciliation, deferred (2026-07-01 progress entry has the addr list).
- **(cleared) T7 — the 0/222 broad-rotation banking.** Two harness bugs (good_sha sha1sum-format + the src/asm/out 077-default) fixed + proven (ov_SC01_000 7/15 byte-identical, check-all 136/136).
## Progress log
- 2026-06-29: **Phase opened at PhaseEnd_Phase22 close (Drew).** Built across this session: cheap-tier A/B (T1, Haiku 4.8×/$), stock-local floor (T2, 0), the LoRA pipeline (T3) + corpus-v2 extern-fix (T4, 6–15 ins 0%→85%), first 4 real open-stub banks (T5, @commit:0320), the `lora_grind` mass-run driver (T6). Calibration run (T7) launched (500 fns) — **18% on ov_SC01_077 but 0/222 broad rotation → #1 debug.** gate_stage commit tag made phase-agnostic. The whole arc + measured numbers: `docs/gen2-mips-matching-model.md`; memory `cheap-tier-ab-validated`. NEXT: T7 debug, then bounded mass-runs + corpus-v3.
- 2026-06-30: **T7 DEBUGGED + FIXED.** 3 Explore scouts (tooling / run-evidence / corpus) + a direct code read (R14 — which resolved a flat contradiction between two scouts) found **two independent bugs** in `lora_grind`'s gate path: **(A)** `good_sha()` passed `"<sha> <name>"` vs harvest_verify's bare `sha1()` → 0 banks for ALL binaries incl. 077 (so 077's "0/12" was a bug artifact); **(B)** `src/asm/out` defaulted to ov_SC01_077 → non-077 drafts dropped at the 077 stub-filter, silently. Fixed `gate_stage.run_gate` (binary-agnostic resolution + bare-hash normalize + a loud negative-control guard) + `lora_grind.good_sha`; byte-neutral (check-all 136/136). ov_SC01_000 spot-run **banked 7/15 (47%) byte-identical** (@commit:0322) → reach-2 `func_8017CE24` propagated ×2. **ROI:** 6/7 reach-1 → broad rotation is high bank-rate / low fleet-% ROI; the fleet lever is **reach≥2 targeting** + corpus-v3. Backlog now correctly classified (4× close=1 = grinder fuel). NEXT: **T8 corpus-v3** (struct types) + **T9 reach≥2 selection** + concurrent grinder.
- 2026-06-30 (cont.): **T9 reach≥2 targeting built + measured.** Added `lora_grind --min-reach N` (lazy sig-based reach oracle == `dedup_propagate`, validated 0-mismatch/60 + the func_8017CE24=2 ground truth; `--min-reach 2` ranks high-reach-first, naturally restricts to overlays). Bounded reach≥2 mass-run: ov_SC01_000's 15 reach≥2 (shared) stubs banked **0/15** (vs the reach-1 spot-run's 7/15) — the model is **weakest on shared code** (corpus skipped the `DEFINE_func` bodies + it's the regalloc/schedule tail). But **5/15 are close≤3 reach-134 near-misses** (3× close=1 = func_8012E27C/BF4C/AD64) → high-value permuter fuel (×134 each). **FINDING: reach≥2 model-only ≠ a fleet lever; the lever is reach≥2-draft → grinder-close (×134)**, which needs `grinder.py`'s per-binary fix (same class as T7) + a backlog `binary` field. (A foreground mass-run hit the 10-min Bash cap mid-2nd-batch; tree recovered clean via `git checkout`, check-all 136/136.) Details: `docs/gen2-mips-matching-model.md` "T9 RESULT". NEXT: the grinder per-binary fix (realize the reach-134 ×134 fuel), then corpus-v3.
- 2026-06-30 (cont.): **Grinder per-binary fix (5-layer) — built + validated; reach≥2 fuel re-characterized.** The grinder/backlog were ov_SC01_077-hardcoded **5 layers deep** — fixed all: `gate_stage` records `binary`, `backlog.FIELDS` keeps it, `backlog.load_best`/`_open_stubs` fleet-aware (stuck-local fns surface via their overlay record), `p16_permute.setup` takes asm-subdir, `grinder` per-binary resolution + grouped gating + None-`nins` allow. Validated: the 3 fresh reach-134 close=1 ov_SC01_000 fns now surface + resolve + gate per-000. **FINDINGS (byte-evidenced):** (a) the reach≥2 close=1 fuel is **MODEL semantic-misses** (7B drafts empty `void f(void){}` for trivial `return 1`/setters), NOT permuter fuel — a corrected draft banks them (**+3 byte-identical** via the fixed gate, @commit:0326); (b) ×reach is **propagation-capped** (inline-matched in `ov_SC01_077_a.c` → dedup "nothing to propagate" → banked ×1). So the reach-134 ×134 payoff routes through **corpus-v3** (leaf variety) + the **dedup-collapse**, NOT the permuter. check-all 136/136 throughout. Details: `docs/gen2-mips-matching-model.md` (grinder fix). NEXT: corpus-v3 + dedup-collapse.
- 2026-07-01 (**A: reach≥2 measurement + the macro-extern-injection lever — fleet 63.82%→64.16%, $0 LLM**): served v3 (`serve_local.py`, ~15s). **(1) Measured v3 on reach≥2** (`lora_grind --min-reach 2`, bounded): banked **6/15 in ov_SC01_005** (+1 in 006) — **v3 banks the reach≥2 band**, vs v2-pre-fix's 0/15 (T9). But the banks are mostly low-reach (the 005≡006 pair = ×2) and re-bank inline per overlay → the fleet-% needs the COLLAPSE. **(2) Scoped the dedup-collapse:** the broad-overlay inline banks are mostly already-propagated (8-hr sweeps) or reach-1; the real ×134 fuel is **29 reach-134 fns matched INLINE in ov_SC01_077 but `dedup_propagate`-skipped as "not self-contained"** (the §19/20 propagation cap; the Phase-20 backlog list). Diagnosed (R14): they fail `compiles_standalone` on **file-scope callee/data externs** (clean lever) or **local types** (fiddly). **(3) Built `tools/inject_capped_externs.py`** (the Phase-20-backlog "macro-extern-injection", never built): source each capped fn's externs from 077's OWN file-scope decls, inject block-scope → self-contained → liftable ×134. Byte-neutral (077 stays `d19c9580…`), fail-safe (only rewrites if then `compiles_standalone`). **Freed 8 (main .c) + 3 (`_a.c`) = 11 reach-134 fns**, propagated ×134, **9 new shared groups (1678→1687)**, fleet **63.82%→64.16%** (+0.34%, +1204 byte-identical), check-all **136/136**, dedup-check 0 failed, **$0 LLM**. cookbook **§28d** + SETUP inventory (R16/R21/R30). **(4) Type-walled subset (14 reach-134 fns) DEFERRED:** the §28b type-lift (`build_engine_types --strip`) is **collision-blocked** (`typedef vec differs between header and source`) — the inline-invented types (MATRIX/Buf/Vec3/Loc/vec/…) need manual name-collision reconciliation; low-ROI/fiddly, byte-gate+collision-guard correctly refused. The 14: `0x8014964C 0x8014A048 0x8014C8C8 0x801678F0 0x8016A8FC 0x801749C8 0x80177EA4` (main) + `0x8012A418 0x8012A464 0x8012C658 0x8012E778 0x8012EA90 0x80134310 0x801395D4` (_a). Blast radius of a lift = overlays ONLY (engine_types.h does NOT reach main/resident). **(5) Launched a bounded broad v3 harvest** (12 batches, propagate-every-4) for $0 corpus growth + incremental banks — running at write time. **NEXT:** verify+commit the harvest; then the strategic fork — (a) v3.1 flywheel retrain on the grown corpus (compounding, ~2h GPU), (b) the type-collision reconciliation (~+0.5%, fiddly), or (c) accept the automated ceiling and head to PhaseEnd. The injection lever now CLOSES the v3-reach≥2→×134 loop for future 077 banks (idempotent re-run). **(6) Saturation-campaign scoping (Drew's "10,000 small funcs" directive — exhaust ≤15 fleet-wide before bigger):** measured the real fuel — **~32,636 open ≤15-ins stub-instances / ~4,597 unique-by-name / 4,087 untried** (NOT saturated — my "saturated" claim from the SC01-only sample was an over-extrapolation Drew caught, R14; SC03 alone has 12,497). Sample shows ~all untried ≤15 are reach-1 (location-unique, ×1) — the shared band was matched/propagated in earlier phases. **THROUGHPUT WALL (the real blocker):** drafting is ~15–18s/func SERIAL (the real draft = full .s input + recovery pipeline), so exhausting ~4k unique funcs is ~a day and the full per-overlay set is multi-day; and the whole-overlay byte-gate (~1s warm) is fast but a big-overlay gate hung 11+ min once. **Fixes applied:** `lora_grind` runs `progress.py --fleet` only on propagate sweeps (was every batch = ~14s overhead dominating small batches); `--min-nins` (target a fresh band w/o re-grinding small failures); `api_draft` max_tokens 4096→512 (the model rambled to 4096 = ~80–130s on no-stop-token drafts). **ARCHITECTURE CONCLUSION (Drew's phase-separation idea = right):** the interleaved draft(GPU)→gate(CPU) per-batch design idles the GPU during every gate (the "low GPU usage" Drew saw). The path: **phase-separate** (bulk-draft GPU → bulk-gate parallel-across-cores CPU → dedupe once) + **pipeline** (draft N+1 while gating N) + **batched inference via vLLM** (the real GPU-saturation lever — Unsloth serve does ~18s/func serial; vLLM continuous-batching is 5–20×). This is a substantial focused build (the proper `bulk_harvest` + a vLLM serving swap) — deferred to a fresh focused effort (R30: capture findings now, the build is continuable). v3 bank-rate on the fresh location code = STILL UNMEASURED (the SC03 measurement hung on the slow gate); measure it first thing (it decides ~8k vs ~2k yield → whether the multi-day campaign is worth it).
- 2026-06-30 (**8-hour autonomous run, Drew away**): prompt-fix + local serving + corpus-v3 + v3 + big batch. **LM Studio ejected** → built **`tools/serve_local.py`** (Unsloth GPU serving as an OpenAI endpoint; the prebuilt llama-cpp-python CUDA wheel SIGILLs on this no-AVX-512 CPU, so the Unsloth/torch path is the reliable one). **PROMPT FIX** (`api_draft.LEAN_SYS` + `format_finetune.SYS`, synced): "translate EVERY instruction, never an empty body" — the v2 empty-leaf overfit, small-leaf band **0/3→2/3**, banked 3 on a fresh ov_SC01_001 batch. **CORPUS-V3**: `export_pairs` now mines the **1623 engine_core.h `DEFINE_func` macros** (the shared setters/return-const the model was blind to — 96.6% of v2 was overlay-unique) + `format_finetune` inlines `engine_types.h` structs → corpus **1312→2891**, trainable **2534+291** (2.5× v2). **v3 trained** (Qwen2.5-Coder-7B QLoRA, loss 1.275→0.085, ~2h), held-out gate-true eval **MATCH 23/40 (57.5%)** (generalizing; v2's mixed-set rate was lower). Grinder concurrent during training = 0 banks (permuter tail exhausted, Phase-22 reality). **Big batch** (v3, broad rotation, 25 binaries, propagate-every-3, **$0 LLM**): banked **~352 fns inline + 45 new shared groups** (1633→1678) → fleet **63.67% → 63.82%** (+502 byte-identical), **136/136 byte-clean**, 25 auto-commits. v3 repeatedly banked the **empty-leaf/setter class v2 couldn't** (func_8012E27C=return 1, 8012AD64/BF4C=setters). **Pipeline validated end-to-end:** a free local fine-tuned model harvests the small/setter bulk at $0, gated identically (G3/P9). Details: `docs/gen2-mips-matching-model.md` ("Corpus-v3 ... 8-hour autonomous run"). **NEXT (fresh session):** (1) **shared/reach≥2 targeting** (`lora_grind --min-reach 2` with v3 — so the setter banks propagate ×134 instead of re-banking inline per binary — the fleet-% lever); (2) **dedup-collapse** the per-binary inline setters; (3) **corpus-v4** = struct-giant types; raise `--max-nins` as the band lifts. Serve: `tools/serve_local.py --adapter models/bfm-match-7b-v3` (R21 / SETUP §Tooling inventory).
- 2026-07-01 (**T10 — the phase-separated + parallel-gate harvester, BUILT + MEASURED**): plan-mode Tier-1, harvester-first (Drew picked it over vLLM-first — the exploration showed parallel-gate is the bigger single lever, no install, and unblocks the measurement). **Built** `tools/bulk_harvest.py` (bulk-draft GPU → `ProcessPoolExecutor` byte-gate over DISTINCT binaries → dedupe-once + one commit; round-robin fuel spread) on back-compat `gate_stage`/`harvest_verify` params (T10.1 — `lock_path`/`verified_out`/`failed_out`/`compute_fleet`, all default to serial; `backlog` redirected per-worker via its module global, no edit). Tooling commit `commit:0379`. **Validated** (count=12/workers=4): 9/12 banked, 136/136 byte-identical, R23-clean commit `commit:0380`, dedup 0 failed. **MEASURED** (count=80/workers=8 fresh SC03 ≤15-ins, commit `commit:0381`): **bank-rate 52/80 = 65%**; gate **0.4s/fn amortized** (8 workers) vs lora_grind's ~30s/fn serial ≈ **75×**; draft **13.5s/fn serial = 97% of wall-clock** → the parallel gate moved the bottleneck ENTIRELY to drafting. Fleet 64.19→64.20% (61 tiny reach-1 banks barely move the byte-weighted % — the ≤15 campaign's value is bank COUNT + completeness, freeing the heavy model for giants; honest P9). check-all 136/136, dedup-check 1687→1715/0. **DECISIONS SETTLED:** (1) campaign is GO (65% over ~4,087 untried uniques ⇒ ~2,650 free banks); (2) drafting is now the only lever; (3) **vLLM is JUSTIFIED** (~15h serial → ~1–3h batched) — deferred to Drew-go (T10.5). Details: `docs/gen2-mips-matching-model.md` "T10 RESULT". Serve still up via `serve_local.py` (v3). NEXT: Drew's call — vLLM install (T10.5) now, or run the ≤15 campaign on the serial path in background chunks, then PhaseEnd (T11). **Drew chose B (serial campaign now, defer vLLM).** **T10.6 LAUNCHED (2026-07-01):** `bulk_harvest --binary-glob 'ov_*' --count 150 --cycles 40 --workers 8` on `serve_local` v3 — the full ≤15-ins overlay saturation ($0), self-terminating on dry fuel (~15h max), **per-cycle commits (crash-safe/resumable via `.run/auto/lora_grind_tried.json`), STOP-able** (`touch .run/auto/STOP`). Monitor: `.run/auto/bulk_harvest_{heartbeat,stats}.json` + `.run/bulk_campaign.log`. Region bank-rate varies (SC03 fresh ~65%, SC01 dregs ~40%). `--cycles` campaign mode is committed (`commit:0384`); the `run_cycle` loop was validated (2-cycle run, per-cycle commits `commit:0385`/`commit:0386`). If a fresh session finds the campaign stopped: re-launch the same command (resumes from `tried`), or `touch .run/auto/STOP` to halt. After the campaign: `make check-all` (R22) + T11 PhaseEnd_Phase23. **Drew's new direction (2026-07-01), the NEXT step after the campaign:** test a HUGE-parameter cloud model via **OpenRouter** (GLM-4.6 / the current big GLM Drew called "glm5.2" — confirm the exact id on OpenRouter at run time) as a DRAFTER — `api_draft.py` already speaks the OpenAI endpoint (OpenRouter is OpenAI-compatible: set `API_BASE`=OpenRouter + key, `MODEL`=the GLM id). A/B vs local v3 on a sample (bank-rate + $/match), focused on the harder band the 7B plateaus on (>15-ins / struct). Byte-gate keeps it correctness-safe (throughput/cost lever only). **PARKED at Drew's direction:** the batching levers — DIY static batching (~3–4×, VRAM-capped on 12 GB) and vLLM (5–20×, needs a separate venv + tighter quant) — both on HOLD; revisit after the OpenRouter test. "We can move forward after that" (Drew). **T10.6 CAMPAIGN RESULT (2026-07-01):** ran **23 cycles → 1,297 fns banked ($0) + 52 propagated groups**, fleet 64.21→**64.6%** (+0.39%), tried 764→**4,214** (**~92% of the ~4,597 ≤15-ins unique fuel saturated**). **check-all 136/136 byte-identical**, dedup 0 failed. **Ended gracefully on a serve_local crash** (bitsandbytes `ops.cu:81` after ~14h continuous serving — the serial 4-bit server's long-run fragility, NOT a code bug; bulk_harvest's per-cycle endpoint-check caught the dead server, committed the partial cycle 23, and exited clean). Commits `commit:0387..commit:0410` + digests `commit:0411`. **Remaining ≤15 tail (resumable, low-value ~28% yield):** ~383 untried + ~132 skipped-marked-`tried` in the crashed cycle 23 (prune those from `.run/auto/lora_grind_tried.json` if resuming). To resume: restart `serve_local.py` then `bulk_harvest --binary-glob 'ov_*' --cycles N`. Wrapped up per Drew → **NEXT: T10.7 OpenRouter** (cloud, no local GPU). Serve_local is DOWN, GPU freed. **T10.7 RESULT (OpenRouter GLM5.2 A/B, 2026-07-01):** on 18 hard-band fns (16-22 ins, P16-walled `ov_SC01_077`), **GLM5.2 10/18 match_one vs v3 1/18** (~10× codegen edge) but **whole-binary banks only 3/18** (2 direct + `func_801577C8` via `fix_arity_callers`) — the **DEF-side loose-typing wall (Phase 16/20) HOLDS** and caps ANY drafter (Fable5 §3c re-test confirmed). 6 hardest fns beyond GLM too (0/6 at MAXTOK=16000). Cost ~$1.02 of $25. Banked 3 GLM fns (`commit:0416`); `api_draft` gained `MAXTOK` env + cost capture (`commit:0414`). Details: `docs/gen2-mips-matching-model.md` "T10.7 RESULT". **DECISION PENDING (Drew):** how to use GLM — (1) direct drafter on def-conflict-free hard fns ($0.03/fn, ~17% wall-capped), (2) corpus-v4 flywheel/retrain (also wall-capped, ~2h GPU, uncertain), or (3) point GLM at the def-side-wall RECONCILIATION (reasoning-shaped unlock, new harness). T10.7c retrain NOT started (wall caps its value → checkpoint for Drew). **T10.7 Option-3 probe DONE (`tools/glm_reconcile.py`, +$0.23):** aimed GLM's REASONING at the def-side wall — expert-grade reasoning (store-width/`sh`-vs-`sw` awareness, K&R promotion, independently derives the §17a-1 cast idiom), captured to `.run/glm_reason/` for idiom mining (R16 → cookbook **§29**). Banked **1/7** (`func_80175184`); mechanical `--any-proto` **0/7**. **VERDICT: the def-side wall is INTRINSIC** (frontier reasoning model + full toolkit cracks ~1/7; residual = narrow-param K&R limit + byte-addressing). **GLM total banks 4/18** vs v3 ~1/18. Tooling: `glm_reconcile.py` (NEW), `api_draft` REASON-capture, `fix_arity_callers --any-proto` (`commit:0418`); cookbook §29 + doc T10.7 RESULT. **GLM's role settled:** (a) $0.03–0.08/fn DIRECT drafter for the def-conflict-FREE hard band (~22%, v3 can't reach), (b) idiom TEACHER — NOT a wall-breaker. Per Fable5 §4.3 the real lever past the wall is the public flip (community), not a bigger model. Total T10.7 spend ~$1.25/$25. **NEXT (Drew's call):** (i) a bounded GLM DIRECT hard-band campaign (bank the conflict-free ~22% at scale, money-open), (ii) the public-flip pivot (Fable5 §6.3), or (iii) close Phase 23 → PhaseEnd (T11).
- 2026-07-01 (**evening — T10.6 ≤15 campaign + T10.7/8/9 the OpenRouter/GLM5.2 frontier-model exploration**): CONTEXT CHECKPOINT (85% ctx, wind-down). **T10.6 DONE:** `bulk_harvest.py` ≤15 saturation campaign ran 23 cycles → **1,297 banks, fleet 64.6%, ~92% of the ≤15 band saturated**, 136/136, crash-graceful exit (serve_local bnb `ops.cu:81` after 14h — not our bug). **T10.7 DONE (GLM5.2 A/B, `z-ai/glm-5.2` via OpenRouter):** on 18 hard-band fns (16-22 ins), **GLM 10/18 match_one vs v3 1/18** (~10× CODEGEN edge) but **whole-binary banks 4/18** — the **DEF-side loose-typing wall (Phase 16/20) is INTRINSIC** (caps ANY drafter; it's overlay-decl-vs-true-sig, a C/ABI limit). Option-3 probe (`glm_reconcile.py`): GLM reasons the wall EXPERTLY (store-width `sh`/`sw`, K&R promotion, independently derives the §17a-1 cast idiom) but banks only 1/7 — reconciliation idioms → cookbook **§29**. **T10.8 DONE (`idiom_hunt.py`, $0.51):** batch-clustered the FAILED backlog by class → GLM re-derives our OWN idioms (pins §17, array-of-struct §18) + confirms walls, **0 new bankable idioms** (the failed residual is the worst discovery material — Drew's catch). **T10.9 IN FLIGHT (bg task `bmv1wxzdz`, `.run/glm_fresh.log`):** Drew's CORRECT reframe — the idiom engine is DEEP hand-solving of FRESH never-tried hard fns (like Opus Max), then mining the byte-correct cracks for NEW idioms (idiom lives in the correct BODY, not the bank). GLM deep-solving **15 fresh 25-118-ins fns** (`ov_SC01_077`/`SC03_001`/`SC02_005`), LEAN, MAXTOK=32k, iters=3, **reasoning captured** (`.run/glm_fresh/*.reasoning.txt`, `REASON=1`). **TOOLING BUILT this session:** `tools/bulk_harvest.py` (phase-separated parallel-gate harvester + `--cycles`), `tools/glm_reconcile.py` (Option-3 wall reconciler), `tools/idiom_hunt.py` (idiom researcher, `--budget`-capped), `api_draft` (`MAXTOK` env + `usage.cost` + `REASON=1` reasoning capture), `fix_arity_callers --any-proto`, `gate_stage.run_gate` parallel params (`lock_path`/`verified_out`/`failed_out`/`compute_fleet`), `harvest_verify --verified-out/--failed-out`. **Commits:** `commit:0379` (T10 tooling)…`commit:0410`/`commit:0412` (≤15 campaign)…`commit:0414`(MAXTOK)…`commit:0418`(glm_reconcile+§29)…`commit:0420`(idiom_hunt+T10.8). **OpenRouter setup:** key in `.env` as `open_router_key` (gitignored, `sk-or-`); GLM5.2 is a REASONING model (needs MAXTOK≥8k or it returns empty content — burns budget on reasoning tokens); ~**$23.6 left of $25** (spent ~$1.90); the key shows `"limit": 20` (possibly a $20 spend cap — raise in the OpenRouter dashboard before a bigger run); no hard concurrency cap (~6 providers serve glm-5.2). **STRATEGIC STATE:** the LLM matching tier is now THOROUGHLY explored — local v3 saturated ≤15 ($0); frontier GLM confirmed the hard-band def-side wall is INTRINSIC (triple-confirmed) + banks a modest ~22% of the conflict-free hard band at ~$0.05/fn; the new-idiom well is DRY on the failed residual (T10.8), with T10.9 testing whether FRESH hand-solving still yields NEW idioms (Drew's hypothesis). **DREW CHOSE: (A) a bounded GLM fresh-hard-band campaign** (which T10.9 pilots). Per the Fable5 strategy review (`docs/fable5-strategy-review-2026-07.md`), the real lever PAST the wall is the **public flip / community labor** (§4.3), not a bigger model. **RESUME (fresh session):** (1) read `.run/glm_fresh.log` + gate the correct bodies (bank the def-conflict-free ones; use `glm_reconcile`/`inject_capped_externs` recovery), (2) MINE `.run/glm_fresh/*.reasoning.txt` for NOVEL idioms (vs cookbook §17-§29) → distill real ones to cookbook, (3) decide the fork: scale the GLM fresh-hard campaign (money-open but modest), pivot to the public flip (Fable5 §6.3, the strategic lever), or PhaseEnd_Phase23 (the LLM tier is well-characterized — a clean Tier-1 close). **T11 (PhaseEnd) remains.** serve_local is DOWN (GPU free). Cost discipline: Drew is cost-conscious — calibrate cheap, `--budget`-cap, no surprise overnight bills.
- 2026-07-02 (**T10.9 RESULT + STOPPING POINT — the fresh-hand-solve idiom test**): Drew's fair test of "does GLM hand-solving FRESH hard funcs teach NEW idioms?" (the mechanism that gave us pins/array-of-struct via Opus-Max). GLM deep-solved **15 fresh 25-118-ins fns** (LEAN, MAXTOK=32k, iters=3, reasoning captured `.run/glm_fresh/*.reasoning.txt`): **4/15 match_one, 2 whole-binary banks** (`func_8017DE28`, `func_8015EEE0`; @`commit:0422`, check-all 136/136), **$2.00**. **IDIOM VERDICT (the point):** GLM's 4 correct bodies reason entirely about **KNOWN gcc mechanics** — delay-slot scheduling, callee-saved `$s0` preservation, reload-after-call aliasing, switch jump tables — all already in cookbook §10/§17/jump-table. **NO new idiom emerged.** GLM is a strong reasoner *applying* mapped idioms, not finding unmapped quirks (expected — 22 phases of Opus-Max mining charted the common gcc-2.7.2 quirks). **So the new-idiom well is DRY, now confirmed from BOTH angles: failed-residual (T10.8) AND fresh-hand-solve (T10.9).** 8/15 were reasoning-truncated empties (giants overrun 32k); near-misses `func_8018119C` (near-4) / `func_80187DAC` (near-5) are grinder/reconcile fodder. **ROBUST STRATEGIC CONCLUSION (multi-angle):** the LLM matching tier is exhausted of cheap+idiom leverage — (1) local v3 saturated ≤15 ($0, ~1300 banks), (2) frontier GLM = a modest ~22% hard-band drafter at ~$0.05-0.13/fn (def-side wall caps banking; wall is INTRINSIC, triple-confirmed), (3) no new idioms from either failed-residual or fresh cracks. Total OpenRouter spend ~$4/$25. **Per the Fable5 strategy review (`docs/fable5-strategy-review-2026-07.md` §4.3/§6.3), the real lever past the wall is the PUBLIC FLIP (community hand-matching), not a bigger model.**
**★ CLEAN STOPPING POINT (fresh session, Tier-1 Max, plan mode):** the LLM-tier exploration is comprehensively concluded. Options for the fresh session, in recommended order: **(A) PhaseEnd_Phase23** — write it (the LLM tier is fully characterized: local grinder + frontier test + idiom hunt, all byte-proven; a clean Tier-1 close), then **(B) plan the PUBLIC FLIP as the next phase** (Fable5 §6.3: R20 clean mirror, rom→decoder, CI, a naming pass on major systems, contributor on-ramp from `docs/backlog.md` — the strategic lever). A bounded GLM fresh-hard-band campaign remains available (money-open, modest fleet%) but is NOT the high-leverage move. Fleet stands at **~64.2%** (function-count) / ~30% byte-weighted, 136/136 byte-identical, 0 NON_MATCHING. All work committed (`commit:0379`…`commit:0422`); Drew pushes (R6). serve_local DOWN, GPU free. `db.*.gbf` churn is R23 restart-noise (do NOT stage).
- 2026-07-02 (**THE GIANT CAMPAIGN — Fable5 wall-break + toolkit generalization; SUPERSEDES the "PhaseEnd next" stopping point above**): Drew redirected from PhaseEnd to hand-solving giants + a Fable5 experiment. **RESULT: 3 giants banked ×134, fleet 64.67% → 64.82%, check-all 136/136, all committed** (`commit:0427`, `commit:0428`, `commit:0429`).
- **Fable5Max (`Agent` model=fable) CRACKED `func_8014EE14`** (248 ins, reach-134) — a §20/§10 store-vs-load **"CONFIRMED unsteerable"** giant — by reading the **gcc-2.7.2 SOURCE** (`tools/reference/gcc-papermario`) + RTL `-da` dumps. **This REFUTES the T10.8/T10.9 "wall intrinsic / idiom well dry" verdict: it was MODEL-RELATIVE** (true for GLM/7B, false for a frontier model reading gcc internals). 3 idioms → cookbook **§30**: (1) store-vs-load is a deterministic `MEM_IN_STRUCT_P`/`/s` aliasing flag (grant via `((struct{s32 f;}*)p)->f` anon-struct member-ref, deny via `*p`), NOT a scheduler tie-break; (2) the def-side return-type wall has a **MACRO escape** (widen a *discarding* caller-macro's `extern void`→`s32`, byte-neutral); (3) the birthing-boost prologue-order re-tie.
- **(a) toolkit pass:** 3 standard **Opus agents** *APPLIED* §30 to 3 backlog "unsteerable" store-vs-load fns → `func_8014F2E0` (66) + `func_80150528` (53) **banked ×134** (`commit:0428`); `func_8014EA4C` 6→3 (partial — abs/frame-fragility residual, permuter/Fable5 tail). +2 idioms → **§30a**: IV-combine single-base-pointer (keep `combine_givs` at N IVs); inline loop-limit → preheader-hoist. **KEY FINDINGS:** standard Opus agents *applying* documented idioms crack leaves at ~80–130k tokens (3–5× cheaper than Fable5 *discovering*); the "unsteerable" backlog is **largely MIS-VERDICTED**; the whole-binary integration is **MECHANICAL** (all 3 giants = the same `void`→`s32` macro-widen) → **TOOLABLE**.
- Also this session: GLM parallel-batch campaign (`tools/glm_parallel.sh`, K concurrent OpenRouter drafters) over the 102 reach-134 hard band → **2 banks/102** (~$10 OpenRouter, def-side wall confirmed at scale; MAXTOK-tuning: reasoning models starve low, 96k overkill, ~48k+iters=2 sweet spot; iter-2 rescues ~29% of near-misses but not past the wall); free v3 grind **+89 fns** (fleet →64.7%) before I stopped it. `type-dependency map`: the 8 compile-blocked giants are **mechanically** unblocked (ghidra-types + externs), **NOT struct-walled**.
- **★ CAMPAIGN REFRAME (the current plan):** bottleneck moved from "can we match" (yes — toolkit + Opus agents) to **THROUGHPUT**. Path: **(1) BUILD the §30 integration recovery tool** (auto macro-widen → every leaf-MATCH auto-banks — the throughput unlock); **(2) Opus-agent WAVE** over the near-miss backlog (~95 "schedule"/80 "regalloc"/31 "hoist" verdicts, mostly mis-labeled) applying the toolkit; **(3) Fable5** for genuinely-novel residuals + the untouched big giants (`func_80144B9C` 770, `func_80132784` 400 HARD-DEFER, region `a`) + the 8 compile-blocked (mechanical unblock + draft). **v4-retrain on the new idioms = LOW ROI** (too few examples, class modest, giants exceed 7B — TOOL the idioms, don't train). **IN FLIGHT:** a 2nd Fable5 on giant `func_801571C4` (198 ins, close=24, region main — prologue-schedule/slot-swap residual). Tasks tracked in the harness list (R28). serve_local DOWN, GPU free. NB `.gitignore`+`docs/backlog.md` had pre-session uncommitted edits (unrelated); `db.*.gbf` = R23 noise.