diff --git a/phase-ends/CURRENT_PHASE.md b/phase-ends/CURRENT_PHASE.md index 1caeb391b..475d64730 100644 --- a/phase-ends/CURRENT_PHASE.md +++ b/phase-ends/CURRENT_PHASE.md @@ -91,3 +91,5 @@ The **whole-binary byte-gate** (`gate_stage`/`harvest_verify`, G3/P9) is the sol - 2026-06-30 (**8-hour autonomous run, Drew away**): prompt-fix + local serving + corpus-v3 + v3 + big batch. **LM Studio ejected** → built **`tools/serve_local.py`** (Unsloth GPU serving as an OpenAI endpoint; the prebuilt llama-cpp-python CUDA wheel SIGILLs on this no-AVX-512 CPU, so the Unsloth/torch path is the reliable one). **PROMPT FIX** (`api_draft.LEAN_SYS` + `format_finetune.SYS`, synced): "translate EVERY instruction, never an empty body" — the v2 empty-leaf overfit, small-leaf band **0/3→2/3**, banked 3 on a fresh ov_SC01_001 batch. **CORPUS-V3**: `export_pairs` now mines the **1623 engine_core.h `DEFINE_func` macros** (the shared setters/return-const the model was blind to — 96.6% of v2 was overlay-unique) + `format_finetune` inlines `engine_types.h` structs → corpus **1312→2891**, trainable **2534+291** (2.5× v2). **v3 trained** (Qwen2.5-Coder-7B QLoRA, loss 1.275→0.085, ~2h), held-out gate-true eval **MATCH 23/40 (57.5%)** (generalizing; v2's mixed-set rate was lower). Grinder concurrent during training = 0 banks (permuter tail exhausted, Phase-22 reality). **Big batch** (v3, broad rotation, 25 binaries, propagate-every-3, **$0 LLM**): banked **~352 fns inline + 45 new shared groups** (1633→1678) → fleet **63.67% → 63.82%** (+502 byte-identical), **136/136 byte-clean**, 25 auto-commits. v3 repeatedly banked the **empty-leaf/setter class v2 couldn't** (func_8012E27C=return 1, 8012AD64/BF4C=setters). **Pipeline validated end-to-end:** a free local fine-tuned model harvests the small/setter bulk at $0, gated identically (G3/P9). Details: `docs/gen2-mips-matching-model.md` ("Corpus-v3 ... 8-hour autonomous run"). **NEXT (fresh session):** (1) **shared/reach≥2 targeting** (`lora_grind --min-reach 2` with v3 — so the setter banks propagate ×134 instead of re-banking inline per binary — the fleet-% lever); (2) **dedup-collapse** the per-binary inline setters; (3) **corpus-v4** = struct-giant types; raise `--max-nins` as the band lifts. Serve: `tools/serve_local.py --adapter models/bfm-match-7b-v3` (R21 / SETUP §Tooling inventory). - 2026-07-01 (**T10 — the phase-separated + parallel-gate harvester, BUILT + MEASURED**): plan-mode Tier-1, harvester-first (Drew picked it over vLLM-first — the exploration showed parallel-gate is the bigger single lever, no install, and unblocks the measurement). **Built** `tools/bulk_harvest.py` (bulk-draft GPU → `ProcessPoolExecutor` byte-gate over DISTINCT binaries → dedupe-once + one commit; round-robin fuel spread) on back-compat `gate_stage`/`harvest_verify` params (T10.1 — `lock_path`/`verified_out`/`failed_out`/`compute_fleet`, all default to serial; `backlog` redirected per-worker via its module global, no edit). Tooling commit `commit:0379`. **Validated** (count=12/workers=4): 9/12 banked, 136/136 byte-identical, R23-clean commit `commit:0380`, dedup 0 failed. **MEASURED** (count=80/workers=8 fresh SC03 ≤15-ins, commit `commit:0381`): **bank-rate 52/80 = 65%**; gate **0.4s/fn amortized** (8 workers) vs lora_grind's ~30s/fn serial ≈ **75×**; draft **13.5s/fn serial = 97% of wall-clock** → the parallel gate moved the bottleneck ENTIRELY to drafting. Fleet 64.19→64.20% (61 tiny reach-1 banks barely move the byte-weighted % — the ≤15 campaign's value is bank COUNT + completeness, freeing the heavy model for giants; honest P9). check-all 136/136, dedup-check 1687→1715/0. **DECISIONS SETTLED:** (1) campaign is GO (65% over ~4,087 untried uniques ⇒ ~2,650 free banks); (2) drafting is now the only lever; (3) **vLLM is JUSTIFIED** (~15h serial → ~1–3h batched) — deferred to Drew-go (T10.5). Details: `docs/gen2-mips-matching-model.md` "T10 RESULT". Serve still up via `serve_local.py` (v3). NEXT: Drew's call — vLLM install (T10.5) now, or run the ≤15 campaign on the serial path in background chunks, then PhaseEnd (T11). **Drew chose B (serial campaign now, defer vLLM).** **T10.6 LAUNCHED (2026-07-01):** `bulk_harvest --binary-glob 'ov_*' --count 150 --cycles 40 --workers 8` on `serve_local` v3 — the full ≤15-ins overlay saturation ($0), self-terminating on dry fuel (~15h max), **per-cycle commits (crash-safe/resumable via `.run/auto/lora_grind_tried.json`), STOP-able** (`touch .run/auto/STOP`). Monitor: `.run/auto/bulk_harvest_{heartbeat,stats}.json` + `.run/bulk_campaign.log`. Region bank-rate varies (SC03 fresh ~65%, SC01 dregs ~40%). `--cycles` campaign mode is committed (`commit:0384`); the `run_cycle` loop was validated (2-cycle run, per-cycle commits `commit:0385`/`commit:0386`). If a fresh session finds the campaign stopped: re-launch the same command (resumes from `tried`), or `touch .run/auto/STOP` to halt. After the campaign: `make check-all` (R22) + T11 PhaseEnd_Phase23. **Drew's new direction (2026-07-01), the NEXT step after the campaign:** test a HUGE-parameter cloud model via **OpenRouter** (GLM-4.6 / the current big GLM Drew called "glm5.2" — confirm the exact id on OpenRouter at run time) as a DRAFTER — `api_draft.py` already speaks the OpenAI endpoint (OpenRouter is OpenAI-compatible: set `API_BASE`=OpenRouter + key, `MODEL`=the GLM id). A/B vs local v3 on a sample (bank-rate + $/match), focused on the harder band the 7B plateaus on (>15-ins / struct). Byte-gate keeps it correctness-safe (throughput/cost lever only). **PARKED at Drew's direction:** the batching levers — DIY static batching (~3–4×, VRAM-capped on 12 GB) and vLLM (5–20×, needs a separate venv + tighter quant) — both on HOLD; revisit after the OpenRouter test. "We can move forward after that" (Drew). **T10.6 CAMPAIGN RESULT (2026-07-01):** ran **23 cycles → 1,297 fns banked ($0) + 52 propagated groups**, fleet 64.21→**64.6%** (+0.39%), tried 764→**4,214** (**~92% of the ~4,597 ≤15-ins unique fuel saturated**). **check-all 136/136 byte-identical**, dedup 0 failed. **Ended gracefully on a serve_local crash** (bitsandbytes `ops.cu:81` after ~14h continuous serving — the serial 4-bit server's long-run fragility, NOT a code bug; bulk_harvest's per-cycle endpoint-check caught the dead server, committed the partial cycle 23, and exited clean). Commits `commit:0387..commit:0410` + digests `commit:0411`. **Remaining ≤15 tail (resumable, low-value ~28% yield):** ~383 untried + ~132 skipped-marked-`tried` in the crashed cycle 23 (prune those from `.run/auto/lora_grind_tried.json` if resuming). To resume: restart `serve_local.py` then `bulk_harvest --binary-glob 'ov_*' --cycles N`. Wrapped up per Drew → **NEXT: T10.7 OpenRouter** (cloud, no local GPU). Serve_local is DOWN, GPU freed. **T10.7 RESULT (OpenRouter GLM5.2 A/B, 2026-07-01):** on 18 hard-band fns (16-22 ins, P16-walled `ov_SC01_077`), **GLM5.2 10/18 match_one vs v3 1/18** (~10× codegen edge) but **whole-binary banks only 3/18** (2 direct + `func_801577C8` via `fix_arity_callers`) — the **DEF-side loose-typing wall (Phase 16/20) HOLDS** and caps ANY drafter (Fable5 §3c re-test confirmed). 6 hardest fns beyond GLM too (0/6 at MAXTOK=16000). Cost ~$1.02 of $25. Banked 3 GLM fns (`commit:0416`); `api_draft` gained `MAXTOK` env + cost capture (`commit:0414`). Details: `docs/gen2-mips-matching-model.md` "T10.7 RESULT". **DECISION PENDING (Drew):** how to use GLM — (1) direct drafter on def-conflict-free hard fns ($0.03/fn, ~17% wall-capped), (2) corpus-v4 flywheel/retrain (also wall-capped, ~2h GPU, uncertain), or (3) point GLM at the def-side-wall RECONCILIATION (reasoning-shaped unlock, new harness). T10.7c retrain NOT started (wall caps its value → checkpoint for Drew). **T10.7 Option-3 probe DONE (`tools/glm_reconcile.py`, +$0.23):** aimed GLM's REASONING at the def-side wall — expert-grade reasoning (store-width/`sh`-vs-`sw` awareness, K&R promotion, independently derives the §17a-1 cast idiom), captured to `.run/glm_reason/` for idiom mining (R16 → cookbook **§29**). Banked **1/7** (`func_80175184`); mechanical `--any-proto` **0/7**. **VERDICT: the def-side wall is INTRINSIC** (frontier reasoning model + full toolkit cracks ~1/7; residual = narrow-param K&R limit + byte-addressing). **GLM total banks 4/18** vs v3 ~1/18. Tooling: `glm_reconcile.py` (NEW), `api_draft` REASON-capture, `fix_arity_callers --any-proto` (`commit:0418`); cookbook §29 + doc T10.7 RESULT. **GLM's role settled:** (a) $0.03–0.08/fn DIRECT drafter for the def-conflict-FREE hard band (~22%, v3 can't reach), (b) idiom TEACHER — NOT a wall-breaker. Per Fable5 §4.3 the real lever past the wall is the public flip (community), not a bigger model. Total T10.7 spend ~$1.25/$25. **NEXT (Drew's call):** (i) a bounded GLM DIRECT hard-band campaign (bank the conflict-free ~22% at scale, money-open), (ii) the public-flip pivot (Fable5 §6.3), or (iii) close Phase 23 → PhaseEnd (T11). - 2026-07-01 (**evening — T10.6 ≤15 campaign + T10.7/8/9 the OpenRouter/GLM5.2 frontier-model exploration**): CONTEXT CHECKPOINT (85% ctx, wind-down). **T10.6 DONE:** `bulk_harvest.py` ≤15 saturation campaign ran 23 cycles → **1,297 banks, fleet 64.6%, ~92% of the ≤15 band saturated**, 136/136, crash-graceful exit (serve_local bnb `ops.cu:81` after 14h — not our bug). **T10.7 DONE (GLM5.2 A/B, `z-ai/glm-5.2` via OpenRouter):** on 18 hard-band fns (16-22 ins), **GLM 10/18 match_one vs v3 1/18** (~10× CODEGEN edge) but **whole-binary banks 4/18** — the **DEF-side loose-typing wall (Phase 16/20) is INTRINSIC** (caps ANY drafter; it's overlay-decl-vs-true-sig, a C/ABI limit). Option-3 probe (`glm_reconcile.py`): GLM reasons the wall EXPERTLY (store-width `sh`/`sw`, K&R promotion, independently derives the §17a-1 cast idiom) but banks only 1/7 — reconciliation idioms → cookbook **§29**. **T10.8 DONE (`idiom_hunt.py`, $0.51):** batch-clustered the FAILED backlog by class → GLM re-derives our OWN idioms (pins §17, array-of-struct §18) + confirms walls, **0 new bankable idioms** (the failed residual is the worst discovery material — Drew's catch). **T10.9 IN FLIGHT (bg task `bmv1wxzdz`, `.run/glm_fresh.log`):** Drew's CORRECT reframe — the idiom engine is DEEP hand-solving of FRESH never-tried hard fns (like Opus Max), then mining the byte-correct cracks for NEW idioms (idiom lives in the correct BODY, not the bank). GLM deep-solving **15 fresh 25-118-ins fns** (`ov_SC01_077`/`SC03_001`/`SC02_005`), LEAN, MAXTOK=32k, iters=3, **reasoning captured** (`.run/glm_fresh/*.reasoning.txt`, `REASON=1`). **TOOLING BUILT this session:** `tools/bulk_harvest.py` (phase-separated parallel-gate harvester + `--cycles`), `tools/glm_reconcile.py` (Option-3 wall reconciler), `tools/idiom_hunt.py` (idiom researcher, `--budget`-capped), `api_draft` (`MAXTOK` env + `usage.cost` + `REASON=1` reasoning capture), `fix_arity_callers --any-proto`, `gate_stage.run_gate` parallel params (`lock_path`/`verified_out`/`failed_out`/`compute_fleet`), `harvest_verify --verified-out/--failed-out`. **Commits:** `commit:0379` (T10 tooling)…`commit:0410`/`commit:0412` (≤15 campaign)…`commit:0414`(MAXTOK)…`commit:0418`(glm_reconcile+§29)…`commit:0420`(idiom_hunt+T10.8). **OpenRouter setup:** key in `.env` as `open_router_key` (gitignored, `sk-or-`); GLM5.2 is a REASONING model (needs MAXTOK≥8k or it returns empty content — burns budget on reasoning tokens); ~**$23.6 left of $25** (spent ~$1.90); the key shows `"limit": 20` (possibly a $20 spend cap — raise in the OpenRouter dashboard before a bigger run); no hard concurrency cap (~6 providers serve glm-5.2). **STRATEGIC STATE:** the LLM matching tier is now THOROUGHLY explored — local v3 saturated ≤15 ($0); frontier GLM confirmed the hard-band def-side wall is INTRINSIC (triple-confirmed) + banks a modest ~22% of the conflict-free hard band at ~$0.05/fn; the new-idiom well is DRY on the failed residual (T10.8), with T10.9 testing whether FRESH hand-solving still yields NEW idioms (Drew's hypothesis). **DREW CHOSE: (A) a bounded GLM fresh-hard-band campaign** (which T10.9 pilots). Per the Fable5 strategy review (`docs/fable5-strategy-review-2026-07.md`), the real lever PAST the wall is the **public flip / community labor** (§4.3), not a bigger model. **RESUME (fresh session):** (1) read `.run/glm_fresh.log` + gate the correct bodies (bank the def-conflict-free ones; use `glm_reconcile`/`inject_capped_externs` recovery), (2) MINE `.run/glm_fresh/*.reasoning.txt` for NOVEL idioms (vs cookbook §17-§29) → distill real ones to cookbook, (3) decide the fork: scale the GLM fresh-hard campaign (money-open but modest), pivot to the public flip (Fable5 §6.3, the strategic lever), or PhaseEnd_Phase23 (the LLM tier is well-characterized — a clean Tier-1 close). **T11 (PhaseEnd) remains.** serve_local is DOWN (GPU free). Cost discipline: Drew is cost-conscious — calibrate cheap, `--budget`-cap, no surprise overnight bills. +- 2026-07-02 (**T10.9 RESULT + STOPPING POINT — the fresh-hand-solve idiom test**): Drew's fair test of "does GLM hand-solving FRESH hard funcs teach NEW idioms?" (the mechanism that gave us pins/array-of-struct via Opus-Max). GLM deep-solved **15 fresh 25-118-ins fns** (LEAN, MAXTOK=32k, iters=3, reasoning captured `.run/glm_fresh/*.reasoning.txt`): **4/15 match_one, 2 whole-binary banks** (`func_8017DE28`, `func_8015EEE0`; @`commit:0422`, check-all 136/136), **$2.00**. **IDIOM VERDICT (the point):** GLM's 4 correct bodies reason entirely about **KNOWN gcc mechanics** — delay-slot scheduling, callee-saved `$s0` preservation, reload-after-call aliasing, switch jump tables — all already in cookbook §10/§17/jump-table. **NO new idiom emerged.** GLM is a strong reasoner *applying* mapped idioms, not finding unmapped quirks (expected — 22 phases of Opus-Max mining charted the common gcc-2.7.2 quirks). **So the new-idiom well is DRY, now confirmed from BOTH angles: failed-residual (T10.8) AND fresh-hand-solve (T10.9).** 8/15 were reasoning-truncated empties (giants overrun 32k); near-misses `func_8018119C` (near-4) / `func_80187DAC` (near-5) are grinder/reconcile fodder. **ROBUST STRATEGIC CONCLUSION (multi-angle):** the LLM matching tier is exhausted of cheap+idiom leverage — (1) local v3 saturated ≤15 ($0, ~1300 banks), (2) frontier GLM = a modest ~22% hard-band drafter at ~$0.05-0.13/fn (def-side wall caps banking; wall is INTRINSIC, triple-confirmed), (3) no new idioms from either failed-residual or fresh cracks. Total OpenRouter spend ~$4/$25. **Per the Fable5 strategy review (`docs/fable5-strategy-review-2026-07.md` §4.3/§6.3), the real lever past the wall is the PUBLIC FLIP (community hand-matching), not a bigger model.** + **★ CLEAN STOPPING POINT (fresh session, Tier-1 Max, plan mode):** the LLM-tier exploration is comprehensively concluded. Options for the fresh session, in recommended order: **(A) PhaseEnd_Phase23** — write it (the LLM tier is fully characterized: local grinder + frontier test + idiom hunt, all byte-proven; a clean Tier-1 close), then **(B) plan the PUBLIC FLIP as the next phase** (Fable5 §6.3: R20 clean mirror, rom→decoder, CI, a naming pass on major systems, contributor on-ramp from `docs/backlog.md` — the strategic lever). A bounded GLM fresh-hard-band campaign remains available (money-open, modest fleet%) but is NOT the high-leverage move. Fleet stands at **~64.2%** (function-count) / ~30% byte-weighted, 136/136 byte-identical, 0 NON_MATCHING. All work committed (`commit:0379`…`commit:0422`); Drew pushes (R6). serve_local DOWN, GPU free. `db.*.gbf` churn is R23 restart-noise (do NOT stage).