diff --git a/phase-ends/CURRENT_PHASE.md b/phase-ends/CURRENT_PHASE.md index 2b6346925..b27cda368 100644 --- a/phase-ends/CURRENT_PHASE.md +++ b/phase-ends/CURRENT_PHASE.md @@ -34,9 +34,11 @@ D. Periodically: `export_pairs → format_finetune → train_lora → redeploy` ## ▶ RESUME HERE (fresh session) **State:** Phase 23 in progress (NOT a phase end). **v3 is the current model** — `bfm-match-7b-v3` (Qwen2.5-Coder-7B QLoRA on **corpus-v3**), adapter at `models/bfm-match-7b-v3` (v2 kept as fallback at `models/bfm-match-7b`). **LM Studio is EJECTED** — serve via **`tools/serve_local.py`** (Unsloth GPU, OpenAI endpoint), NOT LM Studio. The **8-hour autonomous run (2026-06-30)** built local serving + the prompt fix + corpus-v3 + v3 + a production batch → **fleet 63.82%** (+502 byte-identical, $0), 136/136 byte-clean, 27 commits this session (local — **Drew pushes**, R6). Pipeline validated end-to-end: a free local model banks the small/setter bulk, including the empty-leaf class v2 couldn't. Corpus `datasets/match_pairs/` + `.venv-train` gitignored. (Phase 22 close `PhaseEnd_Phase22.md` is committed `commit:0325`.) -**NEXT TASK — pick a path (both documented + set up by the run):** -**(A) More LLM harvesting with v3 — the fleet-% lever.** Run **`lora_grind --min-reach 2` with v3**: now that v3 banks the shared setters, target the **shared (reach≥2)** ones so each bank propagates **×134** instead of re-banking inline per binary (what capped this run's % at +0.15). Pair with the **dedup-collapse** of the per-binary inline setters → shared `engine_core.h` macros, and the **data flywheel** (the ~352 new banks grow the corpus → retrain v3.1). Cheapest, immediate, $0. -**(B) Train a 14B.** The measure-before-investing gate is now **SATISFIED** — v3-on-7B paid (banks the small/setter bulk), so a dense **Qwen2.5-Coder-14B** for the **>15-ins / struct-giant** band is the warranted next investment. `train_lora.py --base unsloth/Qwen2.5-Coder-14B-Instruct-bnb-4bit --rank 32`. Needs ~16 GB → a **cloud A100/H100** (the 3080 Ti is 12 GB; 4-bit + offload locally works but is slow). Do **corpus-v4** first (emit the struct-giant types the >40-ins fns need — the band v3 still compile-fails). The byte-gate makes a wrong 14B a throughput risk only. +**NEXT TASK — DECIDED (Drew, 2026-06-30): do A first, THEN B.** (Sequencing rationale: A is $0 + immediate + realizes the ×134 lever this run set up, and A's new banks enrich the corpus that B trains on — so A-then-B compounds.) + +**① START THE FRESH SESSION HERE — (A) more LLM harvesting with v3, the fleet-% lever.** Run **`lora_grind --min-reach 2` with v3**: now that v3 banks the shared setters, target the **shared (reach≥2)** ones so each bank propagates **×134** instead of re-banking inline per binary (what capped this run's % at +0.15). Pair with the **dedup-collapse** of the per-binary inline setters → shared `engine_core.h` macros, and run the **data flywheel** (the run's ~352 new banks grow the corpus → retrain a v3.1; raise `--max-nins` as the band lifts). Serve v3 via `serve_local.py` (command below). Cheapest, immediate, $0. Close A when its flywheel plateaus. + +**② THEN — (B) train a 14B for the giant band.** The measure-before-investing gate is **SATISFIED** (v3-on-7B paid), so a dense **Qwen2.5-Coder-14B** for the **>15-ins / struct-giant** band is the warranted investment. Do **corpus-v4 first** (emit the struct-giant types the >40-ins fns need — the band v3 still compile-fails), then `train_lora.py --base unsloth/Qwen2.5-Coder-14B-Instruct-bnb-4bit --rank 32`. Needs ~16 GB → a **cloud A100/H100** (the 3080 Ti is 12 GB; 4-bit + offload locally works but slow). The byte-gate makes a wrong 14B a throughput risk only. **Serve + run (when Drew says go):** ```