From 671177aece85eca364d5f7bca5f1e04f59e8fabb Mon Sep 17 00:00:00 2001 From: Drew T <50529377+Druthulu@users.noreply.github.com> Date: Mon, 29 Jun 2026 16:26:29 -0600 Subject: [PATCH] =?UTF-8?q?docs(phase-22):=20capture=20stock-local-model?= =?UTF-8?q?=20floor=20=E2=80=94=200=20reliable=20banks,=20motivates=20the?= =?UTF-8?q?=20LoRA=20specialist?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Qwen3.6-35B-A3B stock drafter: fair harness fixed compiles but model stuck at fixed near-misses (can't refine from diff); full cookbook worse+2.3x slower than curated (dilution). Gets structure, misses gcc-2.7.2 precision — exactly what src-pair LoRA targets. Forward: fine-tune or permuter-seed. --- docs/gen2-mips-matching-model.md | 26 ++++++++++++++++++++++++++ 1 file changed, 26 insertions(+) diff --git a/docs/gen2-mips-matching-model.md b/docs/gen2-mips-matching-model.md index 7377d5ff6..e78efa16c 100644 --- a/docs/gen2-mips-matching-model.md +++ b/docs/gen2-mips-matching-model.md @@ -47,6 +47,32 @@ A stock local model test (in progress) tells us the floor. The fine-tune is wort the specialist's *gate-true* match rate clears the stock model by enough to matter — which the free gate eval settles directly. Don't train blind; train against a target number. +### Stock-model floor — measured 2026-06-29 (the result that motivates this) + +**Qwen3.6-35B-A3B as a stock drafter fails on the byte-match step**, and this is precisely the gap a +fine-tune fills. On the 20 reach1 functions via `tools/api_draft.py`: +- **Blind harness:** 0/13 (run killed early); mostly compile-fails + near-misses. +- **Fair harness** (inline common.h + live cookbook + corpus examples): *fixed* compilation, but the + model **stuck at fixed near-misses** — identical closeness across all 4 diff-feedback iterations + (e.g. func_8013373C = near-4 in blind, curated, AND full cookbook). It cannot act on the + instruction-level diff to refine. +- **Full vs curated cookbook:** full (whole 192k-char file, ~58k-tok prompt) was **worse and 2.3× + slower** than the curated matching-only subset — attention dilution, gate-confirmed. More context + is not the lever. + +Diagnosis: the model gets the *structure* right (correct control flow, field semantics) but misses +gcc-2.7.2 **precision** — element-vs-byte offset scaling, `lh`/`lhu` signedness, an extra `move`, +frame size — and can't self-correct from the diff. That precision is exactly what `src/`-pair LoRA +bakes into weights. **The stock floor is ~0 reliable banks; that is the number to beat.** (Contrast: +Haiku, a frontier *small* model, reliably matched the ≤52-ins bulk — so the gap is capability, not task.) + +Two forward levers besides fine-tuning: +- **Permuter-seed role:** the model's structurally-correct near-misses are good *permuter seeds* — let + the 32-thread permuter brute-force the regalloc/schedule precision the model can't. Plays to its + strength; cheap to test. +- **Cloud cheap tier (Haiku/GLM-5.2)** stays the working low-cost drafter today (the local-free tier + needs the fine-tune or the seed role to be useful). + ## Open questions / notes - **Corpus quality > size.** ~1,700 verified pairs is plenty for LoRA; dedup near-identical reach