Files
BFM-decomp/docs/gen2-mips-matching-model.md
T
Drew T 5ef2d4d2ab docs(phase-25): T4 CLOSEOUT — v4 retrain FAILED (negative), keep v3; T5 handoff
- A/B gate-true: v4 <= v3 (marginally worse on medium, tie easy/hard) -> discard v4, keep v3 (frozen ceiling)
- 'corpus quality > size' confirmed; 7B is capacity-bound (0/5 even on trained fns, not truncation)
- decision-log (R31): the local-7B tier is off the endgame critical path; engine = frontier-crack -> deterministic-propagate -> byte-gate + permuter
- gen2-mips-matching-model: the A/B + the maxlen-2048 truncation flaw (drop-over-length OR grad-checkpointing, NOT CPU offload) for any future retrain
- CURRENT_PHASE: T4 done/failed; NEXT = T5 (Ultracode measure-wave); phase OPEN, no PhaseEnd
2026-07-08 18:36:50 -06:00

379 lines
29 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Gen2 idea — a BFM/PSX-gcc-2.7.2 *matching* specialist model
**Status:** design / not started. Captured 2026-06-29. The high-ceiling idea from the cheap-tier thread.
## The gap, and why it's ours to fill
There is **no matching-decompilation LLM** — for MIPS or anything else. The one RE-specific model
family, **LLM4Decompile**, targets **x86** and optimizes for *readable, recompilable, semantically-
equivalent* C (passes the function's tests), **not byte-identical** machine code under a *specific*
ancient compiler. That's the wrong objective for us. The reason no matching model exists is **data**:
the only place (target-asm ↔ byte-exact gcc-2.7.2 C) pairs exist is inside decomp projects' git
histories. The community therefore runs *general* strong coders + deterministic tools (m2c, permuter,
the byte-gate). The real domain "specialist" today isn't an LLM — it's **m2c** (rule-based MIPS→C).
**But we have the data off-the-shelf models lack:** ~1,300 **gate-verified** (asm ↔ byte-matching
gcc-2.7.2 C) pairs in this repo's banked history (plain `src/` defs; more recoverable from deduped
`DEFINE_func_*` macros) — real, clean, byte-exact, specific to our exact
toolchain and this game's idioms. That is precisely the corpus LLM4Decompile had to *synthesize* for
x86, except ours is correct and on-target. So we can *build* the matching specialist that doesn't
exist — narrow domain + high-quality verified data is the ideal setup for a small local model to
beat a frontier generalist **on this one task** (it won't generalize, and doesn't need to).
This is the *right-data* version of the "Opus-distilled Qwen" models: instead of generic reasoning
traces, bake our cookbook's gcc-2.7.2 quirks (regalloc order §17, %lo array-of-struct fold §18,
schedule/loop forms) into the **weights**, not just the prompt.
## Recipe
1. **Corpus** — `tools/export_pairs.py` mines every banked function into `{fn, region, asm, c}`
JSONL: C from `src/`, the target `.s` recovered from git history (the exact splat format
`api_draft.py` feeds at inference → train/inference format match). Split ~90/10 train/test by fn.
2. **Base model** — start from a strong local coder: **Qwen2.5-Coder-32B** or **Qwen3.6-35B-A3B**
(MoE, cheap to serve). Code-pretrained base + our SFT on top.
3. **Fine-tune** — **LoRA SFT** (not full FT): prompt = the drafter prompt (target asm + ghidra-C +
toolkit, same as `api_draft`/`worker_wave`), completion = the banked C. Rank 16–64, a few epochs.
Trainable on a rented A100/H100 in hours, or slowly/local for a small rank. Keep the prompt format
byte-identical to inference.
4. **Eval — free and rigorous** — run the adapter as an `api_draft` arm over a *held-out* frozen
target set and score with `tools/ab_score.py` (the byte-gate). Compare to the stock base model and
to Haiku. The gate is the arbiter; no eval guesswork. Iterate rank/epochs/data on that number.
5. **Deploy** — the adapter becomes a cheap-tier `local` arm (LM Studio / llama.cpp serve base+LoRA);
`api_draft.py` already targets it. Slots straight into the cost-escalation ladder as Tier-0/1.
## Why measure before training
A stock local model test (in progress) tells us the floor. The fine-tune is worth the effort **iff**
the specialist's *gate-true* match rate clears the stock model by enough to matter — which the free
gate eval settles directly. Don't train blind; train against a target number.
### Stock-model floor — measured 2026-06-29 (the result that motivates this)
**Qwen3.6-35B-A3B as a stock drafter fails on the byte-match step**, and this is precisely the gap a
fine-tune fills. On the 20 reach1 functions via `tools/api_draft.py`:
- **Blind harness:** 0/13 (run killed early); mostly compile-fails + near-misses.
- **Fair harness** (inline common.h + live cookbook + corpus examples): *fixed* compilation, but the
model **stuck at fixed near-misses** — identical closeness across all 4 diff-feedback iterations
(e.g. func_8013373C = near-4 in blind, curated, AND full cookbook). It cannot act on the
instruction-level diff to refine.
- **Full vs curated cookbook:** full (whole 192k-char file, ~58k-tok prompt) was **worse and 2.3×
slower** than the curated matching-only subset — attention dilution, gate-confirmed. More context
is not the lever.
Diagnosis: the model gets the *structure* right (correct control flow, field semantics) but misses
gcc-2.7.2 **precision** — element-vs-byte offset scaling, `lh`/`lhu` signedness, an extra `move`,
frame size — and can't self-correct from the diff. That precision is exactly what `src/`-pair LoRA
bakes into weights. **The stock floor is ~0 reliable banks; that is the number to beat.** (Contrast:
Haiku, a frontier *small* model, reliably matched the ≤52-ins bulk — so the gap is capability, not task.)
### Pilot RESULT — measured 2026-06-29 (NEGATIVE on meaningful functions)
7B QLoRA (Qwen2.5-Coder-7B, 3 epochs) on 638 compile-filtered pairs, evaluated on 75 held-out banked
fns (`tools/eval_lora.py`, gate-true):
- **≤5-ins trivial: 39/41 MATCH (95%)** — memorized the leaf-function pattern.
- **≥6 ins: 0/34. Meaningful (>15 ins): 0/24** — same as stock-local.
- Near-misses are **FAR, not close**: `near N ≈ nins` (all instructions mismatch → structurally-wrong
output), only 1/33 non-trivial within 5 of a match. So it's NOT "one epoch away."
Root cause = the corpus: the compile-filter (needed because bare `src/` defs don't compile standalone)
**threw out the 536 harder functions** (globals/structs), starving the model of non-trivial signal;
638 examples with 297 trivial → a 7B overfit the easy pattern. **The fix that filter requires is the
same one the corpus needs: self-contained completions WITH externs (corpus-v2), which recovers the hard
functions AND makes them trainable.** Until that's done, the fine-tune verdict is unproven, not refuted —
but the far near-misses suggest data quality/coverage (and likely a bigger base) are the real levers,
not epochs. Cheap-cloud (Haiku/GLM) remains the working tier meanwhile.
### Corpus-v2 RESULT — measured 2026-06-29 (POSITIVE: data was the bottleneck)
Fix: `export_pairs` now captures the `extern <type> D_xxx;` block the src declares immediately above
each def (correct byte-verified types) → self-contained completions, standalone-compile 52%→92%,
training set 638→1111 with non-trivial examples 257→813. Retrained the SAME 7B; held-out gate-true eval:
| band | v1 | v2 |
|---|---|---|
| 6–15 ins | 0% | **85% (23/27)** |
| 16–40 ins | 0% | 13% (3/22) |
| >40 ins | 0% | 0% (10 compile-fail = need struct types) |
| non-trivial (>5) | 0/34 | **26/73** |
| meaningful (>15) | 0/24 | 3/46 |
**Conclusion: corpus quality was the bottleneck, not the model or the task.** A free local 7B now
byte-matches trivial + small-medium (≤15 ins) functions at 85–93% — a real Tier-0 for the bulk,
rivaling Haiku on that band at $0. Limits: ≥16 ins falls off (7B capacity), giants compile-fail (the
extern-capture covers globals but not struct *types* → corpus-v3 = also emit the struct defs each fn
needs). Decision gate (staggered plan) = GO: scale to a cloud-trained dense 14–32B to extend the band
upward. Caveat: this eval is held-out BANKED (objdump format); production on OPEN stubs still needs the
.s-format alignment (spimdisasm). Tooling: format_finetune→train_lora→eval_lora; serve via LM Studio
(GPU) — Unsloth's bundled llama.cpp is CPU-only.
Two forward levers besides fine-tuning:
- **Permuter-seed role:** the model's structurally-correct near-misses are good *permuter seeds* — let
the 32-thread permuter brute-force the regalloc/schedule precision the model can't. Plays to its
strength; cheap to test.
- **Cloud cheap tier (Haiku/GLM-5.2)** stays the working low-cost drafter today (the local-free tier
needs the fine-tune or the seed role to be useful).
### T7 RESULT — the broad-rotation gate debugged 2026-06-30 (0/222 was TWO harness bugs, not the model)
The 500-fn calibration run (`lora_grind`) banked **0/222** across ov_SC01_000→ov_SC02_005 while the
model banked ~18% on ov_SC01_077. Root-caused (R14 — by reading the code + the run's own backlog,
which resolved a direct contradiction between two scout agents) to **two independent bugs in
`lora_grind`'s use of `gate_stage.run_gate`**, NOT model quality:
- **Bug A — good_sha format:** `lora_grind.good_sha()` returned the whole sha1sum line `"<sha> <name>"`;
`harvest_verify` compares it against a bare `sha1()` → never equal → **0 banks for EVERY binary
including 077**. So the "077 0/12" in that run was a *bug artifact*, not an exhausted tail (a claim I
nearly enshrined before reading the `.sha` format — the R14 payoff).
- **Bug B — path mis-resolution:** the gate call passed only `binary`+`good_sha`, leaving
`src`/`asm`/`out` at the hardcoded ov_SC01_077 defaults → non-077 drafts dropped at the 077 stub-filter
→ 0 banks, **silently**, and the backlog near-miss classes were contaminated by the mis-resolved asm
(so that run's "199 compile-fail" breakdown was untrustworthy for non-077 fns).
Fix (`tools/gate_stage.py`): `run_gate` now resolves `src`/`asm`/`out`/`good_sha` from `binary` when
unset (binary-agnostic — the Phase-9 "no silent default an overlay inherits" discipline; good_sha
normalized to the bare hash) + a **loud negative-control guard** (warns when 0 drafts are stubs in the
binary's own sources — the silent-0 can never recur). `lora_grind.good_sha` also fixed at source.
Byte-neutral (check-all 136/136); zero `lora_grind` logic change beyond the one-line good_sha fix.
**The first trustworthy on-OPEN-stubs signal (ov_SC01_000 spot-run, 15 smallest ≤15-ins stubs):**
- **Gate-banked 7/15 (47%) byte-identical** (check-all 136/136, auto-committed). Proxy `match_one`:
6 leaf-exact + 1 recovered via the TU-plumbing pipeline. Backlog (now correctly classified): 6
near-misses — **four at closeness=1** (prime grinder/permuter fuel) — + 2 standalone-compile-fails
(struct types → corpus-v3). The 0/222 was 100% the bugs; the model is a strong Tier-0 on the small
open-stub tail **fleet-wide**, the production number the corpus-v2 caveat (above) flagged as unmeasured.
**ROI finding (the sizing input):** of the 7 banks, **6 are reach-1 (overlay-UNIQUE, ×1)** and 1 is
reach-2. The broad rotation's small non-077 stubs are predominantly overlay-unique → **high bank-RATE,
low fleet-% ROI** (each ×1; the fleet % barely moved, +8 fns). The fleet-% levers are therefore
**reach≥2 targeting** (the ×134 multiplier — a `lora_grind`/`wave_targets` `--min-reach` filter, T9) and
the canonical-site (077) harvest, plus **corpus-v3** for the struct compile-fails (T8) — NOT a blind
broad rotation. A broad ≤15-ins run remains worthwhile for per-overlay completeness, corpus growth
(retrain fuel), and seeding the permuter grinder with the close=1 near-misses.
### T9 RESULT — reach≥2 targeting built; the model alone is weakest on shared code (2026-06-30)
Added `--min-reach N` to `lora_grind` (a lazy sig-based reach oracle == `dedup_propagate`'s, validated
0-mismatch over 60 stubs + the `func_8017CE24`=2 ground truth) so the mass-run can prefer SHARED
functions — one bank → ×reach. Bounded reach≥2 run, ov_SC01_000 batch (15 reach≥2 ≤15-ins stubs):
**0/15 banked** — vs the reach-1-heavy spot-run's 7/15. The model is **weakest exactly on reach≥2
(shared) code**, for two compounding reasons: (1) the corpus skipped the shared `DEFINE_func` macro
bodies (`export_pairs` reads only `src/` defs → 96.6% of the corpus is overlay-unique), and (2) the
shared engine functions are the harder regalloc/schedule residuals that survived the whole Phase-21
apparatus. So **reach≥2 model-only is NOT a fleet-% lever by itself** (a real, measured negative).
BUT the reach≥2 drafts are high-value FUEL: of the 15, **5 are close≤3 reach-134 near-misses** — three
at **close=1** (`func_8012E27C/BF4C/AD64`) + two `match_one`-MATCH-but-gate-rejected — each worth ×134
if closed. So the real fleet lever is **reach≥2 draft → permuter-grinder close (×134)** (the
CURRENT_PHASE "synergy"), and/or **corpus-v3-with-shared-bodies** to lift the model's direct banking.
Realizing the grinder path needs the **same per-binary fix T7 applied to `lora_grind`**: `grinder.py`
calls `run_gate` with no `binary` (→ defaults to ov_SC01_077) AND the backlog record stores no `binary`
field — so the permuter can't gate a *non-077* near-miss today. That two-part fix (grinder per-binary
resolution + a backlog `binary` field) is the next concrete step to turn the reach-134 close=1 fuel into
×134 banks. The reach oracle + `--min-reach` are reusable for that and for a corpus-v3 retrain.
### Grinder per-binary fix — built (5 layers); the reach≥2 close=1 fuel is semantic-misses + propagation-capped (2026-06-30)
To grind the reach≥2 close=1 fuel via the permuter, the grinder needed the same binary-agnostic
treatment T7 gave `lora_grind` — and it ran **five layers deep** (the whole grinder/backlog pipeline
was ov_SC01_077-hardcoded): (1) `gate_stage` records the source `binary`; (2) `backlog.FIELDS` keeps it;
(3) `backlog.load_best`/`_open_stubs` is **fleet-aware** (a 077-matched-but-stuck-local fn now surfaces
via its overlay record instead of being dropped as "matched"); (4) `p16_permute.setup` takes the target
binary's asm-subdir; (5) `grinder` resolves per-binary asm + gates **grouped by binary** + allows unknown
`nins`. Validated end-to-end: the 3 fresh reach-134 close=1 ov_SC01_000 fns now surface, resolve to
ov_SC01_000's asm, and gate via ov_SC01_000. Backward-compatible (legacy records → 077).
**Two byte-evidenced findings redirected the fuel strategy:**
- **The reach≥2 close=1 fuel is largely MODEL semantic-misses, not permuter fuel.** Diagnosed by the
byte: `func_8012E27C`'s target is literally `return 1` (2 ins), but the 7B drafted `void f(void){}`
(the corpus's overfit empty-leaf pattern); `func_8012BF4C`/`func_8012AD64` are trivial `sw`/`sh` setters
also drafted empty. The permuter (regalloc/schedule only) can't add a missing return/store — but a
**corrected draft** does: banked all 3 byte-identical via the fixed per-binary gate (the concrete
non-077-banking proof, fleet +3). So this fuel's lever is **a better draft (corpus-v3 leaf variety),
not the permuter.**
- **×reach is propagation-capped for the stuck-local class.** These 3 are already inline-matched in
`ov_SC01_077_a.c` (matched in 077, never propagated — the §19/20 cap), so `dedup_propagate --auto-from`
reports "nothing to propagate" (it can't auto-collapse an already-inline-matched fn into a shared
macro). They banked **×1** (ov_SC01_000 only). Realizing ×reach needs the **dedup-collapse** of the
inline copies into one `engine_core.h` macro — the existing Phase-19/20 lever.
Net: the grinder/gate pipeline is now **fully binary-agnostic** (capability unlocked + validated), but the
reach-134 ×134 payoff routes through **corpus-v3** (better leaf drafts for the semantic-miss fuel) + the
**dedup-collapse** (for stuck-local inline matches) — NOT the permuter, which the bytes show isn't the
closer for this fuel.
### Corpus-v3 + local GPU serving + the prompt fix — 2026-06-30 (8-hour autonomous run)
LM Studio was ejected (Drew freed the GPU), so serving moved in-repo: **`tools/serve_local.py`**
serves base+LoRA via Unsloth (`.venv-train`, cu128) as an OpenAI endpoint that `api_draft`/`lora_grind`
hit unchanged. (The prebuilt `llama-cpp-python` CUDA wheels SIGILL on this no-AVX-512 CPU; the
Unsloth/torch path is reliable and needs no build — `serve_local.py` loads in ~6s.)
**The prompt fix (cheap, no retrain).** The v2 corpus overfit an empty `void f(void){}` leaf pattern —
it drafted trivial `return 1` / setter functions as *empty*. One clause added to the LEAN prompt
("translate EVERY instruction, never an empty body; a `jr`+`addiu` delay slot is `return N`, a `sw/sh`
is a store") took the small-leaf band **0/3 → 2/3** in a prompt test, and banked 3 on a fresh
ov_SC01_001 batch end-to-end. Baked into `api_draft.LEAN_SYS` + `format_finetune.SYS` (kept in sync).
**Corpus-v3 (the root fix).** `export_pairs` now also mines the **1623 `DEFINE_func` macro bodies** in
`engine_core.h` — the shared setters/return-const/dispatchers `extract_defs` never saw (96.6% of v2 was
overlay-unique inline defs → the source of the empty-leaf overfit). `format_finetune` inlines
`engine_types.h` structs in the compile-filter so struct-using bodies are KEPT, not dropped. Corpus
**1312 → 2891** (1312 inline + 1579 macros); trainable **2534 train + 291 test** (2.5× v2's 1111),
97.8% compile.
**v3 trained + eval'd.** Qwen2.5-Coder-7B QLoRA, 3 epochs, loss **1.275 → 0.085** (~2h on the 3080 Ti
at batch 1 / maxlen 2048, VRAM-tight). Held-out gate-true eval (unseen functions): **MATCH 23/40
(57.5%)** — strong and generalizing, well above v2's mixed-set rate and the stock-local ~0 floor. So v3
is the better model (more+better data + the baked-in prompt) and drove the production batch.
**Production harvest (v3, ~3h, broad rotation over ~25 binaries, propagate-every-3, $0 LLM cost):**
the v3 batch banked **~352 functions** (inline, byte-gated) and the propagate sweeps lifted **+45 new
shared groups** (1633→1678) ×reach, taking the fleet **63.67% → 63.82%** (**+502 byte-identical
functions**), **136/136 byte-clean** throughout, across **25 auto-committed batches**. v3 repeatedly
banked the previously-impossible **empty-leaf/setter class** (`func_8012E27C` = `return 1`,
`func_8012AD64`/`8012BF4C` = `sw`/`sh` setters — all drafted *empty* by v2). The modest headline %
is the expected T9 reality (small open stubs are mostly reach-1 ×1; and the shared setters hit the
§19/20 propagation cap, so they re-bank inline per binary rather than collapse once ×134). **The
validated result is the pipeline itself:** a free, local, fine-tuned model now harvests the
small/setter bulk for $0, including the exact class the whole prior apparatus could not draft —
gated identically (G3/P9), with the data flywheel (banks → corpus → retrain) and propagation
(match-once → ×reach) intact. Next levers (unchanged, now higher-yield with v3): **shared/reach≥2
targeting** (so the setter banks propagate ×134 instead of re-banking inline) + the **dedup-collapse**
of the per-binary inline setters; corpus-v4 = struct-giant types; raise `--max-nins` as the band lifts.
### T10 RESULT — bulk_harvest (phase-separated + parallel-gate) built + measured — 2026-07-01
The interleaved `lora_grind` was ~48s/fn (≈18s serial draft + ≈30s recovery/gate) and idled the GPU
during every gate. **`tools/bulk_harvest.py`** phase-separates: **(A)** bulk-draft all K fresh stubs
(GPU, serial via the existing `serve_local`), **(B)** `ProcessPoolExecutor` byte-gate over DISTINCT
binaries (CPU — `build/<bin>/**` is isolated; `gate_stage.run_gate(propagate=False, commit=False,
per-binary lock, per-worker scratch, compute_fleet=False)`), **(C)** dedupe-once + one commit. Round-
robin fuel spread so the batch fans across many overlays. The whole-binary byte-gate stays the sole
arbiter (G3/P9). Back-compat enablers (defaults unchanged, so lora_grind/grinder/orchestrator are
byte-unaffected): `gate_stage.run_gate` gained `lock_path` / `verified_out` / `failed_out` /
`compute_fleet`; `harvest_verify` gained `--verified-out`/`--failed-out`; `backlog` is redirected
per-worker via its module-global path (no edit).
**Measured (v3, fresh SC03 ≤15-ins, serial Unsloth serving):**
- **Bank-rate 52/80 = 65%** (broad, 50 overlays) + 9/12 = 75% (the tiny validation band) — the REAL
whole-binary rate (not the `match_one` proxy, which over-counts ~2×).
- **Gate 0.4s/fn amortized** (30s total / 80 fns, 8 workers) — vs ~30s/fn serial in lora_grind ≈ **75×**.
- **Draft 13.5s/fn serial** — now **97% of wall-clock**; end-to-end ~14s/fn (was ~48s) ≈ 3.4×.
- 136/136 byte-identical (R22), dedup-check 0 failed, fleet 64.19→64.20% (61 tiny reach-1 banks barely
move the byte-weighted %; the ≤15 campaign's value is bank COUNT + completeness — freeing the heavy
model for the giant/struct band — NOT a fleet-% jump; honest, P9).
**Decisions this settles:**
1. **The campaign is GO** — 65% over the ~4,087 untried unique ≤15-ins fns ⇒ ~2,650 free byte-matches
(a minority reach≥2 → propagate ×reach; most reach-1 ×1).
2. **The bottleneck is now DRAFTING, not gating** — the parallel gate made the gate ~free (0.4s/fn); the
only remaining lever is inference speed. The full unique-≤15 campaign is ~15.3h of serial drafting.
3. **⇒ vLLM batched serving is JUSTIFIED** (the "measure before investing" gate passes): continuous
batching (5–20×) turns ~15h serial drafting into ~1–3h. It drops into Phase A unchanged (same
`/v1/chat/completions`). Path: a separate `.venv-vllm`; try bnb-direct + runtime LoRA first, fall
back to merge→fp16→AWQ 4-bit (fits the 12 GB card). Quantization is a throughput risk only, never
correctness (byte-gate). Deferred to Drew's go.
### T10.7 RESULT — GLM5.2 (OpenRouter) A/B on the HARD band — 2026-07-01
Sample: 18 hard-band fns (16–22 ins, untried, `ov_SC01_077` struct-core — the P16-walled band the
≤15 campaign doesn't touch). Drafted with **GLM5.2** (`z-ai/glm-5.2`, $0.93/$3.00 per Mtok, a
**reasoning** model, 1M ctx) via OpenRouter and local **v3**, same LEAN pipeline, iters=2. `api_draft`
gained `MAXTOK` env (reasoning models burn the 512 cap on reasoning tokens → empty content — raised to
8k/16k) + `usage.cost` capture.
| | match_one (CODEGEN) | whole-binary BANK | cost |
|---|---|---|---|
| v3 (local) | **1/18** (6%) | ~1/18 | $0 |
| GLM5.2 | **10/18** (56%) | **3/18** | ~$1.02 total ($0.03–0.08/fn) |
**GLM is ~10× better at hard-band codegen** (bodies): 10/18 vs 1/18 match_one; fns v3 could only reach
"near 10 → fail" GLM one-shot MATCHed. **But whole-binary banks are only 3/18** (2 direct +
`func_801577C8` via `fix_arity_callers`). The 7 stranded match_one-matches are blocked by the
**DEF-side loose-typing wall** (Phase 16/20): the overlay forward-declares the fn (from a banked
caller) with a signature conflicting with the fn's true byte-correct def (`conflicting types for
func_X`). **This caps ANY drafter** — the conflict is overlay-decl-vs-true-sig, independent of who
writes the body. `fix_arity_callers` (relax `(void)`→`()`) recovered only 1; the other 6 have
non-`(void)` conflicting forward-decls or narrow params. **Fable5 review §3c re-test verdict: the
def-side wall HOLDS** — GLM's better bodies don't break it; the wall is the ceiling, not the drafter.
The 6 hardest fns are beyond GLM too (0/6 even at MAXTOK=16000 — runaway reasoning / genuine complexity).
**Strategic read (the honest one):**
- GLM's codegen edge converts to banks only on the **def-conflict-FREE** subset (~3/18 = 17% here); at
$0.03/fn that's a real, immediate lever v3 can't touch, but wall-capped.
- The **flywheel** (distill GLM bodies → retrain v3) is **also wall-capped** — v3-trained-on-GLM hits
the same def-conflicts; it lifts v3's hard-band *codegen* on the conflict-free subset (uncertain
transfer, ~2h GPU), it does not break the wall.
- **The real unlock is the def-side wall itself, not the drafter.** "Loose-typing reconciliation is a
reasoning problem" (Fable5 §4.1): a reasoning model given the *conflicting decls + the body* could
reason out a consistent declaration strategy — a different prompt/harness than body-drafting, and the
highest-ceiling (most experimental) follow-up.
Banked: **3 GLM hard-band fns** (func_8013373C, func_8012F8C8, func_801577C8), byte-gated, check-all
136/136. Total OpenRouter spend **~$1.02 of $25**. Options handed to Drew: (1) GLM as a direct drafter
on def-conflict-free hard fns; (2) corpus-v4 flywheel (wall-capped); (3) point GLM at the def-side-wall
*reconciliation* (the reasoning-shaped unlock).
**Option-3 probe result (`tools/glm_reconcile.py`, +$0.23):** aimed GLM's REASONING directly at the wall
(body + the actual conflicting TU decls + the reconciliation toolkit → a consistent byte-identical decl
set; reasoning captured to `.run/glm_reason/` for R16). GLM's reasoning is **expert-grade** — it reads
store widths (`u16` params emit `sh`; widening to `s32` would emit `sw`), reasons K&R promotion, and
**independently derives the call-site-cast idiom** (§17a-1) for un-relaxable callee *definitions*. But it
banked only **1/7** (`func_80175184`); mechanical `--any-proto` relaxation banked **0/7**. **Verdict: the
def-side wall is INTRINSIC** — a frontier reasoning model with the full toolkit cracks ~1/7; the residual is
the irreducible narrow-param wall (K&R can't no-proto a `u16`/`s16` def) + byte-level addressing subtleties.
Total GLM banks on the sample: **4/18** (2 direct + 1 fix_arity + 1 reconcile) vs v3 ~1/18. **Conclusion:**
GLM's role is (a) a **$0.03–0.08/fn direct drafter for the def-conflict-FREE hard band** (~22%, which v3
can't reach — a real, affordable lever on the byte-weighty functions), and (b) an **idiom TEACHER**
(reasoning captured + distilled → cookbook §29; corpus-v4 optional but wall-capped). It is NOT a
wall-breaker. Per Fable5 review §4.3, the real lever past this wall is community labor (the public flip),
not a bigger model. Total T10.7 spend ~$1.25 of $25. The reconciliation idioms → cookbook §29.
## Open questions / notes
- **Corpus quality > size.** ~1,700 verified pairs is plenty for LoRA; dedup near-identical reach
copies so the model doesn't overfit one idiom. Consider weighting by residual class (cookbook
taxonomy) so under-represented quirks (schedule, IV-combine) aren't drowned by plumbing.
- **Format discipline:** train on the literal `.s` + the C as `api_draft` will present them. Any drift
between train and inference prompt format wastes the specialization.
- **It won't generalize** beyond BFM/gcc-2.7.2 — that's the point. It's a drafter for *this* project.
- **Repo placement:** build it IN this repo (the git history + asm/src are the essential context the
exporter reads). Tooling + this doc on `main`; the dataset JSONL and any GGUF/LoRA weights go in a
**gitignored** dir (`datasets/`, `models/`) per R20 (large binaries stay out of git). No branch or
worktree needed — the work is purely additive.
## Artifacts
- `tools/export_pairs.py` — corpus miner (this doc's step 1)
- `datasets/match_pairs/` — exported JSONL (gitignored)
- eval: reuse `tools/api_draft.py` + `tools/ab_score.py`
### v4 RESULT — retrain on the larger post-giant corpus is a NEGATIVE (2026-07-08, Phase 25 T4)
Retrained v4 (same recipe as v3: Qwen2.5-Coder-7B QLoRA, rank 16, 3 epochs, maxlen 2048, batch 1) on the
re-exported corpus **2,891→3,574 pairs** (+994 medium 16-40 ins + 597 large >40 ins from the giant campaign).
Final loss ~0.082 (converged like v3). Gate-true A/B vs v3 on identical held-out functions, 3 bands:
- **easy 6-14 ins:** v3 5/5, v4 5/5 (tie — no regression).
- **medium 18-40 ins:** v3 0/12 but near-misses closer (one `near-1`), 1 compile-fail; v4 0/12 with 4 compile-fails
and farther near-misses → **v3 better** (closer on 9/12). v4 slightly REGRESSED.
- **hard 45-85 ins:** both 0/10 (tie — the 7B capacity wall).
**Verdict: discard v4, keep v3.** More (harder) data did NOT lift the capacity ceiling — the "corpus quality >
size" note, confirmed. v4 scored 0/5 even on 76-83 ins functions it TRAINED on (verified ~1.4-1.7k tok, inside
maxlen 2048 → genuine capacity, not truncation).
**Two setup findings for any future retrain:**
1. **maxlen-2048 truncates functions >~85 ins** (example = system + asm + C ≈ N×22 + 150 tok). The giant-campaign
corpus has many such functions → they trained on CUT-OFF completions (teaches incomplete C — actively harmful,
the likely source of v4's medium regression). **Fix: DROP over-length examples** (clean, zero VRAM cost) rather
than train on truncated ones; OR maxlen 4096 via **gradient checkpointing** (recompute, ~25% slower, no extra
VRAM — NOT CPU offload, which is 2-4× slower for training / 5-20× for inference over the ~25 GB/s PCIe straw on
this WSL2 box vs ~900 GB/s VRAM). The biggest giants clip even at 4096.
2. **Train/inference maxlen mismatch:** v4 trained at 2048 but serves at 4096 — a big function that fits at
inference was never trained for that context. Train at the context you'll infer at.
**Strategic conclusion (→ decision-log 2026-07-08):** the local-7B tier is capacity-bound and **off the endgame
critical path**. The engine is `frontier-crack → deterministic-propagate (family_remap/dedup_propagate) → byte-gate`
+ permuter-soften. v3 stays as a frozen $0 mop-up for the ≤~15-ins setter/leaf tail; **no more retrains** — a real
capability jump needs a bigger base (14B-4bit fits the 12 GB card; 32B → cloud A100) or the frontier tier, not more
data on the 7B. v4 adapter kept at `models/bfm-match-7b-v4` (gitignored) for reference; the partial GGUF merge was
aborted (unneeded — serve_local runs base+adapter).