mirror of
https://github.com/Druthulu/BFM-decomp
synced 2026-10-01 07:40:42 -04:00
docs(phase-25): T4 CLOSEOUT — v4 retrain FAILED (negative), keep v3; T5 handoff
- A/B gate-true: v4 <= v3 (marginally worse on medium, tie easy/hard) -> discard v4, keep v3 (frozen ceiling) - 'corpus quality > size' confirmed; 7B is capacity-bound (0/5 even on trained fns, not truncation) - decision-log (R31): the local-7B tier is off the endgame critical path; engine = frontier-crack -> deterministic-propagate -> byte-gate + permuter - gen2-mips-matching-model: the A/B + the maxlen-2048 truncation flaw (drop-over-length OR grad-checkpointing, NOT CPU offload) for any future retrain - CURRENT_PHASE: T4 done/failed; NEXT = T5 (Ultracode measure-wave); phase OPEN, no PhaseEnd
This commit is contained in:
@@ -70,3 +70,31 @@ the session that produced it. Route TECHNICAL idioms to the cookbook; this file
|
||||
detect-collisions-first probe would have scoped the safe subset up front instead of discovering it via a
|
||||
failed build. General lesson: an automated bulk transform needs an explicit *soundness boundary*, and the
|
||||
byte-gate (not optimism) is what stops a partial success from masquerading as a full one.
|
||||
|
||||
## 2026-07-08 · Phase 25 — the local-7B tier is capacity-bound and off the endgame critical path (T4)
|
||||
- **Context / belief:** the fine-tuned local drafter (`bfm-match-7b-v3`) was a core cheap tier; retraining **v4**
|
||||
on the much larger post-giant-campaign corpus (2,891→3,574 pairs, +994 medium + 597 large functions v3 never
|
||||
saw) should extend its band upward and make it a stronger drafter for the T5 wave.
|
||||
- **Dead-end:** v4 **did not beat v3** — it was marginally WORSE. Gate-true head-to-head on identical held-out
|
||||
functions: easy 6-14 ins both 5/5; **medium 18-40 ins** v3's near-misses closer (one at `near-1`, permuter fuel)
|
||||
with 1 compile-fail vs v4's 4 — v3 closer on 9/12; **hard 45-85 ins** both 0/10. Crucially v4 scored 0/5 even on
|
||||
the 76-83 ins functions it TRAINED on (verified ~1.4-1.7k tok, well inside maxlen 2048 → NOT truncation → genuine
|
||||
capacity). (Note: a real corpus-prep flaw exists — functions >85 ins WERE truncated at maxlen 2048 → training on
|
||||
cut-off completions, likely the source of v4's slight medium regression — but it doesn't touch the decisive band.)
|
||||
- **Pivot:** discard v4, **keep v3 (the frozen ceiling)**, and stop investing in the local-7B tier. Not retired
|
||||
(still a $0 mop-up for the ≤~15-ins setter/leaf tail), just no longer load-bearing and no more retrains.
|
||||
- **Why:** the byte-gate A/B settled it directly (G3/P9). "**Corpus quality > size**" landed empirically: v2→v3
|
||||
gained from *better* data (the extern-capture fix); v3→v4 was just *more/harder* data and it didn't lift a
|
||||
capacity ceiling. Byte-matching's hard part is compiler-codegen REASONING (scales UP with model size), not
|
||||
language breadth (which a smaller model could shed) — so neither "more data" nor "a smaller RE-specialist" is the
|
||||
lever; the reasoning has to come from a large pretrained base or a frontier model, and the RE-smartness that IS
|
||||
small+deterministic already exists as **m2c** (rules, not weights).
|
||||
- **Hindsight / for the wiki:** **the endgame engine is `frontier-crack → deterministic-propagate → byte-gate`,
|
||||
with the permuter softening near-misses — the local small model is a convenience on the small tail, not a
|
||||
load-bearing part.** For a *matching* decomp you already own the ground-truth compiler + a perfect verifier, so
|
||||
the ML task is candidate-PROPOSAL + search (proposal quality scales with reasoning/size; the check is free). A
|
||||
bespoke small "RE model" founders on data scarcity (the asm↔C-under-a-specific-compiler corpus only exists, tiny,
|
||||
in decomp git histories). The honest tiering: **m2c** for structure, a **frontier reasoner** for the byte-exact
|
||||
precision on the hard/byte-weighty band, the **permuter** for regalloc/schedule search, a **frozen small LoRA**
|
||||
only for the cheap ≤15-ins tail. Don't spend GPU-hours chasing band-extension on a 7B; rent a bigger GPU or use
|
||||
the frontier tier when the hard band is the target.
|
||||
|
||||
@@ -345,3 +345,34 @@ not a bigger model. Total T10.7 spend ~$1.25 of $25. The reconciliation idioms
|
||||
- `tools/export_pairs.py` — corpus miner (this doc's step 1)
|
||||
- `datasets/match_pairs/` — exported JSONL (gitignored)
|
||||
- eval: reuse `tools/api_draft.py` + `tools/ab_score.py`
|
||||
|
||||
### v4 RESULT — retrain on the larger post-giant corpus is a NEGATIVE (2026-07-08, Phase 25 T4)
|
||||
|
||||
Retrained v4 (same recipe as v3: Qwen2.5-Coder-7B QLoRA, rank 16, 3 epochs, maxlen 2048, batch 1) on the
|
||||
re-exported corpus **2,891→3,574 pairs** (+994 medium 16-40 ins + 597 large >40 ins from the giant campaign).
|
||||
Final loss ~0.082 (converged like v3). Gate-true A/B vs v3 on identical held-out functions, 3 bands:
|
||||
- **easy 6-14 ins:** v3 5/5, v4 5/5 (tie — no regression).
|
||||
- **medium 18-40 ins:** v3 0/12 but near-misses closer (one `near-1`), 1 compile-fail; v4 0/12 with 4 compile-fails
|
||||
and farther near-misses → **v3 better** (closer on 9/12). v4 slightly REGRESSED.
|
||||
- **hard 45-85 ins:** both 0/10 (tie — the 7B capacity wall).
|
||||
|
||||
**Verdict: discard v4, keep v3.** More (harder) data did NOT lift the capacity ceiling — the "corpus quality >
|
||||
size" note, confirmed. v4 scored 0/5 even on 76-83 ins functions it TRAINED on (verified ~1.4-1.7k tok, inside
|
||||
maxlen 2048 → genuine capacity, not truncation).
|
||||
|
||||
**Two setup findings for any future retrain:**
|
||||
1. **maxlen-2048 truncates functions >~85 ins** (example = system + asm + C ≈ N×22 + 150 tok). The giant-campaign
|
||||
corpus has many such functions → they trained on CUT-OFF completions (teaches incomplete C — actively harmful,
|
||||
the likely source of v4's medium regression). **Fix: DROP over-length examples** (clean, zero VRAM cost) rather
|
||||
than train on truncated ones; OR maxlen 4096 via **gradient checkpointing** (recompute, ~25% slower, no extra
|
||||
VRAM — NOT CPU offload, which is 2-4× slower for training / 5-20× for inference over the ~25 GB/s PCIe straw on
|
||||
this WSL2 box vs ~900 GB/s VRAM). The biggest giants clip even at 4096.
|
||||
2. **Train/inference maxlen mismatch:** v4 trained at 2048 but serves at 4096 — a big function that fits at
|
||||
inference was never trained for that context. Train at the context you'll infer at.
|
||||
|
||||
**Strategic conclusion (→ decision-log 2026-07-08):** the local-7B tier is capacity-bound and **off the endgame
|
||||
critical path**. The engine is `frontier-crack → deterministic-propagate (family_remap/dedup_propagate) → byte-gate`
|
||||
+ permuter-soften. v3 stays as a frozen $0 mop-up for the ≤~15-ins setter/leaf tail; **no more retrains** — a real
|
||||
capability jump needs a bigger base (14B-4bit fits the 12 GB card; 32B → cloud A100) or the frontier tier, not more
|
||||
data on the 7B. v4 adapter kept at `models/bfm-match-7b-v4` (gitignored) for reference; the partial GGUF merge was
|
||||
aborted (unneeded — serve_local runs base+adapter).
|
||||
|
||||
+26
-18
@@ -35,15 +35,24 @@ idiom curriculum → sweep one exemplar per family (curriculum-ordered), each cr
|
||||
real substance):** crack the **127 draftable family exemplars** curriculum-ordered (T6 order), template/propagate ×134.
|
||||
- [ ] Close — clean-fleet verify · PhaseEnd synthesis · plain-English recap (R25) · Phase-26 backlog *(NOT yet — phase open)*
|
||||
|
||||
## FRESH SESSION — RESUME HERE (T7-MECHANICAL done & pulled ahead; **NEXT = T4 → T5 → T6 → T7-cracking**; phase OPEN, no PhaseEnd)
|
||||
**Clarification (Drew, 2026-07-08):** "do NOT start T4 *yet*" meant *finish the pulled-ahead mechanical T7 sweep
|
||||
first* — **NOT** defer T4 to Phase 26. T4/T5/T6 + T7's exemplar-cracking (the 127 draftable families / 6.7 MB) are
|
||||
the phase's CORE remaining work. The mechanical sweeps were the cheap prelude.
|
||||
## FRESH SESSION — RESUME HERE (**NEXT = T5, the Ultracode measure-wave**; phase OPEN, no PhaseEnd; do NOT close)
|
||||
**Phase is OPEN.** Order: T7-mechanical (done, pulled ahead) → T4 (done, FAILED/negative) → **T5 (NEXT)** → T6 →
|
||||
T7-cracking → Close. The phase's CORE prize is still ahead: **the 127 draftable family exemplars / 6.7 MB**, cracked
|
||||
by the frontier tier and propagated ×134 by deterministic tooling. Do NOT write a PhaseEnd until T5/T6/T7-cracking
|
||||
are done and top-family ROI drops (open-ended milestone, per the plan-of-record). *(A prior session misread "→ Close"
|
||||
in a handoff header and nearly closed early — see `docs/decision-log.md`; reconcile against the plan-of-record.)*
|
||||
|
||||
**T5 — the cheap-tier soften + measure wave (Step A).** Over the T2 exemplars + reach-1 + reach-134-tractable:
|
||||
cheap drafters (**v3** — NOT v4, which failed T4 — + cheap-Opus applying §31) → `match_one` closeness + `klass` per
|
||||
fn; permuter-ILS softens regalloc/schedule seeds. Output = the complete class/closeness **frontier map** +
|
||||
pre-advanced seeds. Bank free wins via `gate_stage` as they land. **This is BREADTH → at session start, PROMPT Drew
|
||||
for `/effort ultracode` and WAIT for the toggle (R27); Claude cannot set effort.** Tools: `wave_targets.py` /
|
||||
`gen_harvest_targets.py`, `tools/workflows/worker_wave.js` + `distill.js`, `permuter_ils.py`, `bulk_harvest.py`.
|
||||
|
||||
**Done + committed:** T0–T3 + T7-part-1 (matched-free sweep, `commit:0476`) + **T7.2 decl-reconcile** (`commit:0479`) +
|
||||
**T7.3 h_exact stragglers** (`commit:0481`) — fleet **66.02% → 70.82% → 71.32% → 71.36%**. T7.2 banked **1,729**
|
||||
(base 532 + `_after` 1,197) via type-lift + mechanical remap; T7.3 banked **~200** via `dedup_propagate`; **R22
|
||||
clean-fleet 136/136** every batch, dedup-check 1813/0.
|
||||
**T7.3 h_exact stragglers** (`commit:0481`) + **T4 (v4 retrain — FAILED, keep v3)** — fleet **66.02% → 70.82% → 71.32%
|
||||
→ 71.36%**. T7.2 banked **1,729** (base 532 + `_after` 1,197) via type-lift + mechanical remap; T7.3 banked **~200**
|
||||
via `dedup_propagate`; **R22 clean-fleet 136/136** every batch, dedup-check 1813/0.
|
||||
|
||||
**The mechanical family method (the phase's engine — T3):** h_norm families are TEMPLATES, not free dedup. Per
|
||||
family: crack ONE exemplar → `family_remap` builds each member's C by positionally substituting the per-overlay
|
||||
@@ -204,17 +213,16 @@ Every batch R22 clean-fleet **136/136**, dedup-check 1813/0. Tools added: `famil
|
||||
`build_engine_types --file/--exclude`. Cookbook **§40a** written (R30). **NEXT = T4** (see the RESUME section above);
|
||||
NOT closing — the curriculum + 127-draftable exemplar-cracking is the phase's core, still ahead.
|
||||
|
||||
### T4 — v4 LoRA retrain + A/B gate ⏳ IN PROGRESS (2026-07-08)
|
||||
- **Corpus re-exported** (`export_pairs`, mines the now-fully-built objects): **2,891 → 3,574 pairs** (v3-era snapshot
|
||||
saved at `datasets/match_pairs/pairs.v3era.jsonl`). Size dist: ≤5=784, 6-15=1199, **16-40=994, >40=597** — the
|
||||
medium/large signal v3 (trained pre-giant-campaign) lacked. `format_finetune` → **3,176 train / 360 test** (99% compile).
|
||||
- **Training LAUNCHED** (`.run/train_v4.log`, pid 282216 @ 2026-07-08): `train_lora --out models/bfm-match-7b-v4
|
||||
--maxlen 2048 --batch 1 --epochs 3 --rank 16` (v3's exact recipe; only the corpus changed → fair A/B). ~2-3h on the
|
||||
3080 Ti. Completion-waiter = bg task; models/ + datasets/ are gitignored (R20) so nothing to commit until the A/B decision.
|
||||
- **NEXT when training ends:** `eval_lora --test datasets/match_pairs/test.jsonl` on BOTH v3 and v4 (same held-out set,
|
||||
gate-true) → A/B. **Keep v4 in the wave only where it beats v3** (else fall back to v3). Then T5 (prompt `/effort ultracode`).
|
||||
- **If context compacts mid-train:** check `.run/train_v4.log` for `train_runtime` (done) or a Traceback (failed);
|
||||
if `models/bfm-match-7b-v4/adapter_model.safetensors` exists, training finished — proceed to eval.
|
||||
### T4 — v4 LoRA retrain + A/B gate ❌ FAILED / NEGATIVE (2026-07-08) — v4 discarded, KEEP v3
|
||||
- Retrained v4 (same recipe as v3, corpus 2,891→3,574) — converged (loss ~0.082). Gate-true A/B vs v3 on identical
|
||||
held-out functions, 3 bands: **easy 6-14** both 5/5; **medium 18-40** v3 closer near-misses + fewer compile-fails
|
||||
(v3 better on 9/12); **hard 45-85** both 0/10. v4 = marginally WORSE. **Verdict: discard v4, keep v3** (frozen
|
||||
ceiling). "Corpus quality > size" confirmed; the 7B is capacity-bound (0/5 even on trained fns, NOT truncation).
|
||||
- **v3 remains the drafter** (`models/bfm-match-7b-v3`). v4 adapter kept at `models/bfm-match-7b-v4` (gitignored) for
|
||||
reference only. Full write-ups: `docs/decision-log.md` (strategic — 7B off the endgame critical path) +
|
||||
`docs/gen2-mips-matching-model.md` (technical — the A/B + the maxlen-2048 truncation flaw for any future retrain).
|
||||
- **Strategic conclusion (Drew-aligned):** the local-7B tier is a $0 mop-up for the ≤~15-ins tail, NOT load-bearing.
|
||||
The endgame engine is **frontier-crack → deterministic-propagate → byte-gate** + permuter-soften. **No more 7B retrains.**
|
||||
|
||||
## RULES ADDED THIS PHASE (transcribe to the PhaseEnd Rules table at close)
|
||||
- **R31 — CONFIRMED by Drew 2026-07-08 (binding now).** Capture the WHY behind strategic pivots in `docs/decision-log.md`, while fresh. At each major
|
||||
|
||||
Reference in New Issue
Block a user