From 366f7e4c7fedf7108b5b94f16adaec63cc092c18 Mon Sep 17 00:00:00 2001 From: Drew T <50529377+Druthulu@users.noreply.github.com> Date: Tue, 30 Jun 2026 00:12:45 -0600 Subject: [PATCH] =?UTF-8?q?feat(phase-23):=20T9=20=E2=80=94=20reach>=3D2?= =?UTF-8?q?=20targeting=20(lora=5Fgrind=20--min-reach);=20shared-code=20is?= =?UTF-8?q?=20the=20model's=20weak=20band?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Add a sig-based reach oracle + --min-reach N to lora_grind so the mass-run can prefer SHARED functions (one bank propagates x reach — the fleet-% multiplier). The oracle reads the same .run/sig.ov_*.jsonl dedup_propagate uses (validated: 0 mismatch over 60 stubs + the func_8017CE24=2 ground truth), so a reach>=N target is exactly one dedup_propagate will stamp x reach after the bank. Bounded reach>=2 mass-run (ov_SC01_000, 15 shared <=15-ins stubs): 0/15 banked, vs the reach-1-heavy spot-run's 7/15. The model is weakest exactly on reach>=2 (shared) code: (1) the corpus skipped the shared DEFINE_func macro bodies (export_pairs reads only src/ defs -> 96.6% of the corpus is overlay-unique), and (2) the shared engine fns are the harder regalloc/schedule residuals. So reach>=2 model-only is NOT a fleet lever by itself. BUT the reach>=2 drafts are high-value FUEL: 5/15 are close<=3 reach-134 near-misses (3x close=1: func_8012E27C/BF4C/AD64) -> x134 each if closed. The real lever is reach>=2 draft -> permuter-grinder close (x134), which needs the SAME per-binary fix T7 applied to lora_grind: grinder.py calls run_gate with no binary (-> 077) and the backlog stores no binary field. That two-part fix is the next step. byte-neutral: check-all 136/136. - docs/gen2-mips-matching-model.md: T9 RESULT - phase-ends/CURRENT_PHASE.md: T9 done; next = grinder per-binary fix, then corpus-v3 --- docs/gen2-mips-matching-model.md | 21 +++++++++++ phase-ends/CURRENT_PHASE.md | 3 +- tools/lora_grind.py | 63 +++++++++++++++++++++++++++++--- 3 files changed, 81 insertions(+), 6 deletions(-) diff --git a/docs/gen2-mips-matching-model.md b/docs/gen2-mips-matching-model.md index bcd7d55a6..dd094d4f0 100644 --- a/docs/gen2-mips-matching-model.md +++ b/docs/gen2-mips-matching-model.md @@ -149,6 +149,27 @@ the canonical-site (077) harvest, plus **corpus-v3** for the struct compile-fail broad rotation. A broad ≤15-ins run remains worthwhile for per-overlay completeness, corpus growth (retrain fuel), and seeding the permuter grinder with the close=1 near-misses. +### T9 RESULT — reach≥2 targeting built; the model alone is weakest on shared code (2026-06-30) + +Added `--min-reach N` to `lora_grind` (a lazy sig-based reach oracle == `dedup_propagate`'s, validated +0-mismatch over 60 stubs + the `func_8017CE24`=2 ground truth) so the mass-run can prefer SHARED +functions — one bank → ×reach. Bounded reach≥2 run, ov_SC01_000 batch (15 reach≥2 ≤15-ins stubs): +**0/15 banked** — vs the reach-1-heavy spot-run's 7/15. The model is **weakest exactly on reach≥2 +(shared) code**, for two compounding reasons: (1) the corpus skipped the shared `DEFINE_func` macro +bodies (`export_pairs` reads only `src/` defs → 96.6% of the corpus is overlay-unique), and (2) the +shared engine functions are the harder regalloc/schedule residuals that survived the whole Phase-21 +apparatus. So **reach≥2 model-only is NOT a fleet-% lever by itself** (a real, measured negative). + +BUT the reach≥2 drafts are high-value FUEL: of the 15, **5 are close≤3 reach-134 near-misses** — three +at **close=1** (`func_8012E27C/BF4C/AD64`) + two `match_one`-MATCH-but-gate-rejected — each worth ×134 +if closed. So the real fleet lever is **reach≥2 draft → permuter-grinder close (×134)** (the +CURRENT_PHASE "synergy"), and/or **corpus-v3-with-shared-bodies** to lift the model's direct banking. +Realizing the grinder path needs the **same per-binary fix T7 applied to `lora_grind`**: `grinder.py` +calls `run_gate` with no `binary` (→ defaults to ov_SC01_077) AND the backlog record stores no `binary` +field — so the permuter can't gate a *non-077* near-miss today. That two-part fix (grinder per-binary +resolution + a backlog `binary` field) is the next concrete step to turn the reach-134 close=1 fuel into +×134 banks. The reach oracle + `--min-reach` are reusable for that and for a corpus-v3 retrain. + ## Open questions / notes - **Corpus quality > size.** ~1,700 verified pairs is plenty for LoRA; dedup near-identical reach diff --git a/phase-ends/CURRENT_PHASE.md b/phase-ends/CURRENT_PHASE.md index ddaaf4763..7c822adc7 100644 --- a/phase-ends/CURRENT_PHASE.md +++ b/phase-ends/CURRENT_PHASE.md @@ -34,7 +34,7 @@ D. Periodically: `export_pairs → format_finetune → train_lora → redeploy` ## ▶ RESUME HERE (fresh session) **State:** Phase 23 in progress (NOT a phase end). Phase 22 closed (`PhaseEnd_Phase22.md`, uncommitted — Drew's gate-2 commit+push). The fine-tuned model **`bfm-match-7b-v2`** (Qwen2.5-Coder-7B QLoRA on corpus-v2) is built; **served by LM Studio** at `http://192.168.1.113:1234/v1` (model id `bfm-match-7b-v2`; GGUF at `models/bfm-match-7b_gguf/`). Corpus `datasets/match_pairs/` + training stack `.venv-train` (gitignored). **T7 FIXED** — the gate banks fleet-wide now (ov_SC01_000 7/15 byte-identical, +1 reach-2 propagated; @commit:0322 + the T7 checkpoint commit). The 0/222 was two harness bugs (good_sha format + src/asm/out 077-default), not the model. -**NEXT TASK — T8: corpus-v3 (struct types).** T7 is closed; a bounded mass-run is now safe to size. The remaining draft losses split into: (a) standalone-compile-fails on **struct types** (→ corpus-v3, THIS task — emit the `struct {...}` defs each fn needs; v2 did globals only) and (b) **reach-1 ROI** (→ T9). Build corpus-v3 → **retest free on the 7B** (does the struct band lift on small/medium?) → only then a dense 14B cloud train if it pays. Then T9 (reach≥2 target selection — the fleet-% lever — a `lora_grind`/`wave_targets` `--min-reach 2` filter) + wire the concurrent permuter grinder on the close=1 near-misses T7 surfaced. +**NEXT TASK — the grinder per-binary fix (to realize the reach≥2 ×134 synergy).** T7 (gate) and T9 reach≥2 targeting (`lora_grind --min-reach`) are DONE. The reach≥2 mass-run's verdict: the model banks **~0 on shared functions directly** (0/15 on ov_SC01_000's reach≥2 batch vs 7/15 reach-1), because the corpus skipped the shared `DEFINE_func` bodies and these are the harder regalloc tail — BUT it generated **3 close=1 reach-134 near-misses** (`func_8012E27C/BF4C/AD64`) = high-value permuter fuel (×134 each). The **permuter grinder** is the tool that closes close=1 regalloc/schedule near-misses — but `grinder.py` has the **same per-binary bug T7 fixed** (`run_gate` with no `binary` → ov_SC01_077; the backlog stores no `binary` field), so it can't gate a non-077 near-miss. Fix = (1) add a `binary` field to the backlog record in `gate_stage`, (2) thread it through `grinder.py`'s `run_gate` call; then run the grinder on the reach-134 close=1 fuel → ×134 banks. **SECOND lever:** corpus-v3 (struct types) to lift the model's DIRECT reach≥2 bank rate (the 2/15 compile-fails + the 0-direct-bank gap). **Run a bounded mass-run (when Drew says go):** ``` @@ -65,3 +65,4 @@ The **whole-binary byte-gate** (`gate_stage`/`harvest_verify`, G3/P9) is the sol ## Progress log - 2026-06-29: **Phase opened at PhaseEnd_Phase22 close (Drew).** Built across this session: cheap-tier A/B (T1, Haiku 4.8×/$), stock-local floor (T2, 0), the LoRA pipeline (T3) + corpus-v2 extern-fix (T4, 6–15 ins 0%→85%), first 4 real open-stub banks (T5, @commit:0320), the `lora_grind` mass-run driver (T6). Calibration run (T7) launched (500 fns) — **18% on ov_SC01_077 but 0/222 broad rotation → #1 debug.** gate_stage commit tag made phase-agnostic. The whole arc + measured numbers: `docs/gen2-mips-matching-model.md`; memory `cheap-tier-ab-validated`. NEXT: T7 debug, then bounded mass-runs + corpus-v3. - 2026-06-30: **T7 DEBUGGED + FIXED.** 3 Explore scouts (tooling / run-evidence / corpus) + a direct code read (R14 — which resolved a flat contradiction between two scouts) found **two independent bugs** in `lora_grind`'s gate path: **(A)** `good_sha()` passed `" "` vs harvest_verify's bare `sha1()` → 0 banks for ALL binaries incl. 077 (so 077's "0/12" was a bug artifact); **(B)** `src/asm/out` defaulted to ov_SC01_077 → non-077 drafts dropped at the 077 stub-filter, silently. Fixed `gate_stage.run_gate` (binary-agnostic resolution + bare-hash normalize + a loud negative-control guard) + `lora_grind.good_sha`; byte-neutral (check-all 136/136). ov_SC01_000 spot-run **banked 7/15 (47%) byte-identical** (@commit:0322) → reach-2 `func_8017CE24` propagated ×2. **ROI:** 6/7 reach-1 → broad rotation is high bank-rate / low fleet-% ROI; the fleet lever is **reach≥2 targeting** + corpus-v3. Backlog now correctly classified (4× close=1 = grinder fuel). NEXT: **T8 corpus-v3** (struct types) + **T9 reach≥2 selection** + concurrent grinder. +- 2026-06-30 (cont.): **T9 reach≥2 targeting built + measured.** Added `lora_grind --min-reach N` (lazy sig-based reach oracle == `dedup_propagate`, validated 0-mismatch/60 + the func_8017CE24=2 ground truth; `--min-reach 2` ranks high-reach-first, naturally restricts to overlays). Bounded reach≥2 mass-run: ov_SC01_000's 15 reach≥2 (shared) stubs banked **0/15** (vs the reach-1 spot-run's 7/15) — the model is **weakest on shared code** (corpus skipped the `DEFINE_func` bodies + it's the regalloc/schedule tail). But **5/15 are close≤3 reach-134 near-misses** (3× close=1 = func_8012E27C/BF4C/AD64) → high-value permuter fuel (×134 each). **FINDING: reach≥2 model-only ≠ a fleet lever; the lever is reach≥2-draft → grinder-close (×134)**, which needs `grinder.py`'s per-binary fix (same class as T7) + a backlog `binary` field. (A foreground mass-run hit the 10-min Bash cap mid-2nd-batch; tree recovered clean via `git checkout`, check-all 136/136.) Details: `docs/gen2-mips-matching-model.md` "T9 RESULT". NEXT: the grinder per-binary fix (realize the reach-134 ×134 fuel), then corpus-v3. diff --git a/tools/lora_grind.py b/tools/lora_grind.py index fc740bd4e..8a9d483cb 100644 --- a/tools/lora_grind.py +++ b/tools/lora_grind.py @@ -44,13 +44,57 @@ def good_sha(b): return open(p).read().split()[0] if os.path.exists(p) else None # bare hash (sha1sum format) +_SIGS = None + + +def _load_sigs(): + """Lazy {overlay -> {addr_int -> h_exact}} from the 134 .run/sig.ov_*.jsonl (make sig-overlays). + The SAME data dedup_propagate reaches from, so a reach>=N fn here is exactly one it will stamp + ×reach after the bank (and a fn the sigs miss wouldn't propagate anyway → correctly excluded).""" + global _SIGS + if _SIGS is None: + _SIGS = {} + for p in glob.glob(os.path.join(REPO, ".run/sig.ov_*.jsonl")): + ov = os.path.basename(p)[4:-6] # sig.ov_SC01_000.jsonl -> ov_SC01_000 + d = {} + for line in open(p): + line = line.strip() + if not line: + continue + r = json.loads(line) + nm = r.get("name", "") + if nm.startswith("func_"): + try: + d[int(nm[5:], 16)] = r.get("h_exact") + except ValueError: + pass + _SIGS[ov] = d + return _SIGS + + +def reach_of(binary, fn): + """#overlays byte-identical (h_exact) to `binary` at fn's addr; None if binary/fn isn't signed. + reach>=2 = a shared fn that propagates ×reach via dedup_propagate after it banks (the fleet lever).""" + sigs = _load_sigs() + try: + addr = int(fn[5:], 16) + except ValueError: + return None + h = sigs.get(binary, {}).get(addr) + if not h: + return None + return sum(1 for ov in sigs if sigs[ov].get(addr) == h) + + def nins(s_path): return sum(1 for l in open(s_path) if re.match(r'\s*/\*\s*[0-9A-Fa-f]+\s+[0-9A-Fa-f]+\s+[0-9A-Fa-f]{8}\s*\*/', l)) -def open_stubs(b, max_nins, tried): - """still-INCLUDE_ASM funcs for binary b (main + split .c) whose .s exists and is <= max_nins ins.""" +def open_stubs(b, max_nins, tried, min_reach=1): + """still-INCLUDE_ASM funcs for binary b (main + split .c) whose .s exists and is <= max_nins ins. + min_reach>1 keeps only fns byte-identical across >= min_reach overlays (the propagation multiplier) + and ranks high-reach-first; min_reach=1 keeps all, smallest-first (no reach cost).""" stubbed = set() for cf in glob.glob(os.path.join(REPO, "src/%s/%s*.c" % (b, b))): stubbed |= set(STUB_RE.findall(open(cf).read())) @@ -61,10 +105,16 @@ def open_stubs(b, max_nins, tried): for sd in glob.glob(os.path.join(REPO, "asm/%s/nonmatchings/*/%s.s" % (b, fn))): n = nins(sd) if 0 < n <= max_nins: - out.append({"name": fn, "addr": "0x" + fn[5:].lower(), "nins": n, + rch = reach_of(b, fn) if min_reach > 1 else None + if min_reach > 1 and (rch is None or rch < min_reach): + break # below the reach threshold (or unsigned) -> skip + out.append({"name": fn, "addr": "0x" + fn[5:].lower(), "nins": n, "reach": rch, "class": "WAVE", "asm": os.path.relpath(sd, REPO), "ghidra_c": ""}) break - out.sort(key=lambda t: t["nins"]) + if min_reach > 1: + out.sort(key=lambda t: (-(t.get("reach") or 0), t["nins"])) # leverage: high-reach, then small + else: + out.sort(key=lambda t: t["nins"]) return out @@ -120,6 +170,9 @@ def near_class_hist(): def main(): ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) ap.add_argument("--max-nins", type=int, default=15) + ap.add_argument("--min-reach", type=int, default=1, + help="only draft fns byte-identical across >= N overlays — the x reach propagation " + "multiplier (the fleet lever); default 1 = all open stubs, smallest-first") ap.add_argument("--batch", type=int, default=12) ap.add_argument("--iters", type=int, default=3) ap.add_argument("--propagate-every", type=int, default=8, help="run a fleet propagate sweep every N batches") @@ -150,7 +203,7 @@ def main(): continue if a.max_batches and batch_i >= a.max_batches: break - stubs = open_stubs(b, a.max_nins, tried) + stubs = open_stubs(b, a.max_nins, tried, a.min_reach) if not stubs: continue did_work = True