From 381b0652a1191d23bad2a6cb1fcfcbee89dc9c28 Mon Sep 17 00:00:00 2001 From: Drew T <50529377+Druthulu@users.noreply.github.com> Date: Thu, 18 Jun 2026 22:53:19 -0600 Subject: [PATCH] =?UTF-8?q?feat(phase-16):=20pipeline=20proven=20=E2=80=94?= =?UTF-8?q?=20known-answer=20ladder=2067%=20m2c-direct=20re-derivation?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - tools/p16_known_answer.py: graduated known-answer validation (Drew's method) — stub already-matched fns -> re-extract .s -> run pipeline -> measure re-derivation, restore safely (never git-checkout mid-harvest, §14c). Auto-picks a difficulty ladder. - RESULT: 8/12 known-answer fns re-derived DIRECTLY by m2c+macros (no permuter, no struct types). Remainder = permuter/struct/typing candidates. Overlay byte-restored (d19c9580). - Infra proven end-to-end: m2c --valid-syntax + common.h macros -> compile -> byte-gate. - CURRENT_PHASE: Drew refinements (graduated ladder, never-stop run, safe-exit sentinel). --- phase-ends/CURRENT_PHASE.md | 13 +++- tools/p16_known_answer.py | 133 ++++++++++++++++++++++++++++++++++++ 2 files changed, 145 insertions(+), 1 deletion(-) create mode 100644 tools/p16_known_answer.py diff --git a/phase-ends/CURRENT_PHASE.md b/phase-ends/CURRENT_PHASE.md index e6cb9e478..1d9f64ef0 100644 --- a/phase-ends/CURRENT_PHASE.md +++ b/phase-ends/CURRENT_PHASE.md @@ -29,7 +29,13 @@ - [ ] **S8** Harvest/propagate sweep + fleet roll-up (xHigh + 1 surgical survey) - [ ] **S9** Close / PhaseEnd (Max) — honest milestone, plain-English recap (R25) -**Current task pointer → S0.** +**Current task pointer → S2/S1 (pipeline proven; building permuter integration + driver).** + +### Live results (2026-06-18 night) +- **Pipeline infra PROVEN end-to-end:** m2c `--valid-syntax` + common.h macros → compiles → harvest_verify byte-gate. `match_one` confirms byte-faithful drafts. +- **Known-answer ladder (Drew's method), `tools/p16_known_answer.py`:** on 12 already-matched fns (known-reachable answers), **m2c-DIRECT re-derivation = 8/12 = 67%** (macro-only, NO permuter, NO struct types). Remainder: 2 near-misses @15 mismatch (permuter), 1 @73 (struct/hand), 1 CC1-fail (typing). Byte-restore safe (overlay back to d19c9580). → strong viability signal; struct types (S1) + permuter lift from here. +- **Unmatched smallest-80 macro-only-no-permuter:** 2/80 byte-gated — expected low (unmatched = the hard residual; no permuter yet). The gap vs 67% confirms the unmatched tail is self-selected hard; GATE-B (permuter+struct on unmatched mediums) measures the real NEW yield. +- **Permuter:** runs (2048 iters/120s @-j8); did NOT close an *unmatched* near-miss (func_8012CB64, score 145 flat — out-of-search-space, §3 class). Next: prove it closes a KNOWN-answer near-miss. ## New tools/files `tools/struct_infer.py`, `tools/m2c_ctx.py`, `src/shared/engine_struct.h` (#ifdef M2C skeleton / #else real layout), `tools/auto_driver.py`, `tools/auto_supervisor.sh`, opt `src/shared/engine_decls.h`. Reused: sig_unify, match_one, harvest_verify, dedup_propagate (patch compiles_standalone += struct header), decompile.py --context, permuter/compile.sh, progress.py --fleet, build_engine_types.py (additive). m2c context MUST be flat directive-free C (rejects #include/#ifndef). @@ -50,6 +56,11 @@ Must be verified + ready to launch unattended by Sun afternoon. Cadence (I own t **Execution-order adaptation (S0 finding):** build the m2c+macros+permuter+byte-gate **pipeline (S2) FIRST** (testable tonight), then add `struct_infer` (S1) as the regalloc/structural enhancer + measure its lift. Same tasks + gates; order adapted to get a known-answer test running fastest. (Within-phase autonomy, P3.) +## Drew refinements (2026-06-18 #2) — graduated validation + never-stop run + safe-exit +- **Graduated known-answer ladder (validation method).** Prove the pipeline on EXISTING decomp across a difficulty ramp: several softball tiny/easy (prove infra) → increasing difficulty → 10+ medium → up to the struct-heavy difficulty we'll actually run. Use already-matched functions (known answers): revert to stub → re-extract its `.s` → run the full pipeline → confirm it re-derives the byte-match we already know. Purpose: tune the methodology to be **stable + competent** before the 5-day run. (Pick struct-heavy matched fns from engine_core.h macros for representativeness; the byte-gate on unmatched mediums is the complementary capability proof = GATE-B.) +- **The 5-day run NEVER STOPS** — loops the FULL worklist to completion (match → propagate → commit → re-derive remaining → escalate permuter effort on the residual), running until ALL gettable work is done or Drew stops it. Not a one-batch run. +- **Safe-exit mechanism (REQUIRED).** Driver checks `.run/auto/STOP` at every function boundary; if present → finish current fn's gate+propagate+commit → final heartbeat "stopped safely" → exit 0; supervisor sees STOP + clean exit → does NOT relaunch. **Trigger:** Drew returns + messages me "exit the run" → I run `tools/auto_stop.sh` (`touch .run/auto/STOP`); or Drew runs the one-liner himself (works with no Claude session). Plus `tools/auto_status.sh` (heartbeat: current fn / banked count / fleet % / last commit) for remote check-in. + ## Blockers - (resolved) Away window = Sun afternoon; S1–S6 must complete by then. diff --git a/tools/p16_known_answer.py b/tools/p16_known_answer.py new file mode 100644 index 000000000..ae60acdf8 --- /dev/null +++ b/tools/p16_known_answer.py @@ -0,0 +1,133 @@ +#!/usr/bin/env python3 +"""p16_known_answer.py — graduated known-answer validation of the Phase-16 struct pipeline. + +Take ALREADY-matched functions (known answers), revert them to INCLUDE_ASM stubs, regenerate +their .s, run the pipeline (m2c --valid-syntax -> draft), and measure how close/whether it +re-derives the byte-match we already know is correct. The byte-gate (harvest_verify) is the +ground truth; match_one gives a fast per-function score for the calibration ramp. + +Restores the overlay .c from a backup afterward (never git-checkout mid-harvest — cookbook §14c). + + python3 tools/p16_known_answer.py --pick 8 # auto-pick a difficulty ladder, calibrate (m2c+match_one) + python3 tools/p16_known_answer.py --funcs func_A,func_B +""" +import argparse, os, re, subprocess, sys, json, random + +REPO = os.path.dirname(os.path.dirname(os.path.abspath(__file__))) +OV = "ov_SC01_077" +SRC = f"src/{OV}/{OV}.c" +ECORE = "src/shared/engine_core.h" +ASM = f"asm/{OV}/nonmatchings/{OV}" +BAK = ".run/p16_ka_backup.c" + + +def sh(cmd, **kw): + return subprocess.run(cmd, capture_output=True, text=True, **kw) + + +def macro_bodies(): + """addr -> instruction-ish line count for every DEFINE_func_ macro in engine_core.h.""" + txt = open(ECORE).read() + out = {} + for m in re.finditer(r"#define DEFINE_func_([0-9A-Fa-f]+)\(\)((?:.*\\\n)*.*\n)", txt): + addr = m.group(1) + body = m.group(2) + # count statements (lines with ; or store/load) as an nins proxy + n = len(re.findall(r"[;}]", body)) + out[addr] = n + return out + + +def instantiated_in_src(addr): + return re.search(rf"DEFINE_func_{addr}\(\)", open(SRC).read()) is not None + + +def main(): + ap = argparse.ArgumentParser() + ap.add_argument("--pick", type=int, default=8, help="auto-pick a difficulty ladder of N matched fns") + ap.add_argument("--funcs", help="comma list of func_ to test instead of auto-pick") + ap.add_argument("--seed", type=int, default=11) + a = ap.parse_args() + os.chdir(REPO) + + # choose the ladder + if a.funcs: + addrs = [f.replace("func_", "") for f in a.funcs.split(",")] + else: + bodies = macro_bodies() + cand = [(addr, n) for addr, n in bodies.items() if instantiated_in_src(addr)] + # ladder: spread across size buckets + random.seed(a.seed) + buckets = {"tiny(1-4)": [], "easy(5-12)": [], "med(13-30)": [], "big(31+)": []} + for addr, n in cand: + if n <= 4: buckets["tiny(1-4)"].append((addr, n)) + elif n <= 12: buckets["easy(5-12)"].append((addr, n)) + elif n <= 30: buckets["med(13-30)"].append((addr, n)) + else: buckets["big(31+)"].append((addr, n)) + per = max(1, a.pick // 4) + addrs = [] + for b in buckets.values(): + random.shuffle(b) + addrs += [addr for addr, n in b[:per]] + funcs = [f"func_{addr}" for addr in addrs] + print(f"ladder ({len(funcs)} known-answer fns): {funcs}") + + # backup, then stub the chosen fns in the overlay .c + src = open(SRC).read() + open(BAK, "w").write(src) + stubbed = [] + for fn in funcs: + addr = fn.replace("func_", "") + pat = rf"^[ \t]*DEFINE_func_{addr}\(\).*$" + repl = f'INCLUDE_ASM("{ASM}", {fn});' + new, k = re.subn(pat, repl, src, flags=re.M) + if k: + src = new; stubbed.append(fn) + open(SRC, "w").write(src) + print(f"stubbed {len(stubbed)} instantiations; re-extracting asm...") + + try: + r = sh(["make", "extract", f"BINARY={OV}"]) + if r.returncode: + print("EXTRACT FAIL:\n" + r.stderr[-800:]); return + # per-fn: m2c -> draft -> match_one score + results = [] + for fn in stubbed: + s = f"{ASM}/{fn}.s" + if not os.path.exists(s): + results.append((fn, "no-asm", None)); continue + out = sh([".venv/bin/python", "tools/m2c/m2c.py", "-t", "mipsel-gcc-c", + "--valid-syntax", "-f", fn, s]).stdout + if not out.strip() or "OSError" in out: + results.append((fn, "m2c-empty", None)); continue + if re.search(r"M2C_ERROR|M2C_BREAK|MULT_HI|\bCLZ\b|M2C_TRAP|GLUE_F64|BSWAP", out): + results.append((fn, "nonfaithful(defer)", None)); continue + dpath = f".run/p16_ka/{fn}.c" + os.makedirs(".run/p16_ka", exist_ok=True) + open(dpath, "w").write(out) + mo = sh([".venv/bin/python", "tools/match_one.py", fn, "--c", dpath, + "--asm-subdir", ASM]).stdout + first = mo.strip().splitlines()[0] if mo.strip() else "?" + if first.startswith("MATCH"): + results.append((fn, "M2C-DIRECT-MATCH", 0)) + elif "mismatched" in first: + mm = re.search(r"(\d+) mismatched", first) + results.append((fn, "near-miss", int(mm.group(1)) if mm else -1)) + else: + results.append((fn, first[:40], -1)) + print("\n=== KNOWN-ANSWER CALIBRATION (m2c direct, no permuter yet) ===") + direct = 0 + for fn, status, score in results: + print(f" {fn:20} {status:22} {'mismatch='+str(score) if score and score>0 else ''}") + if status == "M2C-DIRECT-MATCH": direct += 1 + print(f"\n m2c-DIRECT re-derivation: {direct}/{len(stubbed)} (the rest are permuter candidates)") + json.dump([{"fn": f, "status": s, "score": sc} for f, s, sc in results], + open(".run/p16_ka_results.json", "w"), indent=1) + finally: + # ALWAYS restore the matched .c (never leave a regression) + open(SRC, "w").write(open(BAK).read()) + print(f"\nrestored {SRC} from backup.") + + +if __name__ == "__main__": + main()