diff --git a/docs/SETUP.md b/docs/SETUP.md index 85b6ba5387..056c59564e 100644 --- a/docs/SETUP.md +++ b/docs/SETUP.md @@ -781,6 +781,7 @@ Every script under `tools/` (plus the two report make-targets), grouped by purpo | | **S69 tooling — the TRIAGE LADDER** (P31 S69) | **`tools/triage_ladder.py` (NEW)** — the zero-token pass that answers *"does this target need an agent at all?"* before one is spent. **The PRE/POST split is the point:** `--pre ` runs the TARGET-SIDE tiers only (BANKED · WALL-332 · PARKED) — no draft, no build, milliseconds — and writes `/triage.json` + `triage_exclude.txt`; `--post` runs the full `residual_rules_b` residual routing, which needs a draft and runs `match_one`. `--escalate B:FN` exits 2 on a walled/banked target (the check S68 was missing when it escalated a §332 wall at closeness 8). `--acceptance` is the R39/R32 harness: false-skip over EVERY open stub, recall over sampled matched fns, and a wall negative control over already-banked code — all pure filesystem work, so it runs in seconds over the whole corpus. **It REFUSES on a non-quiescent tree** (`pgrep -af` rows for gate/build lanes + `lane_inflight`): a live gate makes the stub oracle transiently wrong in both directions (§377). Wired into `wave_args.py` (drops walled/parked targets at draw time, reusing `pre_classify` — one implementation, R33) and into `tools/workflows/escalate_fable.js`, which now REFUSES any target that does not carry `triage:'DRAFT'`. **Routing correction (§376):** the `INTEG-STANDALONE-MATCH` / `NOCOMPILE-UNDECLARED-*` tiers are GATE-FIRST candidates, never free banks — S69 gated that class **0/28** raw; the route is `fix_arity_callers --any-proto` then the gate. | | | **S69 tooling — the §378 SELF-CALLER CAST** (P31 S69) | **`tools/cast_self_callers.py` (NEW)** — the mirror of `cast_call_sites.py`. That one fixes the DRAFT calling a conflicting CALLEE; this one fixes the TU's OWN already-banked code calling the function the draft is about to DEFINE, which nothing handled and which is the terminal blocker of the §376 pile. **Never run it alone** — it answers the error that `fix_arity_callers --any-proto` CREATES: no-protoing the conflicting decl makes the draft's definition the prototype in scope, so the TU's own call fails anew with `too few arguments`. Order is `arity -> self-cast -> [--sync-decls] -> gate`. Byte-neutral because gcc-2.7.2 folds a cast of a known function symbol back to a direct `jal`. `--sync-decls` covers the narrow-param case a no-proto decl CANNOT legally reach (C89 requires promotion-stable parameter types when one declaration has no prototype, so `void f(s16)` is illegal against `extern void f();` — which is exactly why `fix_arity_callers` skips it as 'narrow-param'); safe only once the call sites are cast, because a declaration then emits no code. It REFUSES a function whose return type it cannot read off the draft (R43). **Journals every edit, and `--undo-journal --keep ` after the gate is MANDATORY** — a leftover cast made `ov_SC07_000` stop compiling and every later gate verdict on it measured a broken baseline. Wired into `recover_integration.py --stages arity,self-cast` (tier `binary`), prescribed by `residual_rules_b`'s decl-conflict tiers, and in the playbook §4b. Measured: **8 banked of 28** including `main/func_80036D58` at zero agent tokens; generalises to the callee the diagnostic NAMES (banked `main/func_80021D38` that way). | | | **S69 tooling — the NEAR-TWIN BAND** (P31 S69) | **`tools/seed_ref.py --near [--max-d N] [--near-control]` (WIDENED, not a new tool — R33)** — the fleet-wide twin oracle gained an edit-distance tier over reloc-normalized instruction streams, because the exact-hash tier answers only *"is there a byte-identical copy?"* while a frontier needs *"is there anything CLOSE?"*. **Measured widening: 22 of 352 reachable open stubs had a d=0 hash twin; 75 of 352 (21%) have a banked match at d<=25 — 3.4x.** Root cause of the gap is §389: `h_norm`'s normalizer drops its pending lui-hi on an intervening R-type, so indexed-global reloc twins hash differently and are invisible to seed_ref/twin_sweep/dedup/family-maps simultaneously. **Do NOT fix `h_norm`** — every stored map and calibration keys on it; the near tier reads through the hole. Verification built in: sound prefilters that cannot lose a true pair, R32 population assertion, R34 cross-check reproducing all 22 exact twins every run, R39 controls (positive 200/200 at d=0; random-pair base rate 1.17%). Classes emitted: HASH-TWIN · RELOC-ONLY (mechanical — remap via `family_remap` and gate; **8 of 10 banked at ~0 agent tokens on first use**) · NEAR-COUSIN (seeded crack). Known remaining gap: `family_sweep.load_sigs` globs `ov_` only, so 38 md_ + 5 resident + 67 main are structurally invisible to it (a 407-ins md_SC05_026 PURE twin of banked ov_MAIN_012 code was found in the wild). Full audit: `.run/S69_fable/report.md`. | +| | **S69 tooling — the CONTAINED tier + the twin LADDER** (P31 S69) | **`tools/seed_ref.py --contained [--contained-control]` (WIDENED again, R33)** — finds an open stub that is a banked body **plus or minus whole blocks** (any gap size), the class edit-distance ranks badly. Needs branch-offset masking (unmasked offsets veto exactly the target pairs) and a min-side-25 floor (89% of raw hits were prologue/epilogue vacuity). **Ranks by (substitutions+regions, cover), NOT by d** — §390's law: a deletion is free, a substitution is thought; the near tier ranked a d=5 substitution twin ABOVE a same-C-minus-one-statement pair that banked at closeness 0. Verified: planted-deletion positive control 60/60, random-pair base rate 0/397, R32 population assert 346/346, `--near` regression reproduces the stored slice exactly (352/22/28). **Use the ladder in playbook §2a-2** (exact → RELOC-ONLY → CONTAINED → cousin → cold), filtering lookalikes at `r = d/min(nins) >= ~0.3`. **Yield: 9 usable stubs, 1 banked.** And the standing conclusion: three fleet-wide probes past RELOC-ONLY returned 0 new / 9 / 2 — **the scanner well is dry; spend integration effort (§376/§378), not scanner effort.** Audit: `.run/S69_fable2/report.md`. | | | **S68 fixes — six instances of the overlay-layout assumption** (cookbook **§363**) | `dedup_propagate` could not even IMPORT (`os.` at module level in the one module that imports `os as _os`). `seed_ref` offered main's LINKED-subseg DEAD TEXT as bankable twins (43 of 82 hits — a draft there gates GREEN while wrong); now refuses and COUNTS the refusal. `parallel_gate.stage_generated` hard-coded `build//.ld`; now asks the Makefile for `_LD_SCRIPT`/`_UNDEF_SYMS`/`_UNDEF_FUNCS` and REFUSES when absent. `rtu_match` gained **`--tu`** (+ `blocker_probe` passes `stub.path` and `stub.asm_dir`) — it reconstructed `src//.c`, which is the overlay layout; main's sources are LOOSE FILES in `src/`. `gate_stage` no longer synthesises `--out`/`--good-sha` — for main those resolved to a nonexistent path and then **ov_SC01_077's SHA** via `DEF_SHA`. **`psyq_integrate`**: the `*_externals.ld` map is now MONOTONIC — it was re-derived against the CURRENT `.ld`, so `firstfile = 0x80061FA8;` was DROPPED on every incremental relink and main was 2 bytes red before any draft was spliced (**the true identity of the 2026-08-15 'main link defect'**). | | | **S68 — the module-binary -O0 carve route** (cookbook **§371**) | `jr_isolate_all` + `overlay_src_split` + the **Makefile -O0 glob widened to `src/md_*/md_*_o0?.c`** open carving for the single-object `md_*` binaries. Three stacked causes behind one `unaddressable content` message (interior-YAML-comment symbol-list truncation; a trailing verbatim-asm chunk with no region; bare tag forward decls), then the **spimdisasm rodata-migration trap**: migrated rodata follows its function ONLY within the same subseg, so a carve silently drops it and `INCLUDE_RODATA` cannot bring it back — rename the `.rodata` subseg to the object its emitters moved to. **The Makefile hunk MUST be committed with the carve** or a fresh clone loses -O0 on the region and every draft banked there mystery-fails. | | | `tools/recover_rejects.py` | **(P31 S59)** Free recovery of PRE-GATE rejects, wired into the maintenance lane. Two paths exist for a draft that does not bank and only one was recorded: a gate failure gets a backlog row (closeness/class/best draft), while a draft the reloc pre-filter drops reached nothing — **569 of 1,261 drafts over eight waves, 45%**. Of the `MISMATCH?` rejects, **13% carry `shape: MATCH`** — right body, wrong symbol names, i.e. the §171 stale-seed class `aprop_symfix` rebases deterministically. Reads `.run/reloc_rejects.jsonl` (written by `ox_campaign.reloc_filter`), keeps shape-MATCH rows that are STILL open stubs, runs `aprop_symfix --fix`, and STAGES the rebased bodies into `.run/sweep_maint//` for the lane's existing free gate. It never substitutes, gates or commits — a bad recovery can waste a build, never a bank. Tried-once is remembered in `.run/recover_rejects_seen.json`. Zero model tokens. | diff --git a/docs/accelerators.md b/docs/accelerators.md index 921e9d23f5..1940cb09cc 100644 --- a/docs/accelerators.md +++ b/docs/accelerators.md @@ -422,3 +422,38 @@ every function you crack becomes an exemplar for everything within a few instruc immediately. Build it late and you accumulate invisible-singleton debt that costs a whole session to recover — and you will never know how much you left on the floor, because the tool reports a confident, true, useless number. + +## #18 — A CLAIM DERIVED FROM BYTES IS NOT A CLAIM VERIFIED BY A COMPILER (P31 S69) + +**What happened.** A tooling agent reported an open function as "= banked twin minus its final +statement — **resid 0**", listed under "mechanically bankable". Read naturally, `resid 0` means *it +compiles to the target*. It did not: the agent had aligned the two BYTE STREAMS, observed one +contiguous 6-instruction block absent and zero other token differences, and **had compiled nothing**. +Challenged, it said so immediately and cleanly: *"my 'resid 0' was a byte-stream containment fact, not +a compiled draft."* It then produced the draft and the real verification — `{"status": "match", +"closeness": 0, "nins": 98}`. The prediction was correct. **The claim's TYPE was not.** + +**Why this is its own accelerator and not just a wording nit.** Every decomp pipeline mixes claim +types that read identically in a report: + +| claim type | what it proves | what it does not | +|---|---|---| +| stream/hash containment | the bytes relate | that any C produces them | +| compiled standalone (`match_one`) | the BODY is right | that the TU accepts the signature (§376) | +| whole-binary gate green | this binary is byte-identical | anything about the other 212 | +| clean-fleet R22 | the fleet is green NOW | that a config change was re-extracted (§384) | + +A report that says "resid 0" or "verified" without naming which tier it reached invites the reader to +assume the strongest one. Downstream that becomes a bank attempt against a draft that does not exist, +or — worse — a "free win" ledger entry nobody re-checks. + +**The standing rule: every similarity or correctness claim names the tier it reached.** "Contained at +d=6 (bytes, uncompiled)" and "MATCH closeness 0 (compiled standalone)" are different sentences and +should look different. Ask any agent that reports a match: *what command produced that number?* If the +answer is a stream comparison, the work is a PREDICTION — valuable, rankable, not bankable. + +**The corollary that saved this one:** the reader could not reproduce the number, said so plainly +rather than passing it along, and asked for the file and the literal command. The agent then +self-corrected AND diagnosed the reader's failed repro to the instruction (an invented byte-aligned +type, §391). Non-reproduction is a finding; treat it as one instead of assuming your own setup is at +fault. diff --git a/docs/cookbook-index.md b/docs/cookbook-index.md index 8adff57461..d57e8ad2d6 100644 --- a/docs/cookbook-index.md +++ b/docs/cookbook-index.md @@ -2,7 +2,7 @@ > **Generated by `tools/cookbook_index.py` — do not hand-edit** (R33). Regenerate after adding a cookbook section. > -> `docs/matching-cookbook.md` is ~716 KB / 1044 sections. Grepping it blind is how three P30 wave-1 agents each "discovered" an idiom that was already written down. **Start here, then read the section.** A section appears under every symptom it addresses. +> `docs/matching-cookbook.md` is ~716 KB / 1046 sections. Grepping it blind is how three P30 wave-1 agents each "discovered" an idiom that was already written down. **Start here, then read the section.** A section appears under every symptom it addresses. **How to use:** name what you SEE in the diff (a stolen delay slot, an extra `la`, a swapped register pair, a `conflicting types` error), find that symptom below, read those sections first. If nothing fits, THEN grind — and add a section when you win. @@ -413,7 +413,7 @@ - **§385** — ★★★ — THE **SCHED2 PRIORITY-DONOR ASM**: closing the "hoisted-invariant vs IV-init preheader swap" class (P31 S69; byte-proven main/func_80038A58, 347 ins, fable escalation 2 → 0) L32309 - **§386** — ★★★ — A BYTE LOAD ON THE **BIV** BASE WAS BORN IN THE COMBINE PASS: SPELL IT AS A SHIFT-MASK, NEVER A DEREF (P31 S69; byte-proven main/func_80020598, 292 ins, escalation 1 → 0) L32335 -### structs, block moves & memcpy (84) +### structs, block moves & memcpy (85) - **§3-T2** — Source statement order drives instruction scheduling L78 - **§5** — Known hard-residual classes (instruction-identical, one byte-exact blocker) L199 @@ -499,6 +499,7 @@ - **§357** — ONE STRUCT POINTER, NOT TWO: A SECOND SOURCE VARIABLE BUILDS A THIRD IV (P31 S68; byte-proven ov_SC06_029/func_80181DF8, 335 ins, 330 → 13) L31589 - **§364** — ★ — THE libgpu `P_TAG` BITFIELD SPELLING IS **OPT-LEVEL DEPENDENT** (P31 S68; two functions, opposite verdicts, same session) L31734 - **§379** — ★★★ — **MEM_IN_STRUCT_P**: THE SAME LOAD, WRITTEN AS A STRUCT MEMBER, SCHEDULES WHERE A CAST CANNOT (P31 S69; byte-proven main/func_80021284 220 ins and main/func_8002D904 217 ins, found INDEPENDENTLY by two agents) L32164 +- **§391** — ★★ — A BYTE-ALIGNED STRUCT COPIES IN FOUR INSTRUCTIONS, A WORD-ALIGNED ONE IN TWO (P31 S69) L32488 ### types, signedness & load/store width (92) @@ -1305,7 +1306,7 @@ - **§376** — ★★★ — A STANDALONE `match_one` CLOSENESS OF 0 IS A CLAIM ABOUT THE **BODY**, NEVER ABOUT THE **TU** (P31 S69; measured 0/28) L32055 - **§384** — ★★★ — A CARVE-CONFIG BANK IS RED UNTIL YOU RE-EXTRACT, AND THAT LOOKS EXACTLY LIKE A FALSE BANK (P31 S69; measured twice, cost one destroyed match) L32259 -### (unbucketed — title matched no symptom vocabulary) (308) +### (unbucketed — title matched no symptom vocabulary) (309) - **§3-How** — to use this L30 - **§1** — Idiom catalog (asm pattern → C that produces it) L39 @@ -1615,6 +1616,7 @@ - **§380** — ★★★ — **A SECOND SET OF A PSEUDO DISQUALIFIES IT FROM `move_movables`** (P31 S69; main/func_800215F4, 465 ins, closeness 106 → 59 → 39) L32195 - **§382** — TWO FOLD REASSOCIATIONS THAT NEED THEIR OWN STATEMENT (P31 S69) L32232 - **§389** — ★★★ — `h_norm` IS BLIND TO INDEXED-GLOBAL RELOCS, SO FREE WORK BECOMES AN INVISIBLE SINGLETON (P31 S69; 31 stubs / 4,811 ins recovered, 8 banked same day) L32409 +- **§390** — ★★★ — MINIMUM DISTANCE IS NOT MINIMUM WORK; RANK TWIN CANDIDATES BY EFFORT, AND FILTER LOOKALIKES BY RATIO (P31 S69; byte-proven ov_SC01_077/func_80184D50, banked) L32450 ## All sections, in order @@ -2663,6 +2665,8 @@ - **§387** — ★★ — **SPLIT-FOLD DISPATCH CLOBBER**: one switch case needs a reload, another must keep the fold (P31 S69; byte-proven main/func_80030F80, 343 ins, escalation 3 → 0) L32361 - **§388** — ★★★ — THE **-O0 COLOURING ORACLE**: simulate `stupid.c` instead of grinding spellings (P31 S69; main/func_80011380 proved a C-level WALL at 6) L32384 - **§389** — ★★★ — `h_norm` IS BLIND TO INDEXED-GLOBAL RELOCS, SO FREE WORK BECOMES AN INVISIBLE SINGLETON (P31 S69; 31 stubs / 4,811 ins recovered, 8 banked same day) L32409 +- **§390** — ★★★ — MINIMUM DISTANCE IS NOT MINIMUM WORK; RANK TWIN CANDIDATES BY EFFORT, AND FILTER LOOKALIKES BY RATIO (P31 S69; byte-proven ov_SC01_077/func_80184D50, banked) L32450 +- **§391** — ★★ — A BYTE-ALIGNED STRUCT COPIES IN FOUR INSTRUCTIONS, A WORD-ALIGNED ONE IN TWO (P31 S69) L32488 --- @@ -3719,3 +3723,5 @@ Notes routinely quote that as a section id. This table resolves it. Grep bait: ` | L32361 | §387 | ★★ — **SPLIT-FOLD DISPATCH CLOBBER**: one switch case needs a reload, another must keep th | | L32384 | §388 | ★★★ — THE **-O0 COLOURING ORACLE**: simulate `stupid.c` instead of grinding spellings (P31 | | L32409 | §389 | ★★★ — `h_norm` IS BLIND TO INDEXED-GLOBAL RELOCS, SO FREE WORK BECOMES AN INVISIBLE SINGLE | +| L32450 | §390 | ★★★ — MINIMUM DISTANCE IS NOT MINIMUM WORK; RANK TWIN CANDIDATES BY EFFORT, AND FILTER LOO | +| L32488 | §391 | ★★ — A BYTE-ALIGNED STRUCT COPIES IN FOUR INSTRUCTIONS, A WORD-ALIGNED ONE IN TWO (P31 S69 | diff --git a/docs/generic-decomp-package.md b/docs/generic-decomp-package.md index 2d8d8e9883..3764320469 100644 --- a/docs/generic-decomp-package.md +++ b/docs/generic-decomp-package.md @@ -63,6 +63,19 @@ against random pairs for the base rate (R39: 1.17% here). **And audit every hash you own for BOTH questions.** Dedup wants under-matching; a frontier join wants over-matching. One hash cannot serve both error directions, and the failure is silent. +**Rank the band by WORK, not by distance, and stop building scanners once it is dry.** Two findings +that cost a session here and are free to inherit: + +* *A deletion is free, a substitution is thought* (§390). Edit distance ranked a 5-substitution twin + above a pair that was the same C minus one trailing statement — the second banked at closeness 0. + Order candidates by (substitutions + regions, coverage); use distance only as a filter. And filter + LOOKALIKES at `r = d/min(nins) >= ~0.3`: 17 of 30 "cousins" here were two different functions + sharing boilerplate, and a wrong twin is worse than no twin because the agent believes it. +* *Know when to stop.* After the reloc-only class, three fleet-wide probes returned **0 new / 9 / 2**. + The similarity well runs dry fast. In the same session the INTEGRATION levers — making an + already-correct body compile inside its real translation unit (§376/§378) — banked an order of + magnitude more. **Budget accordingly: scanners early, integration forever.** + **3. The differential-oracle harness (accelerators #15) — the one that works at 0%.** Two independent paths per question, disagreement fails loudly, on a schedule. diff --git a/docs/matching-cookbook.md b/docs/matching-cookbook.md index 5e06a5ae42..bfaf8ace37 100644 --- a/docs/matching-cookbook.md +++ b/docs/matching-cookbook.md @@ -32446,3 +32446,55 @@ error directions. **Diff tell:** an open stub your card calls "no banked twin — derive from the .s", whose body is a per-location copy of engine code that exists in a sibling overlay. Run the near tier before believing a singleton verdict. + +## §390 ★★★ — MINIMUM DISTANCE IS NOT MINIMUM WORK; RANK TWIN CANDIDATES BY EFFORT, AND FILTER LOOKALIKES BY RATIO (P31 S69; byte-proven ov_SC01_077/func_80184D50, banked) + +**The measurement.** The near tier offered two banked candidates for one open stub: + +| candidate | distance | what it costs a human | +|---|---|---| +| `ov_SC03_006:0x8018b868` | **d=5** — 5 scattered SUBSTITUTIONS | rewrite an expression: real thought | +| `ov_SC03_007:func_8018283C` | **d≈6** — ONE contiguous block absent | copy the C, **delete one statement** | + +Edit distance ranked the *substitution* twin first, because 5 < 6. The deletion twin was the strictly +cheaper answer and it banked: copy `func_8018283C`'s body, drop its trailing +`*(s32 *)(...) &= 0x7FFFFFFF;`, rename, `match_one` -> **MATCH, closeness 0, 98/98 ins**. + +**The law.** *A deletion is free and a substitution is thought.* Distance counts moved instructions; +it does not count the reasoning needed to convert one body into another. Rank by +**(substitutions + regions, coverage)**, never by raw `d` — `seed_ref --contained` does; `--near` does +not, by design (it is the recall tier). + +**The companion filter — LOOKALIKES.** A widened band promotes coincidence. Of 30 NEAR-COUSIN rows, +only **13 were true cousins**; the other **17 were boilerplate lookalikes**. They separate cleanly on + + r = d / min(nins) true cousins r <= 0.27 lookalikes r >= 0.37 + +Anything at r >= ~0.3 is two functions that merely share prologue/epilogue/dispatch shape. Never +send one to an agent as a "twin" — a wrong twin is worse than no twin, because the agent trusts it. + +**THE MECHANICAL FRONTIER PAST RELOC-ONLY IS SINGLE DIGITS — three fleet-wide nulls, all controlled:** + +| probe | scope | result | +|---|---|---| +| skeleton join (opcodes+registers, immediates dropped) | 91,638 banked × 352 open | **0 new** — hits are exactly the 46 known same-length twins. The "same shape, different constants" class DOES NOT EXIST in this corpus | +| contained / block-indel (± whole blocks, any gap) | fleet | **9 usable**, 4 previously invisible; 89% of raw hits were prologue/epilogue vacuity until a min-side-25 floor was applied | +| past-the-cap cousins (d26-45, ≤30% drift, ≥100 ins) | 178 pairs | **2** — both replace-heavy seeded cracks | + +**Conclusion, and it is a spending decision:** after RELOC-ONLY (§389) the scanner well is dry. +Further similarity tooling buys single digits; the integration levers (§376/§378) bought dozens in +the same session. **Spend integration effort, not scanner effort, from here.** + +## §391 ★★ — A BYTE-ALIGNED STRUCT COPIES IN FOUR INSTRUCTIONS, A WORD-ALIGNED ONE IN TWO (P31 S69) + +`typedef struct { u8 b[8]; } Blk8;` (`src/shared/engine_types.h:497`) is BYTE-aligned, so gcc-2.7.2 +cannot assume word alignment and emits the unaligned quartet **`lwl / lwr / swl / swr`** per copy. +A word-aligned aggregate of the same size (`struct { int a, b; }`) emits **`lw / sw`**. + +**Diff tell: an exact multiple of 2 instructions missing, scaling with the number of struct +assignments.** Measured while reproducing a twin: substituting an invented word-aligned `Blk8` for +the real byte-aligned one lost exactly 8 instructions across two copies (90 vs the target's 98) and +read as a plausible "near, closeness 70" — a wrong TYPE masquerading as a codegen residual. + +**So: never invent an aggregate type to make a draft compile.** Resolve it from +`src/shared/engine_types.h`. An invented type does not fail loudly; it fails as a believable diff. diff --git a/docs/wave-playbook.md b/docs/wave-playbook.md index 4989a40da0..a535f8e89b 100644 --- a/docs/wave-playbook.md +++ b/docs/wave-playbook.md @@ -112,6 +112,29 @@ bodies hash differently and read as singletons. A `RELOC-ONLY` row is mechanical banked exemplar onto the open address, then gate — **8 of 10 banked at ~0 agent tokens on first use**, one 94-ins exemplar serving five open copies. Never send a RELOC-ONLY row to a drafting agent. +### 2a-2. THE TWIN LADDER — take the CHEAPEST tier available, never the closest number (S69) + +Distance is a FILTER, not the ranking key: a deletion is free and a substitution is thought (§390). +Work down this ladder and stop at the first tier that has a row; only widen when the tier above is +empty. We never "go straight to d25" — widening only lets tiers 2-3 SEE candidates that a d=0-only +tool called singletons. + +| tier | detector | cost | measured S69 | +|---|---|---|---| +| 1 exact twin (d=0) | `seed_ref` hash | free — copy the body verbatim | 22 rows | +| 2 RELOC-ONLY (any d) | `seed_ref --near` + `family_remap.classify_member` | mechanical — remap, gate | 31 rows, **8 banked, ~0 tokens** | +| 3 CONTAINED (± whole block) | `seed_ref --contained` | near-mechanical — delete/add statements | 9 usable, **1 banked** | +| 4 true cousin (few substitutions) | `--near`, ratio `r = d/min(nins) <= 0.27` | seeded crack — a cheap agent holding the twin's C | 13 rows | +| 5 no match | — | cold draft, full price | 277 of 352 | + +**Filter lookalikes before handing anything to an agent.** At `r >= ~0.37` the "twin" is two +different functions sharing boilerplate — 17 of 30 NEAR-COUSIN rows were exactly that. A wrong twin +is worse than no twin, because the agent believes it. + +**And do not build more scanners.** Three fleet-wide probes past RELOC-ONLY returned 0 new / 9 / 2 +(§390). The scanner well is dry; the integration levers (§376/§378) out-earned it by an order of +magnitude in the same session. + ### 2b. RUN `neighbor_ref` FOR EVERY CARD — the biggest measured cost lever in the wave ``` diff --git a/tools/seed_ref.py b/tools/seed_ref.py index bd954cc2b0..718c29258c 100644 --- a/tools/seed_ref.py +++ b/tools/seed_ref.py @@ -34,11 +34,24 @@ disables an ENTIRE binary — measured: `main`, where 765 of 1048 stubs carry cu tools/seed_ref.py --binary ov_SC03_107 --fn func_8013DD68 tools/seed_ref.py --all --json .run/seed_refs.json + +NEAR TIER (P31 S69 Fable audit — venv-only, see the block below `for_stub`): + .venv/bin/python tools/seed_ref.py --near [--max-d 5] [--json PATH] + .venv/bin/python tools/seed_ref.py --near-control + +CONTAINED TIER (P31 S69 Fable-2 audit — venv-only, see the block above `contained_scan`): +banked twins 1..4 whole statement-blocks away at ANY gap size (the near tier's cap conflates a +46-ins inserted block with "far"). Advisory seeding only — deletion-direction rows are +mechanically decidable, insertion-direction rows carry the block's exact asm for a seeded crack. + .venv/bin/python tools/seed_ref.py --contained [--min-cover 0.8] [--json PATH] + .venv/bin/python tools/seed_ref.py --contained-control """ import argparse +import collections import functools import json import os +import struct import sys HERE = os.path.dirname(os.path.abspath(__file__)) @@ -177,6 +190,654 @@ def for_stub(binary, fn, include_linked=False): return None +# -------------------------------------------------------------------------------------------- +# NEAR TIER (P31 S69 Fable audit, 2026-09-01) — banked twins at edit distance 1..K, not only +# hash equality. +# -------------------------------------------------------------------------------------------- +# WHY (byte-witnessed; full audit in .run/S69_fable/report.md). `h_norm`'s hi/lo tracker +# (sig_image.norm_stream) DISCARDS the pending lui-hi when an R-type intervenes, so the +# indexed-global triad +# lui $at, %hi(arr); addu $at, $at, idx; lw $r, %lo(arr)($at) +# keeps its %lo IMMEDIATE unmasked in the hash. Two per-overlay copies of the same function +# whose only difference is that array's ADDRESS therefore hash to DIFFERENT h_norm, and every +# h_norm-keyed consumer (for_stub above, twin_sweep, dup_report, the family maps' structural +# tier, config/dedup.us.yaml) reads them as unrelated singletons. Measured 2026-09-01 over the +# 360 reachable open stubs: 22 had a d=0 (hash) twin, and 36 MORE had a banked twin at distance +# 1..5 — 31 of them RELOC-ONLY (every differing word an address-reloc slot under +# family_remap.reloc_indices, which DOES propagate hi through add/addu — the wider, proven +# model behind the 76-88% mechanical-remap rates). Example: `func_80185F6C` is banked in +# ov_SC06_033 and its body serves FIVE open 94-ins copies whose single differing word is +# `lw $a1, %lo(D_xxxx)($at)` — one extern rename each. +# +# THE METRIC: banded Levenshtein over norm_stream 32-bit tokens (registers and true immediates +# KEPT), so d==0 IS h_norm equality — asserted against for_stub on every run (R34/R39: the near +# tier must reproduce every d=0 the hash oracle finds, or die). Prefilters are SOUND lower +# bounds (|dnins| <= K; opclass-histogram L1 <= 2K; token-multiset L1 <= 2K): no true d<=K pair +# can be excluded by them (R32). h_norm ITSELF is deliberately left unchanged — every stored +# map, ledger and calibration keys on it; this tier widens the JOIN, not the hash. +# +# Verdicts per row: RELOC-ONLY (mechanical-remap candidate — route to family_remap; expect the +# h_norm-class 76-88% rate, minus jtbl carve blockers) vs NEAR-COUSIN (real code drift <= K — +# seeded-crack card fuel: copy the twin's C, expect exactly the listed words to differ). + +def _near_mods(): + """Lazy venv-only deps (sig_image needs rabbitizer). Imported here, not at module top, so + the card pipeline's plain-python3 `import seed_ref` keeps working; --near under the wrong + interpreter refuses loudly instead of degrading the metric (R43).""" + try: + import sig_image + except ModuleNotFoundError as e: + raise SystemExit("seed_ref --near needs rabbitizer (via sig_image): run under " + ".venv/bin/python. Refusing rather than degrading (R43).") from e + import family_remap + return sig_image, family_remap + + +def _norm_tokens(SI, words): + raw = struct.pack("<%dI" % len(words), *words) + nb = SI.norm_stream(raw) + return struct.unpack("<%dI" % (len(nb) // 4), nb) + + +def _words_of(FR, b, a, n): + data = FR._img(b) + off = a - FR.vram_of(b) + if off < 0 or off + 4 * n > len(data): + return None + return list(struct.unpack_from("<%dI" % n, data, off)) + + +def _opclass_hist(words): + h = collections.Counter() + for w in words: + op = w >> 26 + h[(op, w & 0x3F if op == 0 else ((w >> 16) & 0x1F if op == 1 else 0))] += 1 + return h + + +def _l1(c1, c2): + d = 0 + for k, v in c1.items(): + d += abs(v - c2.get(k, 0)) + for k, v in c2.items(): + if k not in c1: + d += v + return d + + +def _mid_jr(words): + return any((w >> 26) == 0 and (w & 0x3F) == 0x08 and ((w >> 21) & 0x1F) != 31 for w in words) + + +def _lev(a, b, cut): + """Levenshtein over token sequences, capped: returns min(d, cut+1). Common-affix trim + + banded DP with an early bail when a whole band row exceeds the cap.""" + la, lb = len(a), len(b) + if abs(la - lb) > cut: + return cut + 1 + i = 0 + while i < la and i < lb and a[i] == b[i]: + i += 1 + j = 0 + while j < la - i and j < lb - i and a[la - 1 - j] == b[lb - 1 - j]: + j += 1 + a = a[i:la - j]; b = b[i:lb - j] + la, lb = len(a), len(b) + if la == 0: + return lb + if lb == 0: + return la + if la > lb: + a, b, la, lb = b, a, lb, la + INF = cut + 1 + prev = list(range(la + 1)) + for jj in range(1, lb + 1): + lo = max(1, jj - cut); hi = min(la, jj + cut) + cur = [INF] * (la + 1) + if lo == 1: + cur[0] = jj if jj <= cut else INF + bj = b[jj - 1]; best = INF + for ii in range(lo, hi + 1): + c = prev[ii] + 1 + if cur[ii - 1] + 1 < c: + c = cur[ii - 1] + 1 + c3 = prev[ii - 1] + (a[ii - 1] != bj) + if c3 < c: + c = c3 + cur[ii] = c + if c < best: + best = c + if best > cut: + return cut + 1 + prev = cur + return prev[la] if prev[la] <= cut else cut + 1 + + +def _classify_pair(FR, wo, wb, to, tb, show=6): + """RELOC-ONLY iff every differing token sits in an address-reloc slot on BOTH sides under + family_remap.reloc_indices (the addu-propagating model); else NEAR-COUSIN.""" + import difflib + ro, rb = FR.reloc_indices(wo), FR.reloc_indices(wb) + sm = difflib.SequenceMatcher(a=to, b=tb, autojunk=False) + reloc_only, detail = True, [] + for op, i1, i2, j1, j2 in sm.get_opcodes(): + if op == "equal": + continue + if op != "replace" or (i2 - i1) != (j2 - j1): + reloc_only = False + detail.append("%s open[%d:%d] twin[%d:%d]" % (op, i1, i2, j1, j2)) + continue + for k in range(i2 - i1): + io, ib = i1 + k, j1 + k + if io in ro and ib in rb: + detail.append("reloc@%d" % io) + else: + reloc_only = False + detail.append("code@%d:%08x<->%08x" % (io, wo[io], wb[ib])) + return ("RELOC-ONLY" if reloc_only else "NEAR-COUSIN"), detail[:show] + + +def _main_splat_sig(): + """main's splat-true sig rows ({addr: row}) — the registry sig for main is the (older) + Ghidra dump; .run/sig.main.jsonl carries splat-seeded slices for everything that was still + a stub when `make sig-main` last ran, which includes every CURRENT stub (stubs only shrink).""" + out = {} + p = os.path.join(REPO, ".run/sig.main.jsonl") + if os.path.exists(p): + for ln in open(p): + ln = ln.strip() + if ln: + r = json.loads(ln) + out[int(r["addr"], 16)] = r + return out + + +def _pool_and_open(include_linked=False, caller="near_scan"): + """The shared population + banked pool (one pass per binary; R32 buckets). + Returns (open_rows, banked, linked_n): open_rows = every reachable open stub with its sig + row; banked = {h_exact: (nins, [(bin, addr)])} over every non-IMPORTED matched body (main: + splat-true sig.main rows first, then registry non-stub rows). Extracted verbatim from + near_scan for the --contained tier (S69 Fable-2); the two tiers MUST share one denominator.""" + open_rows, linked_n, oracle_refused = [], 0, [] + banked = {} + + def _add(h, n, bb, aa): + e = banked.setdefault(h, (n, [])) + if len(e[1]) < 3: + e[1].append((bb, aa)) + + msig = _main_splat_sig() + for b in progress.BINARIES: + try: + stubs = corpus.stubs(b) + except Exception as e: + oracle_refused.append((b, repr(e)[:90])) + continue + st_addrs = set(stubs) + if b == "main": + seen = set() + for a2, r in msig.items(): + if a2 in st_addrs: + continue + _add(r["h_exact"], r["nins"], "main", a2); seen.add(a2) + for a2, r in corpus.sig("main").items(): + if a2 in st_addrs or a2 in seen or r.get("src") == "IMPORTED": + continue + _add(r["h_exact"], r["nins"], "main", a2) + else: + for a2, r in corpus.matched(b).items(): + if r.get("src") == "IMPORTED": + continue + _add(r["h_exact"], r["nins"], b, a2) + for s in stubs.values(): + if not include_linked and is_linked_stub(b, s): + linked_n += 1 + continue + row = msig.get(s.addr) if b == "main" else corpus.sig(b).get(s.addr) + if row is None: + raise SystemExit("%s: open stub %s:%s missing from its sig — refusing to " + "scan a population I cannot stream (R32)." % (caller, b, s.symbol)) + open_rows.append(dict(binary=b, addr=s.addr, symbol=s.symbol, nins=row["nins"], + h_exact=row["h_exact"])) + if oracle_refused: + raise SystemExit("%s: corpus oracle refused %s — the denominator would be wrong " + "(R32/R41)." % (caller, oracle_refused)) + return open_rows, banked, linked_n + + +def near_scan(max_d=5, include_linked=False): + """Fleet-wide near-twin scan. Returns {rows, clusters, stats}; every reachable open stub + appears in exactly one row (R32 — the arithmetic is asserted), with its minimum distance to + any banked body (None if > max_d) and, at d<=max_d, the twin + RELOC-ONLY/NEAR-COUSIN + verdict. `clusters` are the open-open connected components at d<=max_d (crack-one-unlock-k).""" + SI, FR = _near_mods() + open_rows, banked, linked_n = _pool_and_open(include_linked, "near_scan") + + # ---- streams + histograms ---- + by_len = collections.defaultdict(list) + bw, unstreamable = {}, 0 + for h, (n, locs) in banked.items(): + for (bb, aa) in locs: + ws = _words_of(FR, bb, aa, n) + if ws is not None: + bw[h] = ws + by_len[n].append(h) + break + else: + unstreamable += 1 + ow = {} + for o in open_rows: + if o["h_exact"] not in ow: + ws = _words_of(FR, o["binary"], o["addr"], o["nins"]) + if ws is None: + raise SystemExit("near_scan: cannot stream open stub %s:%s (R32)." + % (o["binary"], o["symbol"])) + ow[o["h_exact"]] = ws + bh = {h: _opclass_hist(w) for h, w in bw.items()} + oh = {h: _opclass_hist(w) for h, w in ow.items()} + + # ---- sound prefilter -> candidates ---- + cand = {} + for o in open_rows: + n, hx = o["nins"], o["h_exact"] + ho = oh[hx] + cs = [h for m in range(max(1, n - max_d), n + max_d + 1) for h in by_len.get(m, ()) + if _l1(ho, bh[h]) <= 2 * max_d] + cand[(o["binary"], o["addr"])] = cs + + need = set(h for cs in cand.values() for h in cs) + tok = {h: _norm_tokens(SI, bw[h]) for h in need} + for hx, ws in ow.items(): + tok[hx] = _norm_tokens(SI, ws) + cnt = {h: collections.Counter(t) for h, t in tok.items()} + + drawn = set() + dp = os.path.join(REPO, ".run/t5/drawn.json") + if os.path.exists(dp): + try: + drawn = {tuple(k.split(":", 1)) for k in json.load(open(dp))} + except (OSError, ValueError): + pass + + # ---- open vs banked ---- + rows = [] + for o in open_rows: + hx = o["h_exact"] + to, co = tok[hx], cnt[hx] + best, bhx = max_d + 1, None + for h in cand[(o["binary"], o["addr"])]: + if _l1(co, cnt[h]) > 2 * max_d: + continue + d = _lev(to, tok[h], best - 1 if best <= max_d else max_d) + if d < best: + best, bhx = d, h + if best == 0: + break + rec = dict(binary=o["binary"], fn=o["symbol"], addr="0x%08x" % o["addr"], + nins=o["nins"], d=(best if best <= max_d else None), + jtbl=_mid_jr(ow[hx]), drawn=(o["binary"], o["symbol"]) in drawn) + if bhx is not None: + tb, ta = banked[bhx][1][0] + rec["twin"] = "%s:0x%08x" % (tb, ta) + rec["twin_nins"] = banked[bhx][0] + if best > 0: + cls, detail = _classify_pair(FR, ow[hx], bw[bhx], to, tok[bhx]) + rec["cls"], rec["diff"] = cls, detail + else: + rec["cls"] = "HASH-TWIN" + rows.append(rec) + + # ---- open vs open clusters ---- + prs = [] + for i in range(len(open_rows)): + o1 = open_rows[i] + for j in range(i + 1, len(open_rows)): + o2 = open_rows[j] + if abs(o1["nins"] - o2["nins"]) > max_d: + continue + h1, h2 = o1["h_exact"], o2["h_exact"] + if h1 == h2: + prs.append((0, i, j)); continue + if _l1(cnt[h1], cnt[h2]) > 2 * max_d: + continue + d = _lev(tok[h1], tok[h2], max_d) + if d <= max_d: + prs.append((d, i, j)) + par = list(range(len(open_rows))) + + def find(x): + while par[x] != x: + par[x] = par[par[x]]; x = par[x] + return x + for d, i, j in prs: + par[find(i)] = find(j) + comp = collections.defaultdict(list) + for i in range(len(open_rows)): + comp[find(i)].append("%s:%s" % (open_rows[i]["binary"], open_rows[i]["symbol"])) + clusters = sorted((v for v in comp.values() if len(v) > 1), key=len, reverse=True) + + # ---- R34/R39: the near tier must reproduce every d=0 the hash oracle (for_stub) finds ---- + mism, unverifiable = [], [] + dmap = {(r["binary"], r["fn"]): r for r in rows} + for o in open_rows: + sr = for_stub(o["binary"], o["symbol"], include_linked=include_linked) + if not sr: + continue + r = dmap[(o["binary"], o["symbol"])] + if r["d"] == 0: + continue + tb, ta = sr["binary"], int(sr["exemplar_addr"], 16) + ws = _words_of(FR, tb, ta, sr["nins"] or 0) if sr.get("nins") else None + if ws is None: + unverifiable.append((o["binary"], o["symbol"], "twin unstreamable")) + continue + mism.append((o["binary"], o["symbol"], sr["tier"], r["d"])) + if mism: + raise SystemExit("near_scan FAILED its hash-oracle cross-check (R34/R39): for_stub finds " + "a d=0 twin the near tier does not: %s" % mism[:5]) + + n_d0 = sum(1 for r in rows if r["d"] == 0) + n_near = sum(1 for r in rows if r["d"] is not None and r["d"] > 0) + stats = dict(reachable=len(rows), linked_skipped=linked_n, max_d=max_d, + d0=n_d0, near=n_near, + reloc_only=sum(1 for r in rows if r.get("cls") == "RELOC-ONLY"), + near_cousin=sum(1 for r in rows if r.get("cls") == "NEAR-COUSIN"), + no_twin=sum(1 for r in rows if r["d"] is None), + clusters=len(clusters), + cluster_stubs=sum(len(c) for c in clusters), + banked_pool=len(banked), banked_unstreamable=unstreamable, + cross_check="OK (%d hash twins reproduced)" % n_d0, + cross_check_unverifiable=unverifiable) + assert stats["d0"] + stats["near"] + stats["no_twin"] == stats["reachable"], "R32 arithmetic" + return dict(rows=rows, clusters=clusters, stats=stats) + + +def near_control(max_d=5, n_pairs=4000, n_pos=200): + """R39 controls for the near tier. + POSITIVE (must be perfect): pairs of banked bodies with EQUAL h_norm but different h_exact + — the class the hash remap already banks — must all measure d=0. Any other answer means the + metric would refuse work that previously succeeded: FAIL. + BASE RATE (context): random size-matched banked pairs, P(d<=max_d) per size band — how often + the metric fires on arbitrary same-size functions. (The banked corpus contains TRUE twin + families, so this over-states the false-positive rate; treat it as an upper bound.) + + SCOPE: the positive control pairs only rows from sig_is_independent binaries. main's registry + h_norm is GHIDRA's normToken — a different normalizer (sig_image's docstring says so) — so an + equal-Ghidra-h_norm main pair may legitimately measure d>0 here; pairing across hash MODELS + measured the models' disagreement, not this metric (found by this control's own first run: + 2/200 failures, both main, e.g. main:0x80011818 vs 0x80011b7c, Ghidra-equal, sig_image d=6).""" + import random + SI, FR = _near_mods() + random.seed(42) + groups = collections.defaultdict(list) # h_norm -> [(b, addr, nins, h_exact)] + for b in progress.BINARIES: + if not corpus.sig_is_independent(b): + continue # h_norm there is Ghidra's model, not norm_stream's + try: + for a2, r in corpus.matched(b).items(): + if r.get("src") == "IMPORTED" or not r.get("h_norm"): + continue + groups[r["h_norm"]].append((b, a2, r["nins"], r["h_exact"])) + except Exception: + continue + pos = 0; posn = 0 + for hn, mem in groups.items(): + if posn >= n_pos: + break + hexs = {} + for m in mem: + hexs.setdefault(m[3], m) + if len(hexs) < 2: + continue + (b1, a1, n1, _), (b2, a2, n2, _) = list(hexs.values())[:2] + w1, w2 = _words_of(FR, b1, a1, n1), _words_of(FR, b2, a2, n2) + if w1 is None or w2 is None: + continue + posn += 1 + d = _lev(_norm_tokens(SI, w1), _norm_tokens(SI, w2), max_d) + if d == 0: + pos += 1 + print("POSITIVE control (h_norm-equal, h_exact-different banked pairs): %d/%d at d=0" + % (pos, posn)) + if pos != posn: + raise SystemExit("near_control FAILED: %d previously-succeeding twin pairs measure d>0 " + "(R39 zero-false-refusal broken)." % (posn - pos)) + + flat = [m for mem in groups.values() for m in mem] + by_len = collections.defaultdict(list) + for m in flat: + by_len[m[2]].append(m) + band = lambda n: "tiny(<16)" if n < 16 else "small(16-49)" if n < 50 else \ + "mid(50-119)" if n < 120 else "big(>=120)" + ctr = collections.Counter() + tries = 0 + while tries < n_pairs: + m1 = random.choice(flat) + cs = [m for k in range(max(1, m1[2] - max_d), m1[2] + max_d + 1) + for m in by_len.get(k, ()) if m[3] != m1[3]] + if not cs: + continue + m2 = random.choice(cs) + w1, w2 = _words_of(FR, m1[0], m1[1], m1[2]), _words_of(FR, m2[0], m2[1], m2[2]) + if w1 is None or w2 is None: + continue + tries += 1 + d = _lev(_norm_tokens(SI, w1), _norm_tokens(SI, w2), max_d) + ctr[(band(m1[2]), d <= max_d)] += 1 + print("BASE RATE (random size-matched banked pairs, P(d<=%d)) — an UPPER bound on false " + "positives (true twin families inflate it):" % max_d) + for bd in ("tiny(<16)", "small(16-49)", "mid(50-119)", "big(>=120)"): + hit, tot = ctr[(bd, True)], ctr[(bd, True)] + ctr[(bd, False)] + print(" %-13s %4d/%4d = %5.2f%%" % (bd, hit, tot, 100.0 * hit / max(tot, 1))) + + +# -------------------------------------------------------------------------------------------- +# CONTAINED TIER (P31 S69 Fable-2 audit, 2026-09-01) — banked twins one-to-four whole BLOCKS +# away, at ANY total gap size (the near tier's edit-distance cap conflates a 46-ins inserted +# block with "far"; this tier has no cap on gap, only on drift OUTSIDE the blocks). +# -------------------------------------------------------------------------------------------- +# WHY (measured; full audit in .run/S69_fable2/report.md). Levenshtein-K misses the pair +# open = banked body + one contiguous inserted/deleted statement block of >K instructions +# even when EVERYTHING else matches — e.g. ov_SC07_000:func_8017F69C (108 ins) is banked +# func_8017f59c (64 ins) plus a 46-ins state-dispatch preamble, byte-identical elsewhere; it sat +# in the "no twin within 25" pile. Measured 2026-09-01 over 352 reachable stubs / 51,760 banked +# bodies >=25 ins: 8 opens have a contained twin at cover>=0.8 — 3 with no near-tier twin at +# all, and for 2 more the contained twin is STRICTLY BETTER work than their near twin +# (ov_SC01_077:func_80184D50's d=5 substitution twin vs ov_SC03_007 minus ONE C statement, +# resid 0). Law: minimum edit DISTANCE is not minimum WORK — a resid-0 one-block twin beats a +# closer substitution twin. +# +# THE METRIC: difflib alignment over BRANCH-MASKED norm_stream tokens (branch imm16 zeroed — +# a branch that jumps ACROSS an inserted block differs in offset while the C is identical, so +# unmasked branch offsets veto exactly the pairs this tier exists to find; C-level insertion +# regenerates the offsets). Verdict: 1..4 indel regions, <=2 substituted tokens total (as +# equal-length replaces of <=2), and equal-token cover >= min_cover of the SHORTER side. +# min_side=25 on both functions: below that, prologue+epilogue boilerplate alone passes any +# containment test (measured: 173 raw hits at min 12 collapse to 9 real at min 25). +# Candidates via prefix-8 OR suffix-8 reg_fields-skeleton buckets (a gap at one edge still +# leaves the other edge's skeleton intact; a gap at BOTH edges is invisible to this tier — +# stated limitation, not silent). +# +# The verdict is ADVISORY seeding (copy the twin's C; insert/delete exactly the listed block): +# the deletion direction (twin longer) is mechanically decidable — delete the C statements the +# block compiles from and the byte-gate arbitrates; the insertion direction still needs the +# block's C authored, but the card carries its exact disassembly. Nothing here banks by itself. + +_BR_OPS_MASK = frozenset((1, 4, 5, 6, 7)) + + +def _tok_brmask(SI, words): + """norm_stream tokens with branch immediates masked (see the block comment above).""" + ts = _norm_tokens(SI, words) + return [t & 0xFFFF0000 if (t >> 26) in _BR_OPS_MASK else t for t in ts] + + +def _contained_verdict(to, tt, max_k=4, max_sub=2): + """difflib classification of one token-stream pair -> dict(k, gap, nsub, m_equal, regions) + or None. k = indel regions (1..max_k); substitutions allowed only as equal-length replaces + of <=2 tokens, <=max_sub total; regions carry (tag, open_span, twin_span) for the card.""" + import difflib + ops = difflib.SequenceMatcher(a=to, b=tt, autojunk=False).get_opcodes() + k = nsub = gap = m_equal = 0 + regions = [] + for tag, i1, i2, j1, j2 in ops: + if tag == "equal": + m_equal += i2 - i1 + continue + if tag == "replace": + if (i2 - i1) == (j2 - j1) and (i2 - i1) <= 2: + nsub += i2 - i1 + regions.append(("sub", (i1, i2), (j1, j2))) + continue + return None + k += 1 + gap += max(i2 - i1, j2 - j1) + regions.append(("indel", (i1, i2), (j1, j2))) + if k == 0 or k > max_k or nsub > max_sub: + return None + return dict(k=k, gap=gap, nsub=nsub, m_equal=m_equal, regions=regions) + + +def contained_scan(min_cover=0.8, min_side=25, include_linked=False): + """Fleet-wide contained-twin scan. Returns {rows, stats}; one row per open stub that has a + contained twin (best = fewest nsub+k, then highest cover), with `d0_twin` marking stubs the + hash tier already serves (a contained twin there is a downgrade — consumers should prefer + the hash/near row). R32: the population is _pool_and_open's, arithmetic asserted.""" + SI, FR = _near_mods() + open_rows, banked, linked_n = _pool_and_open(include_linked, "contained_scan") + + pre_idx, suf_idx, binfo = {}, {}, {} + for h, (n, locs) in banked.items(): + if n < min_side: + continue + w = None + for (bb, aa) in locs: + w = _words_of(FR, bb, aa, n) + if w is not None: + break + if w is None: + continue + kp = tuple(FR.reg_fields(x) for x in w[:8]) + ks = tuple(FR.reg_fields(x) for x in w[-8:]) + pre_idx.setdefault(kp, []).append(h) + suf_idx.setdefault(ks, []).append(h) + binfo[h] = (n, locs[0][0], locs[0][1], w) + + tokc = {} + + def T(h): + if h not in tokc: + tokc[h] = _tok_brmask(SI, binfo[h][3]) + return tokc[h] + + rows, n_scanned, seen_hx = [], 0, {} + for o in open_rows: + if o["nins"] < min_side: + continue + n_scanned += 1 + hx = o["h_exact"] + if hx in seen_hx: # h_exact-equal opens share the verdict + if seen_hx[hx] is not None: + rows.append(dict(seen_hx[hx], open="%s:%s" % (o["binary"], o["symbol"]), + addr="0x%08x" % o["addr"])) + continue + wo = _words_of(FR, o["binary"], o["addr"], o["nins"]) + if wo is None: + raise SystemExit("contained_scan: cannot stream %s:%s (R32)." + % (o["binary"], o["symbol"])) + cset = set(pre_idx.get(tuple(FR.reg_fields(x) for x in wo[:8]), ())) | \ + set(suf_idx.get(tuple(FR.reg_fields(x) for x in wo[-8:]), ())) + to = _tok_brmask(SI, wo) + best = None + for h in cset: + n, bb, aa, wt = binfo[h] + if n == o["nins"]: + continue # same length = the hash/near tiers' territory + tt = T(h) + m = min(len(to), len(tt)) + lcp = 0 + while lcp < m and to[lcp] == tt[lcp]: + lcp += 1 + lcs = 0 + while lcs < m - lcp and to[-1 - lcs] == tt[-1 - lcs]: + lcs += 1 + if lcp + lcs < 14: # cheap gate; full verdict below is authoritative + continue + v = _contained_verdict(to, tt) + if v is None or v["m_equal"] < min_cover * m: + continue + row = dict(open="%s:%s" % (o["binary"], o["symbol"]), addr="0x%08x" % o["addr"], + nins=o["nins"], twin="%s:0x%08x" % (bb, aa), twin_nins=n, + side=("open" if o["nins"] > n else "twin"), + cover=round(v["m_equal"] / m, 2), **v) + if best is None or (row["nsub"] + row["k"], -row["cover"]) < \ + (best["nsub"] + best["k"], -best["cover"]): + best = row + seen_hx[hx] = best + if best is not None: + best["d0_twin"] = for_stub(o["binary"], o["symbol"], + include_linked=include_linked) is not None + rows.append(best) + + stats = dict(scanned=n_scanned, of_reachable=len(open_rows), linked_skipped=linked_n, + min_cover=min_cover, min_side=min_side, banked_pool=len(binfo), + hits=len(rows)) + assert all(r["cover"] >= min_cover and r["k"] >= 1 for r in rows), "verdict invariant" + return dict(rows=rows, stats=stats) + + +def contained_control(n_pos=60, n_neg=400, min_side=25): + """R39 controls for the contained tier, self-contained so they never go stale. + POSITIVE (must be perfect): plant the pair by construction — take a random banked body of + >=min_side+10 ins, DELETE a random contiguous 5..30-token mid-block, and the classifier + must report k=1, nsub=0, cover 1.0. Any other answer means the tier would refuse a + textbook one-block twin: FAIL. + NEGATIVE (base rate, context): random banked pairs sharing a prefix-8 OR suffix-8 skeleton + bucket with different lengths — how often ARBITRARY same-shape pairs pass the verdict. + True contained families inflate it; treat as an upper bound.""" + import random + SI, FR = _near_mods() + random.seed(42) + _, banked, _ = _pool_and_open(False, "contained_control") + pool = [] + for h, (n, locs) in banked.items(): + if n >= min_side + 10: + w = _words_of(FR, locs[0][0], locs[0][1], n) + if w is not None: + pool.append(w) + if len(pool) >= 4000: + break + ok = 0 + for i in range(n_pos): + w = random.choice(pool) + t = _tok_brmask(SI, w) + glen = random.randint(5, min(30, len(t) - min_side)) + at = random.randint(4, len(t) - glen - 4) + t2 = t[:at] + t[at + glen:] + v = _contained_verdict(t, t2) + if v and v["k"] == 1 and v["nsub"] == 0 and v["m_equal"] == len(t2): + ok += 1 + print("POSITIVE control (planted one-block deletions): %d/%d classified k=1 sub=0 full-cover" + % (ok, n_pos)) + if ok != n_pos: + raise SystemExit("contained_control FAILED: the classifier refuses planted one-block " + "twins (R39).") + hit = tot = 0 + for i in range(n_neg): + w1, w2 = random.choice(pool), random.choice(pool) + if len(w1) == len(w2): + continue + t1, t2 = _tok_brmask(SI, w1), _tok_brmask(SI, w2) + m = min(len(t1), len(t2)) + v = _contained_verdict(t1, t2) + tot += 1 + if v and v["m_equal"] >= 0.8 * m: + hit += 1 + print("BASE RATE (random banked pairs, different lengths): %d/%d = %.2f%% pass " + "cover>=0.8 — an UPPER bound (true contained families inflate it)" + % (hit, tot, 100.0 * hit / max(tot, 1))) + + def main(): ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) ap.add_argument("--binary") @@ -185,8 +846,66 @@ def main(): ap.add_argument("--json") ap.add_argument("--include-linked", action="store_true", help="do NOT refuse targets in LINKED subsegs (dead text; see _linked_segs)") + ap.add_argument("--near", action="store_true", + help="near tier: banked twins at edit distance 1..K too (venv-only)") + ap.add_argument("--max-d", type=int, default=5, help="near tier distance cap K (default 5)") + ap.add_argument("--near-control", action="store_true", + help="R39 controls for the near tier (positive + base rate)") + ap.add_argument("--contained", action="store_true", + help="contained tier: banked twins 1..4 whole blocks away, any gap size " + "(venv-only)") + ap.add_argument("--min-cover", type=float, default=0.8, + help="contained tier: equal-token cover floor over the shorter side") + ap.add_argument("--contained-control", action="store_true", + help="R39 controls for the contained tier (planted positives + base rate)") a = ap.parse_args() + if a.near_control: + near_control(max_d=a.max_d) + return 0 + + if a.contained_control: + contained_control() + return 0 + + if a.contained: + out = contained_scan(min_cover=a.min_cover) + s = out["stats"] + print("contained tier: %d/%d stubs >=%d ins scanned (of %d reachable; +%d LINKED " + "skipped) vs %d banked bodies · %d contained-twin rows (cover>=%.2f)" + % (s["scanned"], s["scanned"], s["min_side"], s["of_reachable"], + s["linked_skipped"], s["banked_pool"], s["hits"], s["min_cover"])) + for r in sorted(out["rows"], key=lambda r: (r["d0_twin"], r["gap"])): + print(" %-30s %3d ins <- %-26s %3d k=%d gap=%3d sub=%d cover=%.2f side=%s%s" + % (r["open"], r["nins"], r["twin"], r["twin_nins"], r["k"], r["gap"], + r["nsub"], r["cover"], r["side"], + " [has d0 hash twin — prefer that]" if r["d0_twin"] else "")) + if a.json: + with open(a.json, "w") as fh: + json.dump(out, fh, indent=1) + print("wrote %s" % a.json) + return 0 + + if a.near: + out = near_scan(max_d=a.max_d, include_linked=a.include_linked) + s = out["stats"] + print("near tier (K=%d): %d reachable open stubs (+%d LINKED skipped) · d=0 %d · " + "d 1..%d %d (RELOC-ONLY %d / NEAR-COUSIN %d) · no twin %d · open-open clusters " + "%d covering %d · cross-check %s" + % (s["max_d"], s["reachable"], s["linked_skipped"], s["d0"], s["max_d"], s["near"], + s["reloc_only"], s["near_cousin"], s["no_twin"], s["clusters"], + s["cluster_stubs"], s["cross_check"])) + for r in sorted((r for r in out["rows"] if r["d"]), key=lambda r: (r["d"], -r["nins"])): + print(" d=%d %-11s %s:%s (%d ins)%s <- %s %s" + % (r["d"], r.get("cls", ""), r["binary"], r["fn"], r["nins"], + " [jtbl]" if r["jtbl"] else "", r.get("twin", "?"), + "; ".join(r.get("diff", []))[:70])) + if a.json: + with open(a.json, "w") as fh: + json.dump(out, fh, indent=1) + print("wrote %s" % a.json) + return 0 + if a.all: out, scanned, dead = [], 0, 0 dead_by_bin = {}