diff --git a/docs/accelerators.md b/docs/accelerators.md index c390391a74..3c60b05304 100644 --- a/docs/accelerators.md +++ b/docs/accelerators.md @@ -292,3 +292,68 @@ function probed "carveable". Re-probed with the planner: **96 of 159 open jtbl f plan-refused**, and a previous session had priced 32 of them as free work on the blind reading. Cost of the fix: eight lines. An optimistic probe does not merely lose opportunities — it manufactures work plans, which is the expensive direction of the error. + +## #15 — THE DIFFERENTIAL-ORACLE HARNESS: run every question down TWO paths and fail on disagreement (P31 S68) + +**The pattern, counted in ONE session.** Ten-plus blockers, every one the same shape: *a tool +computed a TRUE number about a NARROWER WORLD than the one we believed it covered.* Not bugs — +correctly-scoped tools whose population widened underneath them as new idioms landed. + +* `gate_stage` compared `main` against **another binary's SHA** (`build/main/main` and + `config/check.main.sha` do not exist, so `DEF_SHA` = ov_SC01_077 took over). Every main draft read + "near" — for weeks. +* `psyq_integrate` dropped `firstfile` on every INCREMENTAL relink, so main was 2 bytes red before + any draft was spliced. This is the true identity of the long-standing "main link defect". +* `match_one`'s standalone probe called 39 of 43 drafts `cc1-fail` for symbols that ARE in their + real TU. +* `seed_ref` offered 43 `main` LINKED-subseg stubs as bankable twins — dead text where a draft gates + GREEN while wrong. +* `wall_sweep` (written that same day) returned a confident **0 across 1,378 files**: the `.s` lines + carry a `/* … */` prefix and the regex anchored at line start. +* `corpus.stubs` passed while **106 of 213** binaries had no `.s` on disk after a restore that + printed `212 extracted, 0 failed`. +* Four agents returned `NO-DRAFT` after a rate limit while one sat **3 instructions** from a match, + its full candidate history on disk. +* §332 stated "6 functions fleet-wide" and named TWO. **A count without an enumeration cannot drive + a filter**, so the draw kept paying agents to rediscover the class (92k and 289k tokens, twice). + +**What caught every single one: two independent measurements disagreeing.** `rtu_match` vs the +whole-binary gate. The standalone probe vs the real TU. A sweep's zero vs a member already known. A +clean rebuild vs an incremental one. An agent's verdict vs its own scratch dir. + +**Why the RULES were not enough.** R32 (assert your coverage), R34 (a second DISAGREEING oracle) and +R40 (exonerate the instrument) all existed and are correct. They are *rules applied by whoever +writes the tool* — and in S68 the author wrote R34's warning into one tool's docstring and then +**rebuilt the exact defect it warns about, in a different file, an hour later** (an mtime liveness +heuristic, after documenting that "a quiet file mtime is not a completion signal"). + +**THE TOOL: a standing harness that runs the same question down two independent paths on a schedule +and fails loudly on divergence.** The pairs exist in any decomp from day one: + +| question | path A | path B | +|---|---|---| +| is it matched? | `corpus.stubs` | the built binary / dedup registry | +| does it compile? | standalone probe | the real TU | +| is the fleet green? | incremental build | from `make clean` | +| what does this scanner cover? | its own count | an over-approximating candidate set | +| did the agent produce work? | its returned verdict | its scratch dir | +| is this function bankable? | the draw filter | the toolchain-wall oracle | + +**Why build it FIRST, before the cookbook has a single entry.** It is the only accelerator on this +list that works at 0% and compounds. Retrieval needs a matched corpus; triage needs drafts; the +cookbook needs matches. But *two ways to measure the same thing* exist from the first function — and +the value grows with every tool added, because every new tool is a fresh chance to be confidently +wrong about scope. + +**Cost asymmetry that makes it obvious in hindsight.** Each of the diagnoses above cost 90k–290k +tokens as a one-off agent investigation. A nightly disagreement report is minutes of compute. Rough +estimate for S68: **about half the session went to harness defects wearing model-failure costumes**, +and that ratio has probably held, invisibly, for most of the project — because a plausible number +never asks to be checked. + +**The scheduling half, and it is not optional (Drew, S68):** the widening is PERIODIC, not one-off. +Tooling was correct when written and went stale as new idioms revealed populations it could not see. +So at every session/phase close, review the tooling against the idioms learned that phase and ask +*which scanner's denominator just got wider?* — the answer converts new knowledge into free banks. +S68's own §332 sweep is the worked example: one idiom review, ten functions / 1,027 instructions +reclassified from "hard" to "not bankable at all", and one in-flight escalation stopped mid-spend. diff --git a/docs/generic-decomp-package.md b/docs/generic-decomp-package.md new file mode 100644 index 0000000000..e75d0bad25 --- /dev/null +++ b/docs/generic-decomp-package.md @@ -0,0 +1,70 @@ +# The Generic Decomp Package — what a NEW decompilation should inherit on day one + +> **Status: the thesis, recorded S68 (2026-08-31) by Drew, from hindsight over this project.** +> Feeds the endgame deliverables (the retrospective + the public "how to AI-decomp" wiki). +> This is NOT a plan for BFM. It is what the NEXT project starts with instead of starting empty. + +## The claim + +This project spent most of its life **brute-forcing functions and then widening tooling whenever a +new idiom revealed a population the tooling could not see.** In hindsight that order is backwards. +A new decomp should spend its FIRST phases building the wide tooling and seeding the knowledge base, +and only then start cracking — because every tool built early pays on every function afterwards, +while every function cracked early pays once. + +The evidence is this project's own zero-token banks: whole classes (twins, families, siblings, +cousins, `-O0` carves, propagation, stranded boundaries) that cost nothing per function ONCE the +tool existed — and that were invisible until an idiom taught us to look. + +## What the next project inherits, and does BEFORE cracking + +**1. The knowledge base, seeded from sources that exist before any match does.** +* Mine the actual COMPILER SOURCE for the target triple. This project's highest-value late idioms + (§368 reload-remat, §372 copy-capture, §370's `schedule_select` bound, §373's `pri(asm)=1`) came + from reading `gcc-2.7.2`'s own passes — `reload1.c`, `cse.c`, `local-alloc.c`, `sched.c`, + `stmt.c`. **None of that required a single matched function.** It could have been mined in week 1. +* Mine SIBLING PROJECTS on the same compiler (this project used Vagrant Story / sotn-decomp). +* Carry `docs/matching-cookbook.md` (399 sections) + `cookbook-index.md` (the symptom→section table) + across as the starting corpus, adapted for the new triple rather than rebuilt. + +**2. The structural tooling, before the first crack.** +The families/twins/dedup layer is what converts one crack into N banks. In BFM this arrived late and +retroactively harvested thousands of instructions. Port it first: +`corpus` (the coverage oracle) · `seed_ref` / twin join on signature hashes · `family_remap` / +`family_sweep` · `dedup_propagate` (position-locked overlay sharing) · the `-O0`/opt-level carve +chain (`o0_detect`, `o0_subsplit`, `o0_boundary`) · `wall_sweep` (toolchain walls) · the draw +filter · the byte-gate + clean-fleet verifier. + +**3. The differential-oracle harness (accelerators #15) — the one that works at 0%.** +Two independent paths per question, disagreement fails loudly, on a schedule. + +**4. The periodic widening review (Drew's addition, and the part this project did only by accident).** +At every session/phase close: review the tooling against the idioms learned that phase and ask +**"which scanner's denominator just got wider?"** New idioms do not only make the next crack easier — +they retroactively convert already-open functions into free banks, but ONLY if a tool is widened to +see them. S68's §332 sweep is the worked example. + +## The order this implies + + phase 0 compiler-source + sibling-project idiom mining -> seed the cookbook + phase 1 structural tooling: corpus, families, twins, dedup, carves, walls, byte gate + phase 2 the differential-oracle harness + the draw filter + phase 3 FIRST cracks — and from here every crack feeds the widening review + ... every phase close: idioms -> tooling widening -> free banks + +## The honest caveat + +Tooling-first does not remove the hard tail. This project's remaining frontier at S68 was **418 +functions / 62,717 instructions**, of which only ~5% was mechanically free and 947 instructions were +*permanently* unbankable from C (toolchain walls). The structural work — §366 case-label unstacking, +§368's uncolorable local, §358's unreferenced aggregate — needed genuine reasoning and always will. +**Tooling-first makes the cheap half nearly free and stops the waste; it does not shrink the hard +half.** Sell it as that and it is true; sell it as "no hand-cracking" and it is not. + +## Where the pieces live today + +`docs/matching-cookbook.md` + `docs/cookbook-index.md` (the knowledge) · +`docs/accelerators.md` (hindsight tools, #15 is the day-one one) · +`docs/wave-playbook.md` (the running procedure, each guard paired with the measurement that earned it) · +`docs/decision-log.md` (WHY each pivot happened) · `tools/` (the toolset) · +`phase-ends/` (the build history the retrospective is reconstructed from).