mirror of
https://github.com/Druthulu/BFM-decomp
synced 2026-10-01 07:40:42 -04:00
docs: accelerators #15 (the differential-oracle harness) + the generic decomp package thesis
#15 — the tool worth building FIRST in any decomp, because it works at 0% and compounds: run every question down TWO independent paths on a schedule and fail on disagreement. Ten-plus S68 blockers had one shape — a tool computing a TRUE number about a NARROWER world than we believed it covered — and EVERY one was caught by two measurements disagreeing, never by review. R32/R34/R40 already say this and were not enough: they are rules applied by whoever writes the tool, and in S68 I wrote R34's warning into one docstring and rebuilt the exact defect it warns about an hour later in another file. Includes Drew's scheduling half, which this project only ever did by accident: the widening is PERIODIC. Tooling is not wrong when written, it goes STALE as new idioms reveal populations it cannot see. At every phase close ask 'which scanner's denominator just got wider?' — that question converts new knowledge into free banks. The §332 sweep is the worked example: one review, 10 fns / 1,027 ins reclassified, one in-flight escalation stopped mid-spend. generic-decomp-package.md — what a NEW decomp inherits on day one and does BEFORE cracking: mine the COMPILER SOURCE and sibling projects for idioms (this project's best late idioms came from reading gcc-2.7.2's own passes and needed no matched function at all — week-1 work done in month N), port the families/twins/dedup/carve layer first, then the oracle harness, and only then crack. With the honest caveat that tooling-first makes the cheap half free and does NOT shrink the hard tail.
This commit is contained in:
@@ -292,3 +292,68 @@ function probed "carveable". Re-probed with the planner: **96 of 159 open jtbl f
|
||||
plan-refused**, and a previous session had priced 32 of them as free work on the blind reading.
|
||||
Cost of the fix: eight lines. An optimistic probe does not merely lose opportunities — it
|
||||
manufactures work plans, which is the expensive direction of the error.
|
||||
|
||||
## #15 — THE DIFFERENTIAL-ORACLE HARNESS: run every question down TWO paths and fail on disagreement (P31 S68)
|
||||
|
||||
**The pattern, counted in ONE session.** Ten-plus blockers, every one the same shape: *a tool
|
||||
computed a TRUE number about a NARROWER WORLD than the one we believed it covered.* Not bugs —
|
||||
correctly-scoped tools whose population widened underneath them as new idioms landed.
|
||||
|
||||
* `gate_stage` compared `main` against **another binary's SHA** (`build/main/main` and
|
||||
`config/check.main.sha` do not exist, so `DEF_SHA` = ov_SC01_077 took over). Every main draft read
|
||||
"near" — for weeks.
|
||||
* `psyq_integrate` dropped `firstfile` on every INCREMENTAL relink, so main was 2 bytes red before
|
||||
any draft was spliced. This is the true identity of the long-standing "main link defect".
|
||||
* `match_one`'s standalone probe called 39 of 43 drafts `cc1-fail` for symbols that ARE in their
|
||||
real TU.
|
||||
* `seed_ref` offered 43 `main` LINKED-subseg stubs as bankable twins — dead text where a draft gates
|
||||
GREEN while wrong.
|
||||
* `wall_sweep` (written that same day) returned a confident **0 across 1,378 files**: the `.s` lines
|
||||
carry a `/* … */` prefix and the regex anchored at line start.
|
||||
* `corpus.stubs` passed while **106 of 213** binaries had no `.s` on disk after a restore that
|
||||
printed `212 extracted, 0 failed`.
|
||||
* Four agents returned `NO-DRAFT` after a rate limit while one sat **3 instructions** from a match,
|
||||
its full candidate history on disk.
|
||||
* §332 stated "6 functions fleet-wide" and named TWO. **A count without an enumeration cannot drive
|
||||
a filter**, so the draw kept paying agents to rediscover the class (92k and 289k tokens, twice).
|
||||
|
||||
**What caught every single one: two independent measurements disagreeing.** `rtu_match` vs the
|
||||
whole-binary gate. The standalone probe vs the real TU. A sweep's zero vs a member already known. A
|
||||
clean rebuild vs an incremental one. An agent's verdict vs its own scratch dir.
|
||||
|
||||
**Why the RULES were not enough.** R32 (assert your coverage), R34 (a second DISAGREEING oracle) and
|
||||
R40 (exonerate the instrument) all existed and are correct. They are *rules applied by whoever
|
||||
writes the tool* — and in S68 the author wrote R34's warning into one tool's docstring and then
|
||||
**rebuilt the exact defect it warns about, in a different file, an hour later** (an mtime liveness
|
||||
heuristic, after documenting that "a quiet file mtime is not a completion signal").
|
||||
|
||||
**THE TOOL: a standing harness that runs the same question down two independent paths on a schedule
|
||||
and fails loudly on divergence.** The pairs exist in any decomp from day one:
|
||||
|
||||
| question | path A | path B |
|
||||
|---|---|---|
|
||||
| is it matched? | `corpus.stubs` | the built binary / dedup registry |
|
||||
| does it compile? | standalone probe | the real TU |
|
||||
| is the fleet green? | incremental build | from `make clean` |
|
||||
| what does this scanner cover? | its own count | an over-approximating candidate set |
|
||||
| did the agent produce work? | its returned verdict | its scratch dir |
|
||||
| is this function bankable? | the draw filter | the toolchain-wall oracle |
|
||||
|
||||
**Why build it FIRST, before the cookbook has a single entry.** It is the only accelerator on this
|
||||
list that works at 0% and compounds. Retrieval needs a matched corpus; triage needs drafts; the
|
||||
cookbook needs matches. But *two ways to measure the same thing* exist from the first function — and
|
||||
the value grows with every tool added, because every new tool is a fresh chance to be confidently
|
||||
wrong about scope.
|
||||
|
||||
**Cost asymmetry that makes it obvious in hindsight.** Each of the diagnoses above cost 90k–290k
|
||||
tokens as a one-off agent investigation. A nightly disagreement report is minutes of compute. Rough
|
||||
estimate for S68: **about half the session went to harness defects wearing model-failure costumes**,
|
||||
and that ratio has probably held, invisibly, for most of the project — because a plausible number
|
||||
never asks to be checked.
|
||||
|
||||
**The scheduling half, and it is not optional (Drew, S68):** the widening is PERIODIC, not one-off.
|
||||
Tooling was correct when written and went stale as new idioms revealed populations it could not see.
|
||||
So at every session/phase close, review the tooling against the idioms learned that phase and ask
|
||||
*which scanner's denominator just got wider?* — the answer converts new knowledge into free banks.
|
||||
S68's own §332 sweep is the worked example: one idiom review, ten functions / 1,027 instructions
|
||||
reclassified from "hard" to "not bankable at all", and one in-flight escalation stopped mid-spend.
|
||||
|
||||
@@ -0,0 +1,70 @@
|
||||
# The Generic Decomp Package — what a NEW decompilation should inherit on day one
|
||||
|
||||
> **Status: the thesis, recorded S68 (2026-08-31) by Drew, from hindsight over this project.**
|
||||
> Feeds the endgame deliverables (the retrospective + the public "how to AI-decomp" wiki).
|
||||
> This is NOT a plan for BFM. It is what the NEXT project starts with instead of starting empty.
|
||||
|
||||
## The claim
|
||||
|
||||
This project spent most of its life **brute-forcing functions and then widening tooling whenever a
|
||||
new idiom revealed a population the tooling could not see.** In hindsight that order is backwards.
|
||||
A new decomp should spend its FIRST phases building the wide tooling and seeding the knowledge base,
|
||||
and only then start cracking — because every tool built early pays on every function afterwards,
|
||||
while every function cracked early pays once.
|
||||
|
||||
The evidence is this project's own zero-token banks: whole classes (twins, families, siblings,
|
||||
cousins, `-O0` carves, propagation, stranded boundaries) that cost nothing per function ONCE the
|
||||
tool existed — and that were invisible until an idiom taught us to look.
|
||||
|
||||
## What the next project inherits, and does BEFORE cracking
|
||||
|
||||
**1. The knowledge base, seeded from sources that exist before any match does.**
|
||||
* Mine the actual COMPILER SOURCE for the target triple. This project's highest-value late idioms
|
||||
(§368 reload-remat, §372 copy-capture, §370's `schedule_select` bound, §373's `pri(asm)=1`) came
|
||||
from reading `gcc-2.7.2`'s own passes — `reload1.c`, `cse.c`, `local-alloc.c`, `sched.c`,
|
||||
`stmt.c`. **None of that required a single matched function.** It could have been mined in week 1.
|
||||
* Mine SIBLING PROJECTS on the same compiler (this project used Vagrant Story / sotn-decomp).
|
||||
* Carry `docs/matching-cookbook.md` (399 sections) + `cookbook-index.md` (the symptom→section table)
|
||||
across as the starting corpus, adapted for the new triple rather than rebuilt.
|
||||
|
||||
**2. The structural tooling, before the first crack.**
|
||||
The families/twins/dedup layer is what converts one crack into N banks. In BFM this arrived late and
|
||||
retroactively harvested thousands of instructions. Port it first:
|
||||
`corpus` (the coverage oracle) · `seed_ref` / twin join on signature hashes · `family_remap` /
|
||||
`family_sweep` · `dedup_propagate` (position-locked overlay sharing) · the `-O0`/opt-level carve
|
||||
chain (`o0_detect`, `o0_subsplit`, `o0_boundary`) · `wall_sweep` (toolchain walls) · the draw
|
||||
filter · the byte-gate + clean-fleet verifier.
|
||||
|
||||
**3. The differential-oracle harness (accelerators #15) — the one that works at 0%.**
|
||||
Two independent paths per question, disagreement fails loudly, on a schedule.
|
||||
|
||||
**4. The periodic widening review (Drew's addition, and the part this project did only by accident).**
|
||||
At every session/phase close: review the tooling against the idioms learned that phase and ask
|
||||
**"which scanner's denominator just got wider?"** New idioms do not only make the next crack easier —
|
||||
they retroactively convert already-open functions into free banks, but ONLY if a tool is widened to
|
||||
see them. S68's §332 sweep is the worked example.
|
||||
|
||||
## The order this implies
|
||||
|
||||
phase 0 compiler-source + sibling-project idiom mining -> seed the cookbook
|
||||
phase 1 structural tooling: corpus, families, twins, dedup, carves, walls, byte gate
|
||||
phase 2 the differential-oracle harness + the draw filter
|
||||
phase 3 FIRST cracks — and from here every crack feeds the widening review
|
||||
... every phase close: idioms -> tooling widening -> free banks
|
||||
|
||||
## The honest caveat
|
||||
|
||||
Tooling-first does not remove the hard tail. This project's remaining frontier at S68 was **418
|
||||
functions / 62,717 instructions**, of which only ~5% was mechanically free and 947 instructions were
|
||||
*permanently* unbankable from C (toolchain walls). The structural work — §366 case-label unstacking,
|
||||
§368's uncolorable local, §358's unreferenced aggregate — needed genuine reasoning and always will.
|
||||
**Tooling-first makes the cheap half nearly free and stops the waste; it does not shrink the hard
|
||||
half.** Sell it as that and it is true; sell it as "no hand-cracking" and it is not.
|
||||
|
||||
## Where the pieces live today
|
||||
|
||||
`docs/matching-cookbook.md` + `docs/cookbook-index.md` (the knowledge) ·
|
||||
`docs/accelerators.md` (hindsight tools, #15 is the day-one one) ·
|
||||
`docs/wave-playbook.md` (the running procedure, each guard paired with the measurement that earned it) ·
|
||||
`docs/decision-log.md` (WHY each pivot happened) · `tools/` (the toolset) ·
|
||||
`phase-ends/` (the build history the retrospective is reconstructed from).
|
||||
Reference in New Issue
Block a user