mirror of
https://github.com/Druthulu/BFM-decomp
synced 2026-09-30 23:37:39 -04:00
0f73618fa82876e81ca7b015dfabc0dbc242f94c
3 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b3ab5c2253 |
feat(phase-29 Task-13A): deterministic residual classifier — the permuter's problem is TARGETING
The autopsy (hindsight-study §7) assumed the permuter loses for want of a mutation. Measured over the whole open backlog, it loses because it is aimed at work a search-closer provably cannot close. - NEW tools/residual_class.py: decide a near-miss's class FROM THE BYTES. Decodes each mismatching MIPS word -> (op-skeleton, register-fields, immediate); drift FIRST (one inserted insn inflates `closeness` by the tail length), then consistent-injective register map -> REGALLOC-PERM (§31 S11/RC-3), same-multiset-reorder -> SCHEDULE-REORDER, DELAY-SLOT, WIDTH/BRANCH-POLARITY/STRENGTH/ADDRESSING/IMM-OFFSET/IMM-VALUE. Every class routes to a BUCKET = which tool the failure wants. Uncovered opcode -> UNKNOWN, COUNTED (R32). 16 synthetic unit tests (test_residual_class.py). - NEW tools/autopsy.py: `collect` materialises the corpus Task-12's telemetry never filled (1 of 6,169 records had a residual) by recompiling every open draft through the EXISTING match_one path (R33) — 1,752 drafts in 21s at -j12. `report` -> docs/autopsy.md. - NEW corpus.o0_sources()/is_o0(): the opt-level oracle DERIVED from the Makefile's own -O0 rules, coverage-asserted. Scoring an -O0 target at -O2 makes the residual 100% artefact (the trap this phase hit four times). - R34 cross-check baked in: residual_class's closeness vs masked_diff.structured_diff's, asserted per row; 1,673/1,673 agree, 0 classifier errors. FINDING: of the 972 records the grinder's own filter admits, only 75 (7.7%) are permuter-shaped; 547 are structural and 348 are drafts that are not the function at all. ~92% of the daemon's CPU went where it could not win — the byte-grounded explanation of "7 banks all-time, all Phase 21, 0 since" (Phase-22 audit). grinder.candidates() now filters on the measured bucket (1,303 -> 78) and takes its directed profile from the measured class, not the logged label (91% carry none -> it ran on gcc defaults). Degrades to undirected if uncollected and says so; --no-targeting A/Bs it. Two measured corollaries (R14, not projections): - 699 records rank as near-misses at closeness up to 278 purely from a length artefact: un-attempted work misfiled as a backlog of hard functions -> new `redraft` bucket. - a 12-draft gate probe of the `integration` bucket banked 1/12 (11 PLUMBING), so the 306 prices Task 14's reconcile ladder rather than promising free banks. func_80167714 (104 ins, reach-134) banked x1, un-propagated by design (§55b). Two defects fixed forward: - masked_diff._common_typedefs() used ONE shared probe path, so parallel match_one processes clobbered each other: 14 of 1,752 drafts lost in a single 12-way run (0.8%), silently, in every parallel wave ever run. Now per-PID. - gate_stage.match_one_closeness never passed --o0 -> phantom residuals for every -O0 function, written straight into the backlog this autopsy reads. R22 clean-fleet: check-all 140 passed, 0 failed of 140; tools-health OK (dedup 1847/0, C1 234343/234343); 0 NON_MATCHING (G4). Flywheel captured in-session (R30/R31): cookbook §60, decision-log entry, SETUP.md inventory. |
||
|
|
427baba3bf |
feat(phase-27 T10): completion dashboard (main in the weighted metric) + the resident second oracle
The metrics contract (roadmap §1) wants all three metrics WITH main in the denominators, and the
second, independent boundary oracle (R34) extended beyond the overlays. Both had landmines.
10a — main into the weighted metric, safely:
- weighted_metrics off the func_-only src_stubs regex onto corpus.stubs (R33). THE LANDMINE IS
REAL: src_stubs("SLUS_007.26") globs src/SLUS_007.26/*.c -> 0 files -> every row "matched" ->
main 100% + fleet % silently inflates. Routing through corpus.stubs is a PROVEN 0.000pp no-op on
the existing fleet (overlays are all func_) and closes the curated-name leak.
- a SEPARATE "MAIN game-code weighted" line (0.7%): main's Ghidra sig excludes the LINKED PsyQ
objects (Ghidra never analysed them), which is exactly right for a game-code metric (LINKED is
complete, counted in fn-count). Reported un-folded and caveated (month-stale sig, PROVISIONAL) —
folding a stale/incomplete value into the decomp.dev headline would mislead the flip checkpoint.
10b — the resident second oracle:
- make sig-resident: sig_image on the resident flat blob (byte-derived, not Ghidra). corpus.
sig_is_independent now covers resident -> audit-corpus checks its boundaries too. Probed clean
BEFORE wiring (144 fns, all 21 stubs present, 0 phantom), verified 0 phantom + 0 truncated.
- sig-overlays now derives its payload list from config/overlays.mk, not a 0.4.dec glob that
silently dropped the 4 SC07 index-1 overlays (the audit's own silent-skip class). tools-health
regenerates sig-overlays + sig-resident first so the audit never crashes on an absent sig.
10c — main's second oracle: docs/second-oracle.md. sig_image can't sign the PS-X EXE yet (0x800
header offset, interleaved data/linked islands, one text range); seeding from splat would destroy
independence for the PHANTOM class specifically. Honest deferral + scoped design, not a fake oracle.
- docs/progress.fleet.md regenerated: 140 binaries · fn-count 82.16% · instr-weighted 67.0%
(the honest post-T7 drop from 68.9%) · distinct 47.8% · MAIN game-code 0.7% (separate).
- SETUP §6.3 updated (R21).
|
||
|
|
f7b7399ebe |
feat(phase-26a): A3 — tools/corpus.py, ONE derived corpus oracle (+ a second oracle that can disagree)
The 28 surviving audit findings collapse to ONE bug repeated ~10 times: a hand-maintained model of
the corpus layout (a file allowlist, a single-.c assumption, a func_-only symbol regex, a REGION_SUB
dict) sitting on top of a filesystem that already answers the question. The fix is not ten repaired
regexes — it is one DERIVED oracle and ten deleted scanners (R33).
WHAT IT DERIVES FROM
1. THE FILESYSTEM. Which .c files make up a binary, and where a function's .s lives, are FACTS OF
THE TREE THAT SPLAT ITSELF WROTE. The INCLUDE_ASM line is SELF-DESCRIBING — its first argument
IS the asm subdir — so there is nothing to guess and no dict to rot. A dict literal is strictly
worse than the filesystem AND it fails OPEN (silently yields a wrong path) instead of closed.
2. THE PROVEN INVARIANT. INCLUDE_ASM pastes the ORIGINAL asm and the build is byte-identical, so a
function NOT wrapped in it is byte-exact. `matched` is DERIVED as sig - stubs, never re-parsed
from C text. (progress.py learned this the hard way: weighted_metrics() derived and was right;
classify() re-parsed C and inherited a bug.)
VALIDATED against the real corpus:
* ov_SC01_077: 264 stubs across 14 files. The old 3-file allowlist saw 30.
* Fleet: 58,717 stubs vs the allowlist's 1,992 — 56,725 (96.6%) were INVISIBLE.
* Coverage-asserted (R32): every INCLUDE_ASM line must parse, every symbol must resolve (ANY C
identifier — a func_-only regex silently misses the 100 curated listCdBuffer stubs), every stub
must have a .s. A silent skip is a DEFECT, not a no-op.
THE SECOND ORACLE (`make audit-corpus`) — the real lesson of this audit.
The byte-gate is structurally BLIND to a bad function boundary: the .s halves are pasted back
verbatim in original order, so the image stays byte-identical and green. Only an oracle that can
DISAGREE can see it. sig_image is that oracle — Ghidra-free, derived from the ORIGINAL bytes,
independent of splat. corpus.audit() cross-checks the two and reports:
PHANTOM — a stub address the sig does not know: splat INVENTED a function.
TRUNCATED — a stub whose .s length != the sig's: splat MIS-SLICED one.
It reports 193 (96 + 97) — reproducing the A2 audit's number EXACTLY, from an independently written
tool. That is a third confirmation of the listCdBuffer defect (auditor -> skeptic -> this).
AND AN R14 SELF-CATCH, recorded because the near-miss is the lesson.
Run naively over all 136 binaries the same check reports 914 slices — 4.7x the truth. It is noise:
main/resident are signed by the GHIDRA dumper, whose boundaries are shorter than splat's by design
(and which never analysed the linked PsyQ subsegs at all), so the comparison measures GHIDRA'S limits,
not splat's errors. Only the overlays are signed by sig_image, the oracle actually validated at
58,524/58,621. sig_is_independent() now encodes that domain, with the reasoning, so nobody repeats it.
A check applied outside its valid domain does not become more thorough — it becomes noise.
`make audit-corpus` is RED by design until A4 removes the bad symbol line; then it becomes a gate.
|