mirror of
https://github.com/Druthulu/BFM-decomp
synced 2026-10-08 01:34:14 -04:00
df7bca2e65
The 68%->95% lever from §176h.C2, mechanized. Reconciliation belongs INSIDE the wave: a banked
draft's declarations become the TU's, so a sibling clash hardens into a file clash and post-bank
recovery is measurably worse (18 parked drafts still MATCH, only 1 survived after their wave banked
vs 5 before).
AUTO-FIXES, each re-verified with match_one and REVERTED if a byte moves (a declaration change is
a codegen change, §176f):
* COSMETIC-TYPEDEF two names for a structurally identical struct -> adopt the other. Compared by
BODY, never by name (OtBlk_80015498 == OtBlk_80016450; Elem12 != B12). This body comparison is
also the answer to §176h.C's spelled-name limit.
* SIGNEDNESS / ALIAS / ARRAY-VS-SCALAR -> adopt the TU's spelling, fixing the use site.
* DEFPARAMS (NEW LEVER) -> adopt the TU's parameter types on the DEFINITION and re-narrow with a
shadowing local: `void f(s32 a0_p) { s16 a0 = (s16)a0_p; <body unchanged> }`. One textual
insertion instead of rewriting every use site, and the cast emits the same sll/sra pair.
Byte-identical on both cases tried.
REFUSES, with named reasons, because these are decisions and not edits: DIFFERENT-STRUCT (two real
layouts for one symbol), IMMOVABLE-TU-DECL (gate_main reverts src/, so it needs its own commit +
rebuild + R22), DEF-SIDE-RETURN (adopting the TU's return type usually costs the match -- measured
on func_8001ABBC), and BROKE-MATCH for anything its own verification rejects.
Measured on wave P's leftover slate: 6 -> 9 compatible, 3 auto-reconciled, 2 repairs reverted by
the tool's own byte check, 7 named for a human.
build_wave_atlas: --rank mass (main's gate cost is per SLATE, so ranking groups by member count
silently collapses a wide band to the smallest functions -- measured 60 cards/2,604 ins where 46
cards/4,829 ins were available), and the selector no longer counts ITS OWN OUTPUT as already-waved
(re-running with identical filters had been shrinking the pool 60 -> 46).
175 lines
9.8 KiB
Python
175 lines
9.8 KiB
Python
#!/usr/bin/env python3
|
|
"""Build a campaign wave from the FRONTIER ATLAS, optimized for gate throughput.
|
|
|
|
Why this exists (P31, 2026-08-14): the pre-baked adapt/weak card piles are the smallest,
|
|
best-seeded tail (12-42 ins). Drawing from them banks ~1,440 ins/wave against 635,744 open —
|
|
0.011pp of fleet per wave. The atlas knows where the mass actually is (cousin-multi 294k ins,
|
|
cold 183k, main-only 38k) and what lever each group needs.
|
|
|
|
TWO selection principles, both measured:
|
|
1. GATE COST SCALES WITH (binary, TU) GROUPS, NOT DRAFTS. Each group is a whole-binary rebuild.
|
|
Wave C was 35 drafts over 27 groups = 1.3 drafts/rebuild, ~50 min of gate for 32 banks.
|
|
So: CONCENTRATE the wave in few binaries. This is free throughput.
|
|
2. MASS BEATS COUNT for the instruction-weighted metric — prefer bigger functions where a seed
|
|
exists, but keep them inside the model ladder's competence.
|
|
|
|
Selection: open stubs (derived from corpus, R32/R33), not spent in a prior wave, from atlas
|
|
groups whose lever is agent-draftable; ranked by binary concentration then instruction mass.
|
|
|
|
MUST NOT run while a gate is in flight (R35 — corpus.stubs() misreports substituted drafts).
|
|
|
|
Usage: build_wave_atlas.py <out.json> [N] [--max-bins K] [--min-ins M] [--levers a,b,c]
|
|
"""
|
|
import json, os, sys, collections, subprocess, argparse, glob
|
|
sys.path.insert(0, 'tools')
|
|
import corpus
|
|
|
|
ap = argparse.ArgumentParser()
|
|
ap.add_argument('out')
|
|
ap.add_argument('n', nargs='?', type=int, default=96)
|
|
ap.add_argument("--max-bins", type=int, default=12, help="concentrate into this many GATE GROUPS (binary,TU)")
|
|
ap.add_argument('--min-ins', type=int, default=0)
|
|
ap.add_argument('--max-ins', type=int, default=120, help='above this the bulk ladder stops being honest')
|
|
ap.add_argument('--levers', default='head-crack,seeded-crack,redraft,len-vein,integration,family-sweep,tiny-direct',
|
|
help='agent-draftable levers; UNKNOWN/tell/jtbl/o0/cc1 need their own lanes')
|
|
ap.add_argument('--atlas', default='.run/atlas.json')
|
|
ap.add_argument('--exclude-bins', default='',
|
|
help='comma-separated binaries to skip. NOTHING is excluded by default. '
|
|
'(History: main used to be excluded on a "link-resolution defect" — that '
|
|
'diagnosis was REFUTED 2026-08-15 by a null-draft control: the false diff '
|
|
'reproduces with ZERO drafts substituted. main simply cannot be gated '
|
|
'INCREMENTALLY, because its extract runs psyq_integrate/ld_interleave and '
|
|
'rewrites the .ld. Draft main like any binary; gate it with '
|
|
'tools/gate_main.py, never gate_lane/gate_stage.)')
|
|
ap.add_argument('--rank', choices=('groups','mass'), default='groups',
|
|
help="'groups' (default) ranks gate groups by MEMBER COUNT -- right for overlays, "
|
|
"where every (binary,TU) group costs its own rebuild. 'mass' ranks purely by "
|
|
"instruction size across all groups -- right for MAIN, whose gate cost is per "
|
|
"SLATE, not per TU: with 'groups' a wide --min-ins band fills from the "
|
|
"biggest-by-count group, which is the SMALLEST-by-instruction one, and the "
|
|
"wave silently collapses to tiny functions (measured: 60 cards / 2,604 ins "
|
|
"avg 43, when 46 cards / 4,829 ins avg 105 were available).")
|
|
ap.add_argument('--target-ins', type=int, default=0,
|
|
help='size the wave by INSTRUCTION MASS: keep drawing cards until this many '
|
|
'instructions are selected (still capped by n). The public metric is '
|
|
'instruction-weighted, so this is the number that matters -- 6000+ is the '
|
|
'P31 S52 standard (wave O: 6,266 ins in 2 gate groups, 47/49 MATCH).')
|
|
ap.add_argument('--only-bins', default='',
|
|
help='comma-separated allow-list; if set, ONLY these binaries are eligible. '
|
|
'Use --only-bins main for a main wave: gate_main.py rebuilds the whole EXE '
|
|
'once per SLATE, so main has no per-TU gate cost and --max-bins can be large.')
|
|
a = ap.parse_args()
|
|
EXCLUDE = {b for b in a.exclude_bins.split(',') if b}
|
|
ONLY = {b for b in a.only_bins.split(',') if b}
|
|
|
|
busy = subprocess.run(['pgrep', '-f', 'tools/(gate_stage|dedup_propagate|gate_lane)'],
|
|
capture_output=True, text=True)
|
|
if busy.returncode == 0 and busy.stdout.strip():
|
|
sys.exit(f"REFUSING: gate in flight (pids {busy.stdout.split()}) — corpus.stubs() would misreport (R35).")
|
|
|
|
# R32/R33: derive the already-waved set from what is ON DISK, never from a hardcoded wave-letter
|
|
# list (the literal 'a'..'l' silently missed waves m and n and would have re-issued their cards).
|
|
# Exclude OUR OWN output: the glob matches it, so re-running the selector after an aborted or
|
|
# re-tuned build marked the previous attempt's cards as 'already waved' and silently shrank the
|
|
# pool (measured: 46 candidates instead of 60 on a re-run with identical filters).
|
|
PRIORS = [p for p in sorted(glob.glob('.run/wave_*_cards.json'))
|
|
if os.path.abspath(p) != os.path.abspath(a.out)]
|
|
taken = set()
|
|
for p in PRIORS:
|
|
try:
|
|
taken |= {c.get('fn') or c.get('name') for c in json.load(open(p))}
|
|
except (FileNotFoundError, json.JSONDecodeError, TypeError):
|
|
pass
|
|
if not PRIORS:
|
|
print('NOTE: no prior wave card files found — nothing filtered as already-waved', file=sys.stderr)
|
|
|
|
levers = set(a.levers.split(','))
|
|
atlas = json.load(open(a.atlas))
|
|
|
|
_open = {}
|
|
def _stubmap(binary):
|
|
if binary not in _open:
|
|
# corpus.stubs() is addr -> Stub; the NAME lives on the record
|
|
_open[binary] = {st.symbol: st for st in corpus.stubs(binary).values()}
|
|
return _open[binary]
|
|
|
|
def is_open(binary, fn):
|
|
return fn in _stubmap(binary)
|
|
|
|
def home_tu(binary, fn):
|
|
"""The stub's home .c — this is the GATE GROUP KEY (gate_lane groups by (binary, src))."""
|
|
st = _stubmap(binary).get(fn)
|
|
return st.path if st else None
|
|
|
|
def model_for(nins):
|
|
if nins <= 50: return 'haiku'
|
|
if nins <= 120: return 'sonnet'
|
|
return 'opus'
|
|
|
|
cands, skipped = [], collections.Counter()
|
|
for g in atlas['groups']:
|
|
if g['lever'] not in levers:
|
|
skipped['lever-not-in-lane'] += g['inst']; continue
|
|
ex = g.get('exemplar') or {}
|
|
seed = (g.get('seed') or {}).get('norm') or (g.get('seed') or {}).get('raw') or {}
|
|
for m in g.get('members', []):
|
|
fn, b, nins = m.get('name'), m.get('b'), m.get('nins') or 0
|
|
if not fn or not b: skipped['no-name'] += 1; continue
|
|
if b in EXCLUDE: skipped['excluded-binary'] += 1; continue
|
|
if ONLY and b not in ONLY: skipped['not-in-only-bins'] += 1; continue
|
|
if fn in taken: skipped['already-waved'] += 1; continue
|
|
if not (a.min_ins <= nins <= a.max_ins): skipped['out-of-band'] += 1; continue
|
|
if not is_open(b, fn): skipped['already-banked'] += 1; continue
|
|
sub = corpus.asm_path(b, fn)
|
|
if not sub: skipped['no-asm'] += 1; continue
|
|
cands.append({
|
|
'tu': home_tu(b, fn),
|
|
'fn': fn, 'binary': b, 'lane': 'mass', 'model': model_for(nins), 'nins': nins,
|
|
'addr': m.get('a'), 'sub': __import__('os').path.dirname(sub),
|
|
'gid': g['gid'], 'lever': g['lever'], 'confidence': g.get('confidence'),
|
|
'lever_alts': g.get('lever_alts', []),
|
|
'exemplar': {'binary': ex.get('b'), 'fn': ex.get('name'), 'nins': ex.get('nins')},
|
|
'seed_sim': seed.get('sim'),
|
|
})
|
|
|
|
# principle 1: CONCENTRATE ON GATE GROUPS. gate_lane groups by (binary, home .c) and each group
|
|
# is one whole-binary rebuild, so drafts-per-GROUP is the throughput number that matters -- not
|
|
# drafts per binary. Wave D was 42 drafts over 23 groups (1.8/group, ~40 min of gate).
|
|
by_tu = collections.defaultdict(list)
|
|
for c in cands:
|
|
by_tu[(c['binary'], c['tu'])].append(c)
|
|
if a.rank == 'mass':
|
|
ranked = sorted(by_tu, key=lambda k: -sum(c['nins'] for c in by_tu[k]))[:a.max_bins]
|
|
else:
|
|
ranked = sorted(by_tu, key=lambda k: -len(by_tu[k]))[:a.max_bins]
|
|
|
|
# principle 3 (P31 S52): SIZE A WAVE BY INSTRUCTIONS, NOT BY CARDS. The public metric is
|
|
# instruction-weighted, so a wave is worth what its instructions are worth: the 12-42-ins card
|
|
# lanes produced ~1,400 ins/wave (~0.011pp) while wave O carried 6,266 ins for the same gate cost
|
|
# and the same draft rate. --target-ins keeps drawing cards until the instruction budget is met
|
|
# (still capped by n, so a wave can never spawn an unbounded fleet).
|
|
wave, tot_ins = [], 0
|
|
def _full():
|
|
if a.target_ins:
|
|
return tot_ins >= a.target_ins or len(wave) >= a.n
|
|
return len(wave) >= a.n
|
|
for k in ranked: # principle 2: within a group, mass first
|
|
for c in sorted(by_tu[k], key=lambda c: -c['nins']):
|
|
if _full(): break
|
|
wave.append(c); tot_ins += c['nins']
|
|
if _full(): break
|
|
if a.target_ins and tot_ins < a.target_ins:
|
|
print(f"NOTE: only {tot_ins} ins available under these filters (target {a.target_ins}) — "
|
|
f"widen --min-ins/--max-ins/--levers or raise n ({len(wave)} of max {a.n} cards used)")
|
|
|
|
json.dump(wave, open(a.out, 'w'), indent=1)
|
|
tot = sum(c['nins'] for c in wave)
|
|
print(f"candidates {len(cands)} in {len(by_tu)} gate groups (skipped {dict(skipped)})")
|
|
ngroups = len({(c['binary'], c['tu']) for c in wave})
|
|
print(f"-> wave {len(wave)} drafts / {tot} ins across {len({c['binary'] for c in wave})} binaries "
|
|
f"in {ngroups} GATE GROUPS = {len(wave)/max(ngroups,1):.1f} drafts per rebuild")
|
|
if wave:
|
|
print("models:", dict(collections.Counter(c['model'] for c in wave)))
|
|
print("levers:", dict(collections.Counter(c['lever'] for c in wave)))
|
|
print("nins: %d-%d (avg %.0f)" % (min(c['nins'] for c in wave), max(c['nins'] for c in wave), tot/len(wave)))
|