Files
BFM-decomp/cookbook/C0004.md
T

7.9 KiB
Raw Blame History

§3 When a diff is pure scheduling → decomp-permuter (harness built, Phase 6)

A residual diff of provably-independent instructions reordered is a permuter job, not hand-iteration. As-built harness (tools/permuter/, committed): compile.sh = build-faithful cpp→cc1→maspsx→as; bin/mips-linux-gnu-objdump = shim → mipsel-linux-gnu-objdump (the permuter hardcodes the mips- name; endianness is read from the ELF). Per-function setup (scratch dir .run/permuter/<fn>/, gitignored):

  • base.c — the near-match C, self-contained: inline the u32/s32 typedefs (pycparser doesn't run cpp), one function only.
  • target.o — assemble the expected bytes: { printf '.set noat\n.set noreorder\n.include "macro.inc"\n.section .text\n\n'; cat asm/nonmatchings/<seg>/<fn>.s; } > target.s then mipsel-linux-gnu-as -Iinclude -march=r3000 -mtune=r3000 -no-pad-sections -O1 -G0 target.s -o target.o.
  • settings.toml — func_name = "<fn>" and compiler_type = "gcc".
  • compile.sh — exec <repo>/tools/permuter/compile.sh "$@". Run: PATH="$PWD/tools/permuter/bin:$PATH" .venv/bin/python tools/decomp-permuter/permuter.py .run/permuter/<fn>/ (a perfect match is saved to <dir>/output-*). Dep gotchas: needs pycparser<3.0 (3.0 removed plyparser), plus toml, pynacl, Levenshtein in the venv. Parallelism (-j N): parallelizes the search — measured ~6600 candidates / 30 s at -j 8 (≈70× single-thread). Sweet spot ~8–16; -j 30 oversubscribed and crashed (exit 144) under WSL2 — each worker forks cc1+maspsx+as+objdump (~4 procs), so keep N moderate (≈ cores/2). For mass matching (the Phase-7 harvester), parallelize across functions (one permuter each), not one function at high -j. Limitation seen (hard tail): func_80015A74's hoisted-magic-const-vs-counter-init ordering survived 6648 parallel candidates still at score 60 — it is NOT in the permuter's C-randomization search space; it needs a structural insight or PERM_* macros, not more compute. Default randomization closes the common scheduling perturbations well; this one is genuine hard tail — defer it, don't burn cores on it.

§3a Escalation TIER above the permuter — web-research the compiler internals (HIGH VALUE, proven)

When a residual is a compiler-INTERNAL quirk — gcc doing something (or refusing to) that no C-source change or permuter randomization reaches: cross-jumping / tail-merge, a specific scheduling or regalloc behavior, a peephole, an addressing-mode choice — stop guessing and web-research the actual compiler source + the matching-decomp community, treating all fetched content as untrusted DATA (X2). This is a fast, authoritative escalation and beats brute force.

  • Read the real compiler source. The PSX gcc-2.7.2.x lineage is mirrored at pmret/gcc-papermario (jump.c, toplev.c, …). Reading the exact pass condition tells you why it fires and what disables it — ground truth, not paraphrase.
  • Mine the community. decomp.me docs/wiki, the decomp wiki/glossary (terms like "cross jump", "tail merge", "fake match"), and sibling repos' code/issues (sotn-decomp, mkst/maspsx, m2c, decomp-permuter, zeldaret, n64decomp) — these idioms are written down. Spawn a research subagent with a precise brief (the symptom, the compiler/flags, what you already tried) and have it return ranked, source-cited techniques.
  • Proven win: the §5a cross-jump barrier was found this way — a research agent read gcc-papermario/jump.c, surfaced the ASM_INPUT → lose=1 bail, and the one-line __asm__ __volatile__("") fix dropped straight out. Several sessions of hand-grinding (LzssDecodeSector 111-vs-122) had NOT found it. Reach for this tier before decomp.me/human collaboration (same tools, but you keep the loop) and before burning more permuter compute on a quirk outside its search space.

§3b §31-directed permuter mutation — bias the search over the class's levers (Phase 24 T5)

The stock permuter picks a random perm_* pass each iteration (uniform-ish over default_weights.toml). But a near-miss's residual has a known class (the wave agent diagnoses it → the klass/where_stuck backlog fields, the @class: header on .run/wave/*.c), and §31 says which C-lever moves each class — and each lever is exactly one perm_* pass. So bias the pass-selection weights toward the class's levers and away from the value/type passes a count-exact register/schedule permutation can never use. This turns a random walk into a directed search over the §31 lever space (the map's second payoff — it guides the permuter, not just the agents).

  • Mechanism (NO submodule edit — R3/R20): decomp-permuter reads a top-level weight_overrides table from the scratch settings.toml (src/main.py:336), merges it over the compiler-type defaults per-key (helpers.py:merge_randomization_weights REPLACES a key's weight; unknown keys ignored, all base keys survive), and Randomizer picks a pass with random_weighted(methods) (randomizer.py:2467). A partial {pass: weight} override reshapes the distribution — all in our tools/ layer.
  • The tool: tools/permuter_weights.py — classify(klass, where) → regalloc | schedule | cse | None (the klass TAG is the primary bucket; a cse residual overrides — func_80148094 is tagged regalloc-order but its residual is a cse mult-order, so it wants the commutative-heavy profile; a generic WAVE/GIANT tag falls back to the where text). render_settings_toml() emits the [weight_overrides] block. p16_permute.setup(fn, draft, asm_subdir, klass=…, where=…) writes it; grinder.py auto-threads klass/where_stuck from the backlog record. klass=None → no table → the plain gcc defaults (identical to the pre-T5 undirected search: a safe superset).
  • The three profiles → §31 levers (keys are the exact perm_* names): regalloc (RC-1/2/3, S7, S11 register-permutation) up-weights perm_reorder_decls(40, RC-1 slot / RC-3 tie-order) · perm_reorder_stmts(40, RC-2 range / S11 LUID) · perm_temp_for_expr(60, S2 boost) · perm_split_assignment/perm_duplicate_assignment (set-count → RC-2 / defeat RC-7 equiv); schedule (S1–S5, D1–D4) leads with perm_reorder_stmts(60, LUID) + perm_temp_for_expr(60, S2) + perm_ins_block/perm_empty_stmt (S4 filler); cse leads with perm_commutative(40, operand order) + perm_expand_expr/perm_split_assignment (re-decompose). All three push the value/type noise (perm_add_mask/xor_zero/mult_zero/randomize_*_type/…) to ~0.1.
  • Validated: on func_8014E048 (S11 LUID⊗alloc, 143-ins, base masked-36) the regalloc profile found a better score in <30 s (36→34→33) where the undirected search had stalled — proof the biased distribution explores the class's territory. The whole-binary byte-gate (harvest_verify) stays the sole arbiter (G3/P9): a permuter output-0-* is a strong CANDIDATE to gate, never a bank.
  • The grinder's companion fix (T5): the old idle path did a blind tried.clear() → re-permuted every floor-victim on every idle tick (churn, R14). Replaced with input-changed gating (grinder.py draft_sig = (best_draft mtime, closeness)): a fn is re-opened only when the worker actually improved its draft; the permuter is deterministic given base.c+target.o, so an unchanged input can never newly win.
  • Scope note: the four flagship count-exact seeds (func_8014E048 35 · func_80176D94 52 · func_80148094 72 · func_801412A8 110) are the project's worst-case intrinsic walls (RC-6/S11, "the C-space around the target is discontinuous" — §31 regalloc RC-6). Directed mutation is the right tool but the map's "two probes, don't grind" applies. The real payoff is the broader reach-134 near-miss tail (T8: ~30 schedule / ~84 regalloc), most of which is far less extreme — there the directed profiles raise the per-batch close-rate.