mirror of
https://github.com/Druthulu/BFM-decomp
synced 2026-09-29 15:18:24 -04:00
docs(s60): the Fable frontier audit + corrections it forced to my own checkpoint
The audit's headline, measured: THE WALL IS AN INTEGRATION WALL, NOT A CODEGEN WALL. Of the 292 functions the gate has refused 6+ times, 178 (61%) have ALREADY produced a closeness-0 draft — match_one byte-equality, whole-binary gate rejection. The blocker is symbols/decls/TU plumbing, and the fleet keeps re-drafting them: 10,049 reject rows over 574 distinct functions. Highest-EV build is a zero-token integration-resolver lane, not more drafting. CORRECTIONS TO MY OWN NUMBERS, verified against the tree before accepting: * siblings are 1,334 behind 480 multi-member groups, NOT ~3,900. 1,292 groups are SINGLETONS carrying 57% of open instruction mass. I conflated the never-drafted stub count with the sibling count and overstated remap leverage ~3x, in this checkpoint and repeatedly in conversation. * 'everything drawable is gen6+' holds only for the collapsed wave-eligible view; whole-pool generation is 53% gen0/1, 25% gen6+, and only 292 fns are 6+ GATE-refused. * '30-67 min gates at 8% CPU' conflated wall_min (includes drafting/queue) with gate wall (12-31 min healthy). Gate cost is proportional to FAILURES, not drafts: ~3 whole-binary builds per failing draft, so banks/gate-min fell 17.5 -> 0.10 as conversion fell. * the 5,388 closeness<=2 rows de-dupe to ~543 open functions; my own 19:40 re-measure found 290 still open, down from its 470 — the re-gate and grinder are draining that pool now. * campaign_status's 'banked today' undercounts: the stub invariant says ~2,644 net, because the A-prop lane's 357 rode in a chore commit its regex cannot see. One documented counterexample to 'model quality is not a bottleneck': func_80181714, where ox-alpha plateaued at closeness 4 while Opus/GLM/DeepSeek each reached reloc-verified MATCH — argues for a small escalation tier AFTER the resolver drains the fake walls. Taken on trust and flagged as such: the A-prop residual split (169 STRUCT / 121 no-seed-decl / 73 IMM / 12 void) — the refusal mechanisms exist in aprop_autodraft.py but no file carries those counts; re-derive before building the decl-inference tool.
This commit is contained in:
@@ -0,0 +1,350 @@
|
||||
# Frontier analysis — S60 (2026-08-25, evening)
|
||||
|
||||
**What this is.** A read-only analysis of the remaining unbanked population and of where the
|
||||
next banks actually come from, written for a session that has NO memory of today. Every number
|
||||
below carries its denominator and its source; the appendix lists the exact commands so any figure
|
||||
can be re-derived against the tree. Nothing in the tree was modified to produce this document.
|
||||
|
||||
**The one-paragraph conclusion.** The wide-wave drafting machine has finished the job it was
|
||||
built for. Its target population — never-drafted, small, seed-adjacent functions — was consumed
|
||||
today at 35–51% conversion, and what the waves now re-draft converts at 1–6% while costing
|
||||
30 minutes of gate each. But the campaign's own ledgers show that roughly **571 still-open
|
||||
functions have already been drafted correctly** (their instruction stream matched the target at
|
||||
the object level; the whole-binary gate rejected them for symbol/declaration/TU reasons), and the
|
||||
gate architecture spends ~3 whole-binary builds per *failing* draft while a real-TU oracle that
|
||||
needs no builds already exists in the tree (`tools/rtu_match.py`). The smartest path is to stop
|
||||
optimizing the wave and instead (1) re-gate what a config wipe falsely refused today, (2) build a
|
||||
deterministic **integration-resolver lane** over the shape-correct stock, (3) invert the gate so
|
||||
whole-binary builds are proportional to *banks* rather than *drafts*, and (4) re-aim the free
|
||||
fleet at the strata where drafting is genuinely unfinished — never-gated pockets, main, and the
|
||||
3–8-mismatch band — instead of re-drafting functions whose drafts are already correct.
|
||||
|
||||
---
|
||||
|
||||
## 0. Trust ledger — surveyed ground vs hearsay
|
||||
|
||||
**Verified against the tree by me tonight (commands in Appendix A):**
|
||||
atlas composition and sibling counts; the generation histogram; every wave conversion number and
|
||||
gate wall time (from `.run/gater.log` timestamps); the two `config/overlays.mk` wipe windows and
|
||||
their commit timestamps; the gate ladder's build accounting (from reading
|
||||
`tools/gate_stage.py` / `tools/harvest_verify.py` / `tools/sweep_parallel.py` in full); the
|
||||
backlog/reject-ledger stock counts (16,837 + 10,049 rows re-parsed); today's bank attribution
|
||||
(per-commit `INCLUDE_ASM` removal diffs over all 162 commits); the A-prop pool size; one live
|
||||
probe of a stored closeness-0 draft (`func_80184A68`); the live process table during a running
|
||||
gate.
|
||||
|
||||
**Repo-recorded measurements I did not independently re-derive** (they are written into tool
|
||||
docstrings/lane scripts as measured, with dates): the bank-rate-by-size table (S59, in
|
||||
`.run/drafter.sh`); "reloc pre-filter drops 45% of drafts / 13% of MISMATCH? are shape-MATCH"
|
||||
(in `tools/recover_rejects.py`); rtu_second_chance's 43-of-182 / 3-of-42 figures (in
|
||||
`tools/rtu_second_chance.py` + `.run/maintenance.log`); the func_80181714 three-model MATCH note
|
||||
(a backlog row).
|
||||
|
||||
**Taken on trust from the owner's briefing and NOT reproduced:** the A-prop residual split
|
||||
"169 STRUCT / 121 no-seed-decl / 73 IMM-unresolved / 12 void near-0" (I verified the *mechanisms*
|
||||
exist — `aprop_autodraft.py:522` refuses on "no seed decl", `:592` on STRUCT — but could not find
|
||||
a file carrying those exact counts; re-derive them before building against them). The "961 linked
|
||||
PsyQ stubs in main" figure.
|
||||
|
||||
**Owner claims that did NOT survive contact with the tree** — see §2: the ~3,900-sibling figure,
|
||||
the "everything drawable is gen6+" framing as a statement about the pool, the "0–14 of ~220"
|
||||
collapse (partly a harness defect), and the 30–67-min/8%-CPU gate picture (a metric artifact plus
|
||||
an already-landed fix).
|
||||
|
||||
---
|
||||
|
||||
## 1. State of the campaign as measured tonight
|
||||
|
||||
Fleet (docs/progress.fleet.md, regenerated by progress.py):
|
||||
instr-weighted **97.4%** (13,171,872 / 13,523,865); distinct-code **94.6%** (86,568 / 90,929
|
||||
unique fns); **main game-code 23.5%** (18,697 / 79,510 weighted) — main is the largest coherent
|
||||
open mass left. Overlays excluding main: 97.8%.
|
||||
|
||||
Open pool (`.run/atlas.json`, regenerated 17:25 today at head commit:2911): **3,106 open non-main
|
||||
instances** in 1,772 structural groups (2,245 distinct skeletons), plus **313 crackable main
|
||||
stubs** (atlas `main_open`; consistent with the owner's number).
|
||||
|
||||
| category | groups | instances | note |
|
||||
|---|---:|---:|---|
|
||||
| A-prop | 520 | 1,177 | banked seed + mapping exists; deterministic-lane food |
|
||||
| cold | 715 | 715 | all singletons, no seed |
|
||||
| cousin-multi | 185 | 717 | structural cousins; only 106 groups single-skeleton |
|
||||
| main-only | 149 | 149 | main lane's territory |
|
||||
| seeded | 178 | 297 | |
|
||||
| tiny | 25 | 51 | |
|
||||
|
||||
Sibling leverage, measured: **480 multi-instance groups hold 1,814 instances (93,165 ins);
|
||||
siblings = 1,334**. Of those 480 groups only **304 share a single skeleton** (true
|
||||
crack-one-remap-the-rest); 92 are all-distinct cousins, 84 mixed. **Singletons: 1,292 instances /
|
||||
121,853 ins — 57% of the open non-main instruction mass has no family leverage at all.**
|
||||
|
||||
Generation (times a fn appeared in any of the 204 wave card files), over the 3,106 open
|
||||
instances: gen0 **433**, gen1 **1,208**, gen2–5 **675**, gen6+ **790**. Distinct-fn gate
|
||||
*attempts* are much rarer than draws (the reloc pre-filter and draft failures eat the difference):
|
||||
only **292 open fns have ≥6 recorded gate attempts**.
|
||||
|
||||
Today's production (00:00 → HEAD, measured from the `INCLUDE_ASM` invariant, not commit
|
||||
subjects): stubs went **6,657 → 4,013 (−2,644 net; 2,877 gross removals** across 162 commits;
|
||||
carves re-add stubs, hence gross > net). Attribution by diffing every commit:
|
||||
|
||||
| lane | stub removals | note |
|
||||
|---|---:|---|
|
||||
| ox waves (~30 gates) | 2,031 | almost all of it before 13:00 — see the arc below |
|
||||
| deterministic (A-prop 357, −O0 ~102, misc maint/serial) | ~591 | **zero model tokens** |
|
||||
| main lane | 255 | its own clean-rebuild gate |
|
||||
|
||||
`campaign_status.py`'s "today: 2185 banked" is a **regex undercount** — it sums "— N banked"
|
||||
commit subjects, and e.g. the A-prop lane's 357 banks rode in a chore commit (commit:2904, 358
|
||||
removals) that the regex cannot see. Count banks from the stub invariant, not from subjects
|
||||
(the derive-from-invariants rule; this is the R32 "silently narrowed scope" class again).
|
||||
|
||||
### 1a. The arc of today, per wave (banked / gated, from `.run/gater.log` intervals)
|
||||
|
||||
dd **217/422** (12.4 min gate) · de **192/388** · df **124/278** · dg **100/283** · di **81/268**
|
||||
· dj **43/219** — these are the sibling-inclusive draws (`ONE_PER_GID=0` landed this morning)
|
||||
eating the never-drafted pool at 20–51%.
|
||||
|
||||
Then **dk 2/201** — the cliff. **It coincides exactly with the first `config/overlays.mk` wipe**
|
||||
(dk's own gate commit commit:2863 at 12:37 committed the registry as a zero-line file). With no
|
||||
binaries registered, overlay builds cannot succeed; the gate verdicts of that window are
|
||||
measurements of a broken harness, not of the drafts.
|
||||
|
||||
Recovery + mixed period: dl 12/160, dm 3/149 (re-gates), dp 54/233, dr 28/188, dt 9/169,
|
||||
dq 8/167, ds 35/159, du 3/159, dx 21/158, dy 12/127, eb–ee 1–10 each, **eh 129/380** (a fresh
|
||||
625-target draw). Then **ei 1/221, ej 0/215, ek 0/220, el 0/211, em 1/225 — all inside wipe #2**
|
||||
(committed by ei's gate commit ~17:14, restored 17:43:55 = commit:2913, root-caused 17:48:45 =
|
||||
commit:2915: four truncating `open(mk,"w").write()` sites in `jr_isolate_all.py:593` and
|
||||
`jtbl_carve.py:1077/1118/1165`). The giveaway is in the gate walls: those five "gates" ran
|
||||
**3.4–6 min for ~220 drafts each** — instant build failures, not judgments. Post-fix, honest:
|
||||
**en 14/227 (6.2%), eo 3/216 (1.4%), and a clean re-gate of ei banked 34/184 (18.5%) where its
|
||||
first gate banked 1** (re-gate of ej: 0/186 — so re-gating recovers some waves, not all).
|
||||
|
||||
**Banks per gate minute** — the metric that binds: dd ≈ **17.5**, de ≈ 12.6, eh ≈ 6.0,
|
||||
en ≈ 0.69, eo ≈ **0.10**. The A-prop maintenance pass banked **357 in one ~32-min sweep ≈ 11/min
|
||||
at zero tokens** (15:54–16:26, `.run/maintenance.log`).
|
||||
|
||||
---
|
||||
|
||||
## 2. Corrections to the going narrative (each measured tonight)
|
||||
|
||||
1. **"~3,900 siblings behind ~334 skeletons" is stale.** The 17:25 atlas shows **1,334 siblings
|
||||
behind 480 multi-instance groups**, only 304 of them single-skeleton. Today's dd–dj+eh burst
|
||||
consumed the rest. 57% of remaining open instruction mass is **singletons** — "crack a skeleton,
|
||||
unlock many" now describes a minority of the frontier.
|
||||
2. **"The drawable pool is ~all gen6+" is a statement about the collapsed wave-eligible view,
|
||||
not the pool.** Over all open instances: 53% are gen0/gen1, 25% gen6+. And *drawn* ≠ *gated*:
|
||||
only 292 open fns have actually been refused by the gate 6+ times.
|
||||
3. **The five zero-waves were partly manufactured.** ej–em were gated against a wiped registry
|
||||
(§1a). The clean-tree conversion tonight is 1.4–6.2% on first gates and up to 18.5% on
|
||||
re-gates — bad, but not zero, and every conclusion drawn from that window needs the re-gate
|
||||
first (R40: exonerate the instrument).
|
||||
4. **"A gate takes 30–67 min at 8% CPU" conflates three things.** The `GATE … · NN.Nmin` figure
|
||||
in the log is `time.time() − wave_draft_t0` — it includes drafting and queue wait. Real gate
|
||||
walls today: 12–31 min healthy, 3–6 min when broken. The "-j was decorative" serialization
|
||||
(every worker taking the fleet lock exclusively because `GATE_NO_ARITY` was unset) was
|
||||
diagnosed and **fixed at 10:43 today** (commit:2844). What remains is §3's failure-proportional
|
||||
cost plus (inferred, see §3c) cross-gate lock contention.
|
||||
5. **"5,388 backlog rows at closeness ≤2" is a row count, not a workload.** I count 8,596 such
|
||||
rows across the ledgers — but rows are re-attempts of the same functions. The de-duplicated,
|
||||
still-open workload is **~543 functions: 470 that have hit closeness 0 and 73 whose best is
|
||||
1–2**. The 5,010 files in `.run/backlog_drafts/` are the stock behind them.
|
||||
6. **"Model quality is not a bottleneck" has a documented counterexample stratum.** Backlog row
|
||||
func_80181714 (ov_SC02_016, 121 ins): ox-alpha plateaued at closeness 4 over 20 oracle calls
|
||||
while Opus, GLM-5.3 and fueled DeepSeek each reached reloc-verified MATCH. One case, recorded
|
||||
during a deliberate bake-off — enough to justify a small escalation tier (§5, step 7), not a
|
||||
fleet migration.
|
||||
|
||||
---
|
||||
|
||||
## 3. Where the gate minutes actually go (tools read in full)
|
||||
|
||||
The pipeline: `ox_campaign --gater` → reloc_identity pre-filter (drops ~45% of drafts) →
|
||||
`sweep_parallel -j N` (one worker per **binary**; per-binary flock; parallel across binaries
|
||||
only) → per binary, `gate_stage.run_gate` runs a **three-stage ladder**, and each stage calls
|
||||
`harvest_verify --chunk 1`, i.e. **one incremental whole-binary build per draft per stage**:
|
||||
|
||||
- stage 0: raw drafts — D builds; banked drafts exit here;
|
||||
- transforms on the failures (canon_resident_calls, cast_call_sites, reconcile_tu, arity
|
||||
pre-pass — subprocesses, cheap relative to builds);
|
||||
- stage 1: D′ builds; sig_unify; stage 2: D″ builds;
|
||||
- plus one final confirmation build per harvest_verify invocation (3 per binary), plus one
|
||||
`match_one` compile per still-failing draft for backlog closeness.
|
||||
|
||||
So a fully-failing binary with D drafts costs **≈ 3D + 3 builds + D match_one runs**, serial
|
||||
within the binary; a fully-banking one costs ≈ D + 1. **Gate cost is proportional to failures,
|
||||
not drafts** — which is why dd (51% conversion) ran 1.8 s/draft and eo (1.4%) ran 8.6 s/draft.
|
||||
The transforms exist to rescue PLUMBING-class failures; the recent failure-class mix
|
||||
(`.run/auto/bulk/*.failed.classified.txt`, 332 rows: 74 DIFF · 63 CC1-FAIL · 51 PLUMBING ·
|
||||
32 CARVE-REFUSED · 112 unclassified) says a large fraction of ladder re-builds are spent
|
||||
re-gating **DIFF** drafts — ones that already compiled and linked and produced wrong bytes, which
|
||||
no declaration rewrite will change.
|
||||
|
||||
**(3c, inferred)** Cross-gate lock contention: gates dr/dt/dq ran 27–34 min while a manual
|
||||
re-gate lane held overlapping per-binary flocks (13:22–15:00); ds/du, same size and shape, ran
|
||||
15–16 min once it ended. Two gates over the same binaries serialize worker-by-worker and the
|
||||
blocked workers occupy pool slots. Not proven causal — but it costs nothing to schedule re-gates
|
||||
into the gater's own queue instead of beside it.
|
||||
|
||||
**The structural fix is not more parallelism; it is making builds proportional to banks.**
|
||||
`tools/rtu_match.py` already compiles the *real TU* with the candidate spliced in — "no build
|
||||
tree, no locks, parallel-safe" (its own docstring) — and returns MATCH/DIFF/CC1 with real
|
||||
diagnostics. It is currently used only as a second-chance recovery. Inverted, it becomes the
|
||||
gate's stage 0: run rtu_match on all drafts 32-wide (no flocks, no link, no SHA), and spend
|
||||
whole-binary builds **only on rtu-MATCH drafts** plus the PLUMBING-recovery ladder for rtu-CC1
|
||||
failures whose diagnostics name a fixable conflict. At 2% conversion that is roughly a 20×
|
||||
reduction in whole-binary builds; the gate for a 200-draft wave becomes ~2 min of parallel TU
|
||||
compiles plus a handful of builds. G3/P9 is untouched — the whole-binary SHA gate remains the
|
||||
sole arbiter of banking; rtu only decides which drafts get to spend a build.
|
||||
|
||||
**Risk, and the cheap way to bound it:** rtu false-negatives (drafts rtu rejects that the
|
||||
whole-binary gate would have banked). Run rtu_match in **shadow mode** for 2–3 gates first —
|
||||
record its verdict for every gated draft next to the gate's verdict, zero behavioral change. If
|
||||
the false-negative rate is under ~2%, flip the order and keep a nightly sweep-up re-gate of
|
||||
rtu-rejects as the safety net. If it is higher, the shadow data says exactly which class rtu
|
||||
misjudges, and that class stays on the old path. Falsifier for the whole idea: a shadow run
|
||||
showing rtu-MATCH does not predict banking (agreement < ~90%) — then the gate stays as it is and
|
||||
the win must come from §4/§5 alone.
|
||||
|
||||
---
|
||||
|
||||
## 4. The finding that reframes the endgame: the wall is an integration wall
|
||||
|
||||
Of the **292** open functions the gate has refused six or more times, **178 (61%) have already
|
||||
produced a closeness-0 draft** — `match_one` byte-equality at the object level, whole-binary gate
|
||||
rejection. 27 more sit at closeness 1–2, 29 at 3–8, 50 at 9+, 8 unscored. Extending past the
|
||||
6+-refused wall to the whole open pool:
|
||||
|
||||
- **470 open fns** have a closeness-0 row in the backlog, best draft saved on disk;
|
||||
- **151 open fns** appear in `.run/reloc_rejects.jsonl` with `shape: MATCH` (instruction stream
|
||||
already matches; only the symbol names are wrong — the deterministic `aprop_symfix` class);
|
||||
- union: **571 open fns ≈ 18% of the open non-main pool** where *drafting is finished* and what
|
||||
remains is symbol identity, declaration conflicts, destination-TU derivation, or carve state.
|
||||
|
||||
The 888-section cookbook cannot help these — there is no compiler idiom left to learn; the
|
||||
residual lives between the object file and the linked image. Meanwhile the fleet keeps
|
||||
re-drafting them: the reject ledger holds 10,049 rows over just **574 distinct fns**, the same
|
||||
~115 "MISMATCH?" drafts recurring wave after wave. That is the single largest measured waste in
|
||||
the current architecture — free tokens, but each re-draft also re-enters the gate and pays §3's
|
||||
failure-proportional cost.
|
||||
|
||||
**One caveat, probed tonight:** the stock is stale. func_80184A68's stored closeness-0 draft
|
||||
(recorded 2026-07-01) now COMPILE-FAILs even standalone — the fleet's declarations moved under
|
||||
it. So the resolver lane below must **re-verify every item at intake against today's tree, at
|
||||
the real TU** (rtu_match), and treat the ledger numbers as an index, not a promise.
|
||||
|
||||
### The integration-resolver lane (new tool; the highest-EV build)
|
||||
|
||||
Zero tokens, bounded, built almost entirely from existing parts:
|
||||
|
||||
- **Inputs:** backlog rows with closeness 0 for still-open fns (+ their `.run/backlog_drafts/`
|
||||
bodies); `reloc_rejects.jsonl` shape-MATCH rows; optionally the 73 closeness-1–2 fns for the
|
||||
grinder handoff.
|
||||
- **Pipeline per item:** (1) still an open stub? else drop (report, don't skip silently);
|
||||
(2) `rtu_match` against the real TU → MATCH: stage to `.run/sweep_maint/<bin>/`;
|
||||
(3) rtu CC1-FAIL: run the existing recovery chain against the *real* TU (`reconcile_tu`,
|
||||
`cast_call_sites`, arity probe) and retry rtu once; (4) rtu DIFF: run `reloc_identity`; on
|
||||
MISMATCH, `aprop_symfix --fix` rebase and retry rtu once; (5) still DIFF: return the fn to the
|
||||
redraft pool with the *current* closeness (the stored one is stale) — honest demotion.
|
||||
- **Outputs:** staged drafts for the maintenance lane's existing free gate; a per-item verdict
|
||||
ledger (every drop named — R32).
|
||||
- **Refusal conditions:** fn not an open stub; draft file missing; rtu work dir cannot be built
|
||||
(report as harness, not as a function verdict); more than one candidate body for a fn without a
|
||||
match_one tiebreak.
|
||||
- **Expected yield, honestly bounded:** analogous recoveries measured 7% (rtu_second_chance,
|
||||
3/42) to 18.5% (tonight's clean ei re-gate, 34/184) to 100% (aprop_symfix, 4/4 on its exact
|
||||
class). On 571 fns that is **~40–110 banks for zero tokens**, plus it permanently drains the
|
||||
re-draft loop. Success test: banks per gate minute of the staged output vs the wave lane's
|
||||
0.1–0.7. Falsifier: an intake pass where < 5% of items survive to staging — then the stock was
|
||||
stale beyond recovery and the 571 return to the redraft pool with fresh closeness numbers,
|
||||
which is itself worth having.
|
||||
|
||||
---
|
||||
|
||||
## 5. The smartest path to 100%, sequenced
|
||||
|
||||
Each step changes what the next one costs; this order is deliberate.
|
||||
|
||||
1. **Re-gate the wipe-window waves (now; zero tokens, no new code).** ej/ek/el/em drafts still
|
||||
sit in `.run/wave_*/`; the ready queue already re-gates ep–es. Route them through the gater's
|
||||
own queue (not a parallel lane — §3c). This yields the *true* current wave conversion, which
|
||||
step 6 depends on. Evidence it worked: ei-style recoveries (34 banks) or honest zeros on a
|
||||
clean tree.
|
||||
2. **Ship the integration-resolver lane (§4) and run it to exhaustion.** It attacks the
|
||||
largest measured stock (571 fns) at the best measured cost (zero tokens, ~1 build per
|
||||
surviving item) and shrinks every later wave by ending the re-draft loop.
|
||||
3. **Shadow-mode rtu_match across 2–3 gates, then invert the gate (§3).** After inversion the
|
||||
gate stops being scarce, which re-prices everything else (the maintenance lane currently waits
|
||||
for "gater idle"; wave size caps exist only because gates are slow).
|
||||
4. **Regenerate the A-prop cards and re-run the free lane on cadence** (576 families /
|
||||
1,071 members / 52,289 ins remain as of 17:23; PURE 761). It banked 357 today in 32 minutes.
|
||||
Add the **decl-from-use inference tool** to unblock the "no seed decl" refusals
|
||||
(aprop_autodraft.py:522): infer each missing extern from the member's own target `.s` — access
|
||||
widths (lb/lbu/lh/lhu/lw/sw…), sign, index scale (sll before addu), call arity from jal + arg
|
||||
register writes — and emit minimal C89 externs; refuse on conflicting width evidence at the
|
||||
same symbol. Owner sizes this class at 121 members; re-derive the count when building.
|
||||
5. **Keep the family flywheel behind every crack** (regen atlas → family_remap on the 304
|
||||
single-skeleton multi-groups; 1,334 siblings, 93k ins) — but as a *consequence* of cracks,
|
||||
not the strategy. The majority of remaining mass (57%) is singletons.
|
||||
6. **Re-aim the free fleet** (tokens are free; aim is the scarce thing), informed by step 1's
|
||||
clean conversion number:
|
||||
- stop drawing any fn in the resolver's stock (a draw-time filter on backlog closeness 0 /
|
||||
shape-MATCH — one predicate in build_wave_atlas);
|
||||
- draft the never-gated pockets: 433 gen0 instances + whatever step 1 shows still converts;
|
||||
- **main**: 313 crackable stubs, 60k weighted ins open, its own clean-rebuild gate, banked 255
|
||||
today — the largest coherent mass and the least survivor-hardened;
|
||||
- the 3–8-closeness band (~29 wall fns + the band below 6 attempts): these are the genuine
|
||||
near-miss codegen residuals — permuter/grinder food first (73 fns at closeness 1–2 are
|
||||
pure `tools/grinder.py` territory), fleet redrafts only where the grinder stalls.
|
||||
7. **A small escalation tier for documented model plateaus** (§2.6): after the resolver drains
|
||||
the fake walls, take the residual true-DIFF wall (my measurement: ~79 fns at closeness 3+,
|
||||
≥6 refusals), order by unlock value (`group_ins`), and give the top 20–30 to a stronger model
|
||||
with the full §31-style context pack. The bake-off row proves at least part of this stratum is
|
||||
matchable-but-not-by-ox. Quote the result per model; kill the tier if 10 attempts bank 0.
|
||||
8. **Fix the two measurement leaks so the campaign can see itself:** count "banked today" from
|
||||
the INCLUDE_ASM invariant (§1); and make `wall_min` in the gate log measure the gate, with
|
||||
drafting/queue latency reported separately.
|
||||
|
||||
**What should NOT be done:** more wave-side tuning (band/mix/card-count) before steps 1–3 —
|
||||
S58–S60 already spent several sessions there, and every lever moved single-digit percent while
|
||||
the architecture left 18% of the pool re-drafting solved functions and paid 3 builds per failure.
|
||||
|
||||
---
|
||||
|
||||
## 6. Dead ends declared (so the next session doesn't re-walk them)
|
||||
|
||||
- **Lead "the gate is mysteriously slow at 8% CPU":** resolved, not mysterious — a metric
|
||||
artifact (wall_min includes drafting), a real serialization bug already fixed at 10:43
|
||||
(commit:2844), failure-proportional ladder cost (§3), and probable cross-gate lock contention.
|
||||
A raw "make builds parallel within a binary" project is NOT the fix; builds within a binary are
|
||||
serial by design (each mutates the same tree), and the win is not running them at all (§3).
|
||||
- **Lead "the gen6+ wall hides an unnamed cookbook class":** refuted for the majority — 61% of
|
||||
the wall already matched at the object level; their blocker is integration, which no cookbook
|
||||
section can address. The honest codegen wall is ~79 open fns at closeness 3+ (plus 50 at 9+
|
||||
that may include mis-scored artifacts — the S59 lesson that a wrong asm-subdir or opt-level
|
||||
manufactures phantom residuals applies; re-score at intake).
|
||||
- **"Uncollapsed wave eh was an outlier proving siblings should be drafted":** directionally
|
||||
right but the stronger events were dd–dg the same morning (217/422 … 100/283). The lesson is
|
||||
fully priced in already (`ONE_PER_GID=0` landed today); there is no further sibling-drafting
|
||||
lever of that size left — the pool is spent (§2.1).
|
||||
- **Chunked gating (chunk > 1) as a speedup:** rejected in the tree's own history — chunked
|
||||
failures mis-attribute innocent neighbours (cookbook:1568) and the bisect was measured at 1.35
|
||||
builds/draft vs 1.0. The rtu inversion supersedes this entire axis.
|
||||
|
||||
## Appendix A — how every number was derived (all read-only)
|
||||
|
||||
- Wave conversions/walls: `.run/gater.log` (`=== GATE <tag>` to `GATE <tag>: banked` intervals);
|
||||
`.run/ox_campaign_ledger.jsonl` for lane/band/targets.
|
||||
- Wipe windows: `git log/show commit:2863 commit:2913 commit:2915` + `config_sane()` docstring in
|
||||
`tools/ox_campaign.py`.
|
||||
- Bank attribution: for each `git log --since=00:00` commit, `git show --unified=0 -- src/ |
|
||||
grep -c '^-.*INCLUDE_ASM('`; net: `git grep -c 'INCLUDE_ASM(' <midnight-commit>|HEAD -- src/`.
|
||||
- Atlas: `.run/atlas.json` (head commit:2911, 17:25) — groups/cat/inst/n_skel/ins/lever/members.
|
||||
- Generations: fn occurrences across `.run/wave_*_cards.json` (204 files) joined to atlas members.
|
||||
- Stock: `.run/backlog.jsonl` + `.run/auto/bulk/*.backlog.jsonl` (16,837 rows; closeness by fn),
|
||||
`.run/reloc_rejects.jsonl` (10,049 rows; `shape` field), `.run/backlog_drafts/` (5,010 files),
|
||||
joined to atlas open set.
|
||||
- Gate mechanics: `tools/sweep_parallel.py`, `tools/gate_stage.py`, `tools/harvest_verify.py`
|
||||
read in full; live `ps` during the ep/ek gates.
|
||||
- A-prop: `.run/aprop_cards.json` (17:23), `.run/maintenance.log` 15:54–16:26 block.
|
||||
- Probe: `tools/reloc_identity.py --fn func_80184A68 --binary ov_SC01_077 --c
|
||||
.run/backlog_drafts/func_80184A68.c …` → COMPILE-FAIL (decl drift since 07-01).
|
||||
@@ -112,9 +112,14 @@ writes `docs/tool-designs/frontier-analysis-s60.md` — **READ THAT FIRST NEXT S
|
||||
### THE STRATEGIC PICTURE (this is what next session must act on)
|
||||
* **3,062 open crackable.** main is **313** crackable, not 1,274 — 961 of its stubs are LINKED PsyQ
|
||||
segments and data blobs, linked byte-exact, never decompiled.
|
||||
* The drawable pool collapses to **~334 distinct skeletons, ~308 of them generation 6+** (drafted six
|
||||
or more times, refused every time). ~3,900 open functions are SIBLINGS behind those skeletons and
|
||||
bank by mechanical remap once an exemplar cracks.
|
||||
* **CORRECTED 19:45 by the Fable audit, and I had this wrong all session.** The atlas at commit:2911
|
||||
measures: 3,106 open instances · 2,245 open skeletons · 1,772 groups, of which **480 are
|
||||
multi-member holding 1,334 siblings** and **1,292 are SINGLETONS carrying 57% of the open
|
||||
instruction mass**. My "~3,900 siblings behind ~334 skeletons" conflated two different
|
||||
populations — the never-drafted stub count with the sibling count — and overstated remap leverage
|
||||
by ~3x. Most remaining work is singletons that each need their own crack. The "~334 drawable,
|
||||
308 gen6+" figure describes only the COLLAPSED wave-eligible view; whole-pool generation is 53%
|
||||
gen0/1 and 25% gen6+, and only **292 functions are 6+ GATE-refused**.
|
||||
* **Wide waves now convert at 1-5%.** The drafting side is solved; the gate is the bottleneck (30-67
|
||||
min per gate at 1-3 concurrent builds, load 2.6 of 32 cores) and the drafter parks when its queue
|
||||
fills. **Optimise banks per GATE MINUTE, not cards per wave.**
|
||||
@@ -122,8 +127,22 @@ writes `docs/tool-designs/frontier-analysis-s60.md` — **READ THAT FIRST NEXT S
|
||||
correlated with bank rate — the best waves had the MOST truncation (cx 8.7% trunc / 43.9% bank,
|
||||
dd 8.3% / 51.4%) and the dead waves the least (dl 1.3% / 0.5%, ej 0.6% / 0%). The zeros are the
|
||||
registry outages; the slide from 51% to 20% is population generation, not agent budget.
|
||||
* **5,388 backlog rows at closeness <= 2** — the grinder (tools/grinder.py, Phase 21, LLM-free,
|
||||
decomp-permuter) had NEVER been run this campaign and is now on them.
|
||||
* **The wall is an INTEGRATION wall, not a codegen wall** (Fable's headline, and the highest-value
|
||||
finding of the day): of the 292 functions the gate has refused 6+ times, **178 (61%) already
|
||||
produced a closeness-0 draft** — match_one byte-equality, whole-binary gate rejection. The blocker
|
||||
is symbols/decls/TU plumbing, and the fleet keeps re-drafting them (10,049 reject rows over 574
|
||||
distinct functions). Its top recommendation is a ZERO-TOKEN integration-resolver lane
|
||||
(rtu_match -> reconcile/cast/arity -> reloc_identity -> symfix -> stage to the maintenance gate).
|
||||
* **5,388 backlog rows at closeness <= 2** de-duplicate to **~543 open functions** (470 at 0, 73 at
|
||||
1-2) — the grinder (tools/grinder.py, Phase 21, LLM-free, decomp-permuter) had NEVER been run this
|
||||
campaign and is now on them. My own re-measure at 19:40 found 290 still-open closeness-0 functions,
|
||||
down from Fable's 470: the re-gate and grinder are draining exactly this pool.
|
||||
* **Gate cost is proportional to FAILURES, not drafts** (measured mechanics): chunk=1 plus the 3-stage
|
||||
ladder means ~3 whole-binary builds per FAILING draft, serial per binary — dd (51% conversion) ran
|
||||
1.8 s/draft, eo (1.4%) 8.6 s/draft, and banks-per-gate-minute fell 17.5 -> 0.10. My "30-67 min
|
||||
gates at 8% CPU" conflated wall_min (which includes drafting and queue time) with gate wall
|
||||
(12-31 min healthy). tools/rtu_match.py inverted into the gate's stage 0 would make builds
|
||||
proportional to BANKS (~20x fewer at current conversion) — shadow-run it over 2-3 gates first.
|
||||
|
||||
### OPEN THREADS (ranked)
|
||||
1. **Read `docs/tool-designs/frontier-analysis-s60.md`** — the Fable analyst was told we are NOT married
|
||||
|
||||
Reference in New Issue
Block a user