From 07a167516d9bf07866b694513ce4509d5edef2cc Mon Sep 17 00:00:00 2001 From: Drew T <50529377+Druthulu@users.noreply.github.com> Date: Tue, 25 Aug 2026 19:36:51 -0600 Subject: [PATCH] docs(s60): the Fable frontier audit + corrections it forced to my own checkpoint MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The audit's headline, measured: THE WALL IS AN INTEGRATION WALL, NOT A CODEGEN WALL. Of the 292 functions the gate has refused 6+ times, 178 (61%) have ALREADY produced a closeness-0 draft — match_one byte-equality, whole-binary gate rejection. The blocker is symbols/decls/TU plumbing, and the fleet keeps re-drafting them: 10,049 reject rows over 574 distinct functions. Highest-EV build is a zero-token integration-resolver lane, not more drafting. CORRECTIONS TO MY OWN NUMBERS, verified against the tree before accepting: * siblings are 1,334 behind 480 multi-member groups, NOT ~3,900. 1,292 groups are SINGLETONS carrying 57% of open instruction mass. I conflated the never-drafted stub count with the sibling count and overstated remap leverage ~3x, in this checkpoint and repeatedly in conversation. * 'everything drawable is gen6+' holds only for the collapsed wave-eligible view; whole-pool generation is 53% gen0/1, 25% gen6+, and only 292 fns are 6+ GATE-refused. * '30-67 min gates at 8% CPU' conflated wall_min (includes drafting/queue) with gate wall (12-31 min healthy). Gate cost is proportional to FAILURES, not drafts: ~3 whole-binary builds per failing draft, so banks/gate-min fell 17.5 -> 0.10 as conversion fell. * the 5,388 closeness<=2 rows de-dupe to ~543 open functions; my own 19:40 re-measure found 290 still open, down from its 470 — the re-gate and grinder are draining that pool now. * campaign_status's 'banked today' undercounts: the stub invariant says ~2,644 net, because the A-prop lane's 357 rode in a chore commit its regex cannot see. One documented counterexample to 'model quality is not a bottleneck': func_80181714, where ox-alpha plateaued at closeness 4 while Opus/GLM/DeepSeek each reached reloc-verified MATCH — argues for a small escalation tier AFTER the resolver drains the fake walls. Taken on trust and flagged as such: the A-prop residual split (169 STRUCT / 121 no-seed-decl / 73 IMM / 12 void) — the refusal mechanisms exist in aprop_autodraft.py but no file carries those counts; re-derive before building the decl-inference tool. --- docs/tool-designs/frontier-analysis-s60.md | 350 +++++++++++++++++++++ phase-ends/CURRENT_PHASE.md | 29 +- 2 files changed, 374 insertions(+), 5 deletions(-) create mode 100644 docs/tool-designs/frontier-analysis-s60.md diff --git a/docs/tool-designs/frontier-analysis-s60.md b/docs/tool-designs/frontier-analysis-s60.md new file mode 100644 index 000000000..2b7a3de9f --- /dev/null +++ b/docs/tool-designs/frontier-analysis-s60.md @@ -0,0 +1,350 @@ +# Frontier analysis — S60 (2026-08-25, evening) + +**What this is.** A read-only analysis of the remaining unbanked population and of where the +next banks actually come from, written for a session that has NO memory of today. Every number +below carries its denominator and its source; the appendix lists the exact commands so any figure +can be re-derived against the tree. Nothing in the tree was modified to produce this document. + +**The one-paragraph conclusion.** The wide-wave drafting machine has finished the job it was +built for. Its target population — never-drafted, small, seed-adjacent functions — was consumed +today at 35–51% conversion, and what the waves now re-draft converts at 1–6% while costing +30 minutes of gate each. But the campaign's own ledgers show that roughly **571 still-open +functions have already been drafted correctly** (their instruction stream matched the target at +the object level; the whole-binary gate rejected them for symbol/declaration/TU reasons), and the +gate architecture spends ~3 whole-binary builds per *failing* draft while a real-TU oracle that +needs no builds already exists in the tree (`tools/rtu_match.py`). The smartest path is to stop +optimizing the wave and instead (1) re-gate what a config wipe falsely refused today, (2) build a +deterministic **integration-resolver lane** over the shape-correct stock, (3) invert the gate so +whole-binary builds are proportional to *banks* rather than *drafts*, and (4) re-aim the free +fleet at the strata where drafting is genuinely unfinished — never-gated pockets, main, and the +3–8-mismatch band — instead of re-drafting functions whose drafts are already correct. + +--- + +## 0. Trust ledger — surveyed ground vs hearsay + +**Verified against the tree by me tonight (commands in Appendix A):** +atlas composition and sibling counts; the generation histogram; every wave conversion number and +gate wall time (from `.run/gater.log` timestamps); the two `config/overlays.mk` wipe windows and +their commit timestamps; the gate ladder's build accounting (from reading +`tools/gate_stage.py` / `tools/harvest_verify.py` / `tools/sweep_parallel.py` in full); the +backlog/reject-ledger stock counts (16,837 + 10,049 rows re-parsed); today's bank attribution +(per-commit `INCLUDE_ASM` removal diffs over all 162 commits); the A-prop pool size; one live +probe of a stored closeness-0 draft (`func_80184A68`); the live process table during a running +gate. + +**Repo-recorded measurements I did not independently re-derive** (they are written into tool +docstrings/lane scripts as measured, with dates): the bank-rate-by-size table (S59, in +`.run/drafter.sh`); "reloc pre-filter drops 45% of drafts / 13% of MISMATCH? are shape-MATCH" +(in `tools/recover_rejects.py`); rtu_second_chance's 43-of-182 / 3-of-42 figures (in +`tools/rtu_second_chance.py` + `.run/maintenance.log`); the func_80181714 three-model MATCH note +(a backlog row). + +**Taken on trust from the owner's briefing and NOT reproduced:** the A-prop residual split +"169 STRUCT / 121 no-seed-decl / 73 IMM-unresolved / 12 void near-0" (I verified the *mechanisms* +exist — `aprop_autodraft.py:522` refuses on "no seed decl", `:592` on STRUCT — but could not find +a file carrying those exact counts; re-derive them before building against them). The "961 linked +PsyQ stubs in main" figure. + +**Owner claims that did NOT survive contact with the tree** — see §2: the ~3,900-sibling figure, +the "everything drawable is gen6+" framing as a statement about the pool, the "0–14 of ~220" +collapse (partly a harness defect), and the 30–67-min/8%-CPU gate picture (a metric artifact plus +an already-landed fix). + +--- + +## 1. State of the campaign as measured tonight + +Fleet (docs/progress.fleet.md, regenerated by progress.py): +instr-weighted **97.4%** (13,171,872 / 13,523,865); distinct-code **94.6%** (86,568 / 90,929 +unique fns); **main game-code 23.5%** (18,697 / 79,510 weighted) — main is the largest coherent +open mass left. Overlays excluding main: 97.8%. + +Open pool (`.run/atlas.json`, regenerated 17:25 today at head commit:2911): **3,106 open non-main +instances** in 1,772 structural groups (2,245 distinct skeletons), plus **313 crackable main +stubs** (atlas `main_open`; consistent with the owner's number). + +| category | groups | instances | note | +|---|---:|---:|---| +| A-prop | 520 | 1,177 | banked seed + mapping exists; deterministic-lane food | +| cold | 715 | 715 | all singletons, no seed | +| cousin-multi | 185 | 717 | structural cousins; only 106 groups single-skeleton | +| main-only | 149 | 149 | main lane's territory | +| seeded | 178 | 297 | | +| tiny | 25 | 51 | | + +Sibling leverage, measured: **480 multi-instance groups hold 1,814 instances (93,165 ins); +siblings = 1,334**. Of those 480 groups only **304 share a single skeleton** (true +crack-one-remap-the-rest); 92 are all-distinct cousins, 84 mixed. **Singletons: 1,292 instances / +121,853 ins — 57% of the open non-main instruction mass has no family leverage at all.** + +Generation (times a fn appeared in any of the 204 wave card files), over the 3,106 open +instances: gen0 **433**, gen1 **1,208**, gen2–5 **675**, gen6+ **790**. Distinct-fn gate +*attempts* are much rarer than draws (the reloc pre-filter and draft failures eat the difference): +only **292 open fns have ≥6 recorded gate attempts**. + +Today's production (00:00 → HEAD, measured from the `INCLUDE_ASM` invariant, not commit +subjects): stubs went **6,657 → 4,013 (−2,644 net; 2,877 gross removals** across 162 commits; +carves re-add stubs, hence gross > net). Attribution by diffing every commit: + +| lane | stub removals | note | +|---|---:|---| +| ox waves (~30 gates) | 2,031 | almost all of it before 13:00 — see the arc below | +| deterministic (A-prop 357, −O0 ~102, misc maint/serial) | ~591 | **zero model tokens** | +| main lane | 255 | its own clean-rebuild gate | + +`campaign_status.py`'s "today: 2185 banked" is a **regex undercount** — it sums "— N banked" +commit subjects, and e.g. the A-prop lane's 357 banks rode in a chore commit (commit:2904, 358 +removals) that the regex cannot see. Count banks from the stub invariant, not from subjects +(the derive-from-invariants rule; this is the R32 "silently narrowed scope" class again). + +### 1a. The arc of today, per wave (banked / gated, from `.run/gater.log` intervals) + +dd **217/422** (12.4 min gate) · de **192/388** · df **124/278** · dg **100/283** · di **81/268** +· dj **43/219** — these are the sibling-inclusive draws (`ONE_PER_GID=0` landed this morning) +eating the never-drafted pool at 20–51%. + +Then **dk 2/201** — the cliff. **It coincides exactly with the first `config/overlays.mk` wipe** +(dk's own gate commit commit:2863 at 12:37 committed the registry as a zero-line file). With no +binaries registered, overlay builds cannot succeed; the gate verdicts of that window are +measurements of a broken harness, not of the drafts. + +Recovery + mixed period: dl 12/160, dm 3/149 (re-gates), dp 54/233, dr 28/188, dt 9/169, +dq 8/167, ds 35/159, du 3/159, dx 21/158, dy 12/127, eb–ee 1–10 each, **eh 129/380** (a fresh +625-target draw). Then **ei 1/221, ej 0/215, ek 0/220, el 0/211, em 1/225 — all inside wipe #2** +(committed by ei's gate commit ~17:14, restored 17:43:55 = commit:2913, root-caused 17:48:45 = +commit:2915: four truncating `open(mk,"w").write()` sites in `jr_isolate_all.py:593` and +`jtbl_carve.py:1077/1118/1165`). The giveaway is in the gate walls: those five "gates" ran +**3.4–6 min for ~220 drafts each** — instant build failures, not judgments. Post-fix, honest: +**en 14/227 (6.2%), eo 3/216 (1.4%), and a clean re-gate of ei banked 34/184 (18.5%) where its +first gate banked 1** (re-gate of ej: 0/186 — so re-gating recovers some waves, not all). + +**Banks per gate minute** — the metric that binds: dd ≈ **17.5**, de ≈ 12.6, eh ≈ 6.0, +en ≈ 0.69, eo ≈ **0.10**. The A-prop maintenance pass banked **357 in one ~32-min sweep ≈ 11/min +at zero tokens** (15:54–16:26, `.run/maintenance.log`). + +--- + +## 2. Corrections to the going narrative (each measured tonight) + +1. **"~3,900 siblings behind ~334 skeletons" is stale.** The 17:25 atlas shows **1,334 siblings + behind 480 multi-instance groups**, only 304 of them single-skeleton. Today's dd–dj+eh burst + consumed the rest. 57% of remaining open instruction mass is **singletons** — "crack a skeleton, + unlock many" now describes a minority of the frontier. +2. **"The drawable pool is ~all gen6+" is a statement about the collapsed wave-eligible view, + not the pool.** Over all open instances: 53% are gen0/gen1, 25% gen6+. And *drawn* ≠ *gated*: + only 292 open fns have actually been refused by the gate 6+ times. +3. **The five zero-waves were partly manufactured.** ej–em were gated against a wiped registry + (§1a). The clean-tree conversion tonight is 1.4–6.2% on first gates and up to 18.5% on + re-gates — bad, but not zero, and every conclusion drawn from that window needs the re-gate + first (R40: exonerate the instrument). +4. **"A gate takes 30–67 min at 8% CPU" conflates three things.** The `GATE … · NN.Nmin` figure + in the log is `time.time() − wave_draft_t0` — it includes drafting and queue wait. Real gate + walls today: 12–31 min healthy, 3–6 min when broken. The "-j was decorative" serialization + (every worker taking the fleet lock exclusively because `GATE_NO_ARITY` was unset) was + diagnosed and **fixed at 10:43 today** (commit:2844). What remains is §3's failure-proportional + cost plus (inferred, see §3c) cross-gate lock contention. +5. **"5,388 backlog rows at closeness ≤2" is a row count, not a workload.** I count 8,596 such + rows across the ledgers — but rows are re-attempts of the same functions. The de-duplicated, + still-open workload is **~543 functions: 470 that have hit closeness 0 and 73 whose best is + 1–2**. The 5,010 files in `.run/backlog_drafts/` are the stock behind them. +6. **"Model quality is not a bottleneck" has a documented counterexample stratum.** Backlog row + func_80181714 (ov_SC02_016, 121 ins): ox-alpha plateaued at closeness 4 over 20 oracle calls + while Opus, GLM-5.3 and fueled DeepSeek each reached reloc-verified MATCH. One case, recorded + during a deliberate bake-off — enough to justify a small escalation tier (§5, step 7), not a + fleet migration. + +--- + +## 3. Where the gate minutes actually go (tools read in full) + +The pipeline: `ox_campaign --gater` → reloc_identity pre-filter (drops ~45% of drafts) → +`sweep_parallel -j N` (one worker per **binary**; per-binary flock; parallel across binaries +only) → per binary, `gate_stage.run_gate` runs a **three-stage ladder**, and each stage calls +`harvest_verify --chunk 1`, i.e. **one incremental whole-binary build per draft per stage**: + +- stage 0: raw drafts — D builds; banked drafts exit here; +- transforms on the failures (canon_resident_calls, cast_call_sites, reconcile_tu, arity + pre-pass — subprocesses, cheap relative to builds); +- stage 1: D′ builds; sig_unify; stage 2: D″ builds; +- plus one final confirmation build per harvest_verify invocation (3 per binary), plus one + `match_one` compile per still-failing draft for backlog closeness. + +So a fully-failing binary with D drafts costs **≈ 3D + 3 builds + D match_one runs**, serial +within the binary; a fully-banking one costs ≈ D + 1. **Gate cost is proportional to failures, +not drafts** — which is why dd (51% conversion) ran 1.8 s/draft and eo (1.4%) ran 8.6 s/draft. +The transforms exist to rescue PLUMBING-class failures; the recent failure-class mix +(`.run/auto/bulk/*.failed.classified.txt`, 332 rows: 74 DIFF · 63 CC1-FAIL · 51 PLUMBING · +32 CARVE-REFUSED · 112 unclassified) says a large fraction of ladder re-builds are spent +re-gating **DIFF** drafts — ones that already compiled and linked and produced wrong bytes, which +no declaration rewrite will change. + +**(3c, inferred)** Cross-gate lock contention: gates dr/dt/dq ran 27–34 min while a manual +re-gate lane held overlapping per-binary flocks (13:22–15:00); ds/du, same size and shape, ran +15–16 min once it ended. Two gates over the same binaries serialize worker-by-worker and the +blocked workers occupy pool slots. Not proven causal — but it costs nothing to schedule re-gates +into the gater's own queue instead of beside it. + +**The structural fix is not more parallelism; it is making builds proportional to banks.** +`tools/rtu_match.py` already compiles the *real TU* with the candidate spliced in — "no build +tree, no locks, parallel-safe" (its own docstring) — and returns MATCH/DIFF/CC1 with real +diagnostics. It is currently used only as a second-chance recovery. Inverted, it becomes the +gate's stage 0: run rtu_match on all drafts 32-wide (no flocks, no link, no SHA), and spend +whole-binary builds **only on rtu-MATCH drafts** plus the PLUMBING-recovery ladder for rtu-CC1 +failures whose diagnostics name a fixable conflict. At 2% conversion that is roughly a 20× +reduction in whole-binary builds; the gate for a 200-draft wave becomes ~2 min of parallel TU +compiles plus a handful of builds. G3/P9 is untouched — the whole-binary SHA gate remains the +sole arbiter of banking; rtu only decides which drafts get to spend a build. + +**Risk, and the cheap way to bound it:** rtu false-negatives (drafts rtu rejects that the +whole-binary gate would have banked). Run rtu_match in **shadow mode** for 2–3 gates first — +record its verdict for every gated draft next to the gate's verdict, zero behavioral change. If +the false-negative rate is under ~2%, flip the order and keep a nightly sweep-up re-gate of +rtu-rejects as the safety net. If it is higher, the shadow data says exactly which class rtu +misjudges, and that class stays on the old path. Falsifier for the whole idea: a shadow run +showing rtu-MATCH does not predict banking (agreement < ~90%) — then the gate stays as it is and +the win must come from §4/§5 alone. + +--- + +## 4. The finding that reframes the endgame: the wall is an integration wall + +Of the **292** open functions the gate has refused six or more times, **178 (61%) have already +produced a closeness-0 draft** — `match_one` byte-equality at the object level, whole-binary gate +rejection. 27 more sit at closeness 1–2, 29 at 3–8, 50 at 9+, 8 unscored. Extending past the +6+-refused wall to the whole open pool: + +- **470 open fns** have a closeness-0 row in the backlog, best draft saved on disk; +- **151 open fns** appear in `.run/reloc_rejects.jsonl` with `shape: MATCH` (instruction stream + already matches; only the symbol names are wrong — the deterministic `aprop_symfix` class); +- union: **571 open fns ≈ 18% of the open non-main pool** where *drafting is finished* and what + remains is symbol identity, declaration conflicts, destination-TU derivation, or carve state. + +The 888-section cookbook cannot help these — there is no compiler idiom left to learn; the +residual lives between the object file and the linked image. Meanwhile the fleet keeps +re-drafting them: the reject ledger holds 10,049 rows over just **574 distinct fns**, the same +~115 "MISMATCH?" drafts recurring wave after wave. That is the single largest measured waste in +the current architecture — free tokens, but each re-draft also re-enters the gate and pays §3's +failure-proportional cost. + +**One caveat, probed tonight:** the stock is stale. func_80184A68's stored closeness-0 draft +(recorded 2026-07-01) now COMPILE-FAILs even standalone — the fleet's declarations moved under +it. So the resolver lane below must **re-verify every item at intake against today's tree, at +the real TU** (rtu_match), and treat the ledger numbers as an index, not a promise. + +### The integration-resolver lane (new tool; the highest-EV build) + +Zero tokens, bounded, built almost entirely from existing parts: + +- **Inputs:** backlog rows with closeness 0 for still-open fns (+ their `.run/backlog_drafts/` + bodies); `reloc_rejects.jsonl` shape-MATCH rows; optionally the 73 closeness-1–2 fns for the + grinder handoff. +- **Pipeline per item:** (1) still an open stub? else drop (report, don't skip silently); + (2) `rtu_match` against the real TU → MATCH: stage to `.run/sweep_maint//`; + (3) rtu CC1-FAIL: run the existing recovery chain against the *real* TU (`reconcile_tu`, + `cast_call_sites`, arity probe) and retry rtu once; (4) rtu DIFF: run `reloc_identity`; on + MISMATCH, `aprop_symfix --fix` rebase and retry rtu once; (5) still DIFF: return the fn to the + redraft pool with the *current* closeness (the stored one is stale) — honest demotion. +- **Outputs:** staged drafts for the maintenance lane's existing free gate; a per-item verdict + ledger (every drop named — R32). +- **Refusal conditions:** fn not an open stub; draft file missing; rtu work dir cannot be built + (report as harness, not as a function verdict); more than one candidate body for a fn without a + match_one tiebreak. +- **Expected yield, honestly bounded:** analogous recoveries measured 7% (rtu_second_chance, + 3/42) to 18.5% (tonight's clean ei re-gate, 34/184) to 100% (aprop_symfix, 4/4 on its exact + class). On 571 fns that is **~40–110 banks for zero tokens**, plus it permanently drains the + re-draft loop. Success test: banks per gate minute of the staged output vs the wave lane's + 0.1–0.7. Falsifier: an intake pass where < 5% of items survive to staging — then the stock was + stale beyond recovery and the 571 return to the redraft pool with fresh closeness numbers, + which is itself worth having. + +--- + +## 5. The smartest path to 100%, sequenced + +Each step changes what the next one costs; this order is deliberate. + +1. **Re-gate the wipe-window waves (now; zero tokens, no new code).** ej/ek/el/em drafts still + sit in `.run/wave_*/`; the ready queue already re-gates ep–es. Route them through the gater's + own queue (not a parallel lane — §3c). This yields the *true* current wave conversion, which + step 6 depends on. Evidence it worked: ei-style recoveries (34 banks) or honest zeros on a + clean tree. +2. **Ship the integration-resolver lane (§4) and run it to exhaustion.** It attacks the + largest measured stock (571 fns) at the best measured cost (zero tokens, ~1 build per + surviving item) and shrinks every later wave by ending the re-draft loop. +3. **Shadow-mode rtu_match across 2–3 gates, then invert the gate (§3).** After inversion the + gate stops being scarce, which re-prices everything else (the maintenance lane currently waits + for "gater idle"; wave size caps exist only because gates are slow). +4. **Regenerate the A-prop cards and re-run the free lane on cadence** (576 families / + 1,071 members / 52,289 ins remain as of 17:23; PURE 761). It banked 357 today in 32 minutes. + Add the **decl-from-use inference tool** to unblock the "no seed decl" refusals + (aprop_autodraft.py:522): infer each missing extern from the member's own target `.s` — access + widths (lb/lbu/lh/lhu/lw/sw…), sign, index scale (sll before addu), call arity from jal + arg + register writes — and emit minimal C89 externs; refuse on conflicting width evidence at the + same symbol. Owner sizes this class at 121 members; re-derive the count when building. +5. **Keep the family flywheel behind every crack** (regen atlas → family_remap on the 304 + single-skeleton multi-groups; 1,334 siblings, 93k ins) — but as a *consequence* of cracks, + not the strategy. The majority of remaining mass (57%) is singletons. +6. **Re-aim the free fleet** (tokens are free; aim is the scarce thing), informed by step 1's + clean conversion number: + - stop drawing any fn in the resolver's stock (a draw-time filter on backlog closeness 0 / + shape-MATCH — one predicate in build_wave_atlas); + - draft the never-gated pockets: 433 gen0 instances + whatever step 1 shows still converts; + - **main**: 313 crackable stubs, 60k weighted ins open, its own clean-rebuild gate, banked 255 + today — the largest coherent mass and the least survivor-hardened; + - the 3–8-closeness band (~29 wall fns + the band below 6 attempts): these are the genuine + near-miss codegen residuals — permuter/grinder food first (73 fns at closeness 1–2 are + pure `tools/grinder.py` territory), fleet redrafts only where the grinder stalls. +7. **A small escalation tier for documented model plateaus** (§2.6): after the resolver drains + the fake walls, take the residual true-DIFF wall (my measurement: ~79 fns at closeness 3+, + ≥6 refusals), order by unlock value (`group_ins`), and give the top 20–30 to a stronger model + with the full §31-style context pack. The bake-off row proves at least part of this stratum is + matchable-but-not-by-ox. Quote the result per model; kill the tier if 10 attempts bank 0. +8. **Fix the two measurement leaks so the campaign can see itself:** count "banked today" from + the INCLUDE_ASM invariant (§1); and make `wall_min` in the gate log measure the gate, with + drafting/queue latency reported separately. + +**What should NOT be done:** more wave-side tuning (band/mix/card-count) before steps 1–3 — +S58–S60 already spent several sessions there, and every lever moved single-digit percent while +the architecture left 18% of the pool re-drafting solved functions and paid 3 builds per failure. + +--- + +## 6. Dead ends declared (so the next session doesn't re-walk them) + +- **Lead "the gate is mysteriously slow at 8% CPU":** resolved, not mysterious — a metric + artifact (wall_min includes drafting), a real serialization bug already fixed at 10:43 + (commit:2844), failure-proportional ladder cost (§3), and probable cross-gate lock contention. + A raw "make builds parallel within a binary" project is NOT the fix; builds within a binary are + serial by design (each mutates the same tree), and the win is not running them at all (§3). +- **Lead "the gen6+ wall hides an unnamed cookbook class":** refuted for the majority — 61% of + the wall already matched at the object level; their blocker is integration, which no cookbook + section can address. The honest codegen wall is ~79 open fns at closeness 3+ (plus 50 at 9+ + that may include mis-scored artifacts — the S59 lesson that a wrong asm-subdir or opt-level + manufactures phantom residuals applies; re-score at intake). +- **"Uncollapsed wave eh was an outlier proving siblings should be drafted":** directionally + right but the stronger events were dd–dg the same morning (217/422 … 100/283). The lesson is + fully priced in already (`ONE_PER_GID=0` landed today); there is no further sibling-drafting + lever of that size left — the pool is spent (§2.1). +- **Chunked gating (chunk > 1) as a speedup:** rejected in the tree's own history — chunked + failures mis-attribute innocent neighbours (cookbook:1568) and the bisect was measured at 1.35 + builds/draft vs 1.0. The rtu inversion supersedes this entire axis. + +## Appendix A — how every number was derived (all read-only) + +- Wave conversions/walls: `.run/gater.log` (`=== GATE ` to `GATE : banked` intervals); + `.run/ox_campaign_ledger.jsonl` for lane/band/targets. +- Wipe windows: `git log/show commit:2863 commit:2913 commit:2915` + `config_sane()` docstring in + `tools/ox_campaign.py`. +- Bank attribution: for each `git log --since=00:00` commit, `git show --unified=0 -- src/ | + grep -c '^-.*INCLUDE_ASM('`; net: `git grep -c 'INCLUDE_ASM(' |HEAD -- src/`. +- Atlas: `.run/atlas.json` (head commit:2911, 17:25) — groups/cat/inst/n_skel/ins/lever/members. +- Generations: fn occurrences across `.run/wave_*_cards.json` (204 files) joined to atlas members. +- Stock: `.run/backlog.jsonl` + `.run/auto/bulk/*.backlog.jsonl` (16,837 rows; closeness by fn), + `.run/reloc_rejects.jsonl` (10,049 rows; `shape` field), `.run/backlog_drafts/` (5,010 files), + joined to atlas open set. +- Gate mechanics: `tools/sweep_parallel.py`, `tools/gate_stage.py`, `tools/harvest_verify.py` + read in full; live `ps` during the ep/ek gates. +- A-prop: `.run/aprop_cards.json` (17:23), `.run/maintenance.log` 15:54–16:26 block. +- Probe: `tools/reloc_identity.py --fn func_80184A68 --binary ov_SC01_077 --c + .run/backlog_drafts/func_80184A68.c …` → COMPILE-FAIL (decl drift since 07-01). diff --git a/phase-ends/CURRENT_PHASE.md b/phase-ends/CURRENT_PHASE.md index 043b404df..d2ad945f6 100644 --- a/phase-ends/CURRENT_PHASE.md +++ b/phase-ends/CURRENT_PHASE.md @@ -112,9 +112,14 @@ writes `docs/tool-designs/frontier-analysis-s60.md` — **READ THAT FIRST NEXT S ### THE STRATEGIC PICTURE (this is what next session must act on) * **3,062 open crackable.** main is **313** crackable, not 1,274 — 961 of its stubs are LINKED PsyQ segments and data blobs, linked byte-exact, never decompiled. -* The drawable pool collapses to **~334 distinct skeletons, ~308 of them generation 6+** (drafted six - or more times, refused every time). ~3,900 open functions are SIBLINGS behind those skeletons and - bank by mechanical remap once an exemplar cracks. +* **CORRECTED 19:45 by the Fable audit, and I had this wrong all session.** The atlas at commit:2911 + measures: 3,106 open instances · 2,245 open skeletons · 1,772 groups, of which **480 are + multi-member holding 1,334 siblings** and **1,292 are SINGLETONS carrying 57% of the open + instruction mass**. My "~3,900 siblings behind ~334 skeletons" conflated two different + populations — the never-drafted stub count with the sibling count — and overstated remap leverage + by ~3x. Most remaining work is singletons that each need their own crack. The "~334 drawable, + 308 gen6+" figure describes only the COLLAPSED wave-eligible view; whole-pool generation is 53% + gen0/1 and 25% gen6+, and only **292 functions are 6+ GATE-refused**. * **Wide waves now convert at 1-5%.** The drafting side is solved; the gate is the bottleneck (30-67 min per gate at 1-3 concurrent builds, load 2.6 of 32 cores) and the drafter parks when its queue fills. **Optimise banks per GATE MINUTE, not cards per wave.** @@ -122,8 +127,22 @@ writes `docs/tool-designs/frontier-analysis-s60.md` — **READ THAT FIRST NEXT S correlated with bank rate — the best waves had the MOST truncation (cx 8.7% trunc / 43.9% bank, dd 8.3% / 51.4%) and the dead waves the least (dl 1.3% / 0.5%, ej 0.6% / 0%). The zeros are the registry outages; the slide from 51% to 20% is population generation, not agent budget. -* **5,388 backlog rows at closeness <= 2** — the grinder (tools/grinder.py, Phase 21, LLM-free, - decomp-permuter) had NEVER been run this campaign and is now on them. +* **The wall is an INTEGRATION wall, not a codegen wall** (Fable's headline, and the highest-value + finding of the day): of the 292 functions the gate has refused 6+ times, **178 (61%) already + produced a closeness-0 draft** — match_one byte-equality, whole-binary gate rejection. The blocker + is symbols/decls/TU plumbing, and the fleet keeps re-drafting them (10,049 reject rows over 574 + distinct functions). Its top recommendation is a ZERO-TOKEN integration-resolver lane + (rtu_match -> reconcile/cast/arity -> reloc_identity -> symfix -> stage to the maintenance gate). +* **5,388 backlog rows at closeness <= 2** de-duplicate to **~543 open functions** (470 at 0, 73 at + 1-2) — the grinder (tools/grinder.py, Phase 21, LLM-free, decomp-permuter) had NEVER been run this + campaign and is now on them. My own re-measure at 19:40 found 290 still-open closeness-0 functions, + down from Fable's 470: the re-gate and grinder are draining exactly this pool. +* **Gate cost is proportional to FAILURES, not drafts** (measured mechanics): chunk=1 plus the 3-stage + ladder means ~3 whole-binary builds per FAILING draft, serial per binary — dd (51% conversion) ran + 1.8 s/draft, eo (1.4%) 8.6 s/draft, and banks-per-gate-minute fell 17.5 -> 0.10. My "30-67 min + gates at 8% CPU" conflated wall_min (which includes drafting and queue time) with gate wall + (12-31 min healthy). tools/rtu_match.py inverted into the gate's stage 0 would make builds + proportional to BANKS (~20x fewer at current conversion) — shadow-run it over 2-3 gates first. ### OPEN THREADS (ranked) 1. **Read `docs/tool-designs/frontier-analysis-s60.md`** — the Fable analyst was told we are NOT married