Commit Graph

936 Commits

Author SHA1 Message Date
Drew T f31f56ef17 docs(phase-31): the drafting pool emptied; the carve lane replaced it (+5, frontier 169)
Wave 3 drew 1 target - 0 left in pool. Of 174 open: 64 main, and of 110 non-main, 41
drafted this session, 68 excluded, 2 walls, ZERO undrawn. Re-probing the 68 with
jtbl_carve --probe found 17 now reporting `tail`, because tonight's jr_isolate_all
fixes changed their overlays. All 17 already had drafts; 10 scored closeness 0 with no
drafting. The gate banked 5, and they are exactly the five overlays jr-isolated tonight.

An exclude list is a snapshot of what the TOOLING could not do and goes stale the moment
the tooling improves - re-probe it after every tool fix.
2026-09-02 05:34:38 -06:00
Drew T 9bb43dd426 docs(phase-31): S71 gate 9 (+2, frontier 174, 36 banked); §419 density lever recorded 2026-09-02 05:19:10 -06:00
Drew T 49d1e06632 docs(phase-31): S71 gate 8 (+2, frontier 176, 34 banked); two proven walls recorded
Two functions recorded as walls with their refutation lists rather than redrafted:
ov_SC03_105/func_801834A4 (loop.c movable ordering, closeness 6) and
ov_SC06_022/func_8017DF28 (expand_block_move's copy_addr_to_reg pseudo reused by cse,
closeness 2, seven levers measured inert). One MATCH blocked purely on carve state with
its exact prescription queued in .run/S71_carve_todo.txt.
2026-09-02 05:02:09 -06:00
Drew T 8709798738 docs(phase-31): S71 gate 7 (+3, frontier 178, 32 banked) + the measured blocker census
Blocker census read off the 37 gate verdicts on disk: DIFF 18, CARVE 7, PARSE 3,
NO-DIAG 3, CONFLICT 2, ARITY 2, UNDEF 2. I had called carve the dominant remaining
class mid-session on the strength of the last two agents I'd read; it is not. What
remains is mostly genuine codegen, the opposite of the integration-dominated picture
this session opened with.

Also recorded rather than redrafted: ov_SC03_105/func_801834A4 as a proven loop.c
movable-ordering wall (closeness 6, two measured-inert levers), and
ov_SC01_004/func_8017EB30 as MATCH-279/279 blocked purely on §8e carve state.
2026-09-02 04:46:02 -06:00
Drew T 072607df7b docs(phase-31): S71 gate 6 (+4, frontier 181) and cookbook §416
Four new byte-proven levers from the overnight lane, none previously in the cookbook:
re-read the store instead of passing the value (CSE store-forwarding), (&SYM)[3] vs a
pointer local as an ADDRESSING choice, one biv with +0/+2/+4 for combine_givs, and a
local's width choosing lh vs lhu+sll/sra.

Also recorded: the same-address twin hint was false three times tonight (ov_SC06_000,
ov_SC01_080, ov_SC03_030) while the same-TU neighbour was the real fuel in every case.

Three of the night's five post-limit MATCHes recovered a body off disk rather than
re-deriving it - func_80181A60 in 2 minutes instead of 16.
2026-09-02 04:32:08 -06:00
Drew T ab46791e06 docs(phase-31): S71 resumed at 04:11 after the session-limit reset
All five in-flight agents died on the 5-hour limit and returned NO-DRAFT; that is a
harness kill, not a verdict about the targets (R40), so they relaunch unchanged.
Gate 5 banked 2 (commit:3614). launch_check.py added after a stale card burned an agent.
2026-09-02 04:12:25 -06:00
Drew T 4d7c297b12 docs(phase-31): S71 CHECKPOINT — fleet R22 GREEN 213/213, honest frontier 187 (23 banked), main incident recorded 2026-09-02 01:50:33 -06:00
Drew T 0e4c98dc97 docs(phase-31): S71 gate cycle 4 — frontier 210 -> 176, 34 banked; residual-class routing + the R48 name-key exposure 2026-09-02 01:31:57 -06:00
Drew T fecbfac5f8 docs(phase-31): §412 — §323 carve blocker 2 was a regex blind to __attribute__; 5 of 6 carve overlays cleared 2026-09-02 01:25:56 -06:00
Drew T 6c4391abb9 docs(phase-31): S71 gate cycle 3 — 29 banked; journal-notes made permanent; jr_isolate_all unblocked 2026-09-02 01:17:47 -06:00
Drew T 89c3cf831b docs(phase-31): S71 wave 1 — journal-fuelled packs, 100% first-pass MATCH; cookbook 1078
* Every pack carried PAST ATTEMPTS ON THIS EXACT FUNCTION, mined per-function from the
  historical agent journals (52 of 60 targets, 131 notes). Every landed agent returned
  MATCH at closeness 0 on the hardest frontier we have.
* §409 — the wave and the nine laws it produced. Law 1: a relocation-stream
  TRANSPOSITION is invisible to match_one, the permuter scorer and every similarity
  tier (HI16/LO16 masking; the §195-D blind spot for a different reloc class), and it
  retroactively explains "MATCH but the gate rejected it" verdicts.
* §410 — COPY THEN ACCUMULATE ON THE COPY: satisfies the $s2 in-place destination and
  the sched1 birthing boost at once, with the agent's measured refutation list.
2026-09-02 01:04:27 -06:00
Drew T 137c418bc7 docs(phase-31): S71 — the 64 standalone matches priced honestly; 12 banked, 52 in four named lanes
* gate 1 (all 64 across 33 binaries): 12 banked — main 11 + ov_SC07_006 1.
* gate 2 tested "a bad draft kills its binary's good ones" by re-staging only the 25
  that recover_integration --probe-only called MATCH in their real TU: 0 banked.
  An honest null — that probe compiles and diffs bytes but never LINKS or CARVES,
  so it is a third oracle with its own blind spot.
* triage (25/25 accounted): CARVE-REFUSED 10, undefined-reference 4, DIFF 3,
  CC1-FAIL-no-diagnostic 2, PARSE 1; gate 1 adds 7 func-decl / 4 data-decl /
  6 type-decl conflicts.
* R37 probe of the carve class: 6 of 8 are one refusal — a subseg would host
  NON-CONTIGUOUS .rodata carves — whose named remedy is jr_isolate_all (§8b).

tools/restage_matching.py — rebuild a gate plan from probe verdicts.
tools/gate_triage.py — route a gate's verdicts to the lane each one names (R47).
2026-09-02 00:39:18 -06:00
Drew T 8cf0104386 fix(cards): defect 5 — an expired BASELINE-RED claim, without discarding any measurement
The S70 patch was refused by its own adversarial review for sorting rows by recency:
a pair's ledger rows are several PROBES about one draft, alternating between
`closeness 4` and `won't compile standalone`, so max(ts) serves whichever probe ran
last — often the least informative. This form keeps both.

* the ts-newest verdict is still selected (file order made the per-binary bulk ledger
  always win regardless of age: 25 pairs mis-selected),
* AND the best measurement ever taken on the pair rides alongside it, so a later
  uninformative probe can no longer erase an earlier residual: 981 of 2,605 pairs
  gain a line they were previously denied.
* BASELINE-RED is a fact about a binary at a moment (R51), frozen into an append-only
  ledger and replayed forever — 2,676 rows all stamped 2026-08-26. gate_feedback now
  reads the same live red union gate_stage consults, so a pack and the next gate run
  cannot disagree: 173 expired claims retired, 0 binaries currently red.

R39 control 3/3 (expired-when-green, harness-line-when-red, measurement-survives).
2026-09-02 00:32:20 -06:00
Drew T 13a16a15c6 fix(pgate): the merge scope missed main entirely — 11 byte-proven banks were dropped silently
* `git status --porcelain -- src/<binary>/` finds nothing for main, whose TUs are
  src/800.c, src/boot.c, ... — so a main worker returned `files: {}` while the bank
  oracle (the stub disappeared) still counted the banks. parallel_gate printed
  "12 banked across 2 binaries" and committed one of them.
* src_scope() takes the scope from the binary's own stub rows (each names its TU),
  captured BEFORE the gate because a bank deletes the stub that names it, and keeps
  the directory prefix for overlays that have one.
  Negative control: main 0 -> 54 TUs, ov_SC07_006 1 -> 3 (superset, no regression).
* A reused worktree kept the previous job's .run/harvest_failed*.classified.txt, so
  verdicts surfaced under the wrong binary; the worker clears them first.
* tools/gate_triage.py — routes a gate's verdicts to the repair lane each names (R47),
  with the staged-draft denominator asserted (R32/R41).

Re-gated main: 11 banked (commit:3586), main real frontier 64 -> 53.
2026-09-02 00:27:52 -06:00
Drew T 698f2959ae fix(campaign): R48 — reloc_filter resolves a draft's binary per-draft, not by bare name
* `binof = {c["fn"]: c["binary"]}` was last-writer-wins, and `status`, `det` and `subof`
  had the same shape — a draft of a name carried by two binaries was stamped with
  whichever card came last and then reloc-checked against the OTHER binary's symbols.
* Resolve per draft instead: the shard's own target list first
  (`.run/wave_<tag>_targets.<i>.json` = `targets[i::workers]`, each row carrying its
  binary), a unique-name card second, a counted refusal when neither can answer (R43).
* R39 negative control over every historical wave: 42,655 drafts, 0 regressions,
  2,317 (5.4%) previously mis-stamped; 2,107 homonym card names fleet-wide.
  Intra-shard ambiguity: 0 of 50,684 (shard, name) pairs over 302,370 shard files.

docs: §408 — §406 refuted as a sweep (0 MATCH / 14 applied, 0 / 210). The 134-member
census counted main's 960 LINKED library stubs and matched a symmetric SHAPE; derived
from the mine-vs-target residual the addressable set is 15 / 210. Decision-log entry
records the pivot: 64 of 210 (30.5%) already match standalone, so the frontier's
largest lane is §376 integration, not codegen.

tools/weave_sweep.py — the derived-selector sweep (R32 coverage, R41 denominators,
--lever-all ablation control).
2026-09-02 00:19:59 -06:00
Drew T da3bb36c56 docs(phase-31): S70 FINAL-5 — fresh-session checkpoint, §406 sweep first; 145 banked (355->210), R22 green, cookbook 1074 2026-09-01 23:56:13 -06:00
Drew T a1e8c59640 docs(phase-31): card defect 4 shipped, defect 5 refused by adversarial review (would have suppressed 159 real verdicts) 2026-09-01 23:44:56 -06:00
Drew T 1632c6adec docs(phase-31): S70 FINAL-4 — 145 banked (355->210), 115/118 wave MATCH, 5 card defects (3 fixed), cookbook 1074 2026-09-01 23:40:28 -06:00
Drew T c6a0dc8f43 docs(phase-31): S70-T14 — blocker 3 fixed, carve route OPEN (+3, frontier 308); all three blockers were routing errors, not walls 2026-09-01 22:00:26 -06:00
Drew T ba24730d39 docs(phase-31): S70-T13 — carve ownership was a MODEL bug (dry run 7/17->15/17); a third blocker (jtbl_carve mixed-data island) is the real wall 2026-09-01 21:28:14 -06:00
Drew T 57031f01e4 docs(phase-31): S70-T3 carve route — 0 banked, blocker localised to stale committed carve ownership (7 binaries) 2026-09-01 21:09:18 -06:00
Drew T b5370955cc docs(phase-31): S70 FINAL-3 — 44 banked (355->311), R22 green, tools-health 333s, §332b island shipped 2026-09-01 20:46:03 -06:00
Drew T 576d69a529 docs(phase-31): S70-T4 §332b shipped (+3, frontier 311); §404 corrected; tools-health 333s green 2026-09-01 20:37:18 -06:00
Drew T 425cec253c docs(phase-31): the pre-wave tool audit — blast radius REFUTED, work_evidence shipped, cleared for waves 2026-09-01 18:59:32 -06:00
Drew T d23bc7ad56 docs(phase-31): audit the pgate blast radius — REFUTED; 606/668 plan rows absolute, defect never fired before S70 2026-09-01 18:44:31 -06:00
Drew T 16fae589d0 docs(phase-31): S70 FINAL-2 — 41 banked (355->314), R22 green 213/213, 4 tool defects (3 fixed) 2026-09-01 17:47:55 -06:00
Drew T 6b51f0383c docs(phase-31): correct S70 — main banked 0, not 1; harvest_verify alone does not persist a splice 2026-09-01 17:42:18 -06:00
Drew T 1d5f951a6d docs(phase-31): S70 FINAL checkpoint — 22 banked (355->333 real), R22 green 213/213, next-session queue led by the family_remap decl fix 2026-09-01 17:29:57 -06:00
Drew T 832dd8ea25 docs(phase-31): S70-T9 — 22 banked of 86; jtbl-carve probe law; pgate gated nothing at rc=0; both undo-journals corrupt 2026-09-01 17:19:39 -06:00
Drew T b31d5a2cae docs(phase-31): S70 scope call — defer the future-decomp generalization; size the banked label corpus (16,301 / 12,383 with drafts) 2026-09-01 16:10:48 -06:00
Drew T e6056c56df docs(phase-31): S70-T2 — coverage probe VERDICT=build the rules; 86 standalone matches; denominator corrected
- ran residual_rules_b over the WHOLE open frontier (1,312 cases, 0 errors, ~2min, $0)
  instead of a 50-row sample; artifacts in .run/S70_*
- DENOMINATOR (Drew's correction, R41): main's 960 PsyQ LINKED stubs are not
  matching targets; true frontier = 355 (67 main REAL + 288 non-main), partitioned
  with progress.linked_subsegs() rather than a hand-rolled filter (R33)
- discriminating test settles population-vs-coverage: fire rate DOES rise as
  residuals get clean (35.3% at <=8 vs 6.1% at >64) but 57% of the cleanest band
  is still UNKNOWN -> coverage binds where rules are worth writing
- hand-label 4/4 labelable to existing cookbook buckets; WIDTH/lhu!=lh has its
  discriminating sig already computed and still returns top=None
- 86 REAL standalone MATCHES (closeness 0) = 24% of the frontier, blocked on TU
  plumbing only -- outranks the rule work (standalone-match-is-not-bankable)
- logs 4 instrument defects in my own probe, incl. one wrong answer reported to
  Drew before checking: 4 of S68's 10 autodecl MATCH drafts are STILL OPEN
2026-09-01 16:07:18 -06:00
Drew T cf30d6d5fa chore(phase-31): S70-T1 — R22 clean-fleet verify GREEN 213/213; correct the probe's citation and denominator
- tools/r22_verify.sh from a clean tree: clean rc=0, extract-all 212+main rc=0,
  check-all 213 passed / 0 failed of 213 (2m49s). Clears the S69 --no-r22 debt.
- R38 read of the recorded measurement behind the "1-2% ceiling" (S68 eval set +
  .run/rules_b/eval_results.jsonl) before designing the queued probe:
  * citation fix: the design is Fable-1 (.run/S69_fable/report.md:93), not Fable-2 §7.7
  * denominator fix (R41): shape rules can only fire on the 39 near rows, not 113;
    real fire rate 3/39 = 7.7% (5/39 with REDRAFT), and 14 are UNKNOWN
  * the probe as written is unrunnable: backlog has 125 rows / 20 with residual text
    and the UNKNOWN pile is 14 -- sampling 50 would report a narrower world (R41/R32)
2026-09-01 15:56:52 -06:00
Drew T 772f0eb991 docs(phase-31): S69 FINAL-4 — true session close, next-session queue led by the residual-classifier coverage probe 2026-09-01 15:45:37 -06:00
Drew T 923777b344 docs(phase-31): S69 FINAL-3 — session close, R22 green 213/213, canonical metrics 2026-09-01 15:13:14 -06:00
Drew T 40e58ee5b3 docs(phase-31): S69 FINAL-2 checkpoint — ~93 banked, three Fable audits, frontier 348 -> 320 2026-09-01 14:41:35 -06:00
Drew T 9d7b26523c docs: cookbook §384 — a carve-config bank is red until you re-extract; correct the S69 checkpoint
The 'false bank' in the S69 checkpoint was not one. Both instances verify
byte-identical after 'make extract BINARY=<b>'. §384 states the law (verification
must regenerate whatever the gate changed the inputs to), the trap inside it (a
src-only revert of a carve commit produces 'table-count drift vs the carve', which
reads like progress), and the give-away I ignored — the commit diffstat showed
config/overlays.mk and a splat yaml sitting next to the .c.
2026-09-01 11:09:31 -06:00
Drew T f95b895909 docs(phase-31): S69 FINAL checkpoint — 44 MATCH of 84 agents, the §378 lever, four self-inflicted defects measured 2026-09-01 10:45:07 -06:00
Drew T 6bc674ae45 docs(phase-31): S69 checkpoint — 8 banked from the '32 free banks' class, the §378 lever, the triage ladder acceptance-green 2026-09-01 00:11:20 -06:00
Drew T a47981a494 docs(phase-31): S68 FINAL-2 checkpoint — 39 closed (453 -> 414), fleet 213/213, harvest drained 2026-08-31 22:33:02 -06:00
Drew T a731022d96 docs(phase-31): S68 post-reset — module-binary -O0 route opened, twin flywheel closed 2026-08-31 18:18:01 -06:00
Drew T 2a0808e3dc docs(cookbook): §370 — a HARD BOUND from sched.c, plus the reorg slot-steal diagnostic
The third fable escalation did NOT close its function (main/func_8001BC6C,
33 -> 28 over ~45 measured compiles), so the checkpoint's '2 for 2' is corrected
to 2 closed of 3. The failure is banked because a negative result that tells
future agents when to STOP is worth its tokens.

THE BOUND: sched.c schedule_select ALWAYS fronts a ready load over an
equal-priority ALU leaf (potential_hazard), so no C spelling can emit an ALU chain
before loads that are simultaneously-ready same-priority leaves. If a target shows
that order, look for reorg slot-steals, hard-reg dependency walls, or late in-block
consumers BEFORE burning compiles on statement permutations.

Also banked: the reorg fill_simple_delay_slots slot-steal diagnostic and its
split-tree precondition (the accumulator must live outside the $v0-heavy tail to
be eligible), three supporting levers, and three REFUTED ones with measurements --
a dead-init boost-kill is a no-op because cse delete_dead_from_cse removes it
before the final reg_scan, dense-block re-ties cost +4 to +9 because each re-tie
re-anchors its own load, and the -fno-schedule-insns oracle does not discriminate
when the residual is a multi-pass composition.

This run applied §361 CORRECTLY -- it removed the prior agent's pin first and
exonerated it for the head -- which is why its four-pass diagnosis can be trusted
where the previous single-tie claim could not.
2026-08-31 17:29:45 -06:00
Drew T 64230e40ea docs(phase-31): S68 FINAL checkpoint — 23 closed (453 -> 430), fleet 213/213, main unblocked after two stacked harness defects 2026-08-31 17:26:38 -06:00
Drew T 9c96b47ce0 docs(phase-31): S68 progress — whale carve (6 fns/2,547 ins), the main-layout bug class, gater hardening 2026-08-31 16:46:57 -06:00
Drew T ed53a68f18 feat(p31 s68): deferred propagation done honestly (2 banked) + seed_ref was offering DEAD TEXT
The S67 FINAL-3 OPEN item, plus the two defects found while doing it.

* fix(dedup_propagate): the tool could not run AT ALL. S67's -j patch wrote
  `os.environ` at module level in the one module that imports `os as _os`, so
  every invocation died with NameError before doing any work. Propagation was
  not deferred, it was impossible. Import-checked the other 7 -j-patched tools.

* propagation, honestly scoped: the real closable set is 11, not 32, derived two
  independent ways that agree (seed_ref exact+same_addr, and a direct corpus
  derivation). The 3,161-entry --auto-from plan over 53 overlays is dedup
  hygiene over already-matched code and closes almost no open stub.
  Applied: 2 banked byte-green (ov_SC04_018 func_80181270, func_80182AF8);
  3 gate-refused and cleanly reverted; 6 blocked with named blockers
  (3 CARRY-FIXABLE, 3 func_80144B9C not-inline-def -> needs the o0 whale carve).
  R22 clean fleet: extract 212/212, check 213 passed 0 failed of 213, rc 0/0/0.
  Frontier 453 -> 451.

* fix(seed_ref): REFUSE targets in LINKED subsegs. The playbook calls this tool
  "the fleet-wide answer" and it reported 82 open stubs with a banked twin --
  43 of them main stubs whose TUs the linker script never references. Any C
  written there compiles, links and leaves the SHA1 green WHETHER OR NOT IT IS
  CORRECT, so a mechanical twin lane fed from that list could have minted up to
  43 gate-green FALSE matches the byte gate cannot see. draw_waves has refused
  these since S66; this oracle did not. The refusal is counted and printed, not
  silent. NC: guarded 39 subset of raw 82, all 43 dropped are main, the non-main
  population is identical.

* wave drawn: .run/S68o1 (24 opus 187-770 ins) + .run/S68m1 (30 main), cards +
  packs + wave_args asserted, queue of 53. Drafting opened at concurrency 5.
2026-08-31 15:59:01 -06:00
Drew T 746a7cf656 docs(phase-31): S67 FINAL-3 — 77 closed (530 -> 453), fleet 213/213, jtbl parallelised (both halves), propagation deferred 2026-08-31 15:23:25 -06:00
Drew T b8a4eae41a docs(phase-31): S67 FINAL-2 — 29 closed (530 -> 501), fleet 213/213, wave 20/20 MATCH, gate_wave.py replaces the serial loop 2026-08-31 12:48:46 -06:00
Drew T b2256e3700 docs(phase-31): S67 FINAL checkpoint — 4 closed (530 -> 526), fleet 213/213, propagation found unverified 2026-08-31 09:45:55 -06:00
Drew T b4d8f06363 docs(phase-31): S66 FINAL checkpoint — 416 fns closed (946 -> 530), fleet 213/213, 99.2% weighted
Written after the last harvest, per the rule that the checkpoint is always last. Supersedes the
interim S66 block. Machine quiesced, tree clean at commit:3350.

Leads with the three things a fresh session must not re-learn: main is ~94 open not 1,099 (and
drafting into a LINKED subseg would gate GREEN while wrong); the family era is over so integration
is the whole game (147 of 591 open fns already had byte-correct drafts stranded on four blockers);
and 'independent' means a different INSTRUMENT — two refusals from parallel_gate were one instrument
twice, after which the serial gate banked 16/32 of that class.

Also records that ~56 of the 416 banks came from ZERO drafting agents, purely from work already on
disk, and Drew's binding harvest-then-toolify-before-the-next-wave rule.
2026-08-31 00:01:43 -06:00
Drew T 3a2533c602 docs(phase-31): S66 checkpoint — 390 fns closed (946 -> 556), fleet 213/213, 99.2% weighted
Written for a fresh session. Headlines: main is ~100 open not 1,099 (960 stubs are LINKED dead text,
and drafting into them would gate GREEN while wrong); families are spent (84-93% singletons, twin
pool dry); integration is now the whole game (147 of 591 open fns already had byte-correct drafts
stranded on four blockers). Records Drew's binding rule — harvest, then toolify, BEFORE the next
wave — and the F18 retraction: two refusals from parallel_gate were one instrument twice, not two
independent tests; the serial gate then banked 16/32 of that class.
2026-08-30 22:34:34 -06:00
Drew T a9dd3518d8 docs(phase-31): point a fresh session at the LAST checkpoint block
CURRENT_PHASE.md now holds 27 checkpoint blocks and several older ones say 'supersedes every earlier
block' — true when written, false now. The S64 FINAL block sits ~370 lines above the live S65 FINAL-4
one and makes the same claim, so a fresh session reading top-down could anchor on a state the tree has
moved past by 647 banked functions. Banner at the top states the rule: the LAST block is the live one.
2026-08-29 20:50:55 -06:00