- GLM deep-solved 15 fresh 25-118-ins hard fns: 4/15 match_one, 2 whole-binary banks (func_8017DE28,
func_8015EEE0), $2.00. def-side wall caps the other 2 correct bodies.
- IDIOM VERDICT (Drew's fair test): GLM's correct bodies reason about KNOWN gcc mechanics (delay slots,
callee-saved $s0, reload-after-call aliasing, switch jump tables) — cookbook §10/§17/jump-table.
NO new idiom. The quirk space is largely mapped (22 phases of Opus-Max mining). Well DRY confirmed
from BOTH angles: failed-residual (T10.8) AND fresh-hand-solve (T10.9).
- tools/idiom_hunt.py: group backlog near-misses by residual class -> GLM names the reusable idiom +
emits corrected C -> byte-gate to validate; captures reasoning; HARD --budget cap
- CALIBRATION ($0.51 total, 2 classes): struct + regalloc-order -> 0 banks. GLM re-derives our OWN
idioms (register-pin §17, array-of-struct §18, type-width §25) and confirms walls, but banks nothing
new — the backlog near-misses are the residual our idioms already failed on (irreducible/def-side wall).
- VERDICT: the new-idiom well is DRY (Fable5 review confirmed empirically for $0.51, not $300 overnight).
GLM's value stays: direct drafter for FRESH def-conflict-free hard fns (~22%), not an idiom generator.
- cookbook §29: reasoning-model reconciliation idioms (match-pointer-type-to-TU-decl, call-site cast
for value mismatch, cast-a-callee-definition, data-type match) + the narrow-param hard limit
- gen2-mips-matching-model.md + CURRENT_PHASE: Option-3 verdict (GLM reasons the wall expertly but
banks 1/7; wall INTRINSIC, Fable5 §3c triple-confirmed); GLM role = $0.03-0.08/fn hard-band drafter
+ idiom teacher; real lever past the wall = public flip, not a bigger model
- tools/glm_reconcile.py (NEW): aim GLM's reasoning at the DEF-side loose-typing wall (body + conflicting
TU decls + reconciliation toolkit -> consistent buildable byte-identical decls); captures reasoning
(.run/glm_reason/, idiom source R16); relax-in-any-TU-file + crash-robust call
- api_draft: REASON=1 saves the reasoning trace per draft (idiom mining on any GLM run)
- fix_arity_callers: --any-proto (relax any prototype, not just (void))
- RESULT: GLM's reasoning is expert-level (store-width/sh-vs-sw awareness, K&R promotion, independently
derives the cast idiom) but banks only 1/7 reconciliations; mechanical relaxation 0/7. The def-side
wall is INTRINSIC (narrow-param + byte-level addressing defeat reconciliation) — Fable5 §3c re-test
CONFIRMS the wall holds even vs a frontier reasoning model aimed directly at it. func_80175184 banked,
check-all 136/136
- def-side conflict (engine_core.h forward-declared func_801577C8(void) vs GLM's byte-correct
(s32) def) resolved by relaxing the caller decl to no-proto; strip GLM externs + gate. 136/136.
- FINDING: only +1 of 8 stranded recovers mechanically; the other 7 are the intrinsic Phase-16/20
DEF-side loose-typing wall (5 have non-(void) conflicting forward-decls, 1 narrow-param) — the
wall caps ANY drafter, not a v3-tuning artifact (Fable5 review §3c re-test: wall HOLDS)
- MAXTOK env (default 512 = local v3 unchanged); reasoning models (GLM5.2) need a high cap or
they spend the budget on reasoning tokens and return empty content
- accumulate usage.cost from the response -> per-run $ + $/fn readout (OpenRouter reports it)
- scorecard: original 2026-06-10 scope vs 22 phases of byte-verified reality (what held,
what emerged beyond scope, what deviated and should be revisited)
- adversarial pass: P21 no-shortcut + giant scheduler walls HOLD; P16 loose-typing wall has
a TIMESTAMP GAP (declared 06-19, pre-dating cast_call_sites/block-scope-externs/v3) -> re-test
- July-2026 resources: frontier-on-hard-band (T10.7 re-aim), continuous architect-tier judgment,
RE-ELEVATE THE PUBLIC FLIP (community labor = the only lever that scales into the proven tail)
- strategic fork: posture A/B/C on the byte-match goal; recommends dual-metric (B), Drew's call
- ranked recs 1-6 + explicit endorsements of what not to change