From d97f1c28c57724d1d36fdcf94776f167b415015d Mon Sep 17 00:00:00 2001 From: Drew T <50529377+Druthulu@users.noreply.github.com> Date: Sun, 23 Aug 2026 00:15:52 -0600 Subject: [PATCH] =?UTF-8?q?chore(phase-31):=20R40=20+=20R41=20ACCEPTED=20b?= =?UTF-8?q?y=20Drew=20=E2=80=94=20binding=20from=202026-08-23?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- phase-ends/CURRENT_PHASE.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/phase-ends/CURRENT_PHASE.md b/phase-ends/CURRENT_PHASE.md index a56d9bc37..c89ef8fac 100644 --- a/phase-ends/CURRENT_PHASE.md +++ b/phase-ends/CURRENT_PHASE.md @@ -480,12 +480,12 @@ over and never touched the second. drafting queues behind one serial gate. Worktrees would fix it. 5. DEFER: the 245-member jtbl quest (per-family learning cost), more model comparisons. -### RULES PROPOSED THIS SESSION (P10 — Drew accepts / modifies / rejects) +### RULES ADDED THIS SESSION — **ACCEPTED BY DREW 2026-08-23, BINDING FROM NOW** | Rule | Reason | |---|---| -| **R40 — EXONERATE THE INSTRUMENT BEFORE YOU ATTRIBUTE A FAILURE TO ITS SUBJECT.** When a measured subject (a model, a binary, a family, a lane) appears to fail, the harness that produced the reading is a suspect until cleared. Before writing "X failed", check the run for: a truncated reply (`finish=length`), a transport error, a rate limit, an unhandled tool fault, a parameter the provider rejected, a missing input the subject was entitled to, and a loop the harness never bounded. Report the failure only once those are excluded — and when one of them WAS the cause, say so as a correction, not as a footnote. | **Seven instances in one session, every one reported to Drew as a model result first**: a 568k-token prompt read as "GLM returns empty"; an unstripped code fence read as "Qwen writes broken C"; uncapped reasoning read as three separate models "failing"; break-on-`finish=length` read as "DeepSeek gave up with 24 turns unspent"; a `reasoning.max_tokens` that 502s one provider read as "nemotron-ultra can't run"; an `IncompleteRead` read as "near 17 is its ceiling"; a directory read that killed the jtbl run at turn 2 of 50. The models were fine; the instrument was not. Extends **R35** (fix the instrument before trusting the measurement) to the ATTRIBUTION step, which R35 does not cover. | -| **R41 — A COST, RATE OR YIELD NUMBER SHIPS WITH ITS DENOMINATOR.** Never quote a marginal figure where a total is implied, or a success rate without the attempts it excludes. Say which one it is in the same sentence: "$0.09 per solved function, $6.31 spent total", "19/19 on a seeded stratified sample of 19", "12–25 requests per function, so 1,000/day = 40–80 functions". | I reported GLM-5.3 as costing **$0.30** for eight messages. Drew's bill said **$5**. Both were true: $0.30 was the marginal cost of the two runs that matched, $6.31 was the spend, and 95% of the difference was experiments and my own configuration failures. The flattering number was the one I kept repeating. This is the session's own dominant defect class — *a true number about a narrower scope than the reader believes* (memory `silently-narrowed-tool-scope`) — turned on my own reporting, and it extends **P9** (milestone honesty) from outcomes to metrics. | +| **R40 (ACCEPTED) — EXONERATE THE INSTRUMENT BEFORE YOU ATTRIBUTE A FAILURE TO ITS SUBJECT.** When a measured subject (a model, a binary, a family, a lane) appears to fail, the harness that produced the reading is a suspect until cleared. Before writing "X failed", check the run for: a truncated reply (`finish=length`), a transport error, a rate limit, an unhandled tool fault, a parameter the provider rejected, a missing input the subject was entitled to, and a loop the harness never bounded. Report the failure only once those are excluded — and when one of them WAS the cause, say so as a correction, not as a footnote. | **Seven instances in one session, every one reported to Drew as a model result first**: a 568k-token prompt read as "GLM returns empty"; an unstripped code fence read as "Qwen writes broken C"; uncapped reasoning read as three separate models "failing"; break-on-`finish=length` read as "DeepSeek gave up with 24 turns unspent"; a `reasoning.max_tokens` that 502s one provider read as "nemotron-ultra can't run"; an `IncompleteRead` read as "near 17 is its ceiling"; a directory read that killed the jtbl run at turn 2 of 50. The models were fine; the instrument was not. Extends **R35** (fix the instrument before trusting the measurement) to the ATTRIBUTION step, which R35 does not cover. | +| **R41 (ACCEPTED) — A COST, RATE OR YIELD NUMBER SHIPS WITH ITS DENOMINATOR.** Never quote a marginal figure where a total is implied, or a success rate without the attempts it excludes. Say which one it is in the same sentence: "$0.09 per solved function, $6.31 spent total", "19/19 on a seeded stratified sample of 19", "12–25 requests per function, so 1,000/day = 40–80 functions". | I reported GLM-5.3 as costing **$0.30** for eight messages. Drew's bill said **$5**. Both were true: $0.30 was the marginal cost of the two runs that matched, $6.31 was the spend, and 95% of the difference was experiments and my own configuration failures. The flattering number was the one I kept repeating. This is the session's own dominant defect class — *a true number about a narrower scope than the reader believes* (memory `silently-narrowed-tool-scope`) — turned on my own reporting, and it extends **P9** (milestone honesty) from outcomes to metrics. | **Not elevated to rules** (Phase-8+ precedent — techniques to the cookbook, findings to the checkpoint): the card-fuel result (**§206** + the fuel A/B) is a *procedure* already enforced by `--cards`; the