mirror of
https://github.com/Druthulu/BFM-decomp
synced 2026-09-30 23:37:39 -04:00
chore(phase-31): R40 + R41 ACCEPTED by Drew — binding from 2026-08-23
This commit is contained in:
@@ -480,12 +480,12 @@ over and never touched the second.
|
||||
drafting queues behind one serial gate. Worktrees would fix it.
|
||||
5. DEFER: the 245-member jtbl quest (per-family learning cost), more model comparisons.
|
||||
|
||||
### RULES PROPOSED THIS SESSION (P10 — Drew accepts / modifies / rejects)
|
||||
### RULES ADDED THIS SESSION — **ACCEPTED BY DREW 2026-08-23, BINDING FROM NOW**
|
||||
|
||||
| Rule | Reason |
|
||||
|---|---|
|
||||
| **R40 — EXONERATE THE INSTRUMENT BEFORE YOU ATTRIBUTE A FAILURE TO ITS SUBJECT.** When a measured subject (a model, a binary, a family, a lane) appears to fail, the harness that produced the reading is a suspect until cleared. Before writing "X failed", check the run for: a truncated reply (`finish=length`), a transport error, a rate limit, an unhandled tool fault, a parameter the provider rejected, a missing input the subject was entitled to, and a loop the harness never bounded. Report the failure only once those are excluded — and when one of them WAS the cause, say so as a correction, not as a footnote. | **Seven instances in one session, every one reported to Drew as a model result first**: a 568k-token prompt read as "GLM returns empty"; an unstripped code fence read as "Qwen writes broken C"; uncapped reasoning read as three separate models "failing"; break-on-`finish=length` read as "DeepSeek gave up with 24 turns unspent"; a `reasoning.max_tokens` that 502s one provider read as "nemotron-ultra can't run"; an `IncompleteRead` read as "near 17 is its ceiling"; a directory read that killed the jtbl run at turn 2 of 50. The models were fine; the instrument was not. Extends **R35** (fix the instrument before trusting the measurement) to the ATTRIBUTION step, which R35 does not cover. |
|
||||
| **R41 — A COST, RATE OR YIELD NUMBER SHIPS WITH ITS DENOMINATOR.** Never quote a marginal figure where a total is implied, or a success rate without the attempts it excludes. Say which one it is in the same sentence: "$0.09 per solved function, $6.31 spent total", "19/19 on a seeded stratified sample of 19", "12–25 requests per function, so 1,000/day = 40–80 functions". | I reported GLM-5.3 as costing **$0.30** for eight messages. Drew's bill said **$5**. Both were true: $0.30 was the marginal cost of the two runs that matched, $6.31 was the spend, and 95% of the difference was experiments and my own configuration failures. The flattering number was the one I kept repeating. This is the session's own dominant defect class — *a true number about a narrower scope than the reader believes* (memory `silently-narrowed-tool-scope`) — turned on my own reporting, and it extends **P9** (milestone honesty) from outcomes to metrics. |
|
||||
| **R40 (ACCEPTED) — EXONERATE THE INSTRUMENT BEFORE YOU ATTRIBUTE A FAILURE TO ITS SUBJECT.** When a measured subject (a model, a binary, a family, a lane) appears to fail, the harness that produced the reading is a suspect until cleared. Before writing "X failed", check the run for: a truncated reply (`finish=length`), a transport error, a rate limit, an unhandled tool fault, a parameter the provider rejected, a missing input the subject was entitled to, and a loop the harness never bounded. Report the failure only once those are excluded — and when one of them WAS the cause, say so as a correction, not as a footnote. | **Seven instances in one session, every one reported to Drew as a model result first**: a 568k-token prompt read as "GLM returns empty"; an unstripped code fence read as "Qwen writes broken C"; uncapped reasoning read as three separate models "failing"; break-on-`finish=length` read as "DeepSeek gave up with 24 turns unspent"; a `reasoning.max_tokens` that 502s one provider read as "nemotron-ultra can't run"; an `IncompleteRead` read as "near 17 is its ceiling"; a directory read that killed the jtbl run at turn 2 of 50. The models were fine; the instrument was not. Extends **R35** (fix the instrument before trusting the measurement) to the ATTRIBUTION step, which R35 does not cover. |
|
||||
| **R41 (ACCEPTED) — A COST, RATE OR YIELD NUMBER SHIPS WITH ITS DENOMINATOR.** Never quote a marginal figure where a total is implied, or a success rate without the attempts it excludes. Say which one it is in the same sentence: "$0.09 per solved function, $6.31 spent total", "19/19 on a seeded stratified sample of 19", "12–25 requests per function, so 1,000/day = 40–80 functions". | I reported GLM-5.3 as costing **$0.30** for eight messages. Drew's bill said **$5**. Both were true: $0.30 was the marginal cost of the two runs that matched, $6.31 was the spend, and 95% of the difference was experiments and my own configuration failures. The flattering number was the one I kept repeating. This is the session's own dominant defect class — *a true number about a narrower scope than the reader believes* (memory `silently-narrowed-tool-scope`) — turned on my own reporting, and it extends **P9** (milestone honesty) from outcomes to metrics. |
|
||||
|
||||
**Not elevated to rules** (Phase-8+ precedent — techniques to the cookbook, findings to the checkpoint):
|
||||
the card-fuel result (**§206** + the fuel A/B) is a *procedure* already enforced by `--cards`; the
|
||||
|
||||
Reference in New Issue
Block a user