#!/usr/bin/env python """bench.py -- the agent bench: fixtures, plan, run, pull, compare, check. bench.py [--fixtures DIR] fixtures --pin write MANIFEST.json (sha256 per file + tree hash) bench.py [--fixtures DIR] fixtures --check 'fixtures OK ', or the drifted/missing/extra paths, exit 1 bench.py plan > one line of non-zero arms: 'coder 12 · expert 8 · …' bench.py run [light|full|canary|max|] [--suite light|full|canary|max] [--repeat k] [--model M] [--effort E] [--jobs 4] [--force] [--no-stop] [--out JSON] [--work-dir DIR] [--agents-dir DIR] [--no-rubric] [--preset max20|max5|pro] plan and run: --preset > ./.claude/pa.json preset > max20; an arm whose agent the ladder omits at the preset is skipped with one line; run renders each agent via pa.install.ladder into /_agents/ and runs that copy. Arms coder, expert, planner, retriever, router, critic, review in that order; an runs its one arm (role: line, else the default table, else the name prefix) on --suite (default light). --repeat k (default 1) runs each task k times: k >= 2 clones /a/, retained files .a.*; records carry attempt i (results format 2; a task passes when every ran attempt passes). Each task: a git-inited clone of repo/ (+ overlay/files) under //// (work default ~/pa3-bench; never under ~/.claude: a sensitive path, edits denied), the agent file in its .claude/agents/, `claude -p --agent …` with the prompt on stdin, graded by _bench_grade. Cache order: per arm the first task's attempt 1 alone, then <= --jobs attempts in flight. Window: refuses (exit 2, one line) when five-hour pct + estimate > bench.window_max_pct (80) unless --force; re-polls after each arm and after a task when >= 60 s since the last poll, clean-stops over the cap (--no-stop: never). Writes the results JSON (docs/pa3-build/design/bench.md ## Results format; --out, else ./bench/results/.json), the README.md beside a results/ dir, else inside it, then `pa_ledger.py bench import `. Planner, critic and review tasks get a rubric side score after grading (rubric/grader.json's model, tools off, in an empty /_rubric/-/; its cost added to the task's; never the pass); --no-rubric skips it. Scored (suite max): a coder task whose expect has score_cmd (argv; {clone} {fixtures} {id}) is graded by _bench_grade.grade_scored, passing at score >= expect.pass_score (1.0); records carry score; a task's "timeout" (s) overrides TASK_TIMEOUT. bench.py ab [light|full|canary|max] [--repeat k] [--jobs n] [--work-dir D] [--agents-dir D] [--results-dir D] Same role: required (else 'ab: roles differ ( vs )', exit 2, nothing launched). Runs A then B as `run --suite S --repeat k` (results files in D, default ./bench/results; two ledger imports), then per task id ' | A p/k $x t | B p/k $x t' and 'A better a · B better b · tie t · cost A $x / B $y · turns A t / B u'; no significance claim; exit = worst of the two runs. bench.py pull [--source URL|DIR] each format-1/2 results JSON of DIR (or the one at URL) into the ledger as source=public; source default .claude/pa.json bench.source; none, or a missing dir: one line, exit 0. bench.py compare [--agents DIR] [--base DIR] per installed agent (default ./.claude/agents), by name, one line ' [why]': no-series | update | match | run-offered | lightly-altered (docs/pa3-build/design/bench.md ## Comparison). User marks: version '1+u1' or 'user/1'; a marked body within 20 % of the base agent's lines (--base, default package agents/, else pa3-src) is lightly-altered, never overwritten. bench.py check [--results DIR] four kinds of line, DIR default ./bench/results: 'budget OK|FAIL (light L, full F points)' (newest light < 4, full < 10: last w5h_after - first w5h_before; a part is the reason when unmeasured or over); 'graders OK|FAIL (full: g grader, h harness, u unaudited; light: ...)' (failed records of the newest runs by failure_kind, agent not counted; FAIL when any); 'noise ±f/n [(date)]|unknown' per unstopped arm of the newest full (ids whose attempts disagree, from that run's repeats, else the newest doc's, any date); 'calibration in/tiers [(out: ...)]' (arm tiers, by task id, within 50-85 %). A max run adds 'max ' to budget (never a FAIL), 'max: ...' to graders, and 'noise max ±f/n [Δscore d]' per unstopped arm of its newest run. Exit 1 only on a budget or graders FAIL. bench.py audit [ | --latest light|full|canary] [--arm A] [--unaudited] [--show N] [--set / agent|grader|harness --why TEXT] [--results DIR] Every failed task '/ [|unaudited] ' of one results JSON (default: the newest in DIR, default ./bench/results); --show N: the first N lines of the retained answer (//.out.json); ids '/#a' when the doc's repeat >= 2; --set / (every failed attempt) or /#a writes failure_kind + audit_note into the JSON, then re-imports it into the ledger. Ends 'unaudited u · agent a · grader g · harness h'; exit 0, or 2 when the file is missing (docs/pa3-build/design/bench.md ## Audit). bench.py readme [--results DIR] regenerate README.md beside a results/ DIR, else inside DIR (default ./bench/results), from its results JSONs; prints 'readme: ' (the bench Action's publish job, branch bench-results). Fixture root: --fixtures, else