42 KiB
Project Architect 3.0
The methodology that turns Claude Code into a governed project architect: it takes a rough idea, an existing codebase, or an old-methodology project and produces a scoped, phased, documented build — then runs the build phase by phase with a living document system, pinned models, measured costs and hard stops between phases so any fresh session can reconstruct exactly where the project is and how to work on it.
For humans
Two commands install the system. Once per machine, from the package folder:
python pa_install.py --root --yes
Once per project, from the project's own root:
python pa_install.py --project . --yes
After that you open a terminal in the project and type claude. The session reads a seed,
determines where the project is (planning a generation, planning a phase, the work loop, a
review pause) and starts. You interact at a few moments: approving a plan, answering a
review's decisions, responding when the critic says the developer must decide, answering the
close question, allowing a command outside the allow rules, and opening
/discuss [model] [effort] when you want to think alongside the system (Opus 5.5 at medium unless you opt
into Fable or Sonnet, or high or max effort; /thoughts stays an alias for one release). Everything else — spawning experts,
briefing coders, running tests, committing, closing phases — happens without a keypress.
The roles in one paragraph: a planner drafts the plan you approve; a router spawns one expert per task and never does task work itself; the expert thinks, decides and briefs coders (which run the edit-build-test loops) and retrievers (which look things up); a critic judges every plan change an expert proposes; a review agent turns a review pause into numbered decisions for you.
A day in the life: you launch with claude, the router reads the seed and picks the next
task, spawns the expert in the background and waits. The expert reads the plan, spawns a
coder, waits idle in the background until the coder's hand-back wakes it (a per-session warmer
daemon pings idle agents with . before the cache TTL, at 285 s or 3,300 s, relayed by the
router's Monitor; the agent answers .), writes the task log and summary, and returns done. The router marks the task done and moves to
the next. When every task is done the closer verifies the milestone, writes the phase-end
record and archives the phase. You relaunch for the next one.
Where to look: HOW_WE_WORK.md for standing decisions, tools and environment;
phase-ends/current/PHASE_PLAN.md for the live plan; phase-ends/TASK_INDEX.md for what
was done; the cookbook (cookbook/INDEX.md, grep by tag) for techniques; rules/INDEX.md
for every rule in force.
Philosophy
The creed, carried from 1.3.0 and amended by three generations of experience:
- The constitution is permanent. The rules, the cookbook, the standing decisions and the phase records are the living system. Never conflate the layers.
- History is append-only. A rule is never deleted; it is superseded with a note. A phase
record is never rewritten; a correction is a new record. A file is never removed from the
repository; it moves to
docs/retired/. - Rules are living. They are proposed at phase ends, adjudicated by the planner, and added
to
rules/as individual files. The expert writesRule candidate:in the task summary; the phase planner adjudicates. - Every phase milestone is an observable, machine-checkable outcome. The gate, not anyone's say-so, is the arbiter.
- Generations are evolutionary leaps, not version bumps. A new generation only when a fundamentally new capability requires the prior generation stable and proven. Most projects need two or three.
- Knowledge compounds only if captured: consult the cookbook before, feed it after, and write context-dependent findings during the session that produced them.
New in 3.0: one task per fresh context (the expert knows its task, not the whole phase). The
plan is the contract between the developer and the system; plan_edit.py is the only writer.
Nothing enters a context that a pointer could replace: a summary id instead of a pasted body,
a file:line instead of a code block, a report id instead of a retriever's raw output. The
plan's reads: line names what the expert loads; everything else is a retriever dispatch.
The document system
The five layers of 2.0 are replaced by a deeper structure. Precedence runs top to bottom; where wording differs, the higher layer governs.
CLAUDE.md(about 190 tokens): the roles line, pointers to the project-architect skill,HOW_WE_WORK.md, the live plan and the constitution. It is not the rules; it points at them. The rules themselves live in.claude/skills/project-architect/SKILL.mdand each agent file names that skill underskills:. Leaf agents (retrievers, the critic) carryomitClaudeMd: true, which dropsCLAUDE.mdand the auto-memory index from their context; the critic keeps the skill because it must quote the rules.- The project-architect skill: eleven numbered sections (precedence, roles, output, proportionality, git, verification, hygiene, knowledge, the developer, long compute, house style) plus a project-rules section that holds project-specific additions. This is what governs every expert and coder.
HOW_WE_WORK.md: standing decisions, environment, tools, the docs map. The living reference for how this project runs on this machine. Target size about 3,500 tokens.PROJECT_CONTEXT.md: the constitution. Permanent and static; never edited after generation.GENERATION_PLAN.md: the generation-level roadmap (phases, milestones, scope, ordering rationale). Changed only at generation boundaries.phase-ends/current/PHASE_PLAN.mdand itstasks/,logs/,research/,discussions/directories: the live work of the current phase. The plan is frozen at approval and changed only throughplan_edit.py. Task summaries, logs and research reports are written by the agents that do the work. At phase close,phaseend_index.py archivemovescurrent/tophase-ends/phase-<N>/.PhaseEnd_Phase<N>.mdandGenerationEnd_<G>.md: the append-only build history. Assembled by scripts from the phase's task logs and summaries, the milestone verification, and the closer's recap.rules/(one file per rule),cookbook/(one file per technique),docs/ops/(environment and build references).
Auto memory stays on. Claude Code's auto memory keeps one directory per repository under
~/.claude/projects/<slug>/memory/. The first 200 lines or 25 KB of the index load at
session start for every agent that reads CLAUDE.md. PA3 keeps it enabled and curates it at
every generation start: standing facts belong in HOW_WE_WORK.md, rules in rules/,
techniques in the cookbook. The memory index is kept small and current by a curator agent
rather than switched off. The memory directory lives in the repository (.claude-state/memory/) behind a
junction so a fresh machine restores it with pa_install --project.
Roles
| Role | Agent | Model | Effort | Spawns | Tools |
|---|---|---|---|---|---|
| Planner (generation) | planner-gen | Opus 5.5 | medium | retrievers | — |
| Planner (phase) | planner-phase | Opus 5.5 | medium | retrievers | — |
| Router | pa-session | Sonnet 5.5 | medium | experts, planners, critic, review, discuss, curator | — |
| Expert (default) | expert-opus55 | Opus 5.5 | medium | coders, retrievers | — |
Expert (hard, effort: high) |
expert-fable | Fable 5.1 | medium | coders, retrievers | — |
| Coder (default) | coder-opus55 | Opus 5.5 | medium | retriever-code | — |
| Retriever (code) | retriever-code | Sonnet 5.5 | medium | — | Read, Grep, Glob, Write |
| Retriever (digest) | retriever-digest | Sonnet 5.5 | medium | — | Read, Grep, Glob, Write |
| Retriever (web) | retriever-web | Sonnet 5.5 | medium | — | WebSearch, WebFetch, Write |
| Critic | critic | Fable 5.1 | medium | — | — |
| Review | review | Fable 5.1 | medium | retriever-code, retriever-digest | — |
| Discussion | discuss, discuss-high, discuss-max | Opus 5.5 | medium, high, max | — | — |
| Memory curator | memory-curator | Fable 5.1 | medium | — | — |
Models and efforts are pinned per agent file; the developer never touches /effort or
/model. The cost structure follows from the pinning: Opus 5.5 is the default builder and
expert (coders, the planners since 3.9.5 and the medium-effort expert), Fable 5.1 the thinker for hard tasks (
the critic, the hard-tier expert), and
Sonnet the middle ground (the router, all three retrievers). The router runs
on Sonnet because it is purely mechanical: it reads a seed, picks the next task, spawns an
expert and waits. Judgment lives in the subagents. The phase planner marks effort: high only
when a task's done-when rests on a judgment no test can arbitrate (a design decision, a harness
probe, a proof read from evidence), never for size or importance; at most one task in five per
phase plan, each high mark's judgment named in the plan's Rationale.
The maximum spawn depth is three: router (0), expert (1), coder (2), retriever-code (3). The
depth is set explicitly in the environment (CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH=3) so a
harness default change cannot silently break it. A coder's retriever is a leaf by construction
(no Agent tool) so the chain cannot grow further.
The expert delegation rule: a coder for any build or test loop, more than about twenty lines,
or more than one file. The expert may make one edit of twenty lines or fewer itself with one
verification run. Two coder failures with different causes return
blocked.
Every role reads a file over guard.whole_read_chars (20,000 chars by default) by range only:
a Read without offset/limit on such a file is denied, and the denial carries the file's
outline inline (what tools/outline.py <path> prints: header block, symbols with their intent
lines, section comments; capped at 60 lines, Markdown at depth 2) and ends "Read the ranges you
need with offset and limit". Retrievers get the same map without a shell; the served outline
writes the same outline credit note as the script. The audit's whole_reads row mirrors the
same threshold. A retriever's report, when written, lands first at
research/pending/<agent>-<slug>.md; the retriever returns REPORT: pending/<file>, and the file
is adopted into R<N>-<nnn>.md by tools/research_add.py adopt, run at the expert's hand, at
task_log.py finish, and at phaseend_index.py assemble.
Lifecycle
The launcher (tools/launch.py) checks the project's state and determines the mode:
- No
GENERATION_PLAN.md, or all its phases closed: planner-gen mode. - Generation open, no approved phase plan: planner-phase mode.
REPLAN.mdpresent: planner-phase in replan mode.REVIEW.mdpresent: review mode.- Otherwise: the router loop.
In planner mode the router spawns a planner (generation or phase) which explores only through
retrievers, drafts into .run/, and returns a summary. The router puts the summary in front
of the developer through the question picker: approve, show the full plan first, or send
changes back. On approval the router writes the plan with plan_edit.py from-draft, stamps
it with plan_edit.py approve, records ratified rules, and commits. From that point
plan_edit.py is the only writer of the plan file; a PreToolUse hook denies any other
write to it.
In the router loop, at every task boundary the router consumes INBOX.md if the developer
left notes, asks plan_edit.py next for the next runnable task, stamps the statusline and
spawns the expert in the background. Then it ends its turn; the harness wakes it when the
expert returns. While an expert runs, the statusline shows the task, the agent, the context
and the time elapsed. The expert reads the plan's Context, Interfaces and Cookbook sections,
its own task entry, and only the summaries the entry names. It works, spawns coders and
retrievers as needed, and returns its contract.
When an expert returns done, the router marks the task done
and goes to the next task. When it returns question or blocked, the router sends the
question to the critic. A handoff (context crossed the 350,000-token threshold) triggers
a respawn: the expert writes TASK_PROGRESS.md, commits and returns; the router respawns
the same agent with the progress file in its brief. The respawn archives the file to
logs/T<n>.progress<k>.md when done.
When plan_edit.py next prints NONE, the router asks the developer once whether anything
else should come first, then spawns the phase closer (expert-opus55).
The closer runs phaseend_index.py verify, which executes every verified by: clause of the
milestone and prints GREEN or RED per clause. It writes RECAP.md, promotes techniques from
generalizable: lines to the cookbook, commits, and returns MILESTONE: green or red. On
green the router assembles, lints and archives the phase, commits, and gives the developer
the recap with the agent-runs table. The session ends with Relaunch: /clear then go.
The carry audit rides on the close. phaseend_index.py assemble runs pa_ledger.py audit --phase,
which walks the phase's transcripts once (zero tokens, repeatable) and writes its tables into the
PhaseEnd's Audit section: reads by file and role, writes by file, results by kind, whole-plan and
tool-source reads, spilled results read whole, noise, carry per request against the previous phase,
and candidates promised beside measured. Rows past the thresholds become - flag: lines; the seed
then carries Audit flag: and the router spawns the auditor (Fable, medium effort) before the next
planner. The auditor judges tables, never transcripts: one verdict per flagged row (tool, split,
guard, script-fix, leave), project-level fixes applied through coders, harness-level ones written as
Tool candidate: lines in phase-ends/current/AUDIT.md with what the tool does, what the model did
instead, the occurrences, the saving on the credit formula and the build size. The planner copies them
into its return; the developer ratifies each with one picker question at plan approval; a built tool
earns the savings credit automatically and the next audit prints measured beside promised.
At a generation boundary, genend_index.py assemble builds the GenerationEnd from the
phase-end files. The next claude opens in planner-gen mode and the generation-start ceremony
(the memory curator, then the auditor over the closing generation's last PhaseEnd) runs before the
new planner drafts.
Every /discuss [model] [effort] session that proceeds writes a discussion record with a Status: tag
(open, executed, planned:<phase> T<k>, deferred:<phase>, dropped, failed) and an
index line; the phase's discussion lines fold into the PhaseEnd on archive, and cumulative
history lives in phase-ends/DISCUSSION_INDEX.md. A record's ## Deferred lines and an
inbox's later: lines both flow into the PhaseEnd's ## Deferred block, so a later planner
either plans the item as a task or leaves it deferred with one line of reason.
Every configured MCP server adds its instruction block to every session, whether or not the
task at hand needs it; doctor lists the servers found in ~/.claude.json and the project's
.mcp.json, and the standing advice is to disable the ones a project does not use.
Contracts
Every role returns a fixed-format contract of at most 300 tokens. Nothing else. The contracts are the system's only inter-agent communication; they carry status, pointers and one-line recommendations, never pasted content.
Expert brief (router to expert):
TASK: T3 — <title>
PLAN: phase-ends/current/PHASE_PLAN.md
HOW: HOW_WE_WORK.md
LOGS TO READ: phase-ends/current/tasks/T1.md, tasks/T2.md
CODER: opus55 | none EFFORT: high | medium
PROGRESS: <path> (respawn only)
NOTE: <developer text or critic's WHY> (only when present)
DONE WHEN: <from the task entry>
RETURN: the expert contract, ≤300 tokens, nothing else.
Expert return:
STATUS: done | blocked | question | handoff | review
LOG: phase-ends/current/tasks/T3.md
CTX: 142k
COMMIT: abc123 | none
MILESTONE: green | red | n/a (phase-end only)
RECOMMENDED: <one line>
QUESTION: <one paragraph; only if question>
Coder brief (expert to coder):
CODER TASK: T3.c1 — <one line>
CHANGE: <what to implement; files and functions; ≤8 lines>
INTERFACES: <named entries from the plan's ## Interfaces>
CONSTRAINTS: <2–4 lines>
BUILD/TEST: <exact commands; what green means>
DONE WHEN: <machine-checkable>
LOG: phase-ends/current/logs/T3.c1.md
RETURN: the coder contract, ≤300 tokens.
Coder return:
STATUS: done | partial | blocked
CHANGED: path:lines, … (≤6 entries)
VERIFIED: <command -- result>
LOG: phase-ends/current/logs/T3.c1.md
CTX: 96k
COMMIT: def456 | none
DEVIATIONS: <one line | none>
BLOCKER: <one paragraph; only if partial/blocked>
Retriever brief: QUESTION · FOR · SCOPE · CAP · REPORT yes|auto · KNOWN. Return: the
answer only, first line REPORT: R24-007 when a report was written. Code retrievers return
signatures with file:line, call sites one per line, shared state and nearby gotchas, up to
40 lines.
Critic brief: the plan sections, task index, named summaries, the question, the
recommendation, and a tier guess. Return: DECISION (continue | edit | needs-developer) · TIER · EDITS (plan_edit.py lines) · WHY (3 lines) · BRIEF (needs-developer only, ends "Recommended: ...").
Replan tiers and the critic
Changes are graded by scope and routed to the agent with the right authority:
- Approach within the task: the expert decides and notes it under Deviations.
- A developer's inbox item that only adds within the milestone: the router applies the edit and adds a Changes entry.
- Any plan change an expert proposes, additive or not: the critic decides. It returns
continue(carry on),edit(exactplan_edit.pycommands the router runs and commits) orneeds-developer(a one-paragraph brief with a recommendation the router puts to the developer in the question picker). - Structural (milestone or scope change): the developer, through
REPLAN.md. The nextlaunch.py --seed-onlystarts a planner-phase pass that archives the old plan and writes a new one for approval.
The critic exists because the expert's context is deliberately narrow. An expert knows its own task, not that another task tests the same thing from the other side. When an expert hits a wall it will sometimes recommend removing the wall. The critic sees the full task index, the milestone, and the plan's rationale, and can judge whether the recommendation protects the milestone or undermines it. Cost: about $0.50 per invocation.
Discussion, questions, notifications
The developer is never interrogated. Agents give recommendations, not questions; when a decision is genuinely the developer's (risk appetite, product intent, money, scope), the recommendation states consequences and ends with a recommended answer.
Three channels reach the developer:
- The question picker (plan approval, review decisions, the critic's
needs-developer, a planner's Developer-decides items): interactive, in the terminal. - Plain text at the end of a turn (the close question, and anything the router must ask outside the picker).
/discuss [model] [effort](the discuss agent): the developer opens a read-only thinking session with Opus 5.5 at medium (opt-in: fable or sonnet, high or max;/thoughtsan alias for one release) in its own terminal view.EditandWriteare denied while the discussion flag is set. When the developer saysproceed, the agent writes the record underdiscussions/and returns the decisions as exact plan edits for the router to apply.
A note for a running expert is typed into the expert's own view (its context stays); the
router relays nothing and relaunches only on stop: or relaunch:. INBOX.md takes
anything for the next task boundary.
Notifications: at the default toast: waiting level (.claude/pa.json; the others are all
and off) a Windows toast fires only when something waits on the developer: a Stop that ends
on a question (the toast carries its text), a permission prompt or a picker question (Claude
Code reports both as permission_prompt), a new REVIEW.md or REPLAN.md, a finished
discussion, and a turn stopped by an error. Phone pushes come from Claude Code itself: with
Remote Control on and its /config push switches set (the machine install turns them on
unless you set them yourself), permission prompts and picker questions reach the Claude app.
The toast child gets a console without a window (CREATE_NO_WINDOW) so there is no flash and
no focus theft.
Mode 1 intake
3.0 change: Gate 2 no longer prompts for /effort max. Generation happens in planner-gen
at medium effort; the model and effort are pinned by the agent file.
In 3.0 the intake and the constitution happen in a plain session (claude --agent plain) or
in Claude Chat with the chat kit; that session writes the outputs of the next chapter itself.
The governed loop starts at the first bare claude after pa_install.py --project; an
intake agent is a later item.
Conversational, not an interrogation. Weave questions into the discussion; track silently against the intake checklist. Do not generate the constitution during brainstorming, no matter how much detail accrues.
Intake checklist (all 12, silently tracked):
- The elevator pitch — what is this, in two sentences?
- Motivation — why build it; is it a rewrite/successor of something?
- Who is the developer — background, skill level, solo or team? (feeds
who-is-dev) - Tech stack — chosen or open; existing constraints?
- Hard constraints — platform, deadlines, budget, licenses, privacy/closed-source posture?
- Core features — the non-negotiable heart.
- Stretch features — wanted, not required.
- Dream features — someday/maybe (Parking Lot material).
- Known risks and pain points — what worries them; what failed before?
- Definition of success — what does "it works" observably look like? (feeds milestones)
- Multi-user/deployment reality — who else touches it, where does it run?
- Resource management — data, compute, storage, external services, costs.
Gate 1, the structured review (when the developer says "ready"): do not generate yet. Present recommended phase ordering with reasoning, standard features they missed, nice-to-haves worth considering, scope-creep warnings, methodology concerns, architecture suggestions with reasoning, and risk flags. Then ask what to adjust.
Gate 2, the generation confirmation (after they confirm the review): generation is Tier-1
work. In 3.0 the planner-gen agent runs at medium effort by its agent file; the developer is
not asked to toggle anything. The intake session writes everything listed in the next
chapter; the first bare claude afterwards enters planner-gen mode.
Generating the constitution
3.0 change: the outputs list is updated. The registry is replaced by individual files under
rules/; effort and model are pinned per agent file, nothing is mapped per phase; the
cookbook is generated as cookbook/ entries and the ops reference as docs/ops/ topics.
CLAUDE.md is about 190 tokens of pointers with two default fail-safes, written by the
installer (item 3 below).
Phase design principles (you design the phases, not the developer):
- Core before periphery. Data before processing. Foundation before features. Safety before action. Observation before optimization.
- Every phase produces something runnable and testable. No planning-only or models-only phases.
- Every phase ends with an explicit Milestone line: an observable, machine-checkable outcome
("
make checkgreen", "the level loads and the character walks", "round-trips 100 files byte-identical") — the gate that the closer will hold the phase to. - Validation is a dedicated phase, not an afterthought.
- Enhancement layers are toggleable, feature-flagged modules added one at a time and measured before the next.
- Size phases to the developer (see Skill-level handling): beginners get fewer, chunkier phases; advanced developers get fine-grained ones.
Domain phase-ladder starting points (adapt, do not copy):
- Game/game-tool: foundation and skeleton, data formats and I/O, core domain model, core mechanic vertical slice, content layers, UI shell, integration, validation, polish, packaging.
- Backend/API: skeleton and config, data layer, domain model, first vertical endpoint, auth and safety, remaining endpoints, integration, validation, observability, deployment.
- Desktop app: skeleton and DI shell, data/persistence, core engine, minimal UI, feature layers, import/export, validation, polish/packaging.
- Automation/pipeline: skeleton and config, source connectors, transform core, dry-run end-to-end, guarded write path, scheduling, observability, validation, hardening.
- RE/decompilation/format: environment and oracle setup, deterministic extraction pipeline, ground-truth maps, byte-exact verification gate, first verified unit, scale-out with a knowledge flywheel, the hard tail via an escalation ladder, integration/packaging.
Generation design principles:
- Generations are evolutionary leaps, not version numbers. A new generation only when a fundamentally new capability requires the prior generation stable and proven, and there is a hard dependency between them.
- The constitution's roadmap covers Gen 1 in phase detail; later generations get a Generation Map sketch and a Future Generations section, not premature phase lists.
Generation outputs (the complete set):
PROJECT_CONTEXT.mdon the skeleton — every section, with the Quick Reference Card, Key Decisions, Build Roadmap with machine-checkable Milestones, Parking Lot.rules/G*.md— the project's domain rules as individual files with anINDEX.md, seeded from the project's oracles, definition-of-done gate, data constraints and environment constraints.CLAUDE.mdis written by the installer (about 190 tokens of pointers and two default fail-safes: never push, the milestone gate is the arbiter); nothing regenerates it, the developer edits the fail-safe lines by hand if the project'sG*rules justify others.docs/ops/— environment, build/run/test commands, version pins as known at generation, one file per topic with anINDEX.md.cookbook/INDEX.md— the initial techniques, one file per entry, named for the project's recurring craft.HOW_WE_WORK.md— the card, generated from environment detection (pa/install/detect.py environment()), the developer interview and the constitution's opening and Quick Reference Card. It is one file, sliced per role ({{PY}} tools/card.py slice <role>prints only the sections whose heading tags name that role) and capped (card.max_chars, checked by{{PY}} tools/card.py check); the archive enforces the cap at phase end and nothing is ever appended to it — a standing decision or an environment fact is routed to a rule, a product doc, a Tools-table row ordocs/ops/instead.HOW_WE_WORK.md ## Developerfilled from intake item 3 and the skill-level assessment (no memory is written); it is one of the sections therouterandplannerslices carry beyond the shared ones.
Rule detection
3.0 change: "propose at PhaseEnd" becomes "the expert writes Rule candidate: in the task
summary; the phase planner adjudicates."
| Signal | Example | Likely rule type |
|---|---|---|
| The developer corrects a tool/framework behavior you assumed | "maspsx's default is not latest" | Pin/make-explicit rule |
| The same mistake happens twice in one phase | Two stale-build false diffs | Verification rule |
| The developer says "never do X" / "always do Y" | "never push", "always dry-run first" | Process rule |
| A wiring step gets missed | Registration forgotten | Checklist rule |
| An assumption proves wrong in practice | "the archive is per-day" was not | Provenance/verification rule |
| A strong workflow preference emerges | Per-task commits after the log | Process rule + memory seed |
| A costly gotcha with a cheap guard | The enum-zero config trap | Cookbook entry first; rule if it recurs |
When proposing: state it specifically ("always Y when Z", not "be careful"), one-line justification, developer approves or modifies or rejects — never add unilaterally. Some things are deviations to record, not rules. Techniques go to the cookbook, norms of conduct to the rules.
A rule enters the package's seed only when it has recurred across projects and fits in one
sentence of conduct; otherwise it stays a project rule (a G* rule or a per-phase R rule).
A sharper form of an existing seed rule becomes one clause in that rule, never a new id;
techniques never enter the seed.
Migration
Migration replaces 2.0's Mode 4. The alternate-heading sweep rule is kept: when consolidating rules from an old project, sweep every rule-bearing heading, not just the literal "Rules Added This Phase." Older phases carry rules under alternate headings ("Key Rules Confirmed", "Architecture/Design Decisions", "Key technical decisions"); a consolidator keying only on the canonical heading silently misses them.
PA3 migration is one engine (pa/install/migrate.py) and three thin layout modules. The engine
handles git mv from a mapping table, heading-level splits, retiring, generating and committing.
The layouts differ only in their mapping tables and their HOW_WE_WORK sources. The installer
detects the project shape and selects the layout automatically.
The three paths:
-
Stock 2.0 (
pa/install/stock20.py): a project withCLAUDE.md,RULES_REGISTRY.mdandphase-ends/in the standard 2.0 layout. The registry is split by its###headings into individual files underrules/. The cookbook is split by its##headings intocookbook/entries. The ops-setup document is split by its##headings intodocs/ops/(giant sections split again at###if over 400 lines). The effort map, the old templates and the olddocs/project-architect.mdare retired.CURRENT_PHASE.mdbecomes a partial log. PhaseEnds stay in place; aLEGACY_INDEX.mdis generated. The PA2SessionEndbackup hook is removed from settings; no copy is kept, the harness holds transcripts for ten years (cleanupPeriodDays). The interphase sorts after the last closed phase and before the next. -
1.x-shaped (
pa/install/onex.py): a project whoseProject Context Markdowns(or similar) folder holds PhaseEnds with underscore separators, plan/verdict/prereg documents mixed in, a large rules registry under a non-standard name, and no.claude/settings.json(or an empty one). PhaseEnds are renamed in place (_to.between version tokens) andsort -Vchecked. Non-phase documents (plans, verdicts, prereg, checkpoints, audits) move todocs/research-archive/with a generated index, linked fromRESEARCH_INDEX.md's## Legacysection. The rules registry is split from its own headings intorules/. The project context becomesPROJECT_CONTEXT.md(verbatim, static). Settings are written fresh from the project snippet. The interphase label is proposed from the newest phase family;--interphaseoverrides. -
Overlay kit (
pa/install/overlay.py): a project with aDIGEST.md, R-numbered rules above 100, a kit folder with its own agents and commands, and a large cookbook with a separate index. Rules come from the DIGEST's rule section and from the kit'sregistry-E*.mdtemplates. The DIGEST is frozen in place with a header line. The cookbook is streamed intocookbook/entries with index rows taken from the existingcookbook-index.mdwhen its rows match the headings. The kit folder and its agents are untouched; name collisions with PA3 agents fail the install and name the collision. Kit agents that do not collide are listed inHOW_WE_WORKas extra experts the plan may name. Existing hooks in.claude/settings.jsonare appended to, never replaced. The interpreter ispython3when the project runs on POSIX.
Every path follows the same runbook discipline:
- Dry run first.
--dry-runprints the full mapping table (git mvlines, split counts, retire lines, unmapped files) and writes nothing. The developer reviews the table and confirms once. - One confirmation. After the dry run, the install runs with one confirmation prompt per phase of the install. No per-file prompts.
git mvonly. Every move preserves history. An untracked source is moved as a plain file and noted.git log --followon a moved file shows the pre-move commit.docs/retired/for everything leaving the load order. Nothing is deleted.- Never push. The developer pushes when satisfied.
- The migration's first planning step is the curator. When a project already has a memory directory, the memory curator runs before the planner drafts the first phase, demoting generation-specific memories into a linked archive and keeping the index small.
The resume pointer (.run/install-state.json) records the last completed step. A crash or
interrupt is followed by a rerun that resumes after the last good step. A second run with
everything done prints SKIP for every step.
Knowledge
PA3 captures knowledge through four indexed stores, each managed by a script that enforces the format and keeps the index current.
Rules (rules/INDEX.md, one file per rule): tools/rules_add.py add writes the file and
appends the index line. Rules are proposed at phase end (Rule candidate: in a task summary),
adjudicated by the phase planner, and never added unilaterally. A superseded rule keeps its
file with a superseded-by: line; the index marks it.
Cookbook (cookbook/INDEX.md, one file per technique): tools/cookbook_add.sh --title ... --tags ... --file ... writes the entry and appends the index. Planners and experts grep the index
by tag before recurring work; they never read it whole (the index can grow past 400 lines in
a large project). A technique that recurs across projects is stronger evidence that it is
general.
Research (research/INDEX.md, one file per report): tools/research_add.py new ... | index <id> writes the report and appends the index. Retriever findings land here; experts cite
report ids and never read the bodies. The index carries the title, the date and the scope.
Skills (.claude/skills/): tools/skill_add.py writes a skill file from a workflow gotcha.
Skills are named per agent and listed in the agent file's skills: frontmatter. They load
into the system prompt of every agent that names them.
The flywheel: every task ends with a full log (logs/T<n>.md) and a summary
(tasks/T<n>.md, at most 150 lines). The next task reads summaries only. Gotchas are
marked generalizable: (promoted to the cookbook at phase end) or workflow: (promoted to a
skill). Decisions are marked binding: (promoted to HOW_WE_WORK.md at phase end). The
proportionality rule applies: no features, abstractions, options or tests beyond the change.
A one-time fetch is a fetch, not a downloader.
The generation-start ceremony: when a generation closes and the next one opens, a
memory-curator agent (Fable 5.1 at medium effort) runs before the planner drafts. It reads
every memory file, the cookbook and rules indexes, and HOW_WE_WORK.md. It demotes
generation-specific state into archive files: memories into memory/gen<N>.md, cookbook
entries into cookbook/gen<N>/, rules into rules/gen<N>/, and superseded standing decisions
into docs/retired/HOW_WE_WORK.gen<N>.md. It commits the result and returns a recap. The
demotion pattern is always: append content to an archive under a generation directory, remove
the pointer from the main index, and keep one link line in the main index pointing to the
archive. Active demoted rules keep their main-index line, marked demoted:gen<N>.
The four subcommands of tools/curate.py (memory, cookbook, rules, how-we-work) each
support --dry-run. The how-we-work --report subcommand lists trim candidates by date and
token budget; --retire moves them into the generation archive. The pa-session detects the
ceremony via a curate: line in the seed (written by launch.py when a GenerationEnd
exists) and spawns the curator in the foreground before the planner.
The curator's pre-flight is pa_ledger.py doctor --sizes: the sizes of the governed files
(CLAUDE.md, each skill, HOW_WE_WORK.md, the memory index, the cookbook and rules indexes)
and the current expert seed against their targets. Over a target is a WARN the curator acts
on; only a memory index past the harness's own cap fails. The closer's ## Audit section in
every PhaseEnd records the expert seed per phase, its growth over the previous phase, and
the same sizes.
Standing decisions of 3.1 through 3.3 that change how a project is run:
- The measured saving is what a vanilla Claude Code session would have paid to carry what the
helpers kept out: one 1h cache write per result per session, re-reads at the read price, the
carried results rewritten only at a real idle gap over an hour, bounded by the parent model's
window, minus every helper run's seed carry. One rule at every scope (run, session, window,
project, account). The statusline reads it in the subscriber's units (percent of the 5h and
7d windows, weeks of allowance at the user's own live exchange rate, "N% fewer tokens than
vanilla, lasts Mx longer" per scope). Every input is the user's own ledger, nothing is
calibrated from anyone else's runs; the modeled replay stays in
reportas a cross-check. - The meter weighs model families differently: a Fable dollar consumed the 5h window about ten times faster than an Opus dollar (as of 2026-09-20). The fit carries per-family weights and the "percentage of window" figures are summed per family.
- One ledger per machine, written only by its own hooks. Other roots are read in copy mode, never over the WSL file boundary. The fit, the window samples and the window savings read the union on the sampler's change path and in the CLI; the hooks' rebuild and the render stay local.
- A project's memories live in its repository (
.claude-state/memory/) behind a junction from the harness's memory folder, sopa_install --projecton a fresh machine restores them. Never delete a memory; demote it to a linked generation archive. Only the index costs context. - The generation-start ceremony (curator, demotions, trim, seed audit,
doctor --sizes) runs before the planner at every generation boundary.
Skill-level handling
3.0 change: the developer profile lands in HOW_WE_WORK.md under ## Developer rather than
only in a memory. No change to the levels.
- Beginner: explain more, simpler architectures, fewer and chunkier phases, warn plainly when scope is over-ambitious. Expand the explain-before-coding depth.
- Intermediate: explain the non-obvious; standard patterns; normal phase granularity.
- Advanced: brief reasoning, respect their calls, fine-grained phases, argue only when the methodology stakes are real.
- When in doubt, ask. Record the profile in
HOW_WE_WORK.mdat install or generation so every future session starts calibrated.
Sunset list and changelog
Sunset list (mechanisms removed from 2.0)
Mid-phase rules check and re-read every four tasks. Session-start rule recitation. Printing
checkpoint blocks; the blocks themselves. "Checkpoint CURRENT_PHASE every turn" and
CURRENT_PHASE.md with its template. The effort map, the tier prompts, the "drop back down"
reminders, the context buckets and reset bands, and the statusline-injection checkpoint hook.
Token-lean SSP block and digest subagent. Duplicated fail-safe rule blocks. TaskCreate after
approval, task-list rebuild on resume, and the Tasks-tool setting. Feedback memories as
corrections and the memory index. "Explain before coding" (replaced by: one line when
deviating). The "no truncating reads / spend freely" clause (keeping "never state a
percentage, never hurry"). Long commit messages. Inline one-liners. "Present results in chat"
(replaced by the task log and REVIEW.md). Verbatim rule printing.
Changelog 2.0 to 3.0
| Area | 2.0 | 3.0 |
|---|---|---|
| Entry point | SETUP.md path A/B/C |
pa_install.py --root + --project |
| Rules store | One RULES_REGISTRY.md (recited every session) |
rules/ one file per rule; grep the index |
| Session startup | Rule recitation + checkpoint restore | Seed file (launch.py --seed-only) |
| Effort and models | Developer toggles /effort; effort map per phase |
Pinned per agent file; the developer never touches /effort |
| Execution | One task, confirmation, plan-mode + Max | One task per fresh context; two gates per phase, autonomy between |
| State recovery | CURRENT_PHASE.md (one large file, reloaded every turn) |
TASK_PROGRESS.md (written at handoff, archived when done) |
| Knowledge stores | One cookbook file, one registry, one ops file | cookbook/, rules/, docs/ops/ with indexes and flywheel scripts |
| Agent state | .claude-state/ (repo-contained) |
Same, plus the memory junction and the curator |
| Phase close | PhaseEnd (append-only) + the developer closes | phaseend_index.py verify then assemble, lint, archive |
| Cost visibility | None | Usage ledger, statusline with savings and pace, pa_ledger.py report |
| Migration paths | Mode 4 (one protocol) | Three layout modules: stock 2.0, 1.x-shaped, overlay kit |
What is measured (the acceptance metrics): average context per expert request at most 150k
tokens; expert seed at most 40k; retriever seed at most 6k; router context at phase end at
most 40k; zero prefix rewrites not at session start; zero Tasks-tool attachments; cost per
request at most $0.10; developer interruptions limited to the critic's needs-developer plus
the milestone gates.
Appendix: the chat kit
The package still works for brainstorming in Claude Chat without Claude Code. Attach three
files to a new chat: this document (project-architect-3.0.md), the constitution skeleton
(PROJECT_CONTEXT.skeleton.md) and the rules seed (rules-seed/INDEX.seed.md). The AI runs
Mode 1 above and generates the constitution as a markdown artifact with the rules and
protocols embedded. That output is migration-ready: when the project later moves to Claude
Code, the installer's migration path picks it up.
The dispatch section from 2.0 shrinks to one rule: if you are Claude Code and a
.claude/pa.json exists, the installed system governs and this document is the reference
specification. If you are Claude Chat with this document attached, run the modes from this
document. Never mix the postures.