diff --git a/docs/SETUP.md b/docs/SETUP.md index 6ca96adfa6..a7f4027c45 100644 --- a/docs/SETUP.md +++ b/docs/SETUP.md @@ -749,6 +749,8 @@ Every script under `tools/` (plus the two report make-targets), grouped by purpo | | `DefineFunctions.java` | Disassemble + create functions at splat's validated entry points (`.run/_funcs.txt`) — completes a raw-blob program's function set (Phase 10). | | | `ApplySymbols.java` + `tools/ghidra_apply_symbols.sh` | **(P31 S78) The Ghidra MIRROR of the curated symbol file (R15/G6), headless with a real save.** `tools/ghidra_apply_symbols.sh [PROG] [symbols files…]` (defaults `SLUS_007.26 config/symbols.us.txt`; MCP must be STOPPED first) reads `name = 0xADDR;` rows and sets every function/label to its curated name; a name held by another address is moved to that address's own curated name first (`firstfile`/`firstfile2`), else to `__at_`. Idempotent; prints `BFMAPPLY renamed_funcs=… unchanged=…`; R9-verify with `ghidra_mcp_verify.sh`. **Use this, not MCP `rename_symbol`/`batch_rename`, for renames:** S78 observed 47 MCP renames NOT persisting through the sentinel stop ("Save succeeded", DB grew, names gone — R9 caught it; cause not yet isolated), while the postScript path persisted 73/73 on the first run. | | **Public flip / CI** | `.github/workflows/no-rom.yml` | **(P33 B7) The ROM-free CI**: job `audits` (audit_public, audit_text_sources, verbatim_check --strict, cookbook_index --check, ghidra_roster --check, work_evidence --selftest, test_lzss, lint_symbol_refs — ≈45 s of checks) + job `compile-only` (binutils-mipsel + `cpp-mipsel-linux-gnu` from apt, cc1 from the tracked tarball sha256-checked, maspsx submodule; PR scope `main resident ov_SC01_077 md_MAIN_013`; `--all` weekly Mon 06:17 UTC + `workflow_dispatch`). Byte-identity is NOT proven in CI (needs the disc) — `docs/verification.md`. | +| | `tools/public_rewrite/` (P33 C1) | **The history-rewrite package** (`docs/public-flip-runbook.md` §3 is the operating table). `common.py` (shared: the purge rules, the DERIVED content-hash sets, identities from the log, the one hash regex, a persistent `cat-file --batch`) · `hash_dict.py [--write-mailmap]` (every commit OBJECT → `commit:NNNN` / twin / orphan; prefix index 7..40; asserts 0 ambiguous; records content-hash collisions as excluded; writes the scratch mailmap) · `scrub.py --test \| --sample \| --file` (THE scrub: hash tokens, addresses → noreply, trailer lines in messages; 12 known-true cases; the HEAD sample with git's own object lookup as the independent oracle) · `gate_scan.py --all\|--refs … [--worktree] [--expect-fail FIXTURE]` (paths ever touched × purge rules; every reachable blob's content sha1 × the ROM set; 5 byte signatures; 50 MiB; emits `rom_blob_ids.txt` = hits ∪ every blob ever under a purge path; the fixture `expected_offenders.txt` is the R39 negative control) · `run_filter.py [--sample]` (the git-filter-repo 2.47.0 module-API run inside the scratch bare clone; refuses elsewhere) · `verify_rewrite.py --old --new` (the pairwise proof) · `build_commit_map.py [--out]` (`docs/commit-map.tsv`, asserted free of old hashes) · `resolve_tokens.py [--check] [--map]` (tokens → shortest unique ≥9-char new abbreviations at the tip) · `absent_scan.py [--repo] [--tree]` (nothing old anywhere) · `probe_github.sh [--after-flip]` (Drew's purge probe). Scratch (`.run/public_rewrite/`, never committed): `dict.json`, `mailmap`, `rom_blob_ids.txt`, `old-to-new.tsv`, `repo.git`, the bundle. | +| | `.venv/bin/git-filter-repo` 2.47.0 | (P33 C1) `pip install git-filter-repo==2.47.0` (in `requirements-python.txt`); used through its module API by `run_filter.py`. | | | `tools/audit_public.py [--paths …]` | **(P33 B7) The first-push gate**: no tracked file under `tools/public_rewrite/purge_set.txt` (the C1 rewrite's own input, filter-repo syntax), none whose SHA1 is ROM-derived (DERIVED set: every `sha1` in `extracted/retail/manifest.jsonl` + `config/check.*.sha` + the redump Track-1 SHA1; zero-length files exempt — the empty-file SHA1 is also SC04/SC05 `FILE_029/1.6`'s), none > 50 MiB. Names every offender, exits 1. ≈1 s over 6,798 paths. | | | `tools/compile_only.py \| --all [-j N] [--list]` | **(P33 B7)** cpp → cc1 → maspsx → as on every eligible TU with the Makefile's flags PARSED at run time; TUs per binary from `_SRC_DIR` with nested-binary pruning (= the Makefile's `C_SRCS`); skips main's 70 LINKED tiles (`progress._main_linked_segs_from_makefile`) and the 47 `INCLUDE_ASM(`/`INCLUDE_RODATA(` TUs (they `.include` asm/); -O0 TUs (`corpus.o0_sources`) compile at -O0. Coverage line with every denominator. Measured: PR scope 54 of 124 TUs in 1.6 s; fleet 4,170 of 4,287 in 123 s at -j32 (≈50 CPU-min). | | | `tools/public_rewrite/purge_set.txt` | **(P33 B7)** THE purge set (Drew's decisions 3+11): the EXE at both historical paths, `glob:dumps/*.bin`, `ghidra/`, `tools/psyq/`, `session archive/`, `glob:tools/ghidra-ext/*.zip`, `tools/brave-CUE/brave.exe`. Read by audit_public now and by C1's `git filter-repo --paths-from-file` later. | @@ -1130,6 +1132,42 @@ fills fast). Nothing is leaking — but the host does not get the memory back on `tools/ghidra_*.sh` are repo-relative (`BFM_GHIDRA_PROJ` overrides the project dir; `ghidra_mcp_verify.sh [PROG]`); Makefile `GHIDRA_PROJ := $(or $(BFM_GHIDRA_PROJ),$(CURDIR)/ghidra)`. +### P33 C1 (S87, 2026-09-07) — the history-rewrite tooling, measured before the irreversible run +- **Design points.** The scrub replaces a hex token only when the WHOLE token is a prefix (≥ 7) of an old commit hash — + so a 16-char sig hash can never be mistaken for a commit; only 7–8-char tokens carry any false-positive risk (≈1.6e-5 + per 7-char token) and the sample prints every 7-char replacement in context to be READ (R63). Tokens that are also + prefixes of a cited CONTENT hash (1,480: the manifest, `check.*.sha`, the dumps, the sha256 checksum files, the redump + CRC32) are excluded — measured 0 collisions. `rom_blob_ids.txt` is content/signature hits ∪ every blob that ever sat + under a purge path, so `--strip-blobs-with-ids` kills a renamed copy that a path rule would miss. Bare session UUIDs in + checkpoint prose (68 at HEAD) are NOT scrubbed (out of scope; no `claude.ai` URL exists in any HEAD blob) — reported + as INFO by `absent_scan`. The mailmap and the dictionary are scratch: no personal address and no old hash is a + literal in the package. +- **Measured (S87):** dictionary 4,420 commit objects (4,030 main, 339 twins, 51 orphans), 150,280 prefixes, **0 + ambiguous, 0 content-hash collisions**, 2 personal identities → noreply, 4 s. Sample over HEAD: 397 MB of text in + 7.9 s (50 MB/s), 1,238 replacements in 98 files, 731 distinct tokens (7-char 186 · 8-char 326 · 9-char 721 · 40-char + 5), 6 address replacements; **git's own lookup resolves exactly the same 731 tokens** (only-git 0, only-ours 0). + `gate_scan --all --expect-fail`: 112,390 reachable blobs / 16.79 GB in 2 m 25 s; every purge rule named (ghidra/ 42 + paths ever, tools/psyq/ 190, dumps 28, archive 3, zips 2, brave.exe 1, the EXE 1+1), 0 stray content offenders, 52 + content/signature ids + 269 blobs ever under a purge path. `absent_scan` on the current repo → FAIL with 82,362 + offenders in 7 m 24 s (its positive control). `scrub --test` 12/12. +- **Trial rewrite #1 (S87, on a scratch bare clone — the reason a trial exists):** filter 274 s (72,500 hash replacements + over every historical blob version, 60 trailers, 81 address replacements, 192 binary blobs untouched); the purge is + real (archive/ghidra/dump blobs absent from the store; 0 purge paths reachable) but it found TWO defects that would have + corrupted the real run: (1) the EMPTY blob was in `rom_blob_ids.txt` (an empty file once sat under a purge path) and + `--strip-blobs-with-ids` dropped every "file emptied" change in history — 7 files silently kept their previous content + and a restore commit became empty and was pruned (a second zero row); fixed: a blob shared with a non-purge path is + never stripped by id (content/signature hits always are), and `verify_rewrite` now asserts no purge path survives and + that the pruned set equals the derived purge-only set; (2) a commit the rewrite leaves byte-identical keeps its hash + (the noreply-authored "Initial commit") and tripped the map's old-hash assertion — unchanged commits are recorded and + exempted in the map, `absent_scan` and the probe. Also: the post-rewrite `gate_scan` must accept 0 blobs under purge + paths (the guard now fires only when path offenders exist), and filter-repo's own gc leaves a ≈500 MB pack (re-deltaing + 16 GB of scrubbed text) — the purged binaries are gone regardless; C9 repacks aggressively. +- **Gotchas:** `git log --raw` abbreviates blob ids — `--no-abbrev` or the id filter drops everything (caught by the + count "0 ever under a purge path"; the function refuses that result when path offenders exist, R43). git-filter-repo gc's the OLD objects + out of the clone after the run, so `verify_rewrite` reads the old side from the working repo (`--old`) and the new + side from the clone (`--new`). The clone stops being a "fresh clone" once the tag is deleted → `--force` is passed by + `run_filter.py` (the only deviation from filter-repo's defaults). `--prune-empty auto`, never `always`. + ### P33 B7 (S87, 2026-09-06) — the ROM-free CI: `.github/workflows/no-rom.yml`, `tools/audit_public.py`, `tools/compile_only.py` - **What CI proves and what it cannot.** The contract (218 binaries byte-identical) needs the disc, never in CI. CI proves the tree is public-clean and the C still compiles with the pinned toolchain; the byte proof is the local recorded run diff --git a/docs/public-flip-runbook.md b/docs/public-flip-runbook.md index bf294a1007..48775d0b90 100644 --- a/docs/public-flip-runbook.md +++ b/docs/public-flip-runbook.md @@ -60,27 +60,27 @@ valid for this tip: a `--cached` removal changes no tracked-content byte. `.venv/bin/pip install git-filter-repo==2.47.0` (SETUP row, R21). `tools/public_rewrite/`: -| Tool | Does | Measured expectation (S86 audit — re-measure in C2) | +| Tool | Does | Measured (S87; C2 re-measures on the final tree) | |---|---|---| | `purge_set.txt` | the purge paths (exists since B7) | 8 rules | -| `gate_scan.py --refs [--worktree] [--expect-fail EXPECTED.txt]` | the first-push gate over history: scans every blob reachable from the refs (and the worktree) for purge-path prefixes, content SHA1s in the known-ROM set (the EXE, the redump Track 1, every `sha1` in `extracted/retail/manifest.jsonl`, every `config/check.*.sha`), byte signatures (`PS-X EXE` at offset 0; the 2,097,152-byte RAM image with the resident's first words at 0xCEDF8; the EXE entry code; PsyQ `LIB\x01` / `LNK\x02` magics), any blob > 50 MiB; emits `rom_blob_ids.txt`. `--expect-fail` = the R39 negative control: exit 0 only when the scan fails naming exactly the expected offender set | on the current repo: the EXE at both paths, 28 dumps, all `ghidra/**` (41 unique blobs across history), all `tools/psyq/**`, the 3 archive parts, 2 zips, `brave.exe`; signature hits = the EXE blob(s) + 28 RAM images | -| `hash_dict.py` | every commit hash → ordinal (`git rev-list --reverse main`); off-main commits (the tag lineage, 126) → the ordinal of their (tree, author-timestamp, subject) twin, else `orphan-NNN`; all prefixes 7..40; asserts 0 ambiguous prefixes and 0 collisions with the content-hash set | 4,401 commit objects at S86 (now more: P33 commits), 149,634 prefixes, 126 twins | -| `scrub.py` | THE one scrub function: `\b[0-9a-f]{7,40}\b` → dictionary lookup → `commit:NNNN` (non-hex in the first 7 chars, so it can never re-match; no Markdown side effects); NUL-sniff binary skip; idempotent | at HEAD: 711 resolving citations, 0 word-embedded false hits (`func_800D128C` / `0x800d128c` do not match); 125 MB/s | -| `run_filter.py [--sample]` | composes the `git filter-repo` call (below); refuses to run outside a bare repo under `.run/public_rewrite/`; logs versions + wall time; `--sample` first (R37): must reproduce ≈1,202 hits / 93 files / ≈3 s over HEAD's 395 MB and the replaced-token set must equal the citations `git cat-file -e` resolves | — | -| `build_commit_map.py` | public `docs/commit-map.tsv` (`ordinal new_hash author_date committer_date subject`, no old hash anywhere — asserted by running scrub over its own output); private `.run/public_rewrite/old-to-new.tsv` for the probe | 4,0xx rows, exactly one mapped to zeros (the pruned archive-upload commit) | -| `resolve_tokens.py [--check]` | at HEAD of the adopted checkout: `commit:NNNN` → the unique 9-char new abbreviation (asserted by `git cat-file --batch-check`); `--check` asserts zero resolvable tokens remain and lists the orphan residue | ≥1 orphan: `tools/verify_worktree.py` cites a dropped TEMP commit | -| `verify_rewrite.py` | the pairwise proof (C5) | every pair | -| `absent_scan.py [--tree HEAD]` | every blob incl. binaries + every message → 0 dictionary prefixes, 0 `Claude-Session:`, 0 personal-address strings | 0 / 0 / 0 | +| `gate_scan.py --all \| --refs [--worktree] [--expect-fail tools/public_rewrite/expected_offenders.txt]` | the first-push gate over history: scans every blob reachable from the refs (and the worktree) for purge-path prefixes, content SHA1s in the known-ROM set (the EXE, the redump Track 1, every `sha1` in `extracted/retail/manifest.jsonl`, every `config/check.*.sha`), byte signatures (`PS-X EXE` at offset 0; the 2,097,152-byte RAM image with the resident's first words at 0xCEDF8; the EXE entry code; PsyQ `LIB\x01` / `LNK\x02` magics), any blob > 50 MiB; emits `rom_blob_ids.txt`. `--expect-fail FIXTURE` = the R39 negative control: the fixture lists `rulemin offending paths`; exit 0 only when every rule has at least that many AND no content/signature/size offender sits outside the purge rules; `rom_blob_ids.txt` = content hits ∪ every blob ever under a purge path | measured S87 on the current repo: 112,390 blobs / 16.8 GB in 2 m 25 s; paths ever: ghidra/ 42, tools/psyq/ 190, dumps 28, archive 3, zips 2, brave.exe 1, the EXE 1+1; 0 strays; 52 content/signature ids + 269 blobs ever under a purge path | +| `hash_dict.py [--write-mailmap]` | every commit hash → ordinal (`git rev-list --reverse main`); off-main commits (the tag lineage, 126) → the ordinal of their (tree, author-timestamp, subject) twin, else `orphan-NNN`; all prefixes 7..40; asserts 0 ambiguous prefixes and 0 collisions with the content-hash set | measured S87: 4,420 commit objects (4,030 main, 339 twins incl. the 126 tag-lineage re-authorings, 51 orphans), 150,280 prefixes, 0 ambiguous, 0 content collisions | +| `scrub.py` | THE one scrub function: `\b[0-9a-f]{7,40}\b` → dictionary lookup → `commit:NNNN` (non-hex in the first 7 chars, so it can never re-match; no Markdown side effects); NUL-sniff binary skip; idempotent | at HEAD (S87): 731 distinct resolving tokens, 1,238 replacements in 98 files, git's own lookup agrees exactly; 397 MB in 7.9 s | +| `run_filter.py [--sample]` | composes the `git filter-repo` call (below); refuses to run outside a bare repo under `.run/public_rewrite/`; logs versions + wall time; `--sample` first (R37) = `scrub.py --sample`, the independent-oracle check | the trial run's numbers are in phase-ends/CURRENT_PHASE.md (S87 C1) | +| `build_commit_map.py [--out PATH]` | public `docs/commit-map.tsv` (`ordinal new_hash author_date committer_date subject`, no old hash anywhere — asserted by running scrub over its own output); private `.run/public_rewrite/old-to-new.tsv` for the probe | 4,0xx rows, exactly one mapped to zeros (the pruned archive-upload commit) | +| `resolve_tokens.py [--check] [--map PATH]` | at HEAD of the adopted checkout: `commit:NNNN` → the unique 9-char new abbreviation (asserted by `git cat-file --batch-check`); `--check` asserts zero resolvable tokens remain and lists the orphan residue | ≥1 orphan: `tools/verify_worktree.py` cites a dropped TEMP commit | +| `verify_rewrite.py --old ~/bfm-decomp --new .run/public_rewrite/repo.git` | the pairwise proof (C5): old commits from the ORIGINAL repo (the clone gc's them away), new from the clone | every pair | +| `absent_scan.py [--repo PATH] [--tree HEAD]` | every text blob + every message + every ref → 0 old-hash prefixes, 0 personal addresses, 0 session URLs, 0 trailer lines, every identity = noreply, no replace/original/tag refs (≈7 min over all objects; INFO: bare UUID count) | 0 offenders (the current repo: FAIL, 82,362 — its positive control) | | `probe_github.sh` | Drew's post-purge probe (C10) | — | -| `mailmap` (scratch, `.run/public_rewrite/mailmap`, generated from `git log --format='%an <%ae>' \| sort -u`) | the two personal identities → the noreply identity | never committed | +| `mailmap` (scratch, `.run/public_rewrite/mailmap`, written by `hash_dict.py --write-mailmap` from the log's identities) | the two personal identities → the noreply identity | never committed | Budget: regex+lookup ≈2.2 CPU-min over 16.8 GB of blobs; the filter-repo stream dominates (10–30 min). Disk: a bare `--no-local` clone ≈0.6 GB + the bundle ≈0.6 GB; no working-tree copy (the WSL disk is capped at 75 GB, ≈13 GB free). ## 4. C2 — negative control, dictionary, sample (Claude) -1. `gate_scan.py --refs --all --worktree --expect-fail .run/public_rewrite/expected_offenders.txt` on the CURRENT repo → - must FAIL naming exactly the S86 set above (the scan that PASSES in C5 is this same tool). +1. `gate_scan.py --all --worktree --expect-fail tools/public_rewrite/expected_offenders.txt` on the CURRENT repo → must + report PASS (= the scan fails exactly as the fixture says; the scan that PASSES with 0 offenders in C5 is this same tool). 2. `hash_dict.py` → prints the counts; assert 0 ambiguous, 0 collisions. 3. `run_filter.py --sample` → the sample numbers above. @@ -116,8 +116,18 @@ Why each flag: `--prune-empty auto`, never `always` (main carries one pre-existi survive); `--replace-refs delete-no-add` (no `refs/replace/` names may be minted — they would leak old hashes); `--strip-blobs-with-ids` catches the EXE wherever it was renamed; the message callback also strips the 60 remaining `Claude-Session:` trailer lines the S76 scrub missed. -Checks: exit 0; the commit map has (old main count) rows with exactly one mapped to zeros; no `refs/replace`; one pack; -pack size recorded (expect 150–250 MB). +Checks: exit 0; the commit map has (old main count) rows; the rows mapped to zeros are EXACTLY the commits whose every +change was a purge path (`verify_rewrite` derives that set — 1 on this history: "session archive update"); no +`refs/replace`; one pack. Pack size: filter-repo's own gc leaves ≈500 MB (trial #1: 534 → 520 MB — the 16 GB of scrubbed +text history re-deltas poorly; the purged binaries ARE gone: the archive/ghidra/dump blobs are absent from the store); +C9's local gc uses an aggressive repack — measured on trial #2: `git -c pack.threads=16 repack -adf --window=250 +--depth=50` took the 500 MB pack to **80 MB in 166 s**. +**Lesson from trial #1 (S87):** stripping blobs BY ID must never include a blob that also lives under a non-purge path — +the EMPTY blob (an empty file once sat under `ghidra/`) was in the list, and `--strip-blobs-with-ids` then dropped every +"file emptied" change in history: those files silently kept their previous content and a later restore commit became +empty and was pruned. `gate_scan` now excludes shared blobs from `rom_blob_ids.txt` (content/signature hits are always +kept), and `verify_rewrite` asserts both that no purge path survives in any new tree and that the pruned set equals the +derived purge-only set. ## 6. C5 — verification on the rewritten clone (Claude; every check an exit code) @@ -125,8 +135,12 @@ pack size recorded (expect 150–250 MB). scrub(old.message)`; new parents = map(old parents) with the pruned commit spliced out; `git diff-tree -r --no-renames old new`: every `D` is a purge path or a ROM blob id, every `M` satisfies `hash-object(scrub(old_blob)) == new_blob`, any `A` FAILS; prints the pair count. Then `gate_scan.py --refs --all` → PASS (the same tool that failed in C2); `absent_scan.py` -→ 0/0/0; `git rev-list --count main` = old count − 1 (the pruned commit); the `%at %ct` lists match with the pruned commit -removed; the pre-existing empty commit's twin exists. +→ 0/0/0; `git rev-list --count main` = old count − pruned; the `%at %ct` lists match with the pruned commits removed; the +pre-existing empty commit's twin exists; no purge path in any new tree; the pruned set == the purge-only commits. +A commit the rewrite leaves BYTE-IDENTICAL keeps its hash (old == new — the noreply-authored "Initial commit", which cites +no hash and touches no purge path): `build_commit_map` records those in `.run/public_rewrite/unchanged_commits.txt`, +`absent_scan` does not count their prefixes as old hashes, and `probe_github.sh` skips them (they legitimately still +resolve on GitHub). Measured trial #1: filter 274 s, verify 385 s, absent_scan 346 s over 16.2 GB of text. ## 7. C6 — adoption in `~/bfm-decomp` (Claude; NO gc yet) @@ -163,9 +177,10 @@ git push --force origin main git push origin :refs/tags/S76-pre-scrub-backup # if it was git fetch --prune origin && git rev-parse origin/main main # equal ``` -**Claude, after Drew's word:** `git reflog expire --expire=now --all && git gc --prune=now`; checks: `git rev-list --all ---count` = the number the map predicts (old main − 1 pruned + the tip commits); `gate_scan.py --refs --all` PASS; -`absent_scan.py` PASS; `.git` size recorded (expect < 400 MB, from 929 MB). +**Claude, after Drew's word:** `git reflog expire --expire=now --all && git -c pack.threads=16 repack -adf --window=250 +--depth=50 && git prune --expire=now`; checks: `git rev-list --all --count` = old main − pruned + the tip commits (the map +predicts it), `gate_scan.py --all` PASS, `absent_scan.py --repo ~/bfm-decomp` PASS, `.git` size recorded (trial #2: the +rewritten pack is 80 MB after the aggressive repack, from 929 MB). ## 11. C10 — the Support purge, the probe, the flip (Drew; decision Max) diff --git a/phase-ends/CURRENT_PHASE.md b/phase-ends/CURRENT_PHASE.md index ff47775195..253bfb58f7 100644 --- a/phase-ends/CURRENT_PHASE.md +++ b/phase-ends/CURRENT_PHASE.md @@ -62,7 +62,7 @@ one-time snapshot, `CLAUDE.md` gains "never `git clean -x`" (R20 amendment propo - [x] **A5** THE RECORDED RUN (`tools/verify_contract.sh` → `.run/P33/verify/`, SUMMARY all EXIT=0) — run Low, read Max — see Log 2026-09-06 A5 - [x] **B9/C3** The preparatory commit (`git rm --cached` purge set; psyq CHECKSUMS moved; zip sha256s; runbook; decision-log entry) — Max (P5c-class) — see Log 2026-09-06 B9/C3 -- [ ] **C1** filter-repo 2.47.0 + `tools/public_rewrite/` — design Max, execution xHigh +- [x] **C1** filter-repo 2.47.0 + `tools/public_rewrite/` — design Max, execution xHigh — see Log 2026-09-07 C1 - [ ] **C2** Negative control + dictionary + sample — xHigh - [ ] **C4** Bundle + Drew's archive mirror push + bare clone + the rewrite — xHigh - [ ] **C5** Verification suite on the rewritten clone — xHigh @@ -248,21 +248,47 @@ Mid-phase rules check after every 4 completed tasks (P6). Commit banked artifact before this commit, `--all` 4,300; refs = main, origin/main, origin/HEAD, 1 stash, the tag (NO `refs/original/`, 1 stash — not the S86 block's 3 + refs/original; re-measure in C2); `origin/main` is still the P32 close commit — **Drew has not pushed the P33 commits yet (R6)**; `.git` 929 MB. Commit: see below. +- **2026-09-07 (S87, Max) — C1 the rewrite tooling, PROVEN by two trial rewrites on a scratch bare clone.** + `git-filter-repo==2.47.0` in the venv (+ `requirements-python.txt`); `tools/public_rewrite/`: `common.py`, `hash_dict.py` + (4,420 commit objects: 4,030 main / 339 twins / 51 orphans; 150,280 prefixes; **0 ambiguous, 0 collisions** with 1,480 cited + content hashes; the scratch mailmap: 2 personal identities → noreply), `scrub.py` (`--test` 12/12; `--sample` over HEAD: + 397 MB in 7.9 s, 1,238 replacements in 98 files, 731 distinct tokens = EXACTLY the set git's own object lookup resolves; + the 186 seven-char replacements read — all commit ranges in PhaseEnds/cookbook), `gate_scan.py` (+ the tracked fixture + `expected_offenders.txt`; negative control PASS: 16.8 GB / 112,390 blobs in 2 m 25 s, every rule named, 0 strays), + `run_filter.py`, `verify_rewrite.py`, `build_commit_map.py`, `resolve_tokens.py`, `absent_scan.py` (positive control on the + current repo: FAIL with 82,362 offenders, 7 m 24 s), `probe_github.sh`. **Trial #1** (filter 274 s) exposed two defects + that would have corrupted the real run — (1) the EMPTY blob was in the strip list (an empty file once sat under a purge + path) and `--strip-blobs-with-ids` dropped every "file emptied" change in history: 7 files kept their previous content and + a restore commit was pruned as a second empty; (2) the byte-identical "Initial commit" (noreply-authored, no citations, + no purge paths) keeps its hash and tripped the map's old-hash assertion — plus the post-rewrite gate's empty-list guard. + Fixes: shared blobs are never stripped by id (content/signature hits always are); `verify_rewrite` asserts no purge path + survives and that the pruned set == the DERIVED purge-only set; unchanged commits recorded (`unchanged_commits.txt`) and + exempted in the map, absent_scan and the probe; the guard fires only with path offenders. **Trial #2: everything PASS** — + gate fixture PASS (268 ids, the empty blob excluded); filter 269 s, **1 pruned** (exactly "session archive update"), main + 4,030 → 4,029; `verify_rewrite` 4,029 pairs / 0 failures (241,534 changed blobs re-derived, 989,166 removals justified, the + empty commit survived) in 430 s; map 4,030 rows / 1 pruned / 0 old hashes / 1 unchanged; absent_scan on the clone PASS + (0/0/0/0/0/0; INFO 370 bare UUIDs) 409 s; gate on the clone PASS (0 offenders, 60 s); resolve_tokens on a trial checkout: + 1,231 tokens in 97 files, residue 7 (`commit:1712` ×3 = the pruned commit's ordinal, orphan-24 ×2, orphan-26, orphan-35 — + cited commits that exist in no lineage), `--check` 0 remaining; absent_scan --tree HEAD PASS. **Pack size:** filter-repo's + gc leaves 500 MB; `repack -adf --window=250 --depth=50` → **80 MB in 166 s** (C9). SETUP P33 C1 section + rows; runbook + §3/§5/§6/§10 updated with the measured facts and the strip-by-id lesson. Trial scratch deleted (C4 re-clones fresh). + For Drew's decision (not in the plan's scope, reported by absent_scan as INFO): 370 bare session UUIDs across history / 68 + at HEAD in checkpoint prose; no `claude.ai` URL anywhere. Commit: see below. -## 🛑 SESSION CHECKPOINT — A1–A5 ✓, B1–B9/C3 ✓; NEXT = C1 the rewrite tooling (2026-09-07 ~00:10 MDT, written by session fa49faf3 "S87" after the C3 commit; SUPERSEDES the earlier blocks) +## 🛑 SESSION CHECKPOINT — A1–A5 ✓, B1–B9/C3 ✓, C1 ✓; NEXT = C2 (2026-09-07 ~04:45 UTC, written by session fa49faf3 "S87" after the C1 commit; SUPERSEDES the earlier blocks) ### 0. How to use this block You are a FRESH SESSION that has read `PROJECT_CONTEXT.md`, `phase-ends/DIGEST.md`, `PhaseEnd_Phase30/31/32.md` and this file, and nothing else (R64). Replay this block verbatim, state phase / done / NEXT / effort, list the rules from the digest -(R1–R73), then WAIT for Drew. **NEXT = C1** (design Max, execution xHigh — Drew set Max for B9/C3; restate per R7/R27). **The purge set has left the index: NEVER `git clean -x` (CLAUDE.md fail-safe).** The harness task list must be REBUILT (Drew wants to monitor it — -one TaskCreate per plan item A1…G2, 40 items, mark A1–A5 + B1–B9/C3 completed; R28). The SessionStart hook restarts the headless +(R1–R73), then WAIT for Drew. **NEXT = C2** (xHigh: the formal negative control + dictionary + sample on the FINAL pre-rewrite tree — the three commands below, ≈5 min). **The purge set has left the index: NEVER `git clean -x` (CLAUDE.md fail-safe).** The harness task list must be REBUILT (Drew wants to monitor it — +one TaskCreate per plan item A1…G2, 40 items, mark A1–A5 + B1–B9/C3 + C1 completed; R28). The SessionStart hook restarts the headless MCP server when `ghidra/bfm.rep` exists (it did not stay up in S87 — `ss -tln` showed nothing on :8080; harmless): B6/B7/B8 need no Ghidra; run `tools/ghidra_mcp_stop.sh` before any headless step (R23). ### 1. Where we are **Phase 33 — 100% verification + the public flip + Gen2 exit.** Gate 1 approved 2026-09-06 (plan mode, Max). R65–R73 ratified. **Done: A1 (`commit:4012`), A2 (`commit:4013`), A3 (`commit:4014`), A4 (`commit:4015`), B1 (`commit:4016`), B2 (`commit:4017`), B3 -(`commit:4018`), B4 (`commit:4019`), B5 (`commit:4022`), B6 (`commit:4023`), B7 (`commit:4024`), B8 (`commit:4025`), A5 (`commit:4026` prep + `commit:4027`/`commit:4028` fixes + `commit:4029`), B9/C3 (the S87 `chore(phase-33): C3 …` commit — 251 index entries removed).** The approved plan is +(`commit:4018`), B4 (`commit:4019`), B5 (`commit:4022`), B6 (`commit:4023`), B7 (`commit:4024`), B8 (`commit:4025`), A5 (`commit:4026` prep + `commit:4027`/`commit:4028` fixes + `commit:4029`), B9/C3 (`commit:4030` — 251 index entries removed), C1 (the S87 `feat(phase-33): C1 …` commit).** The approved plan is VERBATIM at the end of this file — Blocks A–G give every task's files, commands and verification; "Execution order and why" is the sequence. Effort: Drew ran S87 at **medium** by explicit choice (the plan says Max for B5); the plan's annotations still stand for the tasks ahead — restate them, Drew decides (R7/R27). @@ -282,26 +308,37 @@ still stand for the tasks ahead — restate them, Drew decides (R7/R27). - **Verification state:** **A5 PASS on `commit:4028`** (`.run/P33/verify/SUMMARY.md`: 218/218 clean fleet, sdk-dual both legs, tools-health OK with 0 PHANTOM/TRUNCATED/PAD-TAIL and zero warns, UNCLAIMED 0, report 100.00/100.0/100.0, stubs 0, backlog 0). C8 re-runs it on the adopted (rewritten) tree; a `--cached` removal (B9/C3) changes no tracked-content bytes. -- **Environment:** the purge set is IGNORED-BUT-PRESENT on disk since C3 — never `git clean -x`; `gh` not authenticated; `git filter-repo` NOT installed (C1); disk ≈ 13 GB free of 75 (`.run/` 38 GB — +- **Environment:** the purge set is IGNORED-BUT-PRESENT on disk since C3 — never `git clean -x`; `gh` not authenticated; `git-filter-repo` 2.47.0 in the venv; disk ≈ 13 GB free of 75 (`.run/` 38 GB — `build/ghidra_rebuild/` and `.run/ghidra_rebuild/` are scratch, ~1 GB, safe to delete); no `/tmp` (R12). `.run/ghidra_export/` (139 MB, 129 live exports from S86) is the input for any future re-derivation of a candidate. ### 3. NEXT — in order 0. **Preflight:** `git status --short | grep -v ghidra/` (empty) · `git log -1 --format='%h %s'` · `df -h ~`. -1. **C1 — the rewrite tooling** (design Max, exec xHigh): `.venv/bin/pip install git-filter-repo==2.47.0` (SETUP row R21) - + `tools/public_rewrite/` per `docs/public-flip-runbook.md` §3 (the table IS the spec, with every measured expectation): - `gate_scan.py` (+ `--expect-fail`, the R39 control), `hash_dict.py`, `scrub.py`, `run_filter.py` (+ `--sample`), - `build_commit_map.py`, `resolve_tokens.py`, `verify_rewrite.py`, `absent_scan.py`, `probe_github.sh`; the mailmap is - generated into `.run/public_rewrite/mailmap` (scratch — the addresses never enter a tracked file). Each tool: a - known-true control before it is trusted (R39/R14). `purge_set.txt` exists (B7). Log, tick, refresh, commit. -2. **C2** (negative control, dictionary, sample) → **C4** (Claude: bundle; **Drew**: create `Druthulu/BFM-decomp-archive` - EMPTY + private, `git remote add archive …`, `git push --mirror archive`; Claude checks `for-each-ref` == `ls-remote`; - then the bare clone + the filter) → C5 → C6 → C7 (noreply identity first) → C8 (`tools/verify_contract.sh` again) → C9 - (Drew force-pushes; then gc) → the Support wait (D/E/F land) → C10 flip → C11. - Reminder for Drew at the next hand-off: `origin/main` is still the P32 close — push the P33 commits (R6); the new - `no-rom` workflow's `audits` job goes GREEN from this commit on (it was RED by design before C3). +1. **C2 — negative control + dictionary + sample** (xHigh; every tool already proven on two trial rewrites — this is the + RECORDED run on the final pre-rewrite tree, so it must be re-run after ANY further commit before C4): (a) `.venv/bin/python + tools/public_rewrite/hash_dict.py --write-mailmap` (expect 0 ambiguous, 0 collisions; the count grows by the C1/C2 commits); + (b) `.venv/bin/python tools/public_rewrite/scrub.py --test && … scrub.py --sample` (only-git 0 / only-ours 0); + (c) `.venv/bin/python tools/public_rewrite/gate_scan.py --all --worktree --expect-fail tools/public_rewrite/expected_offenders.txt` + (PASS = fails exactly as the fixture says; `--worktree` = audit_public OK). Log the three verdicts, tick, refresh, commit. + NOTE: C2's dictionary is built from THIS repo and must be built AFTER the last commit that precedes C4 (the C2 commit + itself is fine to leave out of the dictionary? NO — rebuild the dictionary once more right before C4c, after C2's commit, + or make the C2 log entry the last commit before the clone: simplest = C2 commit, then C4a bundle, then hash_dict again + (4 s) before the clone). +2. **C4** — (a) Claude: `git bundle create .run/public_rewrite/pre-rewrite.bundle --all --reflog && git bundle verify …`; + (b) **Drew**: create `Druthulu/BFM-decomp-archive` EMPTY + private, `git remote add archive https://github.com/Druthulu/ + BFM-decomp-archive.git && git push --mirror archive`; Claude checks `git for-each-ref` == `git ls-remote archive`; (c) Claude: + `rm -rf .run/public_rewrite/repo.git && git clone --no-local --bare ~/bfm-decomp .run/public_rewrite/repo.git && git -C + .run/public_rewrite/repo.git tag -d S76-pre-scrub-backup && git rev-parse S76-pre-scrub-backup > .run/public_rewrite/ + old_tag_tip.txt`, then `hash_dict.py --write-mailmap` + `gate_scan.py --all --expect-fail …` (fresh dict + ids for the + final tree), then `run_filter.py` (≈5 min; expect 1 pruned). The trial chains are the template: `.run/public_rewrite/ + trial2.sh` (scratch, not tracked) ran gate → clone → filter → verify → map → absent → gate → resolve → absent-tree → repack. +3. **C5** `verify_rewrite.py` (≈7 min, 0 failures) + `absent_scan.py` (≈7 min) + `gate_scan.py --all --repo .run/public_rewrite/ + repo.git` (≈1 min) → **C6** adopt → **C7** `build_commit_map.py` (writes docs/commit-map.tsv) + `resolve_tokens.py` (noreply + identity first: `git config user.email 50529377+Druthulu@users.noreply.github.com`) + `absent_scan.py --tree HEAD` + the tip + commit → **C8** `tools/verify_contract.sh` → **C9** gate + Drew force-push + aggressive repack (80 MB) → D/E/F during the + Support wait → **C10** flip → **C11**. ### 4. Files S87 touched -B9/C3: 251 index removals (files on disk), `docs/public-flip-runbook.md` (new), `docs/SETUP.md` §2.3/§2.4, `docs/verification.md`, `CLAUDE.md`, `docs/decision-log.md`. A5: `tools/verify_contract.sh` (new), `tools/audit_frontier.py` (derived denominator lines), `.gitignore` (the `.run/P33/verify/` allowlist), `.run/P33/verify/*` (tracked evidence), `docs/verification.md` §2, regenerated `docs/family-hseq.md` + `docs/progress*.md`/`duplicates*.md`. B8: `docs/verification.md` (new), `docs/SETUP.md` (§4.4/§4.6/§4.8/§6.3 pointers, backup posture). B7: `.github/workflows/no-rom.yml`, `tools/audit_public.py`, `tools/compile_only.py`, `tools/public_rewrite/purge_set.txt` (all new), `docs/SETUP.md`. B6: `dumps/CHECKSUMS.sha1` (new), `dumps/INDEX.md`, `docs/memory-map.md`. B5: `tools/ghidra_scripts/ImportAnnotations.java` (3 compile fixes + the `/undefined` resolver), `tools/ghidra_rebuild.sh` +C1: `tools/public_rewrite/{common,hash_dict,scrub,gate_scan,run_filter,verify_rewrite,build_commit_map,resolve_tokens,absent_scan}.py`, `probe_github.sh`, `expected_offenders.txt` (all new), `requirements-python.txt`, `docs/SETUP.md`, `docs/public-flip-runbook.md`. Scratch `.run/public_rewrite/`: `dict.json`, `mailmap`, `rom_blob_ids.txt`, `old_tag_tip.txt`, `trial*.{sh,log}`, `trial_commit-map.tsv`, `trial_old-to-new.tsv`, `unchanged_commits.txt`, `nonpurge_blob_ids.txt`, `filter.log`. B9/C3: 251 index removals (files on disk), `docs/public-flip-runbook.md` (new), `docs/SETUP.md` §2.3/§2.4, `docs/verification.md`, `CLAUDE.md`, `docs/decision-log.md`. A5: `tools/verify_contract.sh` (new), `tools/audit_frontier.py` (derived denominator lines), `.gitignore` (the `.run/P33/verify/` allowlist), `.run/P33/verify/*` (tracked evidence), `docs/verification.md` §2, regenerated `docs/family-hseq.md` + `docs/progress*.md`/`duplicates*.md`. B8: `docs/verification.md` (new), `docs/SETUP.md` (§4.4/§4.6/§4.8/§6.3 pointers, backup posture). B7: `.github/workflows/no-rom.yml`, `tools/audit_public.py`, `tools/compile_only.py`, `tools/public_rewrite/purge_set.txt` (all new), `docs/SETUP.md`. B6: `dumps/CHECKSUMS.sha1` (new), `dumps/INDEX.md`, `docs/memory-map.md`. B5: `tools/ghidra_scripts/ImportAnnotations.java` (3 compile fixes + the `/undefined` resolver), `tools/ghidra_rebuild.sh` (`.proof` markers; dies unless `failed=0`), `tools/ghidra_annotations_delta.py` (the three drift classes), new `tools/ghidra_roster.py`, `tools/ghidra_mcp_start.sh` (silent no-op guard), `.claude/settings.json` (relative hooks), `Makefile` (roster check in tools-health), `docs/SETUP.md` (P33 B5 section, 5 inventory rows, §2.8), `config/ghidra/*` diff --git a/requirements-python.txt b/requirements-python.txt index f3a580ea8e..5c83cdf842 100644 --- a/requirements-python.txt +++ b/requirements-python.txt @@ -41,3 +41,4 @@ pycparser==2.23 toml==0.10.2 PyNaCl==1.6.2 cffi==2.0.0 +git-filter-repo==2.47.0 # P33 C1: the public-flip history rewrite (tools/public_rewrite/); pure Python diff --git a/tools/public_rewrite/absent_scan.py b/tools/public_rewrite/absent_scan.py new file mode 100644 index 0000000000..8906ad2d2e --- /dev/null +++ b/tools/public_rewrite/absent_scan.py @@ -0,0 +1,112 @@ +#!/usr/bin/env python3 +"""absent_scan.py — nothing that had to go is still anywhere: every object + every message + every ref (P33 C5/C9). + + tools/public_rewrite/absent_scan.py [--repo PATH] # every object in the store (blobs, commits), every ref + tools/public_rewrite/absent_scan.py --tree HEAD # only the blobs of one tree (the tip, after C7) + +Offenders (exit 1 on any): a hex token 7..40 that is a prefix of an OLD commit hash (the dictionary; the excluded +content-hash prefixes are not offenders), a personal e-mail address (from the scratch mailmap), a `claude.ai/code/session` +URL, a `Claude-Session:` trailer line in a commit message, an author/committer e-mail that is not the noreply one, a +`refs/replace/*` or `refs/original/*` ref or any tag. INFO only: the count of bare UUIDs (session ids in checkpoint +prose — not in the plan's scope, reported so the owner can decide). On the CURRENT repo this scan must FAIL (its +positive control); on the rewritten clone and the adopted checkout it must pass with 0 / 0 / 0. +""" +import argparse +import re +import sys + +sys.path.insert(0, str(__import__("pathlib").Path(__file__).resolve().parent)) +import common as C # noqa: E402 + +UUID_RE = re.compile(rb"\b[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}\b") + + +def main(argv): + ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) + ap.add_argument("--repo", default=str(C.CLONE_DIR)) + ap.add_argument("--tree") + a = ap.parse_args(argv) + repo = __import__("pathlib").Path(a.repo) + d = C.load_dict() + index = set() + for h in d["commits"]: + index.update(C.prefixes_of(h)) + excluded = set(d["excluded"]) + unch = C.SCRATCH / "unchanged_commits.txt" # written by build_commit_map: commits the rewrite left identical + if unch.exists(): + for h in unch.read_text().split(): + excluded.update(C.prefixes_of(h)) + emails = [e.encode() for e in C.personal_emails_from_mailmap()] + off = {"old_hash": 0, "email": 0, "session_url": 0, "trailer": 0, "identity": 0, "ref": 0} + examples, uuids, n_blobs, n_text, n_commits, nbytes = [], 0, 0, 0, 0, 0 + cf = C.CatFile(repo) + + def scan_text(data, where): + nonlocal uuids + for m in C.HEX_RE.finditer(data): + t = m.group(0).decode() + if t in index and t not in excluded: + off["old_hash"] += 1 + if len(examples) < 30: + a0, b0 = max(0, m.start() - 30), min(len(data), m.end() + 30) + examples.append(f"{where}: old hash {t}: …{data[a0:b0].decode('utf-8', 'replace')!r}…") + for e in emails: + if e in data: + off["email"] += data.count(e) + if len(examples) < 30: + examples.append(f"{where}: personal address") + if C.SESSION_URL_RE.search(data): + off["session_url"] += len(C.SESSION_URL_RE.findall(data)) + if len(examples) < 30: + examples.append(f"{where}: claude.ai session URL") + uuids += len(UUID_RE.findall(data)) + + if a.tree: + ids = [(ln.split()[2], ln.split("\t")[1]) for ln in C.git(["ls-tree", "-r", a.tree], repo).splitlines() if " blob " in ln] + for oid, path in ids: + _, _, data = cf.get(oid) + n_blobs += 1 + if data is None or C.is_binary(data): + continue + n_text += 1; nbytes += len(data) + scan_text(data, path) + else: + for oid, typ, size in C.iter_all_objects(repo, types=("blob", "commit")): + _, _, data = cf.get(oid) + if data is None: + continue + if typ == "blob": + n_blobs += 1 + if C.is_binary(data): + continue + n_text += 1; nbytes += len(data) + scan_text(data, f"blob {oid[:12]}") + else: + n_commits += 1 + c = C.parse_commit(data) + if C.TRAILER_RE.search(c["message"]): + off["trailer"] += 1 + examples.append(f"commit {oid[:12]}: Claude-Session trailer") + scan_text(c["message"], f"commit {oid[:12]} message") + for who in ("author", "committer"): + if c[who][1] != C.NOREPLY_EMAIL: + off["identity"] += 1 + if len(examples) < 30: + examples.append(f"commit {oid[:12]}: {who} e-mail is not the noreply identity") + refs = C.git(["for-each-ref", "--format=%(refname)"], repo).split() + for r in refs: + if r.startswith(("refs/replace/", "refs/original/", "refs/tags/")): + off["ref"] += 1 + examples.append(f"ref {r}") + cf.close() + total = sum(off.values()) + print(f"absent_scan {repo}{' tree ' + a.tree if a.tree else ''}: {n_blobs} blobs ({n_text} text, {nbytes / 1e6:.0f} MB), " + f"{n_commits} commits scanned; offenders: {off}; INFO bare UUIDs: {uuids}") + for e in examples[:30]: + print(" " + e[:200]) + print("absent_scan: " + ("PASS — 0 offenders" if not total else f"FAIL — {total} offenders")) + return 1 if total else 0 + + +if __name__ == "__main__": + sys.exit(main(sys.argv[1:])) diff --git a/tools/public_rewrite/build_commit_map.py b/tools/public_rewrite/build_commit_map.py new file mode 100644 index 0000000000..557ae63af5 --- /dev/null +++ b/tools/public_rewrite/build_commit_map.py @@ -0,0 +1,76 @@ +#!/usr/bin/env python3 +"""build_commit_map.py — the public ordinal → new-hash map, and the private old → new map (P33 C7). + + tools/public_rewrite/build_commit_map.py [--new .run/public_rewrite/repo.git] + +Writes docs/commit-map.tsv (PUBLIC): one row per commit of the OLD main in order — `ordinal new_hash author_date +committer_date subject` — with NO old hash anywhere (asserted: the scrub over the file's own bytes makes 0 +replacements). The pruned purge-only commit keeps its row with `new_hash` = 40 zeros so the ordinals stay dense and +every `commit:NNNN` token in history has a row. Dates/subjects are read from the NEW commits (already scrubbed by the +rewrite); the pruned row's from the old commit, scrubbed here. +Writes .run/public_rewrite/old-to-new.tsv (PRIVATE, for probe_github.sh): `ordinal old_hash new_hash`. +""" +import argparse +import sys + +sys.path.insert(0, str(__import__("pathlib").Path(__file__).resolve().parent)) +import common as C # noqa: E402 +import scrub # noqa: E402 + + +def main(argv): + ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) + ap.add_argument("--new", default=str(C.CLONE_DIR)) + ap.add_argument("--old", default=str(C.REPO)) + ap.add_argument("--out", default=str(C.COMMIT_MAP_PUBLIC), help="the public map path (a scratch path for a trial run)") + ap.add_argument("--private-out", default=str(C.OLD_TO_NEW_FILE)) + a = ap.parse_args(argv) + out_pub, out_priv = __import__("pathlib").Path(a.out), __import__("pathlib").Path(a.private_out) + new_repo = __import__("pathlib").Path(a.new) + import verify_rewrite + cmap = verify_rewrite.commit_map(new_repo) + d = C.load_dict() + old_main = [h for h, e in sorted(d["commits"].items(), key=lambda kv: kv[1].get("ord", 0)) if e["kind"] == "main"] + if len(old_main) != d["main_count"]: + C.die(f"dictionary main count mismatch ({len(old_main)} vs {d['main_count']})") + s = scrub.Scrubber() + rows, priv = [], [] + for h in old_main: + o = d["commits"][h]["ord"] + new = cmap.get(h) + if new is None: + C.die(f"old main commit ordinal {o} ({h[:9]}) is not in the commit-map — the clone did not hold the same main") + if new == C.ZEROS: + ad, cd, subj = C.git(["log", "-1", "--format=%aI%x00%cI%x00%s", h], a.old).rstrip("\n").split("\0") + subj = s.scrub_text(subj.encode()).decode("utf-8", "replace") + else: + ad, cd, subj = C.git(["log", "-1", "--format=%aI%x00%cI%x00%s", new], new_repo).rstrip("\n").split("\0") + rows.append(f"{o:04d}\t{new}\t{ad}\t{cd}\t{subj}") + priv.append(f"{o:04d}\t{h}\t{new}") + header = ("# docs/commit-map.tsv — the public history's commit map (P33 C7). Generated by tools/public_rewrite/build_commit_map.py; never edit.\n" + "# Before the public flip the history was rewritten (ROM-derived paths purged, personal addresses mapped, old commit hashes\n" + "# in every historical doc replaced by the inert token commit:NNNN). NNNN is the ordinal of the commit on main in the ORIGINAL\n" + "# order (1 = the first commit); this table maps it to the rewritten commit. A row of 40 zeros is the one commit that touched\n" + "# only purged paths and was pruned. Documents at the tip cite the new hashes directly (tools/public_rewrite/resolve_tokens.py).\n" + "ordinal\tnew_hash\tauthor_date\tcommitter_date\tsubject\n") + text = header + "\n".join(rows) + "\n" + # a commit the rewrite left byte-identical (old == new: e.g. the noreply-authored "Initial commit", which cites no + # hash and touches no purge path) legitimately keeps its hash — exempt those from the old-hash assertion + unchanged = {o for o, n in cmap.items() if o == n} + d2 = {"commits": {h: e for h, e in d["commits"].items() if h not in unchanged}, "ambiguous": d["ambiguous"], + "excluded": d["excluded"]} + s2 = scrub.Scrubber(d2) + out = s2.scrub_text(text.encode()) + if s2.stats["replaced"] or s2.stats["emails"]: + C.die(f"the public map would carry {s2.stats['replaced']} old-hash tokens / {s2.stats['emails']} addresses — refusing") + (C.SCRATCH / "unchanged_commits.txt").write_text("\n".join(sorted(unchanged)) + ("\n" if unchanged else ""), encoding="utf-8") + out_pub.write_text(text, encoding="utf-8") + out_priv.write_text("ordinal\told_hash\tnew_hash\n" + "\n".join(priv) + "\n", encoding="utf-8") + zeros = sum(1 for r in rows if "\t" + C.ZEROS + "\t" in r) + print(f"build_commit_map: {len(rows)} rows -> {out_pub} ({zeros} pruned rows); private old-to-new -> {out_priv}; " + f"0 old-hash tokens in the public file (asserted; {len(unchanged)} commits unchanged by the rewrite keep their hash)") + return 0 + + +if __name__ == "__main__": + sys.exit(main(sys.argv[1:])) diff --git a/tools/public_rewrite/common.py b/tools/public_rewrite/common.py new file mode 100644 index 0000000000..e43a1c23a4 --- /dev/null +++ b/tools/public_rewrite/common.py @@ -0,0 +1,240 @@ +"""common.py — shared pieces of the public-flip rewrite tooling (P33 C1). Not a CLI. + +Everything ROM-derived, personal or old-hash-bearing lives ONLY under .run/public_rewrite/ (gitignored): the +dictionary of old commit hashes, the mailmap, the blob-id list, the bundle, the bare clone. No personal address and +no old hash is ever a literal in this package — they are derived from the repository at run time. +""" +import fnmatch +import hashlib +import json +import os +import pathlib +import re +import subprocess +import sys + +REPO = pathlib.Path(__file__).resolve().parents[2] +TOOLS = REPO / "tools" +HERE = pathlib.Path(__file__).resolve().parent +SCRATCH = REPO / ".run" / "public_rewrite" +PURGE_FILE = HERE / "purge_set.txt" +DICT_FILE = SCRATCH / "dict.json" +MAILMAP_FILE = SCRATCH / "mailmap" +ROM_IDS_FILE = SCRATCH / "rom_blob_ids.txt" +OLD_TO_NEW_FILE = SCRATCH / "old-to-new.tsv" +CLONE_DIR = SCRATCH / "repo.git" +COMMIT_MAP_PUBLIC = REPO / "docs" / "commit-map.tsv" + +NOREPLY_NAME = "Drew T" +NOREPLY_EMAIL = "50529377+Druthulu@users.noreply.github.com" +ZEROS = "0" * 40 + +# The one hash regex (validated at HEAD, S86: 711 resolving citations, 0 word-embedded false hits; `func_800D128C` and +# `0x800d128c` do not match because \b needs a non-word char on both sides and `_`/`x` are word chars). +HEX_RE = re.compile(rb"\b[0-9a-f]{7,40}\b") +# The inert token: non-hex inside its first 7 chars, so it can never re-match HEX_RE. +TOKEN_RE = re.compile(rb"\bcommit:(?:(?P[0-9]{4,})|orphan-(?P[0-9]+)|amb-(?P[0-9]+))\b") +TRAILER_RE = re.compile(rb"(?m)^Claude-Session:[^\n]*\n?") +SESSION_URL_RE = re.compile(rb"https://claude\.ai/code/session[^\s)>\"']*") +MIN_PREFIX = 7 +SIZE_CAP = 50 * 1024 * 1024 + + +def git(args, repo=REPO, text=True, check=True, input=None): + r = subprocess.run(["git", "-C", str(repo)] + list(args), capture_output=True, text=text, check=False, input=input) + if check and r.returncode != 0: + err = r.stderr if text else r.stderr.decode("utf-8", "replace") + raise RuntimeError(f"git {' '.join(args)} failed in {repo}: {err.strip()[:300]}") + return r.stdout + + +def die(msg, rc=1): + print(f"{pathlib.Path(sys.argv[0]).name}: {msg}", file=sys.stderr) + sys.exit(rc) + + +# ---------------------------------------------------------------- purge rules (the one file, filter-repo syntax) +def purge_rules(): + prefixes, globs = [], [] + for ln in PURGE_FILE.read_text(encoding="utf-8").splitlines(): + ln = ln.strip() + if not ln or ln.startswith("#"): + continue + if ln.startswith("glob:"): + globs.append(ln[5:]) + elif ln.startswith("regex:"): + die(f"regex: rules are not supported by this tooling ({ln})") + else: + prefixes.append(ln) + if not prefixes and not globs: + die(f"{PURGE_FILE} holds no rules — refusing (R43)") + return prefixes, globs + + +def under_purge(path, prefixes, globs): + """The rule a path falls under, or None. `path` is a str; directory prefixes match everything beneath.""" + for p in prefixes: + if path == p or path.startswith(p if p.endswith("/") else p + "/"): + return p + for g in globs: + if fnmatch.fnmatchcase(path, g): + return "glob:" + g + return None + + +# ---------------------------------------------------------------- the content-hash sets (derived, never typed) +def rom_content_sha1s(): + """SHA1s of ROM-derived artifacts the repo itself declares: the manifest, every check.*.sha, the dumps, the disc.""" + out = {} + man = REPO / "extracted" / "retail" / "manifest.jsonl" + if not man.exists(): + die(f"{man} missing — cannot derive the ROM hash set (R32)") + for ln in man.read_text(encoding="utf-8").splitlines(): + if ln.strip(): + o = json.loads(ln) + if o.get("size", 1) > 0: # the zero-length payloads' sha1 is the empty-file sha1 — not ROM bytes + out[o["sha1"].lower()] = "manifest:" + o["path"] + checks = sorted((REPO / "config").glob("check.*.sha")) + if not checks: + die("no config/check.*.sha — cannot derive the binary hashes (R32)") + for c in checks: + for ln in c.read_text(encoding="utf-8").splitlines(): + parts = ln.split() + if len(parts) >= 2 and len(parts[0]) == 40: + out[parts[0].lower()] = f"{c.name}:{parts[1]}" + dumps = REPO / "dumps" / "CHECKSUMS.sha1" + if dumps.exists(): + for ln in dumps.read_text(encoding="utf-8").splitlines(): + parts = ln.split() + if len(parts) >= 2 and len(parts[0]) == 40: + out[parts[0].lower()] = "dumps:" + parts[1] + sys.path.insert(0, str(TOOLS / "bfm_extract")) + from extract_exe import REDUMP_TRACK1_SHA1 # noqa: E402 + out[REDUMP_TRACK1_SHA1.lower()] = "redump:Track 1" + return out + + +def cited_content_hashes(): + """Every hash a doc may legitimately cite by prefix and that must NEVER be mistaken for a commit: the ROM sha1s + above, plus every 64-hex sha256 in the tracked checksum files, plus the redump CRC32 (an 8-hex token).""" + hashes = set(rom_content_sha1s()) + for f in (TOOLS / "bin" / "CHECKSUMS.sha256", TOOLS / "psyq_CHECKSUMS.sha256"): + if f.exists(): + hashes.update(m.group(0).lower() for m in re.finditer(r"\b[0-9a-f]{64}\b", f.read_text(encoding="utf-8"))) + sys.path.insert(0, str(TOOLS / "bfm_extract")) + from extract_exe import REDUMP_TRACK1_CRC32 # noqa: E402 + hashes.add(REDUMP_TRACK1_CRC32.lower()) + return hashes + + +def prefixes_of(h, lo=MIN_PREFIX, hi=40): + return {h[:n] for n in range(lo, min(len(h), hi) + 1)} + + +# ---------------------------------------------------------------- identities (derived from the log, never literal) +def personal_identities(repo=REPO): + """(name, email) pairs in the history whose email is not the noreply one.""" + seen = set() + for ln in git(["log", "--all", "--format=%an%x00%ae%x00%cn%x00%ce"], repo).splitlines(): + an, ae, cn, ce = ln.split("\0") + for n, e in ((an, ae), (cn, ce)): + if e != NOREPLY_EMAIL: + seen.add((n, e)) + return sorted(seen) + + +def write_mailmap(path=MAILMAP_FILE, repo=REPO): + ids = personal_identities(repo) + lines = [f"{NOREPLY_NAME} <{NOREPLY_EMAIL}> <{e}>" for _, e in ids] + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text("\n".join(lines) + "\n", encoding="utf-8") + return ids + + +def personal_emails_from_mailmap(path=MAILMAP_FILE): + if not path.exists(): + die(f"{path} missing — run hash_dict.py --write-mailmap first") + out = [] + for ln in path.read_text(encoding="utf-8").splitlines(): + m = re.match(r".*<([^>]+)>\s*<([^>]+)>\s*$", ln) + if m: + out.append(m.group(2)) + return out + + +# ---------------------------------------------------------------- the dictionary +def load_dict(path=DICT_FILE): + if not path.exists(): + die(f"{path} missing — run hash_dict.py first") + return json.loads(path.read_text(encoding="utf-8")) + + +def is_binary(data): + return b"\0" in data[:8192] + + +def git_blob_id(data): + h = hashlib.sha1() + h.update(b"blob %d\0" % len(data)) + h.update(data) + return h.hexdigest() + + +class CatFile: + """A persistent `git cat-file --batch` over one repo (bytes in, (id, type, data) out).""" + + def __init__(self, repo): + self.p = subprocess.Popen(["git", "-C", str(repo), "cat-file", "--batch"], stdin=subprocess.PIPE, + stdout=subprocess.PIPE, bufsize=0) + + def get(self, ref): + self.p.stdin.write(ref.encode() + b"\n") + self.p.stdin.flush() + hdr = self.p.stdout.readline().decode() + parts = hdr.split() + if len(parts) < 3 or parts[1] == "missing": + return None, None, None + n = int(parts[2]) + data = b"" + while len(data) < n: + chunk = self.p.stdout.read(n - len(data)) + if not chunk: + break + data += chunk + self.p.stdout.read(1) # the trailing newline + return parts[0], parts[1], data + + def close(self): + try: + self.p.stdin.close() + self.p.wait(timeout=10) + except Exception: + self.p.kill() + + +def iter_all_objects(repo, types=("blob",)): + """(id, type, size) for every object in the store (reachable or not), via --batch-all-objects.""" + out = git(["cat-file", "--batch-all-objects", "--batch-check"], repo) + for ln in out.splitlines(): + oid, typ, size = ln.split() + if typ in types: + yield oid, typ, int(size) + + +def parse_commit(raw): + """Parse a raw commit object (bytes) → dict(tree, parents, author, committer, message) with (name, email, ts, tz).""" + head, _, msg = raw.partition(b"\n\n") + d = {"parents": [], "message": msg} + for ln in head.split(b"\n"): + if ln.startswith(b"tree "): + d["tree"] = ln[5:].decode() + elif ln.startswith(b"parent "): + d["parents"].append(ln[7:].decode()) + elif ln.startswith(b"author ") or ln.startswith(b"committer "): + key, _, rest = ln.partition(b" ") + m = re.match(rb"(.*) <([^>]*)> (\d+) ([+-]\d{4})$", rest) + if not m: + raise ValueError(f"unparsable {key.decode()} line: {rest[:80]!r}") + d[key.decode()] = (m.group(1).decode("utf-8", "surrogateescape"), m.group(2).decode(), + int(m.group(3)), m.group(4).decode()) + return d diff --git a/tools/public_rewrite/expected_offenders.txt b/tools/public_rewrite/expected_offenders.txt new file mode 100644 index 0000000000..00203c2900 --- /dev/null +++ b/tools/public_rewrite/expected_offenders.txt @@ -0,0 +1,12 @@ +# tools/public_rewrite/expected_offenders.txt — the R39 NEGATIVE CONTROL fixture for gate_scan.py --expect-fail (P33 C1/C2). +# On the CURRENT (pre-rewrite) history the gate MUST fail naming every purge rule below with at least this many +# offending paths (paths ever touched, `git log --all --name-only`), and MUST report no content/signature/size +# offender outside the purge rules. Measured S87 (2026-09-06) on 4,300 reachable commits. Format: rulemin. +extracted/SLUS_007.26 1 +extracted/retail/SLUS_007.26 1 +glob:dumps/*.bin 28 +ghidra/ 27 +tools/psyq/ 189 +session archive/ 3 +glob:tools/ghidra-ext/*.zip 2 +tools/brave-CUE/brave.exe 1 diff --git a/tools/public_rewrite/gate_scan.py b/tools/public_rewrite/gate_scan.py new file mode 100644 index 0000000000..ed33d3caee --- /dev/null +++ b/tools/public_rewrite/gate_scan.py @@ -0,0 +1,187 @@ +#!/usr/bin/env python3 +"""gate_scan.py — the first-push gate over HISTORY: no ROM-derived blob reachable from the given refs (P33 C1). + + tools/public_rewrite/gate_scan.py --all [--worktree] [--expect-fail EXPECTED.txt] [--emit-ids FILE] [--repo PATH] + tools/public_rewrite/gate_scan.py --refs main [--worktree] + +Four checks, every one derived (R33) from the repo's own declarations: + 1. PATHS — every path ever touched by any commit reachable from the refs (`git log --name-only`) against the purge + rules in purge_set.txt; + 2. CONTENT — every reachable blob's content SHA1 against the ROM set (the manifest's non-empty payloads, every + config/check.*.sha, dumps/CHECKSUMS.sha1, the redump Track-1 SHA1) — a renamed copy is caught by content; + 3. SIGNATURES — `PS-X EXE` at offset 0; a 2,097,152-byte RAM image carrying the resident blob's first 16 bytes at + 0xCEDF8 or the EXE's entry code at 0x10000; a blob beginning with the EXE's entry code (a header-less EXE copy); + PsyQ `LIB\\x01` / `LNK\\x02` magics. The signature bytes come from the extracted payloads on disk (refused if absent + — a scan that silently skips a check is not a gate, R32); + 4. SIZE — any blob > 50 MiB. +`--emit-ids FILE` (default .run/public_rewrite/rom_blob_ids.txt) writes the blob ids that must die everywhere: +content/signature hits ∪ every blob that ever sat under a purge path (`git log --raw`), for `--strip-blobs-with-ids`. +`--expect-fail EXPECTED.txt` is the R39 NEGATIVE CONTROL: lines `\\t`; exit 0 only when +every listed rule has at least that many offending paths AND no content/signature/size offender sits outside the purge +rules (an unexpected ROM blob elsewhere is exactly what the gate exists to find). Without it: exit 1 on any offender. +`--worktree` additionally runs tools/audit_public.py over the tracked files (the tip). +""" +import argparse +import collections +import hashlib +import subprocess +import sys + +sys.path.insert(0, str(__import__("pathlib").Path(__file__).resolve().parent)) +import common as C # noqa: E402 + + +def signatures(): + exe = C.REPO / "extracted" / "retail" / "SLUS_007.26" + res = C.REPO / "extracted" / "retail" / "MAIN.CD.dir" / "FILE_010.dir" / "1.1" + if not exe.exists() or not res.exists(): + C.die(f"signature sources missing ({exe}, {res}) — run make disc-extract; refusing to scan without them (R32)") + with open(exe, "rb") as f: + head = f.read(0x810) + with open(res, "rb") as f: + res16 = f.read(16) + return {"exe_entry": head[0x800:0x810], "resident16": res16} + + +def reachable_blobs(repo, refs): + """[(blob id, first path)] reachable from refs.""" + out = C.git(["rev-list", "--objects"] + refs, repo) + ids, paths = [], {} + for ln in out.splitlines(): + parts = ln.split(" ", 1) + ids.append(parts[0]) + if len(parts) == 2: + paths[parts[0]] = parts[1] + chk = C.git(["cat-file", "--batch-check"], repo, input="\n".join(ids) + "\n") + blobs = [] + for ln in chk.splitlines(): + p = ln.split() + if len(p) >= 3 and p[1] == "blob": + blobs.append((p[0], paths.get(p[0], ""))) + return blobs + + +def paths_ever(repo, refs): + return set(C.git(["log", "--name-only", "--format="] + refs, repo).splitlines()) - {""} + + +def blobs_under_purge(repo, refs, prefixes, globs): + """(under, shared): blob ids of every version of every path that ever sat under a purge rule (`git log --raw`), + and the subset that ALSO appears under a NON-purge path somewhere in history. A shared blob must never be stripped + by id: the trial rewrite (S87) had the EMPTY blob in the list (an empty file once sat under a purge path) and + `--strip-blobs-with-ids` then dropped every "file emptied" change in history — those files silently kept their + previous content and one restore commit became empty and was pruned.""" + out = C.git(["log", "--raw", "--no-renames", "--no-abbrev", "--format="] + refs, repo) # --no-abbrev: full blob ids + under, elsewhere = set(), set() + for ln in out.splitlines(): + if not ln.startswith(":"): + continue + meta, _, path = ln.partition("\t") + f = meta.split() + if len(f) < 4: + continue + target = under if C.under_purge(path, prefixes, globs) else elsewhere + for oid in (f[2], f[3]): + if len(oid) == 40 and set(oid) != {"0"}: + target.add(oid) + return under, under & elsewhere + + +def main(argv): + ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) + g = ap.add_mutually_exclusive_group(required=True) + g.add_argument("--all", action="store_true") + g.add_argument("--refs", nargs="+") + ap.add_argument("--worktree", action="store_true") + ap.add_argument("--expect-fail") + ap.add_argument("--emit-ids", default=str(C.ROM_IDS_FILE)) + ap.add_argument("--repo", default=str(C.REPO)) + a = ap.parse_args(argv) + refs = ["--all"] if a.all else a.refs + prefixes, globs = C.purge_rules() + rom = C.rom_content_sha1s() + sig = signatures() + offenders = [] # (kind, path, detail, blob id or "") + # 1. paths + for p in sorted(paths_ever(a.repo, refs)): + r = C.under_purge(p, prefixes, globs) + if r: + offenders.append(("path", p, f"purge rule {r}", "")) + # 2–4. blobs + blobs = reachable_blobs(a.repo, refs) + cf = C.CatFile(a.repo) + content_hits, nbytes = set(), 0 + for oid, path in blobs: + _, _, data = cf.get(oid) + if data is None: + continue + nbytes += len(data) + why = [] + if len(data) > C.SIZE_CAP: + why.append(f"{len(data):,} bytes > 50 MiB") + h = hashlib.sha1(data).hexdigest() + if h in rom: + why.append(f"content sha1 == {rom[h]}") + if data[:8] == b"PS-X EXE": + why.append("signature PS-X EXE") + if len(data) == 2097152 and (data[0xCEDF8:0xCEDF8 + 16] == sig["resident16"] or data[0x10000:0x10010] == sig["exe_entry"]): + why.append("signature 2 MiB RAM image") + if data[:16] == sig["exe_entry"]: + why.append("signature EXE entry code at 0") + if data[:4] in (b"LIB\x01", b"LNK\x02"): + why.append("signature PsyQ LIB/LNK") + if why: + content_hits.add(oid) + offenders.append(("blob", path, "; ".join(why), oid)) + cf.close() + by_kind = collections.Counter(k for k, *_ in offenders) + under, shared = blobs_under_purge(a.repo, refs, prefixes, globs) + if by_path := sum(1 for k, *_ in offenders if k == "path"): + if not under: + C.die(f"{by_path} purge-path offenders but NO blob under any purge rule — the --raw parser is broken (R43)") + # strip by id: every content/signature hit (wherever it lives) + purge-path blobs that live NOWHERE else + ids = sorted(content_hits | (under - shared)) + with open(a.emit_ids, "w", encoding="utf-8") as f: + f.write("\n".join(ids) + ("\n" if ids else "")) + # report + print(f"gate_scan: refs {refs}: {len(blobs)} reachable blobs ({nbytes / 1e9:.2f} GB) scanned against {len(rom)} ROM " + f"sha1s + 5 signatures + 50 MiB; paths ever touched checked against {len(prefixes) + len(globs)} purge rules; " + f"offenders: paths {by_kind['path']}, blobs {by_kind['blob']} ({len(content_hits)} content/signature ids); " + f"rom_blob_ids: {len(ids)} ({len(under)} ever under a purge path, {len(shared)} of them shared with a non-purge path and " + f"therefore NOT stripped by id: {', '.join(x[:12] for x in sorted(shared)[:5])}) -> {a.emit_ids}") + rc = 0 + if a.expect_fail: + want = {} + for ln in open(a.expect_fail, encoding="utf-8"): + ln = ln.strip() + if ln and not ln.startswith("#"): + rule, _, n = ln.partition("\t") + want[rule.strip()] = int(n or 1) + per_rule = collections.Counter(d.split("purge rule ", 1)[1] for k, p, d, _ in offenders if k == "path") + for rule, n in sorted(want.items()): + ok = per_rule.get(rule, 0) >= n + print(f" {'ok ' if ok else 'BAD'} rule {rule}: {per_rule.get(rule, 0)} offending paths (expected ≥ {n})") + rc |= not ok + stray = [(p, d) for k, p, d, _ in offenders if k == "blob" and not C.under_purge(p, prefixes, globs)] + print(f" {'ok ' if not stray else 'BAD'} content/signature/size offenders outside the purge rules: {len(stray)}") + for p, d in stray[:20]: + print(f" STRAY {p}: {d}") + rc |= bool(stray) + print(f"gate_scan --expect-fail: {'PASS (the scan fails exactly as expected)' if not rc else 'FAIL'}") + else: + for k, p, d, oid in offenders[:40]: + print(f" OFFENDER [{k}] {p} {oid[:12]} {d}") + if len(offenders) > 40: + print(f" … {len(offenders) - 40} more") + rc = 1 if offenders else 0 + print(f"gate_scan: {'PASS — 0 offenders' if not offenders else f'FAIL — {len(offenders)} offenders'}") + if a.worktree: + sys.path.insert(0, str(C.TOOLS)) + import audit_public + wrc = audit_public.main([]) + rc |= wrc + return int(rc) + + +if __name__ == "__main__": + sys.exit(main(sys.argv[1:])) diff --git a/tools/public_rewrite/hash_dict.py b/tools/public_rewrite/hash_dict.py new file mode 100644 index 0000000000..14ab15465f --- /dev/null +++ b/tools/public_rewrite/hash_dict.py @@ -0,0 +1,104 @@ +#!/usr/bin/env python3 +"""hash_dict.py — the dictionary of EVERY old commit hash → its inert token (P33 C1/C2). + + tools/public_rewrite/hash_dict.py [--repo PATH] [--write-mailmap] + +Writes .run/public_rewrite/dict.json (private: it holds every old hash): + * every commit OBJECT in the store — reachable or not (`git cat-file --batch-all-objects`), because a doc may cite a + commit that was later rewritten away (S86: 119 unreachable commit objects; one cited by tools/verify_worktree.py); + * main's commits → ordinal 1..N in `git rev-list --reverse main` order (token `commit:NNNN`, zero-padded to 4); + * every other commit → the ordinal of its TWIN on main — same (tree, author timestamp, subject); the S76 tag lineage + is 126 such re-authored twins — else `orphan-K`; + * the prefix index (every prefix 7..40 of every hash) is rebuilt on load; ambiguous prefixes (two commits sharing + one) are recorded and MUST be 0 (measured 0 at 7 chars over 4,420 objects); prefixes that are also prefixes of a + cited CONTENT hash (the manifest / check.*.sha / dumps sha1s, the sha256 checksums, the redump CRC32) are recorded + as `excluded` — the scrub leaves such tokens untouched (expected < 1; R41 prints the count). +`--write-mailmap` also writes .run/public_rewrite/mailmap from the history's identities (never a literal anywhere). +""" +import argparse +import collections +import datetime as dt +import json +import sys + +sys.path.insert(0, str(__import__("pathlib").Path(__file__).resolve().parent)) +import common as C # noqa: E402 + + +def build(repo): + objs = [oid for oid, typ, _ in C.iter_all_objects(repo, types=("commit",))] + main = C.git(["rev-list", "--reverse", "main"], repo).split() + ordinal = {h: i + 1 for i, h in enumerate(main)} + cf = C.CatFile(repo) + meta = {} + for oid in objs: + _, typ, raw = cf.get(oid) + d = C.parse_commit(raw) + subject = d["message"].split(b"\n", 1)[0] + meta[oid] = (d["tree"], d["author"][2], subject) + cf.close() + twin_key = {} + for h in main: # first main commit with a key wins + twin_key.setdefault(meta[h], ordinal[h]) + commits, n_twin, n_orphan = {}, 0, 0 + orphans = sorted(h for h in objs if h not in ordinal and meta[h] not in twin_key) + orphan_no = {h: i + 1 for i, h in enumerate(orphans)} + for h in objs: + if h in ordinal: + commits[h] = {"kind": "main", "ord": ordinal[h]} + elif meta[h] in twin_key: + commits[h] = {"kind": "twin", "ord": twin_key[meta[h]]} + n_twin += 1 + else: + commits[h] = {"kind": "orphan", "orphan": orphan_no[h]} + n_orphan += 1 + # prefix index → ambiguity + index = {} + ambiguous = set() + for h in objs: + for p in C.prefixes_of(h): + if p in index and index[p] != h: + ambiguous.add(p) + index.setdefault(p, h) + # collisions with cited content hashes + content = C.cited_content_hashes() + excluded = sorted(p for h in content for p in C.prefixes_of(h) if p in index) + return { + "built_from": {"repo": str(repo), "head": C.git(["rev-parse", "HEAD"], repo).strip(), + "date": dt.datetime.now(dt.timezone.utc).isoformat(timespec="seconds")}, + "main_count": len(main), + "commits": commits, + "ambiguous": sorted(ambiguous), + "excluded": excluded, + "stats": {"objects": len(objs), "main": len(main), "twins": n_twin, "orphans": n_orphan, + "prefixes": len(index), "ambiguous": len(ambiguous), "excluded": len(excluded), + "content_hashes": len(content)}, + } + + +def main(argv): + ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) + ap.add_argument("--repo", default=str(C.REPO)) + ap.add_argument("--write-mailmap", action="store_true") + a = ap.parse_args(argv) + C.SCRATCH.mkdir(parents=True, exist_ok=True) + d = build(a.repo) + C.DICT_FILE.write_text(json.dumps(d, indent=0), encoding="utf-8") + s = d["stats"] + print(f"hash_dict: {s['objects']} commit objects ({s['main']} on main, {s['twins']} twins, {s['orphans']} orphans), " + f"{s['prefixes']} prefixes 7..40, {s['ambiguous']} ambiguous, {s['excluded']} collide with {s['content_hashes']} " + f"cited content hashes -> {C.DICT_FILE}") + if d["excluded"]: + print(" excluded (left untouched by the scrub): " + ", ".join(d["excluded"][:20])) + if a.write_mailmap: + ids = C.write_mailmap(repo=a.repo) + print(f"hash_dict: mailmap written ({len(ids)} personal identities -> the noreply identity) -> {C.MAILMAP_FILE}") + if d["ambiguous"]: + print(f"hash_dict: {len(d['ambiguous'])} AMBIGUOUS prefixes (scrub would emit commit:amb-N) — FAIL, investigate", + file=sys.stderr) + return 1 + return 0 + + +if __name__ == "__main__": + sys.exit(main(sys.argv[1:])) diff --git a/tools/public_rewrite/probe_github.sh b/tools/public_rewrite/probe_github.sh new file mode 100644 index 0000000000..7bf055daf5 --- /dev/null +++ b/tools/public_rewrite/probe_github.sh @@ -0,0 +1,46 @@ +#!/usr/bin/env bash +# tools/public_rewrite/probe_github.sh — Drew's post-purge probe: do the OLD hashes still resolve on GitHub? (P33 C10) +# +# tools/public_rewrite/probe_github.sh [--after-flip] [--repo Druthulu/BFM-decomp] [--n 30] +# +# Needs `gh auth login` in the running shell (Drew's) and .run/public_rewrite/old-to-new.tsv (build_commit_map.py). +# Samples N old hashes evenly over the map + the pruned commit's old hash + the old tag tip (if recorded) and, for each: +# * `gh api repos//commits/` must FAIL with 404 (still 200 = GitHub still serves the old object); +# * `git fetch origin ` must FAIL (still fetchable = still on the server, cached views or not). +# Positive control: the current `main` sha MUST succeed both ways (proves the probe can see a live commit). +# --after-flip additionally probes 7-char prefixes UNAUTHENTICATED at https://github.com//commit/<7> (expect 404). +# Exit 0 only when every old sha is gone and the control passes. Re-run daily while anything returns 200. +set -uo pipefail +REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)" +GH_REPO="Druthulu/BFM-decomp"; N=30; AFTER=0 +while [ $# -gt 0 ]; do case "$1" in --after-flip) AFTER=1 ;; --repo) GH_REPO="$2"; shift ;; --n) N="$2"; shift ;; *) echo "unknown arg $1" >&2; exit 2 ;; esac; shift; done +MAP="$REPO/.run/public_rewrite/old-to-new.tsv" +[ -f "$MAP" ] || { echo "probe: $MAP missing (run build_commit_map.py)"; exit 2; } +command -v gh >/dev/null || { echo "probe: gh not on PATH"; exit 2; } +gh auth status >/dev/null 2>&1 || { echo "probe: gh is not authenticated — run: gh auth login"; exit 2; } +cd "$REPO" +mapfile -t OLD < <(awk -F'\t' 'NR>1 && $2 != $3 {print $2}' "$MAP") # skip commits the rewrite left identical (old == new) +total=${#OLD[@]}; step=$(( total / N )); [ "$step" -lt 1 ] && step=1 +SAMPLE=(); for ((i=0; i&1 | grep -oE 'HTTP [0-9]{3}' | head -1)" + api_gone=0; [ -z "$code" ] && code="HTTP 200"; [[ "$code" == *404* ]] && api_gone=1 + if git fetch --quiet origin "$sha" 2>/dev/null; then fetch_gone=0; else fetch_gone=1; fi + if [ $api_gone = 1 ] && [ $fetch_gone = 1 ]; then echo " gone $sha"; else echo " ALIVE $sha (api $code, fetch $([ $fetch_gone = 1 ] && echo fails || echo SUCCEEDS))"; alive=$((alive+1)); fi + if [ $AFTER = 1 ]; then + http="$(curl -s -o /dev/null -w '%{http_code}' "https://github.com/$GH_REPO/commit/${sha:0:7}")" + [ "$http" = 404 ] || { echo " ALIVE unauthenticated /commit/${sha:0:7} -> HTTP $http"; alive=$((alive+1)); } + fi +done +cur="$(git rev-parse main)" +if gh api "repos/$GH_REPO/commits/$cur" --silent >/dev/null 2>&1 && git fetch --quiet origin "$cur" 2>/dev/null; then + echo " control: current main $cur resolves (OK)" +else + echo " control FAILED: current main $cur does not resolve — the probe cannot be trusted"; exit 2 +fi +if [ $alive = 0 ]; then echo "probe: PASS — every sampled old hash is gone from $GH_REPO"; exit 0; fi +echo "probe: $alive still ALIVE — wait for the Support purge / GC and re-run"; exit 1 diff --git a/tools/public_rewrite/resolve_tokens.py b/tools/public_rewrite/resolve_tokens.py new file mode 100644 index 0000000000..8ec2fb2820 --- /dev/null +++ b/tools/public_rewrite/resolve_tokens.py @@ -0,0 +1,102 @@ +#!/usr/bin/env python3 +"""resolve_tokens.py — at the tip of the ADOPTED (rewritten) checkout, turn `commit:NNNN` tokens back into hashes (P33 C7). + + tools/public_rewrite/resolve_tokens.py # rewrites tracked text files in place; prints the residue + tools/public_rewrite/resolve_tokens.py --check # exit 1 if any resolvable token remains at HEAD; lists the residue + +Each `commit:NNNN` (an ordinal of the original main, main or twin) becomes the SHORTEST UNIQUE abbreviation (≥ 9 chars) +of the rewritten commit, from docs/commit-map.tsv — uniqueness asserted with `git rev-parse --verify` in THIS repo. +The residue that stays as tokens, listed for the tip commit's message: ordinals whose row is the pruned commit (40 +zeros), `commit:orphan-K` (a cited commit that exists in no lineage — e.g. tools/verify_worktree.py cites a dropped +TEMP commit) and `commit:amb-K` (an ambiguous old prefix; measured 0). The fixed-point rule: this runs ONLY at the tip, +after the rewrite — a new hash must never be written into a historical blob. +""" +import argparse +import re +import subprocess +import sys + +sys.path.insert(0, str(__import__("pathlib").Path(__file__).resolve().parent)) +import common as C # noqa: E402 + + +def load_map(path=C.COMMIT_MAP_PUBLIC): + if not path.exists(): + C.die(f"{path} missing — run build_commit_map.py first") + m = {} + for ln in path.read_text(encoding="utf-8").splitlines(): + if ln.startswith("#") or ln.startswith("ordinal\t"): + continue + o, new, *_ = ln.split("\t") + m[int(o)] = new + return m + + +_abbr_cache = {} + + +def abbrev(h, repo): + if h in _abbr_cache: + return _abbr_cache[h] + for n in range(9, 41): + r = subprocess.run(["git", "-C", str(repo), "rev-parse", "--verify", "--quiet", h[:n] + "^{commit}"], + capture_output=True, text=True) + if r.returncode == 0 and r.stdout.strip() == h: + _abbr_cache[h] = h[:n] + return h[:n] + C.die(f"{h} does not resolve in {repo} — is this the adopted (rewritten) checkout?") + + +def main(argv): + ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) + ap.add_argument("--check", action="store_true") + ap.add_argument("--repo", default=str(C.REPO)) + ap.add_argument("--map", default=str(C.COMMIT_MAP_PUBLIC), help="the commit map (a scratch path for a trial run)") + a = ap.parse_args(argv) + repo = __import__("pathlib").Path(a.repo) + m = load_map(__import__("pathlib").Path(a.map)) + files = C.git(["ls-files", "-z"], repo, text=False).decode("utf-8", "surrogateescape").split("\0") + n_files, n_resolved, residue, remaining = 0, 0, {}, 0 + for rel in files: + if not rel or rel == "docs/commit-map.tsv": + continue + p = repo / rel + if not p.is_file(): + continue + data = p.read_bytes() + if C.is_binary(data) or not C.TOKEN_RE.search(data): + continue + + def sub(mt): + nonlocal n_resolved, remaining + if mt.group("ord"): + o = int(mt.group("ord")) + new = m.get(o) + if new and new != C.ZEROS: + if a.check: + remaining += 1 + return mt.group(0) + n_resolved += 1 + return abbrev(new, repo).encode() + residue[mt.group(0).decode()] = residue.get(mt.group(0).decode(), 0) + 1 + return mt.group(0) + residue[mt.group(0).decode()] = residue.get(mt.group(0).decode(), 0) + 1 + return mt.group(0) + + out = C.TOKEN_RE.sub(sub, data) + if out != data and not a.check: + p.write_bytes(out) + n_files += 1 + if a.check: + print(f"resolve_tokens --check: {remaining} resolvable tokens remain at HEAD ({'OK' if not remaining else 'FAIL'}); " + f"residue (unresolvable, left as tokens): {sum(residue.values())} in {len(residue)} distinct") + for t, n in sorted(residue.items()): + print(f" residue {t} ×{n}") + return 1 if remaining else 0 + print(f"resolve_tokens: {n_resolved} tokens resolved in {n_files} files; residue {sum(residue.values())} " + f"({', '.join(f'{t}×{n}' for t, n in sorted(residue.items())) or 'none'})") + return 0 + + +if __name__ == "__main__": + sys.exit(main(sys.argv[1:])) diff --git a/tools/public_rewrite/run_filter.py b/tools/public_rewrite/run_filter.py new file mode 100644 index 0000000000..1ddbfd53ad --- /dev/null +++ b/tools/public_rewrite/run_filter.py @@ -0,0 +1,118 @@ +#!/usr/bin/env python3 +"""run_filter.py — compose and run the git-filter-repo rewrite inside the bare clone (P33 C4c). + + tools/public_rewrite/run_filter.py --sample # = scrub.py --sample (R37: measure before the real run) + tools/public_rewrite/run_filter.py [--clone PATH] # the rewrite; refuses outside a bare clone under .run/public_rewrite/ + +The filter, via the git_filter_repo module API (2.47.0), equals: + git filter-repo --force --invert-paths --paths-from-file + --strip-blobs-with-ids .run/public_rewrite/rom_blob_ids.txt --mailmap .run/public_rewrite/mailmap + --prune-empty auto --replace-refs delete-no-add +plus a blob callback (scrub every text blob: old hashes -> tokens, personal addresses -> noreply) and a message +callback (the same + the `Claude-Session:` trailer lines dropped). Why each flag: `--prune-empty auto`, never `always` +(main carries one pre-existing empty commit that must survive); `--replace-refs delete-no-add` (no refs/replace/ +may be minted — they would leak old hashes); `--strip-blobs-with-ids` (the EXE dies wherever it was renamed); +`--force` (the clone is no longer a "fresh clone" once the S76 tag is deleted in it — the only deviation). +Preconditions asserted: the clone is bare, lives under .run/public_rewrite/, has exactly refs/heads/main, one pack, no +loose objects; dict.json, mailmap and rom_blob_ids.txt exist. Logs versions + wall time to .run/public_rewrite/filter.log; +afterwards counts the commit-map rows (and how many map to zeros), asserts no refs/replace, reports packs and size. +""" +import argparse +import os +import subprocess +import sys +import time + +sys.path.insert(0, str(__import__("pathlib").Path(__file__).resolve().parent)) +import common as C # noqa: E402 + + +def precheck(clone): + clone = clone.resolve() + if C.SCRATCH.resolve() not in clone.parents: + C.die(f"{clone} is not under {C.SCRATCH} — refusing (the rewrite runs only on the scratch bare clone)") + if C.git(["rev-parse", "--is-bare-repository"], clone).strip() != "true": + C.die(f"{clone} is not a bare repository") + refs = C.git(["for-each-ref", "--format=%(refname)"], clone).split() + if refs != ["refs/heads/main"]: + C.die(f"the clone must hold exactly refs/heads/main before the rewrite; it holds {refs} (delete the tag/others first)") + co = dict(ln.split(": ") for ln in C.git(["count-objects", "-v"], clone).splitlines()) + if co.get("packs") != "1" or co.get("count") != "0": + C.die(f"the clone must be one pack with zero loose objects; count-objects: {co}") + for f in (C.DICT_FILE, C.MAILMAP_FILE, C.ROM_IDS_FILE): + if not f.exists(): + C.die(f"{f} missing") + return clone + + +def clean_rules_file(): + out = C.SCRATCH / "purge_rules.filter-repo.txt" + lines = [ln.strip() for ln in C.PURGE_FILE.read_text(encoding="utf-8").splitlines()] + lines = [ln for ln in lines if ln and not ln.startswith("#")] + out.write_text("\n".join(lines) + "\n", encoding="utf-8") + return out + + +def main(argv): + ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) + ap.add_argument("--sample", action="store_true") + ap.add_argument("--clone", default=str(C.CLONE_DIR)) + a = ap.parse_args(argv) + if a.sample: + import scrub + return scrub.main(["--sample"]) + clone = precheck(__import__("pathlib").Path(a.clone)) + import git_filter_repo as fr + import scrub + s = scrub.Scrubber() + rules = clean_rules_file() + argv_fr = ["--force", "--invert-paths", "--paths-from-file", str(rules), + "--strip-blobs-with-ids", str(C.ROM_IDS_FILE), "--mailmap", str(C.MAILMAP_FILE), + "--prune-empty", "auto", "--replace-refs", "delete-no-add"] + log = C.SCRATCH / "filter.log" + with open(log, "a", encoding="utf-8") as lf: + lf.write(f"\n=== run_filter {time.strftime('%Y-%m-%dT%H:%M:%SZ', time.gmtime())} clone={clone}\n") + lf.write(f"git: {C.git(['--version']).strip()} git-filter-repo: {getattr(fr, '__version__', '?')} ({fr.__file__})\n") + lf.write(f"args: {' '.join(argv_fr)}\n") + lf.write(f"rules: {rules.read_text()}") + old_head = C.git(["rev-parse", "main"], clone).strip() + n_before = int(C.git(["rev-list", "--count", "main"], clone)) + os.chdir(clone) + t0 = time.time() + + def blob_cb(blob, metadata): + blob.data = s.scrub_text(blob.data) + + def message_cb(message): + return s.scrub_message(message) + + args = fr.FilteringOptions.parse_args(argv_fr) + f = fr.RepoFilter(args, blob_callback=blob_cb, message_callback=message_cb) + f.run() + wall = time.time() - t0 + # post-run facts + cmap = clone / "filter-repo" / "commit-map" + rows = [ln.split() for ln in cmap.read_text().splitlines() if ln and not ln.startswith("old")] + zeros = sum(1 for r in rows if r[1] == C.ZEROS) + refs = C.git(["for-each-ref", "--format=%(refname)"], clone).split() + co = dict(ln.split(": ") for ln in C.git(["count-objects", "-v"], clone).splitlines()) + n_after = int(C.git(["rev-list", "--count", "main"], clone)) + st = dict(s.stats) + summary = (f"run_filter: DONE in {wall:.0f} s; commit-map rows {len(rows)} ({zeros} -> zeros = pruned); main {n_before} -> " + f"{n_after} commits (old tip {old_head[:9]} -> {C.git(['rev-parse', 'main'], clone).strip()[:9]}); refs {refs}; " + f"count-objects {co.get('packs')} packs, {co.get('count')} loose, size-pack {int(co.get('size-pack', 0)) // 1024} MB; " + f"scrub stats {st}") + print(summary) + with open(log, "a", encoding="utf-8") as lf: + lf.write(summary + "\n") + bad = [r for r in refs if r.startswith("refs/replace/")] + if bad: + C.die(f"refs/replace minted: {bad}") + if zeros != 1: + print(f"run_filter: WARNING — expected exactly 1 pruned commit (the purge-only 'session archive update'), got {zeros}", + file=sys.stderr) + return 0 + + +if __name__ == "__main__": + sys.exit(main(sys.argv[1:])) diff --git a/tools/public_rewrite/scrub.py b/tools/public_rewrite/scrub.py new file mode 100644 index 0000000000..7e86ac19c4 --- /dev/null +++ b/tools/public_rewrite/scrub.py @@ -0,0 +1,211 @@ +#!/usr/bin/env python3 +"""scrub.py — THE one scrub function of the public-flip rewrite (P33 C1), and its sample / self-test. + + tools/public_rewrite/scrub.py --test # known-true cases (R39) + tools/public_rewrite/scrub.py --sample [--rev HEAD] # every tracked blob of a tree: hits, files, per-length counts, + # the 7-char replacements in context, an INDEPENDENT oracle + tools/public_rewrite/scrub.py --file IN --out OUT # scrub one file (debugging) + +What it does to a text blob or a commit message: + * every hex token 7..40 chars long that is a prefix of exactly one OLD commit hash (the dictionary) becomes the + inert token `commit:NNNN` (main ordinal), `commit:orphan-K`, or `commit:amb-K` for an ambiguous prefix; tokens + that are also prefixes of a cited CONTENT hash are left alone (the exclusion set); anything else is untouched; + * every personal e-mail address (from the scratch mailmap) becomes the noreply address; + * for MESSAGES only: `Claude-Session: …` trailer lines are dropped (60 on main, 130 over all refs at S87). +Binary blobs (a NUL in the first 8 KiB) are never touched. The function is idempotent: scrub(scrub(x)) == scrub(x). +""" +import argparse +import collections +import re +import sys +import time + +sys.path.insert(0, str(__import__("pathlib").Path(__file__).resolve().parent)) +import common as C # noqa: E402 + + +class Scrubber: + def __init__(self, dictionary=None, personal_emails=None): + d = dictionary if dictionary is not None else C.load_dict() + self.commits = d["commits"] + self.index = {} + for h in self.commits: + for p in C.prefixes_of(h): + self.index.setdefault(p, h) + self.ambiguous = {p: i + 1 for i, p in enumerate(d["ambiguous"])} + self.excluded = set(d["excluded"]) + self.emails = [e.encode() for e in (personal_emails if personal_emails is not None + else C.personal_emails_from_mailmap())] + self.noreply = C.NOREPLY_EMAIL.encode() + self.stats = collections.Counter() + self.replaced = collections.Counter() # token -> count (for the sample) + + def token_for(self, tok): + """The replacement for one hex token (bytes) or None.""" + t = tok.decode() + if t in self.excluded: + self.stats["excluded"] += 1 + return None + if t in self.ambiguous: + self.stats["ambiguous"] += 1 + return b"commit:amb-%d" % self.ambiguous[t] + h = self.index.get(t) + if h is None: + self.stats["unresolved"] += 1 + return None + e = self.commits[h] + if e["kind"] in ("main", "twin"): + return b"commit:%04d" % e["ord"] + return b"commit:orphan-%d" % e["orphan"] + + def _sub(self, m): + r = self.token_for(m.group(0)) + if r is None: + return m.group(0) + self.stats["replaced"] += 1 + self.stats[f"replaced_len{len(m.group(0))}"] += 1 + self.replaced[m.group(0)] += 1 + return r + + def scrub_text(self, data): + if C.is_binary(data): + self.stats["binary"] += 1 + return data + out = C.HEX_RE.sub(self._sub, data) + for e in self.emails: + if e in out: + self.stats["emails"] += out.count(e) + out = out.replace(e, self.noreply) + return out + + def scrub_message(self, msg): + out, n = C.TRAILER_RE.subn(b"", msg) + self.stats["trailers"] += n + out = self.scrub_text(out) + if n: + out = out.rstrip(b"\n") + b"\n" + return out + + +# ---------------------------------------------------------------- self-test (known-true cases) +def selftest(): + fake_main = "0123456789abcdef0123456789abcdef01234567" + fake_twin = "fedcba9876543210fedcba9876543210fedcba98" + fake_orphan = "1111111111111111111111111111111111111111" + content = "abcdef0123456789abcdef0123456789abcdef01" # a "manifest" sha1: prefix abcdef0 collides with nothing + d = {"commits": {fake_main: {"kind": "main", "ord": 12}, fake_twin: {"kind": "twin", "ord": 3}, + fake_orphan: {"kind": "orphan", "orphan": 2}}, + "ambiguous": [], "excluded": sorted(C.prefixes_of(content))} + s = Scrubber(d, personal_emails=["someone@example.com"]) + cases = [ + (b"see 0123456789a for the fix", b"see commit:0012 for the fix", "9-char main prefix"), + (b"(" + fake_main.encode() + b")", b"(commit:0012)", "full 40-char main hash"), + (b"twin fedcba98 here", b"twin commit:0003 here", "8-char twin prefix -> the twin's ordinal"), + (b"lost 1111111 commit", b"lost commit:orphan-2 commit", "orphan"), + (b"func_800D128C and 0x800d128c stay", b"func_800D128C and 0x800d128c stay", "word-embedded hex untouched"), + (b"abcdef0123 is a manifest sha1 prefix", b"abcdef0123 is a manifest sha1 prefix", "excluded content-hash prefix"), + (b"deadbeefcafe is not a commit", b"deadbeefcafe is not a commit", "unknown hex untouched"), + (b"012345 too short", b"012345 too short", "6 chars never match"), + (b"mail someone@example.com now", b"mail " + C.NOREPLY_EMAIL.encode() + b" now", "personal address -> noreply"), + (b"\0binary 0123456789a", b"\0binary 0123456789a", "binary untouched"), + ] + bad = 0 + for src, want, name in cases: + got = s.scrub_text(src) + ok = got == want + bad += not ok + print(f" {'ok ' if ok else 'BAD'} {name}: {got!r}") + msg = b"subject 0123456789a\n\nbody\n\nClaude-Session: https://example/x\n" + got = s.scrub_message(msg) + want = b"subject commit:0012\n\nbody\n" + ok = got == want + bad += not ok + print(f" {'ok ' if ok else 'BAD'} message: trailer dropped + hash tokenized: {got!r}") + idem = s.scrub_text(s.scrub_text(b"x 0123456789a y fedcba98")) == s.scrub_text(b"x 0123456789a y fedcba98") + bad += not idem + print(f" {'ok ' if idem else 'BAD'} idempotent") + print(f"scrub --test: {'OK' if not bad else f'{bad} FAILED'}") + return 1 if bad else 0 + + +# ---------------------------------------------------------------- the sample over one tree (R37: measure first) +def sample(rev, repo): + s = Scrubber() + entries = [ln.split("\t")[1] for ln in C.git(["ls-tree", "-r", rev], repo).splitlines()] + ids = {ln.split("\t")[1]: ln.split()[2] for ln in C.git(["ls-tree", "-r", rev], repo).splitlines()} + cf = C.CatFile(repo) + t0 = time.time() + files_hit, nbytes, ctx7 = 0, 0, [] + all_tokens = collections.Counter() + for path in entries: + _, _, data = cf.get(ids[path]) + if data is None or C.is_binary(data): + continue + nbytes += len(data) + before = s.stats["replaced"] + out = s.scrub_text(data) + if s.stats["replaced"] != before: + files_hit += 1 + for m in C.HEX_RE.finditer(data): + all_tokens[m.group(0)] += 1 + if len(m.group(0)) == 7 and s.token_for(m.group(0)): + a, b = max(0, m.start() - 40), min(len(data), m.end() + 40) + ctx7.append(f"{path}: …{data[a:b].decode('utf-8', 'replace')}…".replace("\n", "⏎")) + cf.close() + wall = time.time() - t0 + # INDEPENDENT oracle: which distinct tokens does GIT ITSELF resolve to a commit object? + distinct = sorted(t.decode() for t in all_tokens) + resolved_by_git = set() + for i in range(0, len(distinct), 2000): + chunk = distinct[i:i + 2000] + out = C.git(["cat-file", "--batch-check"], repo, input="\n".join(f"{t}^{{commit}}" for t in chunk) + "\n", check=False) + for tok, ln in zip(chunk, out.splitlines()): + if " commit " in ln: + resolved_by_git.add(tok) + ours = {t.decode() for t in s.replaced} + only_git = sorted(resolved_by_git - ours) + only_ours = sorted(ours - resolved_by_git) + per_len = {k: v for k, v in sorted(s.stats.items()) if k.startswith("replaced_len")} + print(f"scrub --sample {rev}: {len(entries)} tracked paths, {nbytes / 1e6:.0f} MB of text scanned in {wall:.1f} s " + f"({nbytes / 1e6 / max(wall, 1e-9):.0f} MB/s); {s.stats['replaced']} replacements in {files_hit} files; " + f"{len(ours)} distinct tokens replaced; excluded {s.stats['excluded']}, ambiguous {s.stats['ambiguous']}, " + f"unresolved hex tokens {s.stats['unresolved']}, e-mail replacements {s.stats['emails']}") + print(f" per length: {per_len}") + print(f" independent oracle (git cat-file on every distinct hex token, {len(distinct)} tokens): git resolves " + f"{len(resolved_by_git)} to commits; ours {len(ours)}; only-git {len(only_git)} (expected: the excluded " + f"content-hash prefixes, if any); only-ours {len(only_ours)} (MUST be 0)") + for t in only_git[:10]: + print(f" only-git: {t} {'(excluded)' if t in s.excluded else '(!! dictionary gap)'}") + for t in only_ours[:10]: + print(f" only-ours: {t} !!") + print(f" 7-char replacements in context ({len(ctx7)} — READ them, R63):") + for c in ctx7[:60]: + print(f" {c[:200]}") + return 1 if only_ours or any(t not in s.excluded for t in only_git) else 0 + + +def main(argv): + ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) + ap.add_argument("--test", action="store_true") + ap.add_argument("--sample", action="store_true") + ap.add_argument("--rev", default="HEAD") + ap.add_argument("--repo", default=str(C.REPO)) + ap.add_argument("--file") + ap.add_argument("--out") + a = ap.parse_args(argv) + if a.test: + return selftest() + if a.sample: + return sample(a.rev, a.repo) + if a.file: + s = Scrubber() + data = open(a.file, "rb").read() + out = s.scrub_text(data) + open(a.out or (a.file + ".scrubbed"), "wb").write(out) + print(f"scrub: {dict(s.stats)}") + return 0 + ap.error("one of --test / --sample / --file") + + +if __name__ == "__main__": + sys.exit(main(sys.argv[1:])) diff --git a/tools/public_rewrite/verify_rewrite.py b/tools/public_rewrite/verify_rewrite.py new file mode 100644 index 0000000000..db0929f5be --- /dev/null +++ b/tools/public_rewrite/verify_rewrite.py @@ -0,0 +1,161 @@ +#!/usr/bin/env python3 +"""verify_rewrite.py — the pairwise proof that the rewrite changed ONLY what it was told to (P33 C5). + + tools/public_rewrite/verify_rewrite.py [--old ~/bfm-decomp] [--new .run/public_rewrite/repo.git] + +For every (old, new) pair in the clone's filter-repo commit-map (old commits read from the ORIGINAL repo — the clone +has gc'd them away; new commits from the clone): + * author and committer: name/e-mail equal after the mailmap; BOTH timestamps and zones equal; + * message: new == scrub_message(old) — byte-equal (hash tokens, addresses, the trailer lines); + * parents: new parents == map(old parents) with every pruned commit spliced out (its own parent substituted); + * trees (`git ls-tree -r`): every path REMOVED sits under a purge rule or its blob id is in rom_blob_ids.txt; every + path CHANGED has the same mode and its new blob id == git-hash(scrub_text(old blob)); any path ADDED fails; gitlinks + (submodules) must be identical. +Then: the new main's commit count == old − pruned; the author/committer timestamp sequences match with the pruned +commits removed; the pre-existing EMPTY commit on main (tree == parent's tree) survived. Every failure is named; exit 1 +on any. Cost: ≈4,000 pairs × (2 cat-file + 2 ls-tree) ≈ 3–5 min. +""" +import argparse +import subprocess +import sys + +sys.path.insert(0, str(__import__("pathlib").Path(__file__).resolve().parent)) +import common as C # noqa: E402 +import scrub # noqa: E402 + + +def commit_map(new): + for cand in (new / "filter-repo" / "commit-map", new / ".git" / "filter-repo" / "commit-map"): + if cand.exists(): + rows = [ln.split() for ln in cand.read_text().splitlines() if ln and not ln.startswith("old")] + return {o: n for o, n in rows} + C.die(f"no filter-repo/commit-map under {new}") + + +def ls_tree(repo, rev): + out = C.git(["ls-tree", "-r", "-z", rev], repo, text=False, check=True) + d = {} + for ent in out.split(b"\0"): + if not ent: + continue + meta, _, path = ent.partition(b"\t") + mode, typ, oid = meta.decode().split() + d[path.decode("utf-8", "surrogateescape")] = (mode, typ, oid) + return d + + +def main(argv): + ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) + ap.add_argument("--old", default=str(C.REPO)) + ap.add_argument("--new", default=str(C.CLONE_DIR)) + a = ap.parse_args(argv) + old_repo, new_repo = __import__("pathlib").Path(a.old), __import__("pathlib").Path(a.new) + cmap = commit_map(new_repo) + pruned = {o for o, n in cmap.items() if n == C.ZEROS} + prefixes, globs = C.purge_rules() + rom_ids = set(C.ROM_IDS_FILE.read_text().split()) if C.ROM_IDS_FILE.exists() else set() + personal = set(C.personal_emails_from_mailmap()) + s = scrub.Scrubber() + cf_old, cf_new = C.CatFile(old_repo), C.CatFile(new_repo) + fails, n_pairs, n_changed_blobs, n_removed = [], 0, 0, 0 + old_commits = {} + + def old_commit(h): + if h not in old_commits: + _, _, raw = cf_old.get(h) + old_commits[h] = C.parse_commit(raw) + return old_commits[h] + + def splice(p): + while p in pruned: + par = old_commit(p)["parents"] + if len(par) != 1: + return None + p = par[0] + return cmap.get(p) + + def ident_ok(old_id, new_id): + name, email, ts, tz = old_id + exp = (C.NOREPLY_NAME, C.NOREPLY_EMAIL, ts, tz) if email in personal else old_id + return exp == new_id, exp + + for old, new in cmap.items(): + if new == C.ZEROS: + continue + n_pairs += 1 + o = old_commit(old) + _, typ, raw = cf_new.get(new) + if raw is None: + fails.append(f"{old[:9]} -> {new[:9]}: new commit missing"); continue + n = C.parse_commit(raw) + for who in ("author", "committer"): + ok, exp = ident_ok(o[who], n[who]) + if not ok: + fails.append(f"{old[:9]}: {who} {n[who]} != expected {exp}") + want_msg = s.scrub_message(o["message"]) + if want_msg != n["message"]: + fails.append(f"{old[:9]}: message differs from scrub(old): {n['message'][:80]!r} vs {want_msg[:80]!r}") + want_par = [splice(p) for p in o["parents"]] + if want_par != n["parents"]: + fails.append(f"{old[:9]}: parents {n['parents']} != expected {want_par}") + if o["tree"] != n["tree"]: + to, tn = ls_tree(old_repo, o["tree"]), ls_tree(new_repo, n["tree"]) + for path in tn: + if C.under_purge(path, prefixes, globs): + fails.append(f"{old[:9]}: purge path SURVIVES in the new tree: {path}") + for path in set(to) - set(tn): + n_removed += 1 + mode, typ, oid = to[path] + if not C.under_purge(path, prefixes, globs) and oid not in rom_ids: + fails.append(f"{old[:9]}: path REMOVED outside the purge set: {path} ({oid[:12]})") + for path in set(tn) - set(to): + fails.append(f"{old[:9]}: path ADDED by the rewrite: {path}") + for path in set(to) & set(tn): + (mo, tyo, ido), (mn, tyn, idn) = to[path], tn[path] + if ido == idn and mo == mn: + continue + if mo != mn or tyo != tyn or tyo != "blob": + fails.append(f"{old[:9]}: {path}: mode/type changed {mo}/{tyo} -> {mn}/{tyn}"); continue + n_changed_blobs += 1 + _, _, data = cf_old.get(ido) + want = C.git_blob_id(s.scrub_text(data)) + if want != idn: + fails.append(f"{old[:9]}: {path}: new blob {idn[:12]} != git-hash(scrub(old {ido[:12]})) {want[:12]}") + if len(fails) > 200: + fails.append("… stopping after 200 failures"); break + cf_old.close(); cf_new.close() + # whole-line checks + old_main = C.git(["rev-list", "--reverse", "main"], old_repo).split() + new_main = C.git(["rev-list", "--reverse", "main"], new_repo).split() + pruned_on_main = [h for h in old_main if h in pruned] + if len(new_main) != len(old_main) - len(pruned_on_main): + fails.append(f"main count: new {len(new_main)} != old {len(old_main)} - pruned {len(pruned_on_main)}") + old_ts = [ln for h, ln in zip(old_main, C.git(["log", "--reverse", "--format=%at %ct", "main"], old_repo).splitlines()) if h not in pruned] + new_ts = C.git(["log", "--reverse", "--format=%at %ct", "main"], new_repo).splitlines() + if old_ts != new_ts: + fails.append("author/committer timestamp sequences differ (after removing the pruned commits)") + # the pruned set must be exactly the commits whose every change was a purge path (derived, not "1") + expected_pruned = set() + for h in pruned | set(old_main): + raw = C.git(["show", "--raw", "--no-renames", "--format=", h], old_repo).splitlines() + changed = [ln.partition("\t")[2] for ln in raw if ln.startswith(":")] + if changed and all(C.under_purge(pth, prefixes, globs) for pth in changed): + expected_pruned.add(h) + if pruned != expected_pruned: + fails.append(f"pruned set {sorted(x[:9] for x in pruned)} != commits whose every change is a purge path " + f"{sorted(x[:9] for x in expected_pruned)}") + empties = [h for h in old_main if len(old_commit(h)["parents"]) == 1 and old_commit(old_commit(h)["parents"][0])["tree"] == old_commit(h)["tree"]] + for h in empties: + if cmap.get(h, C.ZEROS) == C.ZEROS: + fails.append(f"pre-existing empty commit {h[:9]} was pruned (use --prune-empty auto, never always)") + print(f"verify_rewrite: {n_pairs} pairs checked ({len(pruned)} pruned = the {len(expected_pruned)} purge-only commits, {len(pruned_on_main)} on main); {n_changed_blobs} changed " + f"blobs re-derived, {n_removed} path removals justified; empty commits on old main: {len(empties)}; " + f"new main {len(new_main)} commits; failures {len(fails)}") + for f in fails[:60]: + print(" FAIL " + f) + print("verify_rewrite: " + ("PASS" if not fails else "FAIL")) + return 1 if fails else 0 + + +if __name__ == "__main__": + sys.exit(main(sys.argv[1:]))