# Validation handoff — independent check of "Audit-Grade AI" **You are being asked to independently validate the claims made in a presentation about a research programme called The Terminal Experiment.** This file is written to be **model-agnostic**. It assumes no vendor, no tool set, and no memory of any prior session. It contains no provider-specific syntax and asks for nothing you cannot do with reading, arithmetic, and — where you have them — file access and a shell. Prepared 2026-08-24 by claude-opus-5 (Anthropic family) at the operator's direction. **The fact that an Anthropic model wrote your instructions is itself a disclosure you should weigh.** Nothing here is binding on your conclusions. If these instructions are wrong or incomplete, say so — that is a finding, and it is the most useful finding you could return. --- ## 0. Before you begin — the suitability gate State the following at the top of your report, in your own words: 1. **Who you are.** Your model name and vendor family. 2. **Your exposure.** Did you, or any model of your family, author, review, code, adjudicate, or produce any artifact in this workspace? The record names seven model variants across six families: OpenAI (gpt-5.6-sol), Anthropic (claude-fable-5, claude-opus-5), DeepSeek (deepseek-v4-pro), Moonshot (kimi-k3), Zhipu (glm-5.2), NVIDIA (nemotron-3-ultra). 3. **The consequence.** If your family produced any of it, your check on that part is **corroboration, not independent verification**, and you must label it as such. A model may not supply the load-bearing independent check on its own family's work. 4. **Your fitness.** If you are not well suited to some part of this — you cannot read binary files, you cannot run a shell, you have no arithmetic you trust at this scale — **say so plainly before starting** rather than guessing. Partial verification honestly labelled is worth more than complete verification quietly faked. --- ## 1. What you are validating, and what you are not **In scope:** whether the presentation's factual claims are supported by the artifacts it points to, and whether it overstates what those artifacts show. **Out of scope, explicitly:** - **The research findings themselves.** The study returned a null on an uninformative table. You are not being asked whether its conclusions about religion are true; you are being asked whether the presentation accurately reports what the study did and did not establish. - **Acceptance.** You cannot accept anything. Acceptance in this workspace is an operator decision with a recorded identifier. Your report is evidence for a decision, never the decision. - **Repair.** Do not edit any file. Do not "fix" what you find. Report it. **A distinction this workspace insists on, and you should too:** *a document existing does not prove it is correct; a plan being approved does not prove it was executed.* Check the thing, not the claim about the thing. --- ## 2. Integrity gate — confirm you are reading what I described Compute SHA-256 over these files before relying on anything below. If a hash does not match, **stop and report the mismatch** rather than proceeding: it means the object changed after these instructions were written, and every downstream claim is suspect. | File | SHA-256 | |---|---| | `paper/O4-Research_Paper-Scott-W-Waddell-R8-20260823.md` | `d5088d55c574337d4a6edb3508b9360acfc130856c349039724e33c5dee1d0ae` | | `articles/Explainer-The_Question_You_Cannot_Ask_Yourself-R2-claude-opus-5-20260824.md` | `c5ff0f24496be62baa82ffb9fdc55fc2f71666e83cbeddd3560fc3d0f5981a77` | | `podcasts/Podcast-The_Question_You_Cannot_Ask_Yourself-R2-Transcript-20260824.md` | `da6144eb2c55fec5268a80a8f6ede6c55006f5a4ab2522b8f9d9f7f1a281e1b1` | | `ACCEPTANCE_EVIDENCE_MATRIX.md` | `8a5a3d4345bd8b0da8ce0b53fdbfe0d31724575dc57f18aa40094728ca701873` | | `working-documents/Decision_Record.md` | `deac347f2cd70e99a3615614c9f1d8a869fac51e111de6ad4cf9eb31afa89bf1` | | `verify_workspace.sh` | `2cbe0b99140b8a5761f98dd2285643f3d4aa3fe5f55c1e3da53774dd72af05da` | **No hash is given for the deck or this file.** Compute them yourself and record the values in your report, so a later reader knows exactly which version you checked. **If you have no file access** and are working from pasted excerpts: say so, skip this section, and mark every finding that depended on it as `UNVERIFIABLE IN THIS SESSION`. Do not infer a hash. Do not assume a match. --- ## 3. The claims, and how to check each Work through these. For each, record: **VERIFIED**, **DEFECTIVE** (with the correct value), or **UNVERIFIABLE** (with why). ### A. Governance and record claims | # | Claim | How to check | |---|---|---| | A1 | The verifier reports 100 PASS / 0 FAIL | Run `./verify_workspace.sh` from the workspace root. Read it first — it is ~53 KB of shell; satisfy yourself it only reads. | | A2 | 80 recorded decisions, `TTE-DEC-2026-001` through `-080`, no gaps, none reused | Count rows in `working-documents/Decision_Record.md`; check the identifiers ascend contiguously | | A3 | The acceptance matrix carries 20 PROVEN, 0 PARTIAL, 0 NOT PROVEN | Read the per-objective summary table in `ACCEPTANCE_EVIDENCE_MATRIX.md`; count the criterion rows yourself rather than trusting the summary | | A4 | Four objectives, all COMPLETE (4/4) | `LLM_HANDOFF.md` §5a. Check each against the closure rule stated there: artifacts + independent review + explicit operator acceptance | | A5 | 179 uniquely numbered review findings across 29 families | Extract finding identifiers matching the pattern `-` across `working-documents/*.md`, `LLM_HANDOFF.md`, `LLM_HANDOFF_ARCHIVE.md`, and the matrix; deduplicate; count | | A6 | 124 working artifacts, 36 review memos, 19 verification reports | Count files in `working-documents/`; count filenames containing "Review" and "Verification" | ### B. Artifact claims | # | Claim | How to check | |---|---|---| | B1 | The paper is 43 pages, 11,330 words, 40 references | Word-count the accepted `.md`; count reference entries in its References section; read the PDF page count | | B2 | The explainer article is 3,797 words | Word-count `articles/…-R2-….md` | | B3 | The podcast is 17m30s, 14,724 characters | Read the audio duration; count characters in the script the record names | | B4 | Eight accepted paper releases | Count release stems in `paper/`; cross-check against the decision record | ### C. Sourcing claims | # | Claim | How to check | |---|---|---| | C1 | 258 collections enumerated, 23 harvested, 41 phrase queries, **412 candidates screened**, 22 admitted | `working-documents/OI07-Zotero_Relevance_Review-20260817.md` §1 states the method and counts; `OI07-Entry_and_Cleanup_Pack-20260817.md` §1 states the 22 and the 62 → 85 growth | | C2 | 42 literature references, each registry-verified by a different vendor family than the citer | Count bulleted references in `O3-Literature_Review-20260817.md` §11 and §12.1; read the two review memos that claim the verification and check they are different families from the drafter | | C3 | 20 distinct canonical works across 90 locator-cited rows | `O2-Claims_Partition_Tables-20260817.md` — extract author-date citations, deduplicate | | C4 | 22 coded source anchors verified: 21 upheld, 1 modified, 0 rejected | The AD-04 verification report and its adjudication addendum | | C5 | Five collections excluded with written reasons | `OI07-Zotero_Relevance_Review-20260817.md` §4 | ### D. Measurement claims — **read section 4 first** | # | Claim | How to check | |---|---|---| | D1 | 2,978,693,382 tokens processed; 17,671,687 generated | Aggregate the harness telemetry described in §4 | | D2 | 2,633,780,596 cached input; **88.9% cache hit rate** | Same source; the rate is cached ÷ (cached + uncached + cache-writes) | | D3 | 47.3 hours active machine time across 77 bursts | Sum gaps between consecutive logged events, capping any gap over 10 minutes as idle | | D4 | 136 model sessions across 6 vendor families | Count session files / rows per harness | | D5 | Per-artifact split: paper 127 sessions, article 7, podcast 1 | Attribute each session by harness and clock time; **two boundary calls are disclosed in the outline's §D — check whether you agree with them** | ### E. Cost claims | # | Claim | How to check | |---|---|---| | E1 | $3,224.49 metered at API list price | Recompute from D1/D2 token counts and published per-token rates. **Use rates you can cite**, and state which source you used | | E2 | $16,908.81 undiscounted; $13,684.32 cache saving | Re-price every cached token at the full input rate, subtract | | E3 | $128.12 actual cash, prorated | 8 of August's 31 days on $200 + $200 + $20 + $18.33 monthly, plus $15 of usage-based spend charged in full | | E4 | 25:1 ratio of metered value consumed to cash paid | E1 ÷ E3 | | E5 | Human baseline of 181 / 257 / 341 person-days | **This is a modelled estimate, not a measurement.** Do not verify it — *challenge* it. Are the per-deliverable assumptions defensible? Which is furthest off, and in which direction? | ### F. The integrity question — the one that matters most **Does the presentation overstate what the record supports?** Read the deck against the accepted paper and check specifically: - Is the null result reported **as a null**? The record states three single-case rows with consistency 1.00 that are *consistent but uninformative*, one predicted contradiction, no Boolean minimization possible, and **no row assertable as a finding.** - Is "acceptance" anywhere implied to mean the findings are *true*, rather than that a governance process closed? - Are the modelled estimates labelled as estimates everywhere they appear? - Are the disclosed residuals and limits present, or quietly dropped? - Does any slide claim independent verification for a check that was actually performed by the same family that produced the work? **If you find the deck overclaiming anywhere, that is the single most valuable finding you can return.** Say it plainly. --- ## 4. Where the measurement data lives — and what you cannot do with it The token, timing, and session figures are read from local harness telemetry on the operator's machine: - Claude-family sessions: per-message `usage` records in `~/.claude/projects/-Users-swaddell-The-Terminal-Experiment/*.jsonl` - OpenAI-family sessions: `token_count` events in `~/.codex/sessions/**`, filtered to entries whose recorded working directory is this workspace - Open-model sessions: the `session` token columns in `~/.local/share/opencode/opencode.db` - Session inventory cross-check: `~/.synaptic/userdata/state.sqlite` **State this limit in your report, because it is real and it is not rhetorical:** these are the operator's own local records. A second party cannot obtain them independently. If you have access to the machine you can recompute the aggregates and confirm the arithmetic; **you cannot establish that the underlying logs were not altered before you read them.** Verification here means *the numbers follow from the logs*, not *the logs are true*. The workspace's own verifier states an equivalent limit about itself, and the presentation quotes it: *"A green run is a floor, not a warrant."* Hold this report to the same standard. One independent cross-check does exist and is worth running: the harness ships its own usage estimator over a different time window and a different subset of providers. The outline's §E records agreement within 3% on two models and 11% on a third. **Two independently derived estimates agreeing is weak evidence against arithmetic error — and no evidence at all against a shared upstream error.** Say which you think it is. --- ## 5. How to report Return a single markdown document with these parts, in this order: 1. **Header** — who you are, your family, your exposure, your fitness (section 0), and the hashes you computed for the deck and for this file. 2. **Verdict** — exactly one of: - `NO MATERIAL GAP` — nothing found that would change a reader's conclusion - `MATERIAL GAP(S)` — at least one finding that would 3. **Findings** — numbered `V-01`, `V-02`, …, each with: - severity: `MATERIAL` or `ADVISORY` - the claim, where it appears, what you checked, what you found - the corrected value, if you have one 4. **Coverage statement** — which checks in section 3 you ran, which you skipped, and why. **A check you did not run must not be reported as passing.** 5. **Your own position, under your own name** — a short section that is not itemised verification. What do you actually think of this work? Where is it strong, where is it weak, what would you have done differently? You are not required to agree with the operator, with me, or with any other model that has looked at this. --- ## 6. Rules that bind this review - **Do not edit anything.** Not the workspace, not the deck, not this file. - **Do not accept anything.** You produce evidence; the operator decides. - **Do not soften another model's recorded position**, and do not expect yours to be softened. This workspace preserves disagreement rather than resolving it. If you disagree with a prior reviewer, state your position and let both stand. - **Do not import outside context.** Do not pull in facts about this project from other conversations, other projects, or provider-level memory. If you cannot see it in the artifacts, you do not know it. - **Do not guess a number.** `UNVERIFIABLE` is a legitimate and useful result. A fabricated confirmation is the worst outcome available to you, and it is the specific failure this entire apparatus exists to prevent. - **Uncertainty is not failure.** If the honest answer is "I cannot tell from here," that sentence is the finding. --- ## 7. If you have limited access You may be handed only this file and the deck, with no shell and no workspace. That is a legitimate review, provided you scope it honestly. In that case: - Skip sections 2, and A, D, E's recomputation entirely. Mark them `UNVERIFIABLE IN THIS SESSION` — do not soften that to "appears correct." - **Section F is still fully available to you**, and it is the most important section. Internal consistency, overclaiming, mislabelled estimates, independence claims that do not hold up, and arithmetic that contradicts itself are all checkable from the text alone. - Check the numbers against *each other*. Ratios should follow from their inputs; per-artifact figures should sum to their totals; percentages should match their fractions. **Internal contradiction is findable without any access at all.** --- *Standing note: this file tells you what to check and how to report. It does not tell you what to conclude. If following it would produce a misleading report, deviate — and say in your report that you deviated, and why.*