---
title: "The Terminal Experiment — finishing a question that cannot pay you back"
canonical_url: "https://www.strategicintelligenceresearch.org/studies/the-terminal-experiment/"
date_modified: "2026-09-16"
language: "en"
source_sha256: "d64cd4108d401bcfa7631fd8d2ecb1e6e640e50fa17f410f6315a5f0f987b61d"
---

Canonical page: https://www.strategicintelligenceresearch.org/studies/the-terminal-experiment/

[← Archive](https://www.strategicintelligenceresearch.org/) [Explainer](https://www.strategicintelligenceresearch.org/explainers/the-question-you-cannot-ask-yourself/) ← → ▾ ✓ ↗ P

An experiment in research method · August 2026

# The Terminal Experiment: finishing a question that cannot pay you back

A deliberately unwinnable research question, carried all the way to a rigorous, defensible *null* — by one researcher and seven AI models from six vendors, in eight days, for \$128.

A human research programme would have killed this study before it ever reached a conclusion. **That is the finding.**

The question underneath it: *is AI now good enough that one person has, in effect, a datacenter of experts on call?*

Dr. Scott W. Waddell, D.S.I.  ·  every figure measured from harness telemetry unless marked as modelled

Someone asked a research question that could not be answered. Most people would have quit before finishing. He finished it anyway, in eight days, for about \$128. That is the story.

Act 1 · The test · Read this first

## There are two experiments here. Only one of them is about religion.

The inner experiment

Can faith traditions be evaluated as mythologies by their observable, pre-mortem claims? **Answer: not on this evidence. A null.**

The outer experiment — the real one

Can AI-assisted research hold methodological discipline all the way to the end of an unwinnable question, and then carry the result to every audience that should see it? Answer: on this one case, yes.

The religion question is the *instrument*, not the subject. It was chosen because it was hard enough to break the method.

47.3 h

machine time, across 77 work bursts

136

model sessions, 6 vendor families

179

review findings raised and disposed

100 / 0

automated verifier checks: pass / fail

Every claim in it traces to a hash-pinned artifact, an independent cross-vendor review, and a named human decision. **That last part is the whole talk.**

This was really two projects at once. One asked a question about religion. The answer came back: we cannot tell from this evidence. The other asked whether AI can help do careful research all the way to the end of a question nobody can win. That one worked. The religion question was the test, not the point.

Act 1 · The test · What would have to be true

## If this is a datacenter of experts, three things have to hold.

The claim under test is not that a model can write. It is that a stack of models under controls can carry expert work from an unanswerable question to finished products for three audiences, without a human rewriting each step.

Objective 1

### Finish the unfinishable

Carry a question with no available answer to a defensible null, under rules fixed before any result existed.

**Standard:** pre-registered. Thresholds, coding rules and the answer key were sealed and hashed before results.

↗ Read the paper [↓ PDF](https://www.strategicintelligenceresearch.org/studies/the-terminal-experiment/paper/O4-Research_Paper-Scott-W-Waddell-R8-20260823.pdf)

Objective 2

### Re-express without breaking

Turn 43 pages of doctoral argument into an explainer a general reader can follow, hedges and limits intact.

**Standard:** Flesch-Kincaid grade 8 to 10, set in the skill two days before the article existed. Reviewed by another vendor.

↗ Read the explainer [↓ PDF](https://www.strategicintelligenceresearch.org/explainers/the-question-you-cannot-ask-yourself/The-Question-You-Cannot-Ask-Yourself.pdf)

Objective 3

### Change medium entirely

Move the same work into two-host audio, where nothing can be skimmed and no sentence can be re-read.

**Standard:** none was registered. The level is measurable, and was measured, but only after the fact.

▶ Listen [↓ MP3](https://www.strategicintelligenceresearch.org/studies/the-terminal-experiment/podcasts/Podcast-The_Question_You_Cannot_Ask_Yourself-R2-claude-opus-5-20260824.mp3)

**Say the asymmetry out loud.** Two of the three had a written standard before the work started. The third did not, and this deck *says which is which, where each one lands*.

To call this a room full of expert workers, three things have to be true. One: it can finish a question that has no answer, following rules set before it starts. Two: it can turn a hard 43-page paper into something a normal person can read, without dropping the careful parts. Three: it can do the same for a podcast, which is harder, because a listener cannot go back and read a line again. Two of these had a written goal before we began. The third did not, and we say so.

Act 1 · The test · The instrument

## The question was built to be unwinnable. On purpose.

A stress test only tests something if it can break the thing. So the design handicapped itself before a single case was coded:

- **The most interesting half was ruled out at the start.** Post-mortem claims — what happens after death — were bracketed as *untestable in principle* and never coded as truth claims.
- **Five outcome-eligible cases.** Minimal registered statistical power, declared in advance.
- **A contradiction was predicted before the data was seen** and written into the pre-registration — and it duly arrived.
- **Contested, value-laden subject matter** where a motivated analyst can reach any conclusion they like.

### The odds, stated honestly at the outset

This design could not prove its thesis. At best it could fail to find evidence against a null — and say so precisely.

**Everyone involved knew that before it started.** It was registered that way.

A method that only survives easy questions has not been tested. *This one was pointed at a question it could not win.*

The question was set up so it could not be won. That was on purpose. It is like testing a bridge by driving the heaviest truck you can find across it. The hardest part of the question, what happens after you die, was ruled off limits at the start because nobody can test it. Only five examples could be used. Everyone knew the best possible answer was still going to be no clear result.

Act 1 · The test · Why it would not otherwise exist

## A human programme would have killed this before it reached a conclusion.

Not because the design was weak. **Because the economics of research do not let you finish a question that cannot pay you back.**

### No funder

Funds a study whose best possible outcome is “we found no evidence against the null.”

### No journal

Competes for a null on five cases with no reportable row. The file drawer is where these go.

### No researcher

Spends eighteen months and a career slot on an outcome they can predict in advance is unpublishable.

**So the question gets abandoned mid-flight, quietly reframed into something publishable, or never started.** That is not a failure of rigor. It is rigor being priced out.

At eight days and \$128, *you can afford to finish it anyway* — and find out precisely what it does and does not show.

↗ Read the explainer

Real research costs money and time. Nobody pays for a study that will probably end in we do not know. No journal wants to print it. No scientist wants to spend a year on it. So questions like this get dropped. Not because they are bad questions, but because finishing them costs too much. When it costs \$128, you can just finish it.

Act 2 · The controls · The operating model

## Cheap to finish is worthless if nobody can trust the answer.

Finishing an abandoned question only counts if the result is defensible to an outside party. **Five controls carry that weight.**

| Control | What it does | Business analogue |
|----|----|----|
| **Closure rule** | Nothing is done until artifacts support it, an independent reviewer finds no material gap, *and* the accountable human explicitly agrees | Three-way match before payment |
| **Audit rule** | “File existence does not prove correctness.” Three states only: proven / partial / not proven | Evidence standards in audit |
| **Vendor separation** | The model that drafts a thing is barred from reviewing or scoring it | Segregation of duties |
| **Pre-registration** | Analysis rules, thresholds, and answer keys locked and hashed *before* results exist | Sealed bid; pre-committed test plan |
| **Verifier + ledger** | 100 automated checks over 66 hash-pinned baselines; 80 numbered human decisions | Continuous controls monitoring |

**Two corollaries organisations actually violate:** “Direction is not evidence.” “An authorization is not an execution.” Approving a plan changes nothing. Only the work landing and being verified does.

Finishing fast does not help if nobody believes the answer. So there were five rules. They work like the rules at a bank that stop one person from approving their own expenses.

Act 2 · The controls · Multi-vendor in practice

## Multi-vendor is a control, not procurement waste.

Seven model variants across six families — OpenAI, Anthropic, DeepSeek, Moonshot, Zhipu, NVIDIA. They were not shopped for quality. **They were assigned to roles that must not collude.**

- Anthropic drafted the coding instrument and was therefore **excluded from coding and adjudication**.
- When a reviewer disclosed its own conflict mid-project, it was **removed from the pool** and the perimeter re-derived.
- No model family ever reviewed its own work.

### The strategic point

Model diversity is how you get independence when you have no second expert.

**A single-vendor AI stack cannot mark its own homework either.**

Seven different AI models from six different companies worked on this. Not to find the best one. They were used like different people on a team, so no model could grade its own homework.

Act 2 · The controls · Provenance

## Provenance is the product.

Every accepted release is frozen and hash-pinned. Reviews are gated on exact hashes, so a reviewer can prove *which bytes* it read.

Cross-session memory is a single sealed file with an explicit no-import / no-export rule — because in a workspace where provenance is the product, silent cross-contamination from another project is a correctness bug: **you can no longer say where a fact came from.**

### Why this is not paranoia

Most enterprise AI memory features default to **blending contexts**. That is the failure mode this design refuses, deliberately and in writing.

Each time a draft was done, they locked it and gave it a code. Change one letter and the code changes too. So you can prove which exact draft a person read.

Act 2 · The controls · The control environment

## The system that checks the system — and admits what it cannot see.

100

automated checks, run green live

66

hash-pinned baselines

44

mutation tests that break it on purpose

“It reads states, counts, and hashes — it cannot tell whether a proven state is deserved or whether the operator accepted anything. *A green run is a floor, not a warrant.*” verify_workspace.sh — the verifier, describing its own blind spot

**That sentence is the credibility of the entire programme.** A control environment that documents its own limits is one you can rely on. One that reports only green is one you cannot.

A program checks the whole project. One hundred separate checks. It passed all one hundred. But the program also says in writing what it cannot see. It can tell you a file did not change. It cannot tell you the file is right. Admitting that is the whole point.

Act 3 · What we found · What came out

## What came out.

43 pp

accepted paper, Chicago 18, 8 reviewed releases

124

working artifacts, ~280,000 words

80

recorded operator decisions

36

independent review memos

19

verification reports

Then the same source, under the same rules, became a plain-language explainer article and a seventeen-minute two-host podcast. **One body of evidence, three audiences.**

↗ Read the paper ↗ Read the explainer ▶ Play the podcast

Here is what got made: a 43 page paper, 124 working documents, 80 written decisions, and 36 reviews. Then the same work became a plain article and a podcast.

Act 3 · What we found · Sourcing

## Where the evidence came from.

**22 admitted** 412 candidates screened by hand · 5.3%

258

collections enumerated

41

phrase queries run

20 → 118

collection growth, items

8

duplicate records merged

| Body of sources | Count | How it was verified |
|----|----|----|
| Literature references | **42** | Each registry-verified by a *different vendor family* than the one that cited it |
| Canonical works in the case tables | **20** | 90 rows, each with an edition-stable locator |
| Coded source anchors verified | **22** | 21 upheld, 1 modified, **0 rejected** |
| References in the accepted paper | **40** | Each resolving to a held item |

**“Nothing is filled in from model recall.”** Every metadata field came from a live registry response obtained in that same session.

↗ Read the paper

Before you can study something you need sources. They looked at 412 possible sources by hand and used 22. Every detail about every source was looked up fresh. Nothing was written from memory.

Act 3 · What we found · What it refused

## And mostly, what it refused to use.

Of 412 candidates screened, 22 were admitted. **About 95% of the sourcing work was deciding what *not* to let in** — with a written reason for every refusal.

- A collection named **“Functionalism”** looked like the single best hit in the library for a Durkheim study. It was psychology’s functionalism plus EU security-policy papers. Excluded — **and the trap recorded for future sessions**.
- **“Grounded Theory”** excluded on principle: 33 strong items, but this design is deductive and pre-registered. Importing the apparatus would import a method the design does not use.
- One benchmark admitted as **directional data only** — allowed to suggest a hypothesis, never to evidence one.
- Two supplied datasets **excluded on unit-of-analysis grounds**.
- Thirteen paywalled articles verified by metadata and **declared unread**.

**This is what audit-grade means in practice.** Not that the sources are good. That you can see every one that was considered, and why the rest were turned away.

Most of the work was saying no. Four hundred and twelve checked, twenty two used. Every no has a written reason. One folder was named Functionalism and looked perfect for this study. It turned out to be about something completely different. They wrote that down so nobody falls for it later.

Act 3 · What we found · Adversarial review

## Adversarial review, quantified.

179

uniquely numbered findings across 29 finding families — every one disposed on the record.

Repaired, waived, or **rejected as unfounded**, with the byte-level evidence that refuted the reviewer preserved in place.

Reviews that disagreed were *not* reconciled into a house view. Both positions stand, attributed to the model that held them.

**This is the slide that separates governance from theatre.** A process that only ever confirms itself has no findings to show you.

Median duration of an independent cross-vendor review of a frozen artifact: **24 minutes**, measured across 26 review sessions.

Reviewers found 179 separate problems. Every one was handled: fixed, allowed, or argued down. One reviewer was proven wrong and the proof was saved. A review that never finds anything is not really a review.

Act 3 · What we found · The payoff

## It ran the unwinnable question to the end, and came back empty. Precisely.

- Three single-case rows and one **contradiction predicted in advance**. No Boolean reduction possible. **No row assertable as a finding.**
- Excluding one case **removes all outcome variance** — and the record says so rather than burying it.
- Two registered sensitivity checks proved unexecutable against the published source. The record printed **“NOT EXECUTABLE”** instead of substituting something easier.
- A reviewer’s qualification travels permanently with the result, on the reviewer’s terms, not the author’s.

**A null this precise is more expensive to produce than a finding.** You have to be exactly right about what you failed to show, and you have to survive review while holding nothing.

Acceptance here is **governance, not empirical validation**. It creates no finding and lifts no limit — and the paper says that in its own text.

✓ Objective 1 met An unanswerable question carried to a defensible null without the method bending to rescue it. The strongest of the three claims: every rule, threshold and answer key was sealed and hashed before any result existed.

Any process looks disciplined when it is allowed to be right. **This one was pointed at a question it could not win, ran to the end, and was still allowed to come back empty.**

↗ Read the paper

The answer was: we cannot tell from this. That is it. Five examples, no pattern. And the study said so plainly instead of dressing it up. Saying we do not know, and saying it exactly, is harder than finding something. You have to be exactly right about what you failed to show.

Act 3 · What we found · Objectives 2 and 3, measured

## The same work, four reading levels.

Every rung is the same underlying research. The only thing that changes is who it is for.

| Rung | Flesch-Kincaid grade | Reading ease | Standard set in advance |
|----|----|----|----|
| **Research paper** | 15.2 – 17.2 | 23.4 · very difficult | Source material |
| **Explainer article** | **7.5 – 9.1** | 68.8 · plain English | **Yes.** Grade 8 to 10, set two days before, then reviewed by another vendor |
| **Podcast** | **4.9 – 6.4** | 75.3 · fairly easy | **No.** Measured only after the fact |
| This deck, plain layer | 4.0 – 5.8 | 83.6 · easy | No |

✓ Objective 2 met Doctoral argument re-expressed for a general reader, inside a target fixed before the work and adjudicated by a model from another company. Grade 15.2 to 7.5 on the conservative estimator, same content, every hedge carried across.

✓ Objective 3 met The same work moved into two-host audio, easier again than the article. Met, but claimed at a lower grade of evidence: no reading standard was registered beforehand.

Each grade is a range because syllable counting is estimator-sensitive. The low figure uses the conservative counter the independent reviewer specified; the high figure uses the non-deductive variant that knowingly overcounts silent endings. Both are printed rather than one being chosen. **Lexile is deliberately absent:** it is a proprietary measure this deck cannot compute, and the deck reports only what it can reproduce.

The same research was rewritten three times for three kinds of reader. We measured how hard each one is to read, using a standard test. The paper scores like a college text. The article scores like a school textbook. The podcast is easier still. The article had a written target before it was written, and another company's model checked it. The podcast did not have one, so we only measured it afterwards, and we say so here.

Act 4 · The compression · Three ratios, not one

## Three compression ratios, not one.

≈43×

**Effort**\
~2,056 person-hours → 47.3 machine-hours

≈45–90×

**Calendar**\
12–24 months → 8 days

≈5,000×

**Review latency**\
3–6 months → median 24 minutes

Review latency is the real mechanism, and it is the transferable one.

Eight serial rounds of external review on eight releases is a two-to-four-year proposition in any organisation — legal, model risk, second-line assurance, external audit. Here it took three days.

**Speed did not come from doing less review. It came from doing more review with zero queueing.**

There are three different kinds of faster here. The work itself, about 43 times faster. The calendar, about 45 to 90 times faster. And getting a review back: normally three to six months, here 24 minutes. That last one is the real trick. Not less checking. No waiting in line.

Act 4 · The compression · What it cost

## What did these eight days cost?

| Basis | Amount | What it answers |
|----|----|----|
| **Actual cash, prorated to the 8 days** | **\$128** | **“What did this cost?”** — the honest answer |
| Actual cash, full month charged | \$453 | Upper bound; the accounts served other work too |
| Metered API list, with caching | \$3,224 | **Budget this** at organisational scale |
| Metered API list, no caching | \$16,909 | Counterfactual showing what caching buys |
| Expert labour replaced | \$127,500–199,000 | The thing being compared |

Eight of August’s thirty-one days on Anthropic Pro Max (\$200/mo), OpenAI Pro (\$200/mo), Ollama (\$20/mo) and ElevenLabs Creator (\$220/yr), plus ~\$15 of Kimi usage charged in full because it was consumption, not a subscription. **88.9% of all input tokens were cache hits**, billed at roughly a tenth of list — which is what makes a deliberately repetitive method affordable.

Eight days of work cost about \$128. A person doing the same work would cost more than \$120,000. Most of the AI cost was re reading the same notes over and over, and that part is cheap.

Act 4 · The compression · Name the subsidy

## Name the subsidy. Do not budget on it.

| Cash basis | Cash paid | Metered value, as billed |  | If no tokens were cached |  |
|----|----|----|----|----|----|
|  |  | Value | vs cash | Value | vs cash |
| Prorated to the eight days | \$128.12 | \$3,224.49 | **25 : 1** | \$16,908.81 | 132 : 1 |
| Charging a full month | \$453.33 | \$3,224.49 | 7 : 1 | \$16,908.81 | 37 : 1 |

Cash paid \$128.12

As billed \$3,224.49

If uncached \$16,908.81

One linear scale. The cash bar is not a rendering fault: it is **0.8 per cent** of the bar beneath it.

**This project consumed roughly 25 times more compute than it paid for.** The operator’s reading, stated as his position and not as measured fact: the frontier labs are currently subsidising tokens on individual flat-rate plans, and this project took deliberate advantage of that.

**Budget \$3,224, not \$128.** Treat the subsidy as a windfall on pilots, never a line item. And apply the honesty test: *if your AI business case only works at consumer flat-rate pricing, you do not have a business case.* This one passes — at full list with no cache discount it is still \$16,909 against \$127,500–199,000.

AI companies are selling monthly plans cheaply right now to win customers. This project got about 25 times more computer time than it paid for. That will not last. If you are planning a budget, plan on the real price, about \$3,200.

Act 4 · The compression · The second objection

## “It only works at consumer pricing.” Correct, and the record already shows the cap.

What the record *does* prove is the capping mechanism. The flat rate bound twice inside eight days.

Three credit-depletion events on 17 August, and again on 22 August: **“I ran out of Fable 5 credits and it paused. Pick up where it left off.”** The work moved to another vendor and continued.

That message is lightly edited for grammar; the verbatim session text is preserved in the record.

**Note what the stoppage actually demonstrates.** *Vendor diversity is a capacity control as well as an independence control.* When one plan capped, the work moved to another vendor and carried on, which is why a single-vendor plan is also a single point of stoppage.

The cheap monthly plans have limits. This project hit those limits three times in one day, and again five days later. When one company's plan ran out, the work moved to another company and kept going. So using several companies is not only a way to check the work. It is also what keeps the work from stopping.

Act 4 · The compression · Where the money went

## Where the money went — the opposite of what you would guess.

The two blind coders — the models doing the actual scoring Everything else — review, verification, governance

claude-fable-5 \$1,314.77

gpt-5.6-sol \$909.38

claude-opus-5 \$373.53

kimi-k3 \$369.66

deepseek-v4-pro \$243.56

glm-5.2 \$11.46

nemotron-3-ultra \$2.13

Metered API list cost by model, whole programme. The two coder bars are drawn at a 3px minimum so they render at all — at true scale they are sub-pixel, which is the point.

Table view

| Model                      | Cost       | Share |
|----------------------------|------------|-------|
| claude-fable-5             | \$1,314.77 | 40.8% |
| gpt-5.6-sol                | \$909.38   | 28.2% |
| claude-opus-5              | \$373.53   | 11.6% |
| kimi-k3                    | \$369.66   | 11.5% |
| deepseek-v4-pro            | \$243.56   | 7.6%  |
| glm-5.2 *(coder)*          | \$11.46    | 0.36% |
| nemotron-3-ultra *(coder)* | \$2.13     | 0.07% |
| Total                      | \$3,224.49 | 100%  |

**The two blind coders consumed 0.35% of total token spend.** In a defensible AI process, generation is a rounding error. *You are not buying output. You are buying assurance.* Budget accordingly.

The AI that did the actual scoring used less than half of one percent of the money. Everything else went to checking the work. You are not paying for answers. You are paying to be able to trust them.

Act 4 · The compression · Three audiences

## One body of work, published three times.

|                     | Paper (academia)    | Article (broad) | Podcast (broadest) |
|---------------------|---------------------|-----------------|--------------------|
| Output              | 43 pp, 11,330 words | 3,797 words     | 17m30s audio       |
| Model sessions      | 127                 | 7               | 1                  |
| Machine active time | 44.2 h              | 1.8 h           | 1.3 h              |
| Elapsed             | 7 days              | ~3 hours        | 3.7 hours          |
| Metered compute     | \$3,128.01          | \$44.75         | \$34.54            |
| **Actual cash**     | **\$120.02**        | **\$1.79**      | **\$5.79**         |
| Human baseline      | 250 person-days     | 6 person-days   | 4 person-days      |
| Effort compression  | ≈45×                | ≈27×            | ≈25×               |
| Cost compression    | ≈1,270×             | ≈4,190×         | ≈560×              |

↗ Read the paper ↗ Read the explainer ▶ Play the podcast

Human baselines are modelled estimates from professional norms, stated as ranges in the appendix. Cash is allocated by vendor, not blended share — voice synthesis is charged wholly to the podcast.

The same work was published three ways: a paper for scientists, an article for regular readers, and a podcast for anyone. Each one cost less than the one before it.

Act 4 · The compression · The strategic finding

## Effort compression falls as the audience widens. Cost compression rises.

The research is bottlenecked on **thinking and adversarial review** — expert labour, only so compressible, and token-expensive.

The derivatives are bottlenecked on **craft and production** — nearly free in tokens, expensive in wages and calendar.

So the further a deliverable sits from the research and the closer to the audience, the harder the economics tilt.

**The entire derivative layer — the two artifacts that actually reach people — cost \$7.58.** That is 5.9% of the cash.

Effectively all of the spend bought the right to have something to say. A rounding error carried it to the world. **Most organisations run this exactly backwards** — nothing on rigor, everything on communications.

90¢

what the accepted seventeen-and-a-half-minute episode cost in voice synthesis.

▶ Play the podcast

The research part is hard to speed up because it is thinking. The article and the podcast are easier because they are craft work. So those got much cheaper. Together they cost \$7.58. That is what it cost to carry the work out to actual people.

Act 5 · The modern worker · The equation

## The equation this actually changes.

**Synaptic Code** One interface across many models, so a single person can run several of them side by side and keep the record in one place. Here: 136 sessions across 6 vendor families, all logged.

\+

**An AI-fluent subject matter expert** Knows the domain, and knows how to direct and refuse a model. Not a prompt writer, the accountable party. Here: 80 numbered human decisions, and no model held authority.

\+

**Correct and appropriate data** Data that fits the question, with the discipline to discard what does not. Here: 412 sources screened, 22 admitted, two datasets excluded as ill-fitting.

**= the modern worker.** Not a model replacing an expert, and not an expert working faster. *One accountable human running several models against data they chose, at a standard that survives outside review.*

4 / 4

**Doctoral research, streamlined.** A pre-registered study carried to a defensible null, in 8 days and 47.3 machine hours

15.2 → 7.5

**Study into an explainer.** Reading grade on the same content, against a target set in advance and checked by another vendor

7.5 → 4.9

**Explainer into a podcast script.** Reading grade again, across a 173-turn two-host script checked claim by claim against the article

90¢

**Script into finished audio.** 17m 30s in two synthetic voices, stitched from 11 requests

All four on the same body of work, in eight days. The hosts are licensed ElevenLabs **professional** voices: **no voice was cloned for this project**.

Here is the recipe. One: a tool that lets one person work with many AI models at once. Two: a person who knows the subject and knows how to work with AI. Three: the right data, and the honesty to throw out data that does not fit. Put those three together and one person can do what used to take a team. Here that meant a full research paper, a plain-English article, a podcast script, and a finished podcast read by computer voices.

Act 6 · Limits and the ask · What transfers

## What transfers — and what does not.

### Transfers

To any high-stakes knowledge work — diligence, regulatory submissions, competitive analysis, model-risk documentation:

- Closure and audit rules
- Segregation of duties across vendors
- Pre-registration before analysis
- Hash-gated freezes
- A decision ledger with permanent IDs
- A machine verifier that knows its limits

### Does not survive contact with procurement: *the price*

\$128 depends on individual flat-rate plans priced far below metered list. Plan on **\$3,224**; treat anything better as a windfall.

### Not yet established: *that this generalises*

This is **one case**. It is an existence proof — it shows the method *can* carry an unwinnable question to a defensible end, not that it will on the next one. The outer experiment was never pre-registered the way the inner one was. **Second and third cases are the work.**

### Does not transfer: *the human*

Someone had to tell a real methodological objection from a plausible-sounding one, 179 times, and decline to inflate a null result when it would have made a better story. **The doctoral training is the load-bearing component. The models are leverage on it.**

The rules work anywhere careful work matters. The cheap price will not last. The person cannot be replaced. Someone had to decide, 179 times, whether a complaint was real. And this happened once. Once is not always.

Act 6 · Limits and the ask · The ask

## Find the question you abandoned because it could not pay you back. Finish it.

Every organisation has them: the diligence thread nobody chased down, the hypothesis that died in a steering meeting, the analysis whose likely answer was “inconclusive” so it was never commissioned. **Those are not bad questions. They are questions that could not clear the cost of being answered.**

1

abandoned question, run to a defensible end

5

controls, applied without exception

\$3.2k

budget on metered pricing, not the subsidy

The cost of finishing just fell by two orders of magnitude. **What becomes worth knowing now?**

\[Operator to replace this slide’s ask with the specific pilot, owner, and decision date you want out of the room.\]

Think of a question your team gave up on because it was not worth the cost. Finishing just got a lot cheaper. So what is worth knowing now?

Act 7 · The record · Read it, listen to it

## The work itself — open it here.

Academia

### The research paper

43 pages · 11,330 words · Chicago 18 · 40 references · 4 figures\
Accepted eighth release, 23 August 2026

Read the PDF [Download PDF](https://www.strategicintelligenceresearch.org/studies/the-terminal-experiment/paper/O4-Research_Paper-Scott-W-Waddell-R8-20260823.pdf)

General reader

### The explainer article

3,797 words · 8 pages · plain language\
Revision 2, independently reviewed, 24 August 2026

Read the article [Download PDF](https://www.strategicintelligenceresearch.org/explainers/the-question-you-cannot-ask-yourself/The-Question-You-Cannot-Ask-Yourself.pdf)

Anyone with a commute

### The podcast

17m 30s · two hosts · 14,724 characters synthesised\
Revision 2, accepted 24 August 2026

Play [MP3](https://www.strategicintelligenceresearch.org/studies/the-terminal-experiment/podcasts/Podcast-The_Question_You_Cannot_Ask_Yourself-R2-claude-opus-5-20260824.mp3) [Transcript](https://www.strategicintelligenceresearch.org/studies/the-terminal-experiment/podcasts/Podcast-The_Question_You_Cannot_Ask_Yourself-R2-Transcript-20260824.md)

One body of evidence, three audiences — the paper for academia, the article for a general reader, the podcast for anyone with a commute. Assets load from the workspace alongside this file. **The paper and the article are PDFs**, read in the browser here or downloadable. If a PDF will not display inline the reader says so and offers the file directly; serving the folder over HTTP rather than opening it from disk fixes most cases.

You can read or listen to the actual work right here. The paper, the plain article, and the podcast.

Act 7 · The record · Provenance of these numbers

## Where these figures came from.

- **Measured, not estimated:** token, session, timing and cost figures are read from harness telemetry for this workspace — 29 Claude Code transcripts, 81 Codex rollouts, 26 OpenCode sessions, cross-checked against 75 recorded threads.
- **Independently cross-checked:** unit rates agree with the harness’s own 30-day usage estimator within 3% for two models and 11% for a third.
- **Conservative by construction:** four models run through a harness that reports no cache split, so their input is priced at full rate and their savings counted as zero. Every cost figure is an upper bound.
- **Modelled, and labelled as such:** the human baselines and loaded labour rates are estimates built from professional norms, shown as ranges. Substitute your own.
- **Disclosed boundary calls:** two sessions straddled artifacts; both attributions are stated in the appendix rather than buried.

The same standard the study held itself to applies to the slides about it. **If a number here is an estimate, it says so.**

These numbers came from the computer's own records, not from anyone's memory. Where a number is an educated guess, it says so.

Act 7 · The record · For the model checking this

## Don’t take our word for it. Here are the instructions to check it.

The argument is that claims should be independently verifiable — and that applies to *this presentation* too. So there is a vendor-neutral instruction set for any model asked to validate it. **What it requires:**

- **Declare itself.** Model, vendor family, and any exposure to the work. If its own family produced part of this, its check on that part is *corroboration, not independent verification* — and it must say so.
- **Gate on hashes.** Six SHA-256 values it must match before relying on anything. Mismatch means stop and report, not proceed.
- **Recompute, not re-read.** Every headline number comes with how to derive it, and the modelled estimates are marked *challenge this* rather than verify it.
- **Report what it skipped.** A check not run must never be reported as passing. `UNVERIFIABLE` is an accepted answer.

[↓ Download the validation handoff](https://www.strategicintelligenceresearch.org/studies/the-terminal-experiment/Validation-Handoff-claude-opus-5-20260824.md) ↗ Read it here

### The two instructions that matter most

If you find the deck overclaiming anywhere, that is the single most valuable finding you can return. §3F

A fabricated confirmation is the worst outcome available to you, and it is the specific failure this entire apparatus exists to prevent. §6

**And the limit it forces the validator to state:** the telemetry is the operator’s own local record. A second party can confirm the numbers follow from the logs. *It cannot establish that the logs are true.*

Works with no file access too: internal contradiction, mislabelled estimates, and independence claims that do not hold up are checkable from the deck alone.

We wrote instructions for a different AI to check this work. It has to say which company made it, prove it is reading the right files, and admit what it did not check. And it is told the most useful thing it can do is catch us exaggerating.

Act 7 · The record · Run it yourself

## Take the method. Leave this study behind.

The transferable asset is not a prompt transcript. It is a controlled sequence that keeps the subject matter expert in charge:

01

**Frame.** Name the question, decision, scope, and evidence that could change your mind.

02

**Freeze.** Set the source and analysis rules before the answer is visible.

03

**Challenge.** Separate creation, criticism, verification, and human acceptance.

04

**Translate.** Carry one accepted evidence base into every audience format.

Your first fifteen minutes

### Build the project before asking for the answer.

1.  **Open the method page.** Add your question, decision, evidence, outputs, and assurance level.
2.  **Download the tailored starter.** Put the Markdown file in a new, empty project folder.
3.  **Open that folder in Synaptic Code, start a chat, and paste the kickoff prompt.** The agent scaffolds the workspace, proposes the design and role split, then stops for your approval before research begins.

[Open the method](https://www.strategicintelligenceresearch.org/studies/the-terminal-experiment/method/index.html) [Download blank starter](https://www.strategicintelligenceresearch.org/studies/the-terminal-experiment/method/Research-Project-Starter.md) Copy kickoff prompt

**The download carries the operating method, not this project.** No research fact, source, conclusion, decision, or project memory is imported. The human remains the decision authority.

Running the method does not reproduce this study’s quality, speed, cost, or result. It creates an inspectable process for a new question.

If you want to try this on your own question, do not copy this study or ask one AI to do everything. Open the method website. It helps you describe the question and decision, choose how much checking you need, and download one file for a new project folder. Open that folder in Synaptic Code and paste the kickoff prompt. Synaptic builds the workspace first and stops for your approval before it starts the research.


## Structured metadata

```json
[
  {
    "@context": "https://schema.org",
    "@type": "WebPage",
    "url": "https://www.strategicintelligenceresearch.org/studies/the-terminal-experiment/",
    "inLanguage": "en",
    "isPartOf": {
      "@type": "WebSite",
      "name": "Strategic Intelligence Research",
      "url": "https://www.strategicintelligenceresearch.org/"
    }
  }
]
```
