The Terminal Experiment
1 / 29
An experiment in research method · August 2026

The Terminal Experiment:
finishing a question that cannot pay you back

A deliberately unwinnable research question, carried all the way to a rigorous, defensible null — by one researcher and seven AI models from six vendors, in eight days, for $128.

A human research programme would have killed this study before it ever reached a conclusion. That is the finding.

The question underneath it: is AI now good enough that one person has, in effect, a datacenter of experts on call?

Dr. Scott W. Waddell, D.S.I.  ·  every figure measured from harness telemetry unless marked as modelled

Act 1 · The test · Read this first

There are two experiments here. Only one of them is about religion.

The inner experiment

Can faith traditions be evaluated as mythologies by their observable, pre-mortem claims? Answer: not on this evidence. A null.

The outer experiment — the real one

Can AI-assisted research hold methodological discipline all the way to the end of an unwinnable question, and then carry the result to every audience that should see it? Answer: on this one case, yes.

The religion question is the instrument, not the subject. It was chosen because it was hard enough to break the method.

47.3 h
machine time, across 77 work bursts
136
model sessions, 6 vendor families
179
review findings raised and disposed
100 / 0
automated verifier checks: pass / fail

Every claim in it traces to a hash-pinned artifact, an independent cross-vendor review, and a named human decision. That last part is the whole talk.

Act 1 · The test · What would have to be true

If this is a datacenter of experts, three things have to hold.

The claim under test is not that a model can write. It is that a stack of models under controls can carry expert work from an unanswerable question to finished products for three audiences, without a human rewriting each step.

Objective 1

Finish the unfinishable

Carry a question with no available answer to a defensible null, under rules fixed before any result existed.

Standard: pre-registered. Thresholds, coding rules and the answer key were sealed and hashed before results.

PDF
Objective 2

Re-express without breaking

Turn 43 pages of doctoral argument into an explainer a general reader can follow, hedges and limits intact.

Standard: Flesch-Kincaid grade 8 to 10, set in the skill two days before the article existed. Reviewed by another vendor.

PDF
Objective 3

Change medium entirely

Move the same work into two-host audio, where nothing can be skimmed and no sentence can be re-read.

Standard: none was registered. The level is measurable, and was measured, but only after the fact.

MP3

Say the asymmetry out loud. Two of the three had a written standard before the work started. The third did not, and this deck says which is which, where each one lands.

Act 1 · The test · The instrument

The question was built to be unwinnable. On purpose.

A stress test only tests something if it can break the thing. So the design handicapped itself before a single case was coded:

  • The most interesting half was ruled out at the start. Post-mortem claims — what happens after death — were bracketed as untestable in principle and never coded as truth claims.
  • Five outcome-eligible cases. Minimal registered statistical power, declared in advance.
  • A contradiction was predicted before the data was seen and written into the pre-registration — and it duly arrived.
  • Contested, value-laden subject matter where a motivated analyst can reach any conclusion they like.

The odds, stated honestly at the outset

This design could not prove its thesis. At best it could fail to find evidence against a null — and say so precisely.

Everyone involved knew that before it started. It was registered that way.

A method that only survives easy questions has not been tested. This one was pointed at a question it could not win.

Act 1 · The test · Why it would not otherwise exist

A human programme would have killed this before it reached a conclusion.

Not because the design was weak. Because the economics of research do not let you finish a question that cannot pay you back.

No funder

Funds a study whose best possible outcome is “we found no evidence against the null.”

No journal

Competes for a null on five cases with no reportable row. The file drawer is where these go.

No researcher

Spends eighteen months and a career slot on an outcome they can predict in advance is unpublishable.

So the question gets abandoned mid-flight, quietly reframed into something publishable, or never started. That is not a failure of rigor. It is rigor being priced out.

At eight days and $128, you can afford to finish it anyway — and find out precisely what it does and does not show.

Act 2 · The controls · The operating model

Cheap to finish is worthless if nobody can trust the answer.

Finishing an abandoned question only counts if the result is defensible to an outside party. Five controls carry that weight.

ControlWhat it doesBusiness analogue
Closure ruleNothing is done until artifacts support it, an independent reviewer finds no material gap, and the accountable human explicitly agreesThree-way match before payment
Audit rule“File existence does not prove correctness.” Three states only: proven / partial / not provenEvidence standards in audit
Vendor separationThe model that drafts a thing is barred from reviewing or scoring itSegregation of duties
Pre-registrationAnalysis rules, thresholds, and answer keys locked and hashed before results existSealed bid; pre-committed test plan
Verifier + ledger100 automated checks over 66 hash-pinned baselines; 80 numbered human decisionsContinuous controls monitoring

Two corollaries organisations actually violate: “Direction is not evidence.” “An authorization is not an execution.” Approving a plan changes nothing. Only the work landing and being verified does.

Act 2 · The controls · Multi-vendor in practice

Multi-vendor is a control, not procurement waste.

Seven model variants across six families — OpenAI, Anthropic, DeepSeek, Moonshot, Zhipu, NVIDIA. They were not shopped for quality. They were assigned to roles that must not collude.

  • Anthropic drafted the coding instrument and was therefore excluded from coding and adjudication.
  • When a reviewer disclosed its own conflict mid-project, it was removed from the pool and the perimeter re-derived.
  • No model family ever reviewed its own work.

The strategic point

Model diversity is how you get independence when you have no second expert.

A single-vendor AI stack cannot mark its own homework either.

Act 2 · The controls · Provenance

Provenance is the product.

Every accepted release is frozen and hash-pinned. Reviews are gated on exact hashes, so a reviewer can prove which bytes it read.

Cross-session memory is a single sealed file with an explicit no-import / no-export rule — because in a workspace where provenance is the product, silent cross-contamination from another project is a correctness bug: you can no longer say where a fact came from.

Why this is not paranoia

Most enterprise AI memory features default to blending contexts. That is the failure mode this design refuses, deliberately and in writing.

Act 2 · The controls · The control environment

The system that checks the system — and admits what it cannot see.

100
automated checks, run green live
66
hash-pinned baselines
44
mutation tests that break it on purpose
“It reads states, counts, and hashes — it cannot tell whether a proven state is deserved or whether the operator accepted anything. A green run is a floor, not a warrant.verify_workspace.sh — the verifier, describing its own blind spot

That sentence is the credibility of the entire programme. A control environment that documents its own limits is one you can rely on. One that reports only green is one you cannot.

Act 3 · What we found · What came out

What came out.

43 pp
accepted paper, Chicago 18, 8 reviewed releases
124
working artifacts, ~280,000 words
80
recorded operator decisions
36
independent review memos
19
verification reports

Then the same source, under the same rules, became a plain-language explainer article and a seventeen-minute two-host podcast. One body of evidence, three audiences.

Act 3 · What we found · Sourcing

Where the evidence came from.

22 admitted 412 candidates screened by hand · 5.3%
258
collections enumerated
41
phrase queries run
20 → 118
collection growth, items
8
duplicate records merged
Body of sourcesCountHow it was verified
Literature references42Each registry-verified by a different vendor family than the one that cited it
Canonical works in the case tables2090 rows, each with an edition-stable locator
Coded source anchors verified2221 upheld, 1 modified, 0 rejected
References in the accepted paper40Each resolving to a held item

“Nothing is filled in from model recall.” Every metadata field came from a live registry response obtained in that same session.

Act 3 · What we found · What it refused

And mostly, what it refused to use.

Of 412 candidates screened, 22 were admitted. About 95% of the sourcing work was deciding what not to let in — with a written reason for every refusal.

This is what audit-grade means in practice. Not that the sources are good. That you can see every one that was considered, and why the rest were turned away.

Act 3 · What we found · Adversarial review

Adversarial review, quantified.

179

uniquely numbered findings across 29 finding families — every one disposed on the record.

Repaired, waived, or rejected as unfounded, with the byte-level evidence that refuted the reviewer preserved in place.

Reviews that disagreed were not reconciled into a house view. Both positions stand, attributed to the model that held them.

This is the slide that separates governance from theatre. A process that only ever confirms itself has no findings to show you.

Median duration of an independent cross-vendor review of a frozen artifact: 24 minutes, measured across 26 review sessions.

Act 3 · What we found · The payoff

It ran the unwinnable question to the end, and came back empty. Precisely.

  • Three single-case rows and one contradiction predicted in advance. No Boolean reduction possible. No row assertable as a finding.
  • Excluding one case removes all outcome variance — and the record says so rather than burying it.
  • Two registered sensitivity checks proved unexecutable against the published source. The record printed “NOT EXECUTABLE” instead of substituting something easier.
  • A reviewer’s qualification travels permanently with the result, on the reviewer’s terms, not the author’s.

A null this precise is more expensive to produce than a finding. You have to be exactly right about what you failed to show, and you have to survive review while holding nothing.

Acceptance here is governance, not empirical validation. It creates no finding and lifts no limit — and the paper says that in its own text.

Objective 1 met An unanswerable question carried to a defensible null without the method bending to rescue it. The strongest of the three claims: every rule, threshold and answer key was sealed and hashed before any result existed.

Any process looks disciplined when it is allowed to be right. This one was pointed at a question it could not win, ran to the end, and was still allowed to come back empty.

Act 3 · What we found · Objectives 2 and 3, measured

The same work, four reading levels.

Every rung is the same underlying research. The only thing that changes is who it is for.

RungFlesch-Kincaid grade Reading easeStandard set in advance
Research paper15.2 – 17.2 23.4 · very difficultSource material
Explainer article 7.5 – 9.1 68.8 · plain EnglishYes. Grade 8 to 10, set two days before, then reviewed by another vendor
Podcast 4.9 – 6.4 75.3 · fairly easyNo. Measured only after the fact
This deck, plain layer4.0 – 5.8 83.6 · easyNo
Objective 2 met Doctoral argument re-expressed for a general reader, inside a target fixed before the work and adjudicated by a model from another company. Grade 15.2 to 7.5 on the conservative estimator, same content, every hedge carried across.
Objective 3 met The same work moved into two-host audio, easier again than the article. Met, but claimed at a lower grade of evidence: no reading standard was registered beforehand.

Each grade is a range because syllable counting is estimator-sensitive. The low figure uses the conservative counter the independent reviewer specified; the high figure uses the non-deductive variant that knowingly overcounts silent endings. Both are printed rather than one being chosen. Lexile is deliberately absent: it is a proprietary measure this deck cannot compute, and the deck reports only what it can reproduce.

Act 4 · The compression · Three ratios, not one

Three compression ratios, not one.

≈43×
Effort
~2,056 person-hours → 47.3 machine-hours
≈45–90×
Calendar
12–24 months → 8 days
≈5,000×
Review latency
3–6 months → median 24 minutes

Review latency is the real mechanism, and it is the transferable one.

Eight serial rounds of external review on eight releases is a two-to-four-year proposition in any organisation — legal, model risk, second-line assurance, external audit. Here it took three days.

Speed did not come from doing less review. It came from doing more review with zero queueing.

Act 4 · The compression · What it cost

What did these eight days cost?

BasisAmountWhat it answers
Actual cash, prorated to the 8 days$128“What did this cost?” — the honest answer
Actual cash, full month charged$453Upper bound; the accounts served other work too
Metered API list, with caching$3,224Budget this at organisational scale
Metered API list, no caching$16,909Counterfactual showing what caching buys
Expert labour replaced$127,500–199,000The thing being compared

Eight of August’s thirty-one days on Anthropic Pro Max ($200/mo), OpenAI Pro ($200/mo), Ollama ($20/mo) and ElevenLabs Creator ($220/yr), plus ~$15 of Kimi usage charged in full because it was consumption, not a subscription. 88.9% of all input tokens were cache hits, billed at roughly a tenth of list — which is what makes a deliberately repetitive method affordable.

Act 4 · The compression · Name the subsidy

Name the subsidy. Do not budget on it.

Cash basisCash paid Metered value, as billed If no tokens were cached
Valuevs cashValuevs cash
Prorated to the eight days$128.12 $3,224.4925 : 1 $16,908.81132 : 1
Charging a full month$453.33 $3,224.497 : 1 $16,908.8137 : 1
Cash paid$128.12
As billed$3,224.49
If uncached$16,908.81

One linear scale. The cash bar is not a rendering fault: it is 0.8 per cent of the bar beneath it.

This project consumed roughly 25 times more compute than it paid for. The operator’s reading, stated as his position and not as measured fact: the frontier labs are currently subsidising tokens on individual flat-rate plans, and this project took deliberate advantage of that.

Budget $3,224, not $128. Treat the subsidy as a windfall on pilots, never a line item. And apply the honesty test: if your AI business case only works at consumer flat-rate pricing, you do not have a business case. This one passes — at full list with no cache discount it is still $16,909 against $127,500–199,000.

Act 4 · The compression · The second objection

“It only works at consumer pricing.” Correct, and the record already shows the cap.

What the record does prove is the capping mechanism. The flat rate bound twice inside eight days.

Three credit-depletion events on 17 August, and again on 22 August: “I ran out of Fable 5 credits and it paused. Pick up where it left off.” The work moved to another vendor and continued.

That message is lightly edited for grammar; the verbatim session text is preserved in the record.

Note what the stoppage actually demonstrates. Vendor diversity is a capacity control as well as an independence control. When one plan capped, the work moved to another vendor and carried on, which is why a single-vendor plan is also a single point of stoppage.

Act 4 · The compression · Where the money went

Where the money went — the opposite of what you would guess.

The two blind coders — the models doing the actual scoring Everything else — review, verification, governance
claude-fable-5$1,314.77
gpt-5.6-sol$909.38
claude-opus-5$373.53
kimi-k3$369.66
deepseek-v4-pro$243.56
glm-5.2$11.46
nemotron-3-ultra$2.13

Metered API list cost by model, whole programme. The two coder bars are drawn at a 3px minimum so they render at all — at true scale they are sub-pixel, which is the point.

Table view
ModelCostShare
claude-fable-5$1,314.7740.8%
gpt-5.6-sol$909.3828.2%
claude-opus-5$373.5311.6%
kimi-k3$369.6611.5%
deepseek-v4-pro$243.567.6%
glm-5.2 (coder)$11.460.36%
nemotron-3-ultra (coder)$2.130.07%
Total$3,224.49100%

The two blind coders consumed 0.35% of total token spend. In a defensible AI process, generation is a rounding error. You are not buying output. You are buying assurance. Budget accordingly.

Act 4 · The compression · Three audiences

One body of work, published three times.

Paper (academia)Article (broad)Podcast (broadest)
Output43 pp, 11,330 words3,797 words17m30s audio
Model sessions12771
Machine active time44.2 h1.8 h1.3 h
Elapsed7 days~3 hours3.7 hours
Metered compute$3,128.01$44.75$34.54
Actual cash$120.02$1.79$5.79
Human baseline250 person-days6 person-days4 person-days
Effort compression≈45×≈27×≈25×
Cost compression≈1,270×≈4,190×≈560×

Human baselines are modelled estimates from professional norms, stated as ranges in the appendix. Cash is allocated by vendor, not blended share — voice synthesis is charged wholly to the podcast.

Act 4 · The compression · The strategic finding

Effort compression falls as the audience widens. Cost compression rises.

The research is bottlenecked on thinking and adversarial review — expert labour, only so compressible, and token-expensive.

The derivatives are bottlenecked on craft and production — nearly free in tokens, expensive in wages and calendar.

So the further a deliverable sits from the research and the closer to the audience, the harder the economics tilt.

The entire derivative layer — the two artifacts that actually reach people — cost $7.58. That is 5.9% of the cash.

Effectively all of the spend bought the right to have something to say. A rounding error carried it to the world. Most organisations run this exactly backwards — nothing on rigor, everything on communications.

90¢

what the accepted seventeen-and-a-half-minute episode cost in voice synthesis.

Act 5 · The modern worker · The equation

The equation this actually changes.

Synaptic CodeOne interface across many models, so a single person can run several of them side by side and keep the record in one place. Here: 136 sessions across 6 vendor families, all logged.
+
An AI-fluent subject matter expertKnows the domain, and knows how to direct and refuse a model. Not a prompt writer, the accountable party. Here: 80 numbered human decisions, and no model held authority.
+
Correct and appropriate dataData that fits the question, with the discipline to discard what does not. Here: 412 sources screened, 22 admitted, two datasets excluded as ill-fitting.

= the modern worker. Not a model replacing an expert, and not an expert working faster. One accountable human running several models against data they chose, at a standard that survives outside review.

4 / 4
Doctoral research, streamlined. A pre-registered study carried to a defensible null, in 8 days and 47.3 machine hours
15.2 → 7.5
Study into an explainer. Reading grade on the same content, against a target set in advance and checked by another vendor
7.5 → 4.9
Explainer into a podcast script. Reading grade again, across a 173-turn two-host script checked claim by claim against the article
90¢
Script into finished audio. 17m 30s in two synthetic voices, stitched from 11 requests

All four on the same body of work, in eight days. The hosts are licensed ElevenLabs professional voices: no voice was cloned for this project.

Act 6 · Limits and the ask · What transfers

What transfers — and what does not.

Transfers

To any high-stakes knowledge work — diligence, regulatory submissions, competitive analysis, model-risk documentation:

  • Closure and audit rules
  • Segregation of duties across vendors
  • Pre-registration before analysis
  • Hash-gated freezes
  • A decision ledger with permanent IDs
  • A machine verifier that knows its limits

Does not survive contact with procurement: the price

$128 depends on individual flat-rate plans priced far below metered list. Plan on $3,224; treat anything better as a windfall.

Not yet established: that this generalises

This is one case. It is an existence proof — it shows the method can carry an unwinnable question to a defensible end, not that it will on the next one. The outer experiment was never pre-registered the way the inner one was. Second and third cases are the work.

Does not transfer: the human

Someone had to tell a real methodological objection from a plausible-sounding one, 179 times, and decline to inflate a null result when it would have made a better story. The doctoral training is the load-bearing component. The models are leverage on it.

Act 6 · Limits and the ask · The ask

Find the question you abandoned because it could not pay you back. Finish it.

Every organisation has them: the diligence thread nobody chased down, the hypothesis that died in a steering meeting, the analysis whose likely answer was “inconclusive” so it was never commissioned. Those are not bad questions. They are questions that could not clear the cost of being answered.

1
abandoned question, run to a defensible end
5
controls, applied without exception
$3.2k
budget on metered pricing, not the subsidy

The cost of finishing just fell by two orders of magnitude. What becomes worth knowing now?

[Operator to replace this slide’s ask with the specific pilot, owner, and decision date you want out of the room.]

Act 7 · The record · Read it, listen to it

The work itself — open it here.

Academia

The research paper

43 pages · 11,330 words · Chicago 18 · 40 references · 4 figures
Accepted eighth release, 23 August 2026

Download PDF
General reader

The explainer article

3,797 words · 8 pages · plain language
Revision 2, independently reviewed, 24 August 2026

Download PDF
Anyone with a commute

The podcast

17m 30s · two hosts · 14,724 characters synthesised
Revision 2, accepted 24 August 2026

One body of evidence, three audiences — the paper for academia, the article for a general reader, the podcast for anyone with a commute. Assets load from the workspace alongside this file. The paper and the article are PDFs, read in the browser here or downloadable. If a PDF will not display inline the reader says so and offers the file directly; serving the folder over HTTP rather than opening it from disk fixes most cases.

Speaking plainly

The podcast · 17m30s
Act 7 · The record · Provenance of these numbers

Where these figures came from.

The same standard the study held itself to applies to the slides about it. If a number here is an estimate, it says so.

Act 7 · The record · For the model checking this

Don’t take our word for it. Here are the instructions to check it.

The argument is that claims should be independently verifiable — and that applies to this presentation too. So there is a vendor-neutral instruction set for any model asked to validate it. What it requires:

  • Declare itself. Model, vendor family, and any exposure to the work. If its own family produced part of this, its check on that part is corroboration, not independent verification — and it must say so.
  • Gate on hashes. Six SHA-256 values it must match before relying on anything. Mismatch means stop and report, not proceed.
  • Recompute, not re-read. Every headline number comes with how to derive it, and the modelled estimates are marked challenge this rather than verify it.
  • Report what it skipped. A check not run must never be reported as passing. UNVERIFIABLE is an accepted answer.

The two instructions that matter most

If you find the deck overclaiming anywhere, that is the single most valuable finding you can return. §3F
A fabricated confirmation is the worst outcome available to you, and it is the specific failure this entire apparatus exists to prevent. §6

And the limit it forces the validator to state: the telemetry is the operator’s own local record. A second party can confirm the numbers follow from the logs. It cannot establish that the logs are true.

Works with no file access too: internal contradiction, mislabelled estimates, and independence claims that do not hold up are checkable from the deck alone.

Act 7 · The record · Run it yourself

Take the method. Leave this study behind.

The transferable asset is not a prompt transcript. It is a controlled sequence that keeps the subject matter expert in charge:

01
Frame. Name the question, decision, scope, and evidence that could change your mind.
02
Freeze. Set the source and analysis rules before the answer is visible.
03
Challenge. Separate creation, criticism, verification, and human acceptance.
04
Translate. Carry one accepted evidence base into every audience format.
Your first fifteen minutes

Build the project before asking for the answer.

  1. Open the method page. Add your question, decision, evidence, outputs, and assurance level.
  2. Download the tailored starter. Put the Markdown file in a new, empty project folder.
  3. Open that folder in Synaptic Code, start a chat, and paste the kickoff prompt. The agent scaffolds the workspace, proposes the design and role split, then stops for your approval before research begins.

The download carries the operating method, not this project. No research fact, source, conclusion, decision, or project memory is imported. The human remains the decision authority.

Running the method does not reproduce this study’s quality, speed, cost, or result. It creates an inspectable process for a new question.