A deliberately unwinnable research question, carried all the way to a rigorous, defensible null — by one researcher and seven AI models from six vendors, in eight days, for $128.
A human research programme would have killed this study before it ever reached a conclusion. That is the finding.
The question underneath it: is AI now good enough that one person has, in effect, a datacenter of experts on call?
Dr. Scott W. Waddell, D.S.I. · every figure measured from harness telemetry unless marked as modelled
Can faith traditions be evaluated as mythologies by their observable, pre-mortem claims? Answer: not on this evidence. A null.
Can AI-assisted research hold methodological discipline all the way to the end of an unwinnable question, and then carry the result to every audience that should see it? Answer: on this one case, yes.
The religion question is the instrument, not the subject. It was chosen because it was hard enough to break the method.
Every claim in it traces to a hash-pinned artifact, an independent cross-vendor review, and a named human decision. That last part is the whole talk.
The claim under test is not that a model can write. It is that a stack of models under controls can carry expert work from an unanswerable question to finished products for three audiences, without a human rewriting each step.
Carry a question with no available answer to a defensible null, under rules fixed before any result existed.
Standard: pre-registered. Thresholds, coding rules and the answer key were sealed and hashed before results.
Turn 43 pages of doctoral argument into an explainer a general reader can follow, hedges and limits intact.
Standard: Flesch-Kincaid grade 8 to 10, set in the skill two days before the article existed. Reviewed by another vendor.
Move the same work into two-host audio, where nothing can be skimmed and no sentence can be re-read.
Standard: none was registered. The level is measurable, and was measured, but only after the fact.
Say the asymmetry out loud. Two of the three had a written standard before the work started. The third did not, and this deck says which is which, where each one lands.
A stress test only tests something if it can break the thing. So the design handicapped itself before a single case was coded:
This design could not prove its thesis. At best it could fail to find evidence against a null — and say so precisely.
Everyone involved knew that before it started. It was registered that way.
A method that only survives easy questions has not been tested. This one was pointed at a question it could not win.
Not because the design was weak. Because the economics of research do not let you finish a question that cannot pay you back.
Funds a study whose best possible outcome is “we found no evidence against the null.”
Competes for a null on five cases with no reportable row. The file drawer is where these go.
Spends eighteen months and a career slot on an outcome they can predict in advance is unpublishable.
So the question gets abandoned mid-flight, quietly reframed into something publishable, or never started. That is not a failure of rigor. It is rigor being priced out.
At eight days and $128, you can afford to finish it anyway — and find out precisely what it does and does not show.
Finishing an abandoned question only counts if the result is defensible to an outside party. Five controls carry that weight.
| Control | What it does | Business analogue |
|---|---|---|
| Closure rule | Nothing is done until artifacts support it, an independent reviewer finds no material gap, and the accountable human explicitly agrees | Three-way match before payment |
| Audit rule | “File existence does not prove correctness.” Three states only: proven / partial / not proven | Evidence standards in audit |
| Vendor separation | The model that drafts a thing is barred from reviewing or scoring it | Segregation of duties |
| Pre-registration | Analysis rules, thresholds, and answer keys locked and hashed before results exist | Sealed bid; pre-committed test plan |
| Verifier + ledger | 100 automated checks over 66 hash-pinned baselines; 80 numbered human decisions | Continuous controls monitoring |
Two corollaries organisations actually violate: “Direction is not evidence.” “An authorization is not an execution.” Approving a plan changes nothing. Only the work landing and being verified does.
Seven model variants across six families — OpenAI, Anthropic, DeepSeek, Moonshot, Zhipu, NVIDIA. They were not shopped for quality. They were assigned to roles that must not collude.
Model diversity is how you get independence when you have no second expert.
A single-vendor AI stack cannot mark its own homework either.
Every accepted release is frozen and hash-pinned. Reviews are gated on exact hashes, so a reviewer can prove which bytes it read.
Cross-session memory is a single sealed file with an explicit no-import / no-export rule — because in a workspace where provenance is the product, silent cross-contamination from another project is a correctness bug: you can no longer say where a fact came from.
Most enterprise AI memory features default to blending contexts. That is the failure mode this design refuses, deliberately and in writing.
That sentence is the credibility of the entire programme. A control environment that documents its own limits is one you can rely on. One that reports only green is one you cannot.
Then the same source, under the same rules, became a plain-language explainer article and a seventeen-minute two-host podcast. One body of evidence, three audiences.
| Body of sources | Count | How it was verified |
|---|---|---|
| Literature references | 42 | Each registry-verified by a different vendor family than the one that cited it |
| Canonical works in the case tables | 20 | 90 rows, each with an edition-stable locator |
| Coded source anchors verified | 22 | 21 upheld, 1 modified, 0 rejected |
| References in the accepted paper | 40 | Each resolving to a held item |
“Nothing is filled in from model recall.” Every metadata field came from a live registry response obtained in that same session.
Of 412 candidates screened, 22 were admitted. About 95% of the sourcing work was deciding what not to let in — with a written reason for every refusal.
This is what audit-grade means in practice. Not that the sources are good. That you can see every one that was considered, and why the rest were turned away.
uniquely numbered findings across 29 finding families — every one disposed on the record.
Repaired, waived, or rejected as unfounded, with the byte-level evidence that refuted the reviewer preserved in place.
Reviews that disagreed were not reconciled into a house view. Both positions stand, attributed to the model that held them.
This is the slide that separates governance from theatre. A process that only ever confirms itself has no findings to show you.
Median duration of an independent cross-vendor review of a frozen artifact: 24 minutes, measured across 26 review sessions.
A null this precise is more expensive to produce than a finding. You have to be exactly right about what you failed to show, and you have to survive review while holding nothing.
Acceptance here is governance, not empirical validation. It creates no finding and lifts no limit — and the paper says that in its own text.
Any process looks disciplined when it is allowed to be right. This one was pointed at a question it could not win, ran to the end, and was still allowed to come back empty.
Every rung is the same underlying research. The only thing that changes is who it is for.
| Rung | Flesch-Kincaid grade | Reading ease | Standard set in advance |
|---|---|---|---|
| Research paper | 15.2 – 17.2 | 23.4 · very difficult | Source material |
| Explainer article | 7.5 – 9.1 | 68.8 · plain English | Yes. Grade 8 to 10, set two days before, then reviewed by another vendor |
| Podcast | 4.9 – 6.4 | 75.3 · fairly easy | No. Measured only after the fact |
| This deck, plain layer | 4.0 – 5.8 | 83.6 · easy | No |
Each grade is a range because syllable counting is estimator-sensitive. The low figure uses the conservative counter the independent reviewer specified; the high figure uses the non-deductive variant that knowingly overcounts silent endings. Both are printed rather than one being chosen. Lexile is deliberately absent: it is a proprietary measure this deck cannot compute, and the deck reports only what it can reproduce.
Review latency is the real mechanism, and it is the transferable one.
Eight serial rounds of external review on eight releases is a two-to-four-year proposition in any organisation — legal, model risk, second-line assurance, external audit. Here it took three days.
Speed did not come from doing less review. It came from doing more review with zero queueing.
| Basis | Amount | What it answers |
|---|---|---|
| Actual cash, prorated to the 8 days | $128 | “What did this cost?” — the honest answer |
| Actual cash, full month charged | $453 | Upper bound; the accounts served other work too |
| Metered API list, with caching | $3,224 | Budget this at organisational scale |
| Metered API list, no caching | $16,909 | Counterfactual showing what caching buys |
| Expert labour replaced | $127,500–199,000 | The thing being compared |
Eight of August’s thirty-one days on Anthropic Pro Max ($200/mo), OpenAI Pro ($200/mo), Ollama ($20/mo) and ElevenLabs Creator ($220/yr), plus ~$15 of Kimi usage charged in full because it was consumption, not a subscription. 88.9% of all input tokens were cache hits, billed at roughly a tenth of list — which is what makes a deliberately repetitive method affordable.
| Cash basis | Cash paid | Metered value, as billed | If no tokens were cached | ||
|---|---|---|---|---|---|
| Value | vs cash | Value | vs cash | ||
| Prorated to the eight days | $128.12 | $3,224.49 | 25 : 1 | $16,908.81 | 132 : 1 |
| Charging a full month | $453.33 | $3,224.49 | 7 : 1 | $16,908.81 | 37 : 1 |
One linear scale. The cash bar is not a rendering fault: it is 0.8 per cent of the bar beneath it.
This project consumed roughly 25 times more compute than it paid for. The operator’s reading, stated as his position and not as measured fact: the frontier labs are currently subsidising tokens on individual flat-rate plans, and this project took deliberate advantage of that.
Budget $3,224, not $128. Treat the subsidy as a windfall on pilots, never a line item. And apply the honesty test: if your AI business case only works at consumer flat-rate pricing, you do not have a business case. This one passes — at full list with no cache discount it is still $16,909 against $127,500–199,000.
What the record does prove is the capping mechanism. The flat rate bound twice inside eight days.
Three credit-depletion events on 17 August, and again on 22 August: “I ran out of Fable 5 credits and it paused. Pick up where it left off.” The work moved to another vendor and continued.
That message is lightly edited for grammar; the verbatim session text is preserved in the record.
Note what the stoppage actually demonstrates. Vendor diversity is a capacity control as well as an independence control. When one plan capped, the work moved to another vendor and carried on, which is why a single-vendor plan is also a single point of stoppage.
Metered API list cost by model, whole programme. The two coder bars are drawn at a 3px minimum so they render at all — at true scale they are sub-pixel, which is the point.
| Model | Cost | Share |
|---|---|---|
| claude-fable-5 | $1,314.77 | 40.8% |
| gpt-5.6-sol | $909.38 | 28.2% |
| claude-opus-5 | $373.53 | 11.6% |
| kimi-k3 | $369.66 | 11.5% |
| deepseek-v4-pro | $243.56 | 7.6% |
| glm-5.2 (coder) | $11.46 | 0.36% |
| nemotron-3-ultra (coder) | $2.13 | 0.07% |
| Total | $3,224.49 | 100% |
The two blind coders consumed 0.35% of total token spend. In a defensible AI process, generation is a rounding error. You are not buying output. You are buying assurance. Budget accordingly.
| Paper (academia) | Article (broad) | Podcast (broadest) | |
|---|---|---|---|
| Output | 43 pp, 11,330 words | 3,797 words | 17m30s audio |
| Model sessions | 127 | 7 | 1 |
| Machine active time | 44.2 h | 1.8 h | 1.3 h |
| Elapsed | 7 days | ~3 hours | 3.7 hours |
| Metered compute | $3,128.01 | $44.75 | $34.54 |
| Actual cash | $120.02 | $1.79 | $5.79 |
| Human baseline | 250 person-days | 6 person-days | 4 person-days |
| Effort compression | ≈45× | ≈27× | ≈25× |
| Cost compression | ≈1,270× | ≈4,190× | ≈560× |
Human baselines are modelled estimates from professional norms, stated as ranges in the appendix. Cash is allocated by vendor, not blended share — voice synthesis is charged wholly to the podcast.
The research is bottlenecked on thinking and adversarial review — expert labour, only so compressible, and token-expensive.
The derivatives are bottlenecked on craft and production — nearly free in tokens, expensive in wages and calendar.
So the further a deliverable sits from the research and the closer to the audience, the harder the economics tilt.
The entire derivative layer — the two artifacts that actually reach people — cost $7.58. That is 5.9% of the cash.
Effectively all of the spend bought the right to have something to say. A rounding error carried it to the world. Most organisations run this exactly backwards — nothing on rigor, everything on communications.
what the accepted seventeen-and-a-half-minute episode cost in voice synthesis.
= the modern worker. Not a model replacing an expert, and not an expert working faster. One accountable human running several models against data they chose, at a standard that survives outside review.
All four on the same body of work, in eight days. The hosts are licensed ElevenLabs professional voices: no voice was cloned for this project.
To any high-stakes knowledge work — diligence, regulatory submissions, competitive analysis, model-risk documentation:
$128 depends on individual flat-rate plans priced far below metered list. Plan on $3,224; treat anything better as a windfall.
This is one case. It is an existence proof — it shows the method can carry an unwinnable question to a defensible end, not that it will on the next one. The outer experiment was never pre-registered the way the inner one was. Second and third cases are the work.
Someone had to tell a real methodological objection from a plausible-sounding one, 179 times, and decline to inflate a null result when it would have made a better story. The doctoral training is the load-bearing component. The models are leverage on it.
Every organisation has them: the diligence thread nobody chased down, the hypothesis that died in a steering meeting, the analysis whose likely answer was “inconclusive” so it was never commissioned. Those are not bad questions. They are questions that could not clear the cost of being answered.
The cost of finishing just fell by two orders of magnitude. What becomes worth knowing now?
[Operator to replace this slide’s ask with the specific pilot, owner, and decision date you want out of the room.]
43 pages · 11,330 words · Chicago 18 · 40 references · 4 figures
Accepted eighth release, 23 August 2026
3,797 words · 8 pages · plain language
Revision 2, independently reviewed, 24 August 2026
17m 30s · two hosts · 14,724 characters synthesised
Revision 2, accepted 24 August 2026
One body of evidence, three audiences — the paper for academia, the article for a general reader, the podcast for anyone with a commute. Assets load from the workspace alongside this file. The paper and the article are PDFs, read in the browser here or downloadable. If a PDF will not display inline the reader says so and offers the file directly; serving the folder over HTTP rather than opening it from disk fixes most cases.
The same standard the study held itself to applies to the slides about it. If a number here is an estimate, it says so.
The argument is that claims should be independently verifiable — and that applies to this presentation too. So there is a vendor-neutral instruction set for any model asked to validate it. What it requires:
UNVERIFIABLE
is an accepted answer.And the limit it forces the validator to state: the telemetry is the operator’s own local record. A second party can confirm the numbers follow from the logs. It cannot establish that the logs are true.
Works with no file access too: internal contradiction, mislabelled estimates, and independence claims that do not hold up are checkable from the deck alone.
The transferable asset is not a prompt transcript. It is a controlled sequence that keeps the subject matter expert in charge:
The download carries the operating method, not this project. No research fact, source, conclusion, decision, or project memory is imported. The human remains the decision authority.
Running the method does not reproduce this study’s quality, speed, cost, or result. It creates an inspectable process for a new question.