How do teachings in selected faith traditions relate to affiliation retention, and how do their texts respond to uncertainty about ultimate meaning?
Research question, paraphrased from the published explainer
A pilot comparison of faith traditions and philosophical reference cases, with attention to AI-assisted coding, preregistered rules, independent checking, and the limits of the evidence.
In brief
The pilot did not support the author’s prediction about teachings and retention. A separate textual analysis left two traditions undecided and found one philosophical counterexample. The registered question joining the two analyses was not tested.
Keep in mind This small pilot does not refute the author’s belief, establish a causal effect, or show that its AI-assisted workflow outperforms ordinary research.
What it found ↗Published explainer section: “What it found.”
The public explainer describes two analyses. Neither supplies an answer to the study’s joint question.
01Study findings
No support for the predicted retention pattern
The pilot found no evidence against the strict Durkheimian null and no support, at this scale, for the prediction linking the selected teachings to affiliation retention.
Limits of this finding
Five outcome-eligible cases came from one country and one survey wave. Four relied on proxy categories. Removing Theravada Buddhism removes all outcome variation; that case also used the loosest proxy. No individual row can be reported as a finding, and the registered design had minimal power. A null result does not establish that the prediction is false.
What it found ↗Published explainer section: “What it found.”
In the explainer’s account, the texts of the Latter-day Saints, Nicene Christianity and Sunni Islam closed the gap. Hinduism and Theravada Buddhism were undecided because the necessary texts were missing. Camus, a philosophical comparison case, was the only one of nine cases that kept it open.
Limits of this finding
The Camus result was described as “tautology-adjacent”: his writing was judged against his own standard. The universal claim had one counterexample, but the religious-only claim remained unrefuted. These are judgments about the supplied texts, not measurements of believers’ experiences; missing evidence remained undecided.
A 2026 pilot described through its public explainer.
Cases and study design
The study used a nine-case comparative pilot with blind double-coding and cross-family adjudication. Preregistered crisp-set qualitative comparative analysis (QCA) covered five outcome-eligible cases.
Scope limit The author retained control over the questions, framework, features, cases, reading lists and scoring rules. Independence was task-specific, not global.
Two analytical scopes
9Cases read and codedFaith traditions and philosophical reference cases
5Outcome-eligible casesRetention analysis · One country · One survey wave
Separate analyses; the joint question was not tested
The five outcome-eligible cases are a subset of the nine cases. They are not additional cases.
Research controlsSettle who is disqualified before you decide who is best.
First, settle who is disqualified before you decide who is best. Before asking which system was strongest for a job, Waddell asked which ones were ineligible for it. Had it helped write the thing it would be judging? Had it seen the answer key? Had it already reviewed this material once before?
The order matters more than it sounds. Ask who is best first, and you will find a reason why your favorite also happens to be eligible. The company whose systems helped design the study and draft the manuscript was barred from the coding and the judging.
What the design could not do was eliminate every conflict, so it disclosed the rest instead. One of the two coders had already seen the accepted design, and the record says so on the page where its scores appear. Independence here was task by task, never global, and the paper uses those words.
Research controlsFix the standard before you measure anything against it.
Second, fix the standard before you measure anything against it. A set of practice cases and their correct answers was sealed before any coder was chosen.
An answer key written afterward is worthless, because by then it has been shaped by what you already saw. That is not a claim about anyone’s honesty. It is a claim about what a person can no longer un-see. The seal was broken once, to correct an error, and only after both practice submissions were locked and could not be changed. Then it was sealed again, and the detour went into the record.
Research controlsTwo observers who cannot talk are worth more than one careful observer.
Third, two observers who cannot talk are worth more than one careful observer. Two systems, from two different companies, read the nine cases and scored them separately. Neither could see the key, and neither could see the other’s work.
The reasoning is simple. If two readers who cannot consult each other arrive at the same reading, that reading is less likely to be an artifact of whoever happened to do it. It is the reason two radiologists are asked to read the same scan.
They matched exactly on about two of every three scores, and on roughly five of six once you count the places where they said the same thing in different words. That sounds strong, and the paper is the one that tells you to hold back. Both readers had been handed the same reading list, so their agreement shows they read the same books the same way. It says nothing about whether those were the right books, and Waddell lists a proper audit of those lists by human scholars as work he did not do.
Research controlsState your prediction before you look at the evidence.
Fourth, state your prediction before you look at the evidence. Before anyone opened the survey data, the decisions were written down and locked: which measure would count as holding onto people, exactly where the cutoff line would sit, how to check that the traditions were even comparable, and five different ways of re-running everything if the first way looked suspicious.
A prediction made afterward is not a prediction at all, it is a description. Set your cutoff after seeing the numbers and you will set it where the numbers look good, and you will feel entirely reasonable while you do it. That is the entire reason the line goes down first.
Research controlsVerification means somebody else does it over.
Fifth, verification means somebody else does it over. When the analysis was finished, two more systems went over it. One recalculated every figure from scratch, and the other went back to the original survey and looked up every published number again, one at a time. They found nothing that conflicted.
Neither of them had worked on the analysis, though one had reviewed the study plan earlier, which the record discloses rather than hides.
Sixth, preserve a disagreement rather than resolving it. When a reviewer objected to one of the conclusions, the objection was not negotiated away. It was written into the paper in the reviewer’s own words and permanently attached to the finding it complicates.
A smoothed-over objection is invisible to the reader, who is then in no position to weigh it. This paper carries its objections around like scars.
Can belonging and open uncertainty vary independently?
The joint question was registered but not analyzed. The study’s two analyses were never brought together, so they support no claim of a trade-off, independence or connection between belonging and open uncertainty.
The workflow was not compared with an ordinary single-researcher study. Its recorded checks and file hashes establish an inspectable process, not the truth of its claims or superiority over other methods.