Psychedelic research measures the dose to two decimal places and the preparation to the nearest hour. When somebody finally asked what the preparation was doing, the record could only answer in hours, and hours said nothing.


The least count

Every measuring instrument has a least count. It is the smallest division the instrument can resolve, and it is fixed before any measurement is taken.

A steel rule graduated in millimetres cannot report four tenths of a millimetre. Better light will not help. A steadier hand will not help. A more careful machinist will read the same rule and report the same millimetre. The information is not obscured, it was never captured, and no amount of later care recovers it.

This has a consequence that matters more than it sounds. Suppose you record a workshop's output with a millimetre rule for ten years, and then somebody asks whether variations of two tenths of a millimetre affected the finished parts. The answer will come back negative. It will come back negative whether or not two tenths mattered, because the archive cannot distinguish a part that differed by two tenths from one that did not. The finding will look like a fact about engineering. It is a fact about the rule.

Now open a modern psychedelic trial protocol.

The dose is specified to a precision the rest of the document does not attempt. Twenty-five milligrams of a synthetic, characterised, batch-controlled formulation. One milligram in the comparator arm. Purity stated, stability data on file, administration timed to the minute.

Then look for the preparation.

You will find a number of hours. Sometimes a number of sessions. Occasionally the name of a manual. Almost never a description sufficient for another team to reproduce what took place, or for a reader to judge whether one trial's preparation resembles another's in any respect beyond duration.

Two decimal places on one side. A round number on the other.

In 2025 somebody finally asked the archive what the preparation was doing. The archive answered in hours, and hours said nothing. That answer is now circulating as a fact about psychedelic therapy, and it is about to be load-bearing in a regulatory decision. It is a fact about the rule.

A steel rule and a micrometer beside a set of pills, suggesting two different standards of precision applied to one treatment.
A steel rule and a micrometer beside a set of pills, suggesting two different standards of precision applied to one treatment.

The number in the middle

Start with the trial that did something almost nobody does.

In July 2026 the Center for Psychedelic Drug Research and Education at Ohio State published the first American clinical study of psilocybin-assisted therapy in veterans with PTSD. Twelve participants. An eleven-week protocol. Roughly eight hours of preparatory therapy across four visits, then two supervised sessions at 15 mg and 25 mg, then six to eight hours of integration. Sixteen hours of skilled clinical contact wrapped around two drug days.

The trial measured PTSD severity three times rather than twice. At baseline. At the end of preparation, before any psilocybin. And a month after the second dose.

The middle measurement moved. By the fourth preparatory session, before a milligram of drug, clinician-rated severity had fallen 6.33 points, an effect size of 0.87. How much a participant improved during preparation predicted how much they improved a month after dosing.

I have written about this trial in detail and will not relitigate it. Two caveats belong in any honest account. The movement appeared on the clinician-administered measure and not significantly on self-report. And some portion of it is regression to the mean, the ordinary tendency of people recruited at their worst to look better when measured again. Neither caveat makes the observation disappear. It makes it a hypothesis, which is what a twelve-person open-label trial is for.

The point here is narrower than the finding. Almost no other trial can tell us whether its own middle number moved, because almost no other trial took the measurement. Ohio State chose a finer graduation and then reported what it read.


The null that is not a null

In 2025 Jennie Hultgren and colleagues at Stockholm University published the first meta-analysis to test whether the amount of therapy in psilocybin-assisted treatment predicts outcome. Sixteen studies, nineteen reports.

The treatment effects were very large: Cohen's d of 1.69 in the short term and 2.10 at longer follow-up.

The therapy predicted nothing. Total hours were unrelated to effect size in the short term, b = −0.05, p = .327, and in the long term, b = −0.07, p = .340. Modelled separately, preparation came out at b = −0.01, p = .912.

That result has begun to circulate as evidence that the therapy in psychedelic-assisted therapy may not be doing much. It will be cited that way in the run-up to approval, because it is a convenient thing for a sponsor to be able to say.

Read what the authors say about why they found nothing.

The predictor available to them was hours. Not content. Not sequence. Not fidelity. The arithmetic sum of preparation and integration time, across a literature whose entire observed range ran from about four and a half to eighteen hours. They describe the reporting of the therapeutic component as severely insufficient, and they are specific about it: optional hours reported without uptake figures, wide ranges given without particulars, different reports of the same trial supplying different numbers. Manuals in use are loose. Fidelity to them is rarely assessed. It is not clear the interventions delivered meet established definitions of psychotherapy at all. Their recommendations are that the therapeutic component be standardised and reported with the rigour applied to the pharmacological one, and that preparation and integration be separated so their contributions can be distinguished.

They were reading a millimetre rule and being asked about tenths.

Consider what entered their dataset. Two trials each reporting eight hours of preparation enter the model as identical observations. One may have delivered manualised psychoeducation with rehearsal of difficult experiences and explicit correction of optimistic expectations. The other may have delivered three unstructured conversations whose content varied by therapist. No field in the dataset distinguishes them, because no field in the source papers does.

An underexposed region of an archive does not announce itself. It reports zero, and zero is a number, and numbers get cited.

This is the general form of the problem, and it is worth stating plainly because it will recur. A reporting standard does not merely describe research. It fixes the least count of the record, and therefore determines which questions can be asked of the record afterwards. Unrecorded variables do not stay merely unknown. They become unrecoverable, and their absence is eventually read as evidence of their unimportance.

The null on therapy hours is the first large instance of that misreading in this field. It will not be the last.


What fixed the least count

I have set out elsewhere, at length, how the field's consensus reporting standard distributes its attention. Here is the compressed version.

The ReSPCT guidelines, published in Nature Medicine in June 2025, are the product of a four-round international Delphi process: eighty-nine experts, seventeen countries, 770 initial free-text nominations reduced to thirty items clearing a seventy per cent consensus threshold. They are serious, valuable work, and they are an enormous improvement on the previous standard, which was a sentence in the methods section and the reader's imagination.

The thirty items sit in four sections. Seven cover the physical environment. Ten cover the dosing session procedure. Eight cover the therapeutic framework and protocol. Five cover participants' subjective experiences.

The dosing day is graduated finely. Ambiance. Lighting. Objects and decorations. Access to nature. Bathroom privacy. The number and roles of people present. The relative positioning of bodies in the room. Whether attention is directed inward or outward. Music. Interpersonal interventions and how consent for them was obtained.

The content of everything before the dosing day is item 21, which asks for the activities performed during preparation sessions. Item 20 asks how many sessions there were and how long they ran.

That is the graduation. One item for the content of a phase that can run to eighteen hours across several weeks.

Two pieces of external evidence bracket it.

The first is compliance. A 2026 analysis in European Neuropsychopharmacology applied the ReSPCT framework retrospectively to thirteen registered psilocybin protocols for major depressive disorder and treatment-resistant depression, eleven Phase II and two Phase III. Procedural safeguards were well documented. Contextual domains were not. Of the thirteen, 84.6 per cent contained no information on cultural competence and safety, 92.3 per cent did not describe objects or decorations, and 84.6 per cent did not report access to nature. Those figures concern the dosing environment, the part of the checklist graduated most finely. If the best-specified section is being satisfied at those rates, an optimistic estimate of preparation reporting is not available.

The second is what one item cannot hold, which is visible wherever the underlying document is public.

The German EPIsoDE trial, published in JAMA Psychiatry in March 2026, released its therapist manual openly. This is rare and creditable. The manual specifies a preparatory session seven days before the first dose running about a hundred minutes to a timed checklist. Further sessions the day before each dose, including guided mindfulness and explicit discussion of difficult experiences and how to meet them. Both therapists present at every preparatory session. Sessions held in the dosing room wherever possible, so the patient becomes familiar with the space before entering an altered state in it. Everything videotaped. And communication between therapists and patient outside scheduled sessions discouraged, documented where it occurs, and routed through the study coordinator where possible.

The publication adds detail the manual does not: a mixed-sex therapist dyad in accordance with published safety guidelines, fourteen hours across eight in-person visits, eight weekly safety calls, and mandatory discontinuation of both monoaminergic medication and any ongoing psychotherapy before entry.

Almost none of that survives compression into "activities performed during the preparation sessions."

The contact prohibition is the sharpest case, because its entire content concerns the space between sessions. No item scoped to sessions will ever surface it, at any level of compliance.

The most instructive detail is the room. A team decided that patients should become familiar with the physical space before taking a strong psychedelic in it. That is a deliberate intervention on the participant's experience, plausibly more consequential than the lighting, and lighting has an item of its own.

None of this is a criticism of the ReSPCT authors. Their own limitations section says further empirical work is needed on the actual importance of variables both included and excluded. And when their panel was asked in the final round where setting begins and ends, the answers were striking: about 68 per cent said it runs from the start of recruitment to the end of follow-up, 16 per cent confined it to the treatment phase, and 10 per cent to the dosing session.

Ninety per cent of the panel located setting outside the dosing session. The instrument they built concentrates its resolution inside it.


What "set" was supposed to mean

The concept that should house this question has been in the field for sixty-five years and has been losing resolution for most of them.

Ido Hartogsohn's history traces it from the Club des Hashischins through 1950s psychotomimetic work on non-drug determinants of drug response and the extra-drug techniques of the psychedelic therapists of that decade, to Timothy Leary's formulation of the pair as a unit, presented to the American Psychological Association in September 1961 under the title Drugs, Set and Suggestibility. Leary and his collaborators restated the hypothesis through the decade. The claim was strong: that set and setting are the principal determinants of what a psychedelic experience contains.

Note what "set" meant there. Not the mood a person happens to be in when the drug is handed over. Everything they bring: personality, expectation, preparation, intention, prior experience, and beliefs about what is about to happen. In the original framing it was largely a made thing, and making it was the therapist's job.

Two things then happened.

The concept escaped the clinic. Compressed to something like "good time, good place, good people," it became genuinely useful harm-reduction advice and correspondingly imprecise. It became a thing everyone agreed about, which in scientific practice is often indistinguishable from a thing nobody investigates.

Then, when clinical research resumed, the half that could be photographed swallowed the half that could not. You can photograph a room, a playlist, a set of eyeshades, a number of people present. You cannot photograph a mind. Minds have to be measured, and measuring them requires instruments the field did not have. So the literature accumulated increasingly detailed descriptions of couches and increasingly perfunctory descriptions of what patients believed.

This is not an inference about how researchers think. It is visible in their own consensus-building. Asked whether set and setting can be studied separately, the ReSPCT panel split almost evenly: 37 per cent agreeing or strongly agreeing, 42 per cent disagreeing or strongly disagreeing. Asked to estimate the conceptual overlap, answers spread across the whole range, averaging around 49 per cent. Sixty-eight per cent reported that their own understanding of setting had changed during the study. Patch and Smith made the same diagnosis directly in the Journal of Psychopharmacology in 2025, arguing that the vagueness of the terms is itself an obstacle to research and practice.

That is a field discovering, while writing down what it believes, that it lacks a shared definition of its founding construct. The candour is admirable. The ambiguity still sits upstream of every attempt to measure what preparation does.

“Set” once meant everything a person brought to the experience. As science learned to measure the room, it became much worse at measuring the mind.
“Set” once meant everything a person brought to the experience. As science learned to measure the room, it became much worse at measuring the mind.

One further point, from the biology rather than the concepts.

Sixty years of set-and-setting claims have never been tied convincingly to a durable biological measurement. When a Copenhagen group went looking in July 2026 for a synaptic signature that might correspond to a context effect, the pre-registered primary outcome was null, the pre-registered exploratory test linking mystical-experience scores to synaptic change was also null, and the finding that reached the title was an unregistered, non-randomised comparison of five people against ten.

The relevance here is simple. There is no biomarker of context. There is no biomarker of preparation. There is unlikely to be one soon. Which means the pre-dose phase can only be measured behaviourally and procedurally, and that means it can only be measured by being written down. The absence of a biological handle raises the stakes on the reporting standard. It does not lower them.


Preparation is not expectancy

Any argument about the pre-dose period runs into the expectancy literature, and the two get conflated in ways that damage the argument in both directions.

The expectancy problem is severe and well characterised. In macrodose trials, participants almost always know their arm. The figures are stark. In an LSD study reported by Holze and colleagues, one participant of twenty mistook the drug for placebo. Bogenschutz and colleagues reported ninety of ninety-five correctly identifying allocation. EPIsoDE, one of the few trials to assess this properly, found 86 per cent correctly identifying the 25 mg condition. Muthukumaraswamy, Forsyth and Lumley concluded that effect sizes in psychedelic randomised trials are likely inflated by de-blinding and response expectancy. Szigeti and Heifets have since developed the analysis of how activated expectancy propagates through such trials.

This was the substance of the FDA advisory committee's difficulty with MDMA-assisted therapy in June 2024, and it is why the agency's final guidance of July 2026 asks sponsors to measure expectancy directly, to consider low doses of the same compound as comparators, and to use central raters blinded to allocation.

Expectancy and preparation overlap. They are not the same variable. Treating them as one produces two errors.

The deflationary error runs: preparation is a mechanism for generating expectancy, so any pre-dose improvement is a confound, and the correct response is less preparation delivered more neutrally. This is roughly the direction of travel in the commercial programmes, where the preferred term is psychological support rather than psychotherapy. There is an excellent regulatory reason for that preference. The FDA does not regulate psychotherapy, and a sponsor would rather not have its drug's efficacy entangled with a co-intervention that cannot be approved.

The inflationary error runs: preparation is obviously therapeutic, so any pre-dose improvement vindicates the therapy model, and the correct response is more of it. This skips the question of which components are active.

Both errors follow from the same failure to distinguish. Expectancy is a belief about outcome. Preparation includes at least four things that are not beliefs about outcome: procedural knowledge of what will happen, rehearsed strategies for managing difficulty, a relationship with the person who will be in the room, and a set of commitments made in order to be there at all.

The state of measurement is worse than the state of theory.

Ohio State captured expectancy with a single item from the Credibility/Expectancy Questionnaire. Participants rated it 6.67 out of 9, arriving confident. The relationship with outcome ran in the expected direction and did not reach significance. In an unblinded trial of twelve people, that is a study with almost no power to detect the effect in either direction, using an instrument too blunt to find it if it were there.

EPIsoDE, with 144 patients, a manualised fourteen-hour therapeutic programme, and a sample 89 per cent of whom had never taken a psychedelic, did not measure patient expectancy at all. Its limitations include the absence of adherence and therapy-quality ratings, therapeutic alliance among them. It states plainly that the contribution of expectancy to its results cannot be determined.

Nothing in the current reporting standard would have prompted either measurement.


What the evidence for preparation actually is

Set against the null on hours, the positive evidence is suggestive, indirect, and thinner than the confidence with which the field asserts that context matters. Here it is in full.

Pre-dose clinical movement. The Ohio State result. Uncontrolled, small, on one of two measures, partly regression to the mean, and the only recent trial to have looked.

Therapeutic alliance. Levin and colleagues, in a psilocybin trial for major depressive disorder, found the alliance between participants and facilitators associated with both acute effects and clinical outcomes. Murphy and colleagues traced a chain in a separate depression trial: alliance going into the first dosing session predicted pre-session rapport, which contributed to emotional breakthrough, which predicted symptom reduction at six weeks, with emotional breakthrough raising alliance going into the second session in turn. Alliance is built before dosing. This is the closest the field has to a mediational account running from the pre-dose period to outcome, and it rests on small samples and the ordinary fragilities of mediation analysis.

Prospective naturalistic prediction. Haijen and colleagues collected data at five time points from people who had independently planned to take a psychedelic, samples falling from 654 to 212. Comfort in the setting, including comfort with the people present, predicted higher wellbeing two weeks later. Intentions oriented to spiritual, therapeutic or nature-connection purposes were associated with better outcomes. In the fuller model, trait factors were more predictive than the acute experience itself. Self-selected, unblinded, self-reported, using drugs of unverified content, and still one of the largest prospective datasets bearing on the question.

A validated instrument. Rosalind McAlpine, George Blackburne and Sunjeev Kamboj developed the Psychedelic Preparedness Scale, published in Scientific Reports in 2024. Twenty items, four factors: Knowledge-Expectations, Intention-Preparation, Psychophysical-Readiness, Support-Planning. Validated in online samples of 516 and 716, then administered before and after a five to seven day psilocybin retreat with 46 participants. Reliability excellent. Participants scoring high before the experience differed from low scorers on wellbeing measures afterwards, in both prospective and retrospective versions.

This is the most important development in the area and it deserves more attention than it has had. It demonstrates that preparedness is measurable, psychometrically respectable, and carries predictive signal. Its authors are appropriately careful, describing the role of preparedness as critical but unproven and framing the scale as an enabling instrument rather than a finding.

An intervention trial. The same group took the obvious next step. The Digital Intervention for Psychedelic Preparation is a twenty-one day self-guided programme in four modules, tested in a randomised feasibility trial against a structurally identical music-listening control, forty healthy volunteers, a supervised 25 mg session at University College London, follow-up to nine months, with feasibility and adherence as primary outcomes.

That is the whole evidence base. One instrument, one feasibility trial, two alliance analyses, one naturalistic survey, and one small uncontrolled trial that happened to measure at the right moment.

The asymmetry between that and the meta-analysis is not what it appears. The null is well powered on a variable that cannot answer the question. The positive results are underpowered on variables that can. The correct response is not to average them. It is to run the studies that would produce a well-powered result on the right variable.

Notice also what every positive result on that list has in common. Each one came from somebody bringing a finer instrument. The PPS is a finer instrument than a count of hours. An alliance measure is a finer instrument than a therapist's name. A third measurement timepoint is a finer instrument than two. The signal appears wherever the graduation appears, which is what you would expect if the effect were real and also, admittedly, what you would expect if researchers who bothered to look were the ones who wanted to find something. That ambiguity is exactly what a properly designed trial resolves and a reporting standard cannot.


Four things called preparation

The conceptual obstacle is that "preparation" names at least four kinds of activity with different mechanisms, different evidential requirements, and different regulatory statuses. Studying them as one variable guarantees uninterpretable results, and it is a large part of why the analysis of hours came back flat.

Preparation as safety procedure. Screening, medical clearance, medication tapering, transport arrangements, agreement on what happens if the participant becomes distressed. The purpose is risk reduction. The mechanism is logistical. Nobody expects it to be therapeutic and its absence is a protocol violation.

Preparation as psychotherapy. Case formulation, work on presenting problems, therapeutic relationship, sometimes explicit trauma work. A clinical intervention with its own evidence base and its own effect size independent of any drug. Also, from a sponsor's perspective, the most dangerous kind to admit to, because it invites the question the advisory committee asked about MDMA.

Preparation as ritual. Practices whose function is not conveyed by their content. Dietary restriction, abstinence, isolation, fasting, the marking of a threshold, a commitment made before witnesses. These carry no mechanistic story in a clinical framework and are usually removed as cultural rather than active.

Preparation as active therapeutic intervention in its own right. Psychoeducation, expectation calibration, intention formation, rehearsal for difficult experiences, training in specific strategies for the acute state. This is the category the literature gestures at most and defines least. It is what the PPS was built to measure and what DIPP was built to deliver.

The distinction is not academic. Suppose a dismantling trial found that removing preparation reduced effect size. Which of the four was removed? If the psychotherapy, the finding is about therapy. If the safety procedure, it may be about adverse events rather than efficacy. If the psychoeducation, it is about the component that is cheapest to scale and easiest to deliver digitally.

Those are three different conclusions with three different consequences for how psychedelic treatment gets built and paid for. No current reporting standard can tell them apart, which means no dismantling trial run against the current standard can either.


The traditions, and what they are and are not evidence of

I should declare an interest. I spent six months living with a Bwiti community in Gabon and completed a full iboga initiation, and seven years in Ecuador before that.

Whatever else that provides, it provides an observation. In the traditions I know directly, whether preparation is separable from the medicine is not an unanswered question. It is an unaskable one. The days beforehand, the restrictions, the specific people present, the ordering of events, the relationship with the person guiding you: none of these are understood as delivery mechanisms for an alkaloid. They are understood as the treatment, of which the alkaloid is one component.

This is not evidence, and it is worth saying so at length, because this publication's position has been consistently against treating traditional practice as a source of conclusions rather than hypotheses.

Two cautions apply.

Ceremonial confidence about mechanism has been wrong before. A practice can be elaborated, coherent, carefully transmitted across generations, and mistaken about why it works. Bloodletting had all of those properties. A detailed indigenous theory of preparation is evidence that practitioners believe preparation is active. It is not evidence that it is.

The second caution cuts against a story the Western psychedelic world tells about itself. The elaborate pre-ceremony diet that is standard at commercial retreats is not a pan-Amazonian universal. Glenn Shepard's ethnographic work on the Matsigenka records ayahuasca used in flexible, often domestic contexts without rigid ritual or dietary preparation. In a number of Amazonian communities the brew is drunk collectively without extended preparatory restriction. The genuinely strict regime of isolation, bland food, sexual abstinence and silence is the dieta proper: the extended master-plant apprenticeship undertaken over weeks or months, typically by people training to heal rather than people seeking healing.

Conflating those two, the short preparatory diet and the long apprenticeship, is substantially a product of the retreat economy, and it has been exported back into Western discourse as though it were ancient and universal. This publication's standing position is that ayahuasca is not one thing, and that brew, tradition, region and lineage must be specified where recoverable. The same discipline applies here. "Indigenous preparation" is not a protocol, and anyone invoking it as a corrective to clinical practice owes the reader a lineage.

What survives both cautions is narrower and more useful. Across a range of traditions the pre-dose period is elaborated rather than minimised, and elaborated in structurally similar ways: restriction, instruction, commitment, relationship, and the marking of a boundary between ordinary time and ceremonial time. Some of that has an obvious pharmacological rationale, notably tyramine restriction where an MAO inhibitor is involved. Most does not. The convergence is a reason to generate hypotheses, not to accept them.

There is a small sign that the clinical field half-knows this. When the ReSPCT panel voted, five candidate items failed to reach whole-group consensus: odours and scents, temperature, facilitator demographics, sociocultural context, and social determinants of health. The panel included a subgroup of four people whose primary expertise was categorised as plant medicine, and that subgroup returned unanimous votes for four of the five. Four out of eighty-nine is 4.5 per cent, and a seventy per cent threshold is a blunt instrument against a minority that size.


Removing components before establishing what they do

Clinical psychedelic research has, over fifteen years, built its protocols largely by subtraction from ceremonial practice. Eyeshades, curated music, the two-person dyad, the couch, the instruction to go inwards and accept what arises: these are recognisably descendants of ritual forms, retained where they seemed defensible and removed where they seemed superstitious.

The removals were not preceded by testing. There is no literature establishing that dietary restriction before psilocybin is inert. There is no literature establishing that a commitment made in front of family functions differently from a consent form signed alone. There is no literature establishing that the absence of a threshold marker changes nothing about how the day is encoded.

These components were removed because they did not fit the framework, which is an ordinary thing for a research programme to do, and is not the same as having shown they do nothing.

The asymmetry is worth naming. When a trial adds a component resembling ritual, curated music being the obvious case, that component eventually acquires its own literature and its own trials. When a trial removes one, the removal generates no literature at all. There is no graduation for absence. A field can therefore accumulate substantial evidence about what it kept while remaining permanently silent about what it discarded, and the silence will be read as absence of effect, in exactly the way the null on hours is now being read.

I do not want to overstate this. Some removed components are almost certainly inert, and some are worse than inert. Practices that increase dependence on a practitioner, or that establish obligations a participant cannot refuse, are harms rather than active ingredients, and the field is right to have left them behind.

The claim is not that traditional preparation should be reinstated. It is that "we removed it and the treatment still worked" is not a finding, because nobody measured what it would have done had it stayed.


How you would actually research this

Complaining that a variable is unmeasured is worthless on its own. What follows is what to measure and how to find out.

A functional taxonomy

The unit of analysis should not be the session or the hour. It should be the function the preparatory activity performs. Function generalises across a clinic in Ohio, a retreat in Peru and an initiation in Gabon, and function is what a dismantling design can manipulate.

What follows is a working taxonomy of eleven domains. The claim is not that these are the correct eleven. It is that a taxonomy of roughly this shape is what the field needs and lacks, and that each domain can be specified precisely enough to be reported, scored, and in principle experimentally removed.

D1. Procedural knowledge. What the participant was told will physically happen: timings, dose, route, duration, who will be present, what the room contains, what they may and may not do. Reportable as content, source and repetition.

D2. Phenomenological priming. What the participant was told the experience will be like, and by whom. The domain most likely to shape the acute experience and least likely to be documented. It includes vocabulary: whether the materials speak of mystical experience, emotional breakthrough, ego dissolution, or nothing at all. A trial whose framing anticipates mystical experience and whose primary mediator is a mystical experience questionnaire has a circularity problem that nobody currently has to disclose.

D3. Expectation calibration. The direction and confidence of stated outcome beliefs, whether the team corrected optimistic expectations, and whether accurate base rates were communicated. Distinct from D2 in concerning outcome rather than experience.

D4. Intention formation. Whether a purpose was articulated, in what form, with what specificity, whether it was recorded, and whether it was revisited.

D5. Difficulty rehearsal. Explicit preparation for challenging states: instructions for what to do if frightened, agreed signals, rehearsed strategies, prior framing of difficulty as potentially therapeutic rather than as failure. EPIsoDE manualised this. Most trials do not report whether they did it at all.

D6. Relational preparation. Alliance and its measurement, contact hours with the specific individuals who will be present at dosing, whether the dosing personnel are the preparing personnel, and the composition of the dyad.

D7. Commitment and cost signalling. What the participant gave up, submitted to, or publicly undertook in order to be there. Dietary restriction, abstinence, travel, time, financial cost, disclosure to family, medication washout, and, as in EPIsoDE, the requirement to discontinue existing psychotherapy. This is where most traditional preparation clusters, and where clinical protocols are near-silent about function even when they impose the requirement. The hypothesis is that costly, effortful, publicly witnessed commitment does work that a consent form does not. It is testable.

D8. Psychophysical readiness. Sleep, substance abstinence, medication status, washout duration, physical state, and any contemplative practice in the preparatory window. Partly captured by the PPS.

D9. Boundary and containment. How the transition into the treatment period was marked, and what rules governed contact outside scheduled sessions. EPIsoDE discouraged between-session contact and documented it where it occurred, a deliberate and consequential decision that no session-scoped item can surface. The absence of a marker, and the absence of a contact rule, are themselves reportable facts.

D10. Support architecture. What exists on the other side. Who knows, who will be present afterwards, what integration is scheduled and with whom, and whether the participant knew this in advance. Anticipated support is a pre-dose variable even though the support itself is not.

D11. Provenance and governance disclosure. What the participant was told about where the compound came from, who owns it, who profits, and, where a plant-derived or traditionally-derived compound is involved, whether benefit-sharing obligations were addressed. This looks like an ethics item and is also a fidelity item. A participant who has been told that the molecule descends from a tradition whose holders were compensated is in a materially different state from one who has not. Whether that difference matters clinically is unknown, which is the point. It has never been measured, and it is trivially cheap to record. The same logic applies to technique. A preparatory practice adopted from a named lineage with practitioner involvement and the same practice adopted from a secondary summary are not the same intervention, and they will not be delivered the same way.

Eleven domains, each specifiable in two or three sentences, each scoreable as present, absent or partial. None of this requires a micrometer. It requires graduations where there are currently none.

Instruments

Two exist and should be standard. The Psychedelic Preparedness Scale, administered before preparation begins and again immediately before dosing, giving a change score rather than a single value. And a validated multi-item expectancy instrument at the same two points, so that preparedness and expectancy can be modelled separately rather than confounded.

To these should be added a preparation fidelity record: a structured form completed by the preparing clinician, indicating which domains were addressed, for how long, and by whom. This is the psychedelic equivalent of the adherence checklists that are routine in psychotherapy trials and, as Hultgren and colleagues note, largely absent here.

Designs

Four, in ascending order of cost.

Measurement only. Add an assessment of the primary outcome at the end of preparation, before dosing, to every trial. This is nearly free. It is what Ohio State did. If it were standard, within three years there would be a meta-analysable dataset on pre-dose movement across dozens of trials and thousands of participants. Nothing else on this list matters as much, because it costs almost nothing and no argument against it survives contact with the Ohio State result.

Preparation dose-response. Randomise participants to two, four or eight hours of the same manualised preparation, holding drug dose constant. This is the study Hultgren and colleagues could not assemble from the literature, because across trials the number of hours is confounded with everything the hours contained. Within a single protocol it is not. If preparation is doing work, there should be a gradient.

Factorial dismantling. A therapy-only arm delivering the same hours of skilled clinical contact without the drug, alongside drug-plus-therapy and drug-with-minimal-support arms. Expensive, commercially unattractive, and the study the field has avoided for a decade. The FDA's endorsement of low-dose comparators in July 2026 makes a version of it more feasible than it was.

Component removal. Within a fixed protocol, randomise the presence or absence of specific domains. D5 and D7 are the best candidates. Both are cleanly manipulable, both are ethically defensible to withhold, and both have clear predicted directions.

What would falsify this

An argument should say what would sink it.

If a well-powered within-protocol dose-response trial found no gradient across two, four and eight hours; if pre-dose movement across a large pooled sample proved fully attributable to regression to the mean once modelled properly; and if PPS change scores failed to predict outcome in clinical rather than retreat populations, then the reasonable conclusion would be that preparation functions largely as a safety and consent procedure with modest expectancy effects, and that the field's current allocation of attention is roughly correct.

I do not expect that result. I would accept it.


Why the timing is not academic

Compass Pathways reported a positive Phase 3 result for COMP360 in treatment-resistant depression in June 2025 and a second in February 2026, with a new drug application indicated for the fourth quarter of 2026. If approved, psilocybin would be the first classic psychedelic authorised for a psychiatric indication.

An approval is a decision about a label. The label will determine what insurers pay for, what clinics deliver, and what the treatment becomes at scale.

There are three ways this goes wrong, and they are not symmetrical.

If preparation is doing substantial work and the label does not require it, the treatment will be delivered stripped, will underperform the trial data, and the underperformance will be attributed to the drug. That is the failure mode that ends renaissances. It is also the cheapest configuration commercially, which means market pressure selects for it. And the null on hours, read carelessly, is precisely the citation that would justify it.

If preparation is doing little and the label requires a great deal of it, access is restricted and cost inflated for no clinical return, and the people excluded are those with the least time and money. Equity is not a secondary consideration here.

And if the preparation component is entangled with the drug effect in a way nobody has measured, the label is written on a category error that every subsequent comparative trial inherits.

There is a fourth consideration, belonging to this publication's particular preoccupations. If preparation turns out to be substantially active, then a good deal of what the clinical model discarded as culture may turn out to have been mechanism, and the compounds' relationship to the traditions they came from becomes something more than historical courtesy. That is a large conditional and I am not asserting it. I am noting that it is currently unanswerable, and that the reason it is unanswerable is a reporting standard.


The rule and the part

Psychedelic research is attempting to establish that a drug causes an outcome in patients whom it deliberately alters for weeks beforehand, using a record that describes the alteration as a number of hours. It has now published a meta-analysis showing that the number of hours predicts nothing.

The second sentence will be quoted. The first is why it should not be believed.

None of this requires anyone to reopen ReSPCT, and nothing here is an argument that they should. Reporting standards handle exactly this situation with extensions, and the precedents sit in the guidelines' own reference list: CONSORT has one for social and psychological interventions, and there is an established standard for reporting implementation studies of complex interventions. The mechanism for adding graduations to one subdomain without disturbing the parent instrument already exists and is well understood.

The case for using it here does not have to be made against the ReSPCT authors either. They have made most of it themselves. Their limitations section says further empirical work is needed on the actual importance of the variables both included in and excluded from the guidelines. Ninety per cent of their expert panel located setting outside the dosing session. And a checklist built, correctly, to be parsimonious enough that people would complete it was never going to carry the preparatory phase at the resolution the phase turns out to need. The gap is the consequence of a good design decision, not the failure of one.

What an extension would require is unglamorous and specific. An audit establishing what items 20 and 21 leave unrecoverable across a defined body of registered protocols. A functional taxonomy of the preparatory phase capable of being reported without adding meaningfully to anyone's burden. Then the ordinary machinery of consensus. The first of those is a piece of work somebody could complete this year.

In the meantime there is one change that costs almost nothing. Measure the primary outcome once more, at the end of preparation, before the first dose.


Sources

Pronovost-Morgan, C., Greenway, K. T., Roseman, L., & The ReSPCT Experts. (2025). An international Delphi consensus for reporting of setting in psychedelic clinical trials. Nature Medicine, 31, 2186–2195. doi:10.1038/s41591-025-03685-9

Hultgren, J., Hafsteinsson, M. H., & Gruneau Brulin, J. (2025). A dose of therapy with psilocybin: a meta-analysis of the relationship between the amount of therapy hours and treatment outcomes in psychedelic-assisted therapy. General Hospital Psychiatry, 96, 234–243. doi:10.1016/j.genhosppsych.2025.07.020

Armstrong, S. B., Levin, A. W., Sepeda, N. D., et al. (2026). Safety, feasibility, and preliminary clinical outcomes of psilocybin-assisted therapy for veterans with severe, treatment-resistant PTSD: an open-label pilot clinical trial. Communications Medicine, 6, 411. doi:10.1038/s43856-026-01767-4

Mertens, L. J., Koslowski, M., Betzler, F., et al. (2026). Efficacy and safety of psilocybin in treatment-resistant major depression: the EPIsoDE randomized clinical trial. JAMA Psychiatry, 83(5), 448–460. doi:10.1001/jamapsychiatry.2026.0132. With the openly published therapist manual, version 3.71.

McAlpine, R. G., Blackburne, G., & Kamboj, S. K. (2024). Development and psychometric validation of a novel scale for measuring 'psychedelic preparedness'. Scientific Reports, 14, 3280. doi:10.1038/s41598-024-53829-z

Hartogsohn, I. (2017). Constructing drug effects: a history of set and setting. Drug Science, Policy and Law, 3, 1–17. doi:10.1177/2050324516683325

Also referenced

Levin, A. W., Lancelotta, R., Sepeda, N. D., et al. (2024). The therapeutic alliance between study participants and intervention facilitators is associated with acute effects and clinical outcomes in a psilocybin-assisted therapy trial for major depressive disorder. PLOS ONE, 19(3), e0300501.

Haijen, E. C. H. M., Kaelen, M., Roseman, L., et al. (2018). Predicting responses to psychedelics: a prospective study. Frontiers in Pharmacology, 9, 897.

Muthukumaraswamy, S. D., Forsyth, A., & Lumley, T. (2021). Blinding and expectancy confounds in psychedelic randomized controlled trials. Expert Review of Clinical Pharmacology, 14(9), 1133–1152.

Szigeti, B., & Heifets, B. D. (2024). Expectancy effects in psychedelic trials. Biological Psychiatry: Cognitive Neuroscience and Neuroimaging.

Patch, K., & Smith, W. R. (2025). What are set and setting: reducing vagueness to improve research and clinical practice. Journal of Psychopharmacology, 39(9), 900–909.

Brennan, W., Kelman, A. R., & Belser, A. B. (2023). A systematic review of reporting practices in psychedelic clinical trials: psychological support, therapy, and psychosocial interventions. Psychedelic Medicine, 1(4), 218–229.

Bridging the reporting gap: application of the ReSPCT guidelines in psilocybin clinical trial protocols. (2026). European Neuropsychopharmacology.

Digital Intervention for Psychedelic Preparation (DIPP). ClinicalTrials.gov NCT06815653.

FDA, Psychedelic Drugs: Considerations for Clinical Investigations, final guidance, July 2026. Docket FDA-2023-D-1987.


Previously on ARDMT: Before the First Dose, on the Ohio State veterans trial and the measurement in the middle; Seventeen to One, on what a reporting standard decides can be discovered; The Room and the Record, on a null result and the title it acquired.