Ethics · Research


The field has begun to suspect that preparation is doing some of the work. Its consensus reporting standard gives the preparatory phase a single item.

Last week this publication looked at a trial in Ohio that measured PTSD severity three times instead of twice, and found that the middle measurement moved. Twelve veterans, eight hours of preparatory therapy, not a milligram of psilocybin, and clinician-rated severity already down 6.33 points with an effect size of 0.87. How much a participant improved during preparation predicted how much they improved a month after dosing.

The obvious next question is what those eight hours contained.

The answer, for that trial and for nearly every other, is that we largely cannot say. Not because the researchers were careless — the Ohio State team were unusually careful, and I will come back to what they did right — but because the field's agreed standard for describing the contexts of psychedelic treatment allocates the preparatory phase one item out of thirty.

This is a story about a checklist. It turns out to be a story about what a checklist decides can be discovered.


The instrument

In June 2025, Nature Medicine published the Reporting of Setting in Psychedelic Clinical Trials guidelines. ReSPCT is a serious piece of work: an international Delphi consensus, eighty-nine experts across seventeen countries and fifty institutions, four iterative rounds of surveys and facilitated debate, thirty percent attrition by the final round. Round one asked each expert to name the ten setting variables they considered most important to report. That produced 770 free-text responses, synthesised into forty-nine candidate items, rated and debated and merged down to the thirty that cleared a pre-specified threshold of seventy percent.

The thirty items sit in four sections. Items one to seven cover the physical environment. Items eight to seventeen cover the dosing session procedure. Items eighteen to twenty-five cover the therapeutic framework and protocol. Items twenty-six to thirty cover participants' subjective experiences.

The dosing day is specified in detail that is genuinely impressive. Ambiance. Access to nature. Objects and decorations. Lighting. Sensory reduction devices. Bathroom accessibility and privacy. The number and roles of people present. The relative positioning of bodies in the room. Whether attention is directed inward or outward. Music. Verbal and physical interpersonal interventions, and how consent for them was obtained. How much control the participant had over their own environment. Dosing regimen. Concurrent medical procedures. What happened immediately before and after dosing. Disturbances and interruptions.

Preparation gets item twenty, which asks for the number and length of preparation, dosing and integration sessions, and item twenty-one, which asks for <cite index="4-1">"activities performed during the preparation sessions."</cite> Integration gets the equivalent single item.

I want to be fair about this, because the crude version of the complaint is wrong. Other items do touch the preparatory phase obliquely. Item eighteen asks for the therapeutic approach and its manual. Item nineteen asks how the study team framed the intervention, including the short- and long-term drug effects — which is, functionally, a question about expectancy formation. Items twenty-four and twenty-five ask about personnel qualifications and cultural competence. None of these are phase-specific. None ask when the framing happened, who delivered it, in what words, or whether the participant was told anything different at screening than they were told the day before dosing. They describe the study's general character rather than the sequence a patient actually moved through.

So the honest statement of the ratio is this: seventeen items describe the room and the day inside it, at the resolution of bathroom privacy. One item describes the content of everything that happened before.


What one item costs

You can watch the cost being incurred.

In 2025, Jennie Hultgren and colleagues at Stockholm University published the first meta-analysis to ask whether the amount of therapy in psilocybin-assisted treatment predicts outcome. Sixteen studies, nineteen reports, a systematic search of PubMed and PsycINFO. The overall treatment effects were very large: Cohen's d of 1.69 in the short term, 2.10 at longer follow-up.

The amount of therapy predicted nothing. Total therapy hours were unrelated to effect size in the short term, b = −0.05, p = .327, and in the long term, b = −0.07, p = .340. When they modelled preparation and integration separately, preparation came out at b = −0.01, p = .912. Integration, b = −0.11, p = .170.

That result has already begun to circulate as evidence that the therapy in psychedelic-assisted therapy may not be doing much. It is worth reading what the authors themselves say about why they could not find anything.

The predictor they tested was hours. Not content, not sequence, not what the hours contained — the arithmetic sum of preparation and integration time, within a range that ran from 4.5 to 18 hours across the entire literature. They describe the reporting of the therapeutic component as <cite index="1-1">severely insufficient</cite>, and they mean it specifically: studies reporting optional therapy hours without recording how many participants took them up, studies reporting wide ranges without specifics, different reports of the same trial giving different numbers. They note that the manuals in use are typically loose, that fidelity and adherence to them are rarely assessed, and that it is not clear the interventions delivered meet established definitions of psychotherapy at all. Their recommendations are that the therapeutic component be standardised and reported with the rigour applied to the pharmacological one, and that preparation and integration be separated so their respective contributions can be distinguished.

A null on hours is not a finding about preparation. It is a finding about the record. The only variable the published literature could supply was a count of minutes, and a count of minutes is not a description of a practice. If two trials both report eight hours, and one spent them on psychoeducation and rapport while the other spent them on breath training and intention-setting, the meta-analytic dataset records them as identical.


The manual and the publication

The clearest way to see what item twenty-one does not hold is to find a trial where the underlying document is public, and read the two side by side.

The German EPIsoDE trial published in JAMA Psychiatry in March 2026: two centres, triple-blind, 144 patients with treatment-resistant depression, psilocybin 25 mg against psilocybin 5 mg and nicotinamide. Its therapist manual has been released openly. This is a rare and valuable thing, and the team deserve credit for it.

The manual specifies a general preparatory session seven days before the first dose, running approximately 100 minutes, to a checklist: about ten minutes on introductions and an overview, thirty to forty-five on the topics relevant to the patient's depression, ten on previous treatments, ten to twenty-five on the patient's expectations and hopes regarding psilocybin, five on the agreements necessary for participation, ten to hand over written information about dosing sessions. Further sessions the day before each dose include a guided mindfulness exercise and discussion of what commonly happens during dosing, including difficult experiences and how to meet them. Both therapists must be present at every preparatory session. The sessions should be held in the dosing room wherever possible, explicitly so the patient becomes familiar with the room and its equipment beforehand. All of it is videotaped. And communication between therapists and patient outside scheduled sessions is discouraged; where it happens it must be documented and, if possible, routed through the study coordinator, so as not to confound the study.

The publication adds detail the manual does not. Patients were assigned two therapists, one female and one male, in accordance with published safety guidelines. The psychotherapeutic programme comprised seven two-hour sessions, fourteen hours in total, across eight in-person visits, with eight weekly therapist safety calls. Entry required discontinuation of ongoing psychotherapy as well as of monoaminergic medication, with a minimum abstinence of two weeks and five weeks for fluoxetine. Eighty-nine percent of the sample had never taken a psychedelic.

Now ask which of that is recoverable from "activities performed during the preparation sessions."

The contact prohibition is the sharpest case, because its entire content concerns the space between sessions. No degree of compliance with an item scoped to sessions will ever surface it. Nor will the washout, nor the requirement to stop existing psychotherapy, nor the eight safety calls, nor the mixed-sex dyad, nor the decision to hold preparation in the room where the drug would later be taken.

That last one is worth sitting with. A team decided that patients should become familiar with the physical space before entering an altered state in it. That is a considered intervention on the participant's experience, arguably a more consequential one than the lighting in the room, and lighting has its own item.

There is a second thing about EPIsoDE that belongs here. It assessed functional unblinding in patients and therapists, which its authors rightly identify as a strength rarely present in this literature. Eighty-six percent of participants correctly identified the 25 mg condition. But the trial did not measure patients' expectations at all, and lists among its limitations the absence of adherence and therapy-quality ratings, therapeutic alliance included. So a trial that manualised fourteen hours of therapeutic contact down to checklist level, in a predominantly psychedelic-naive treatment-resistant sample, cannot say what its patients expected of the treatment or how well they got on with the people preparing them. It states plainly that the contribution of expectancy to its results cannot be determined.

Nothing in ReSPCT would have prompted either measurement.


The four panellists

There is a table in the ReSPCT supplementary material that I have not seen discussed anywhere, and it is the part of the study I find hardest to stop thinking about.

Five candidate items failed to reach whole-group consensus. Odours and scents, at fifty-eight percent. Temperature, fifty-eight. Facilitators' demographics and cultural backgrounds, sixty. Sociocultural context, fifty-six. Social determinants of health, sixty-eight, which is to say two percentage points short.

The published subgroup breakdown shows how those votes distributed. The panel had a subgroup whose primary expertise was categorised as plant medicine. It comprised four people out of eighty-nine. That subgroup returned four out of four — unanimous — on odours, on temperature, on sociocultural context, and on social determinants of health. It also returned four out of four on both halves of the cultural competence and safety item, which entered the debate round at whole-group ratings of fifty-two and forty-seven percent and was rescued to seventy-three after experts argued for it, becoming item twenty-five.

Precision matters here, so: the plant medicine subgroup did not endorse the facilitator demographics item, which was carried in other subgroups. Four of the five, not five of five.

Still. Four of the five variables the panel declined to include concern the sensory, embodied and socially situated features of a healing encounter — what the space smells like, what temperature the body is held at, the wider social and legal world the person came from and returns to. And the smallest expertise group on the panel voted unanimously for all four while the whole group did not.

The generous reading of this is also the correct one. The ReSPCT authors ran subgroup analyses specifically because subgroup analyses are not standard in Delphi work and they wanted to surface minority positions that would otherwise vanish into an aggregate. They then put the divergent items into a facilitated debate round rather than simply dropping them, and one of those items survived and made the final instrument. The mechanism worked. It is visible in their paper only because they built it and reported it.

But four of eighty-nine is 4.5 percent, and a seventy-percent threshold is a blunt instrument against a minority that size. The paper's own limitations say that the focus on clinical trials excluded broader, non-clinical perspectives on the variables shaping psychedelic experiences, and that further empirical work is needed on the actual importance of the variables both included and excluded.


Where the panel thought setting was

One more result from the final round, which I think is the most quietly interesting number in the paper.

Asked about the temporal boundaries of setting, sixty-eight percent of the experts agreed that it runs from the start of recruitment to the end of follow-up. Sixteen percent confined it to the treatment phase. Ten percent confined it to the dosing session.

Ninety percent of the panel located setting outside the dosing session. The instrument they produced concentrates its resolution inside it.

That is not hypocrisy and it is not oversight. It is what happens when you build something parsimonious enough to be adopted. The authors state that they aimed for guidelines that were straightforward to implement, and they were right to: a sixty-item checklist that nobody completes documents nothing at all. Thirty items that get used are worth vastly more than sixty that don't, and ReSPCT is an enormous improvement on the previous standard, which was a sentence in the methods section and the reader's imagination.

The gap between stated scope and operationalised scope is simply the next piece of work. Reporting guidelines handle this with extensions — CONSORT has one for social and psychological interventions, and there is a standard for implementation studies of complex interventions. Both are cited in ReSPCT's own reference list. The mechanism for adding resolution to a subdomain without reopening the parent instrument already exists, and its authors have already told us where they think the resolution is missing.


Back to Ohio

Which brings me back to the trial I started with, and to a detail I left out.

The Ohio State team reported their setting against the ReSPCT guidelines. Of all the things that trial did right — pre-registration before recruitment opened, an independent assessor scoring the primary outcome, participant-level data published, a candid limitations section stating that small uncontrolled trials produce inflated effect sizes — reporting against the consensus standard is the one most directly relevant here.

And a trial that did everything right, reporting fully against the standard, was still not obliged to say what its eight hours of preparation consisted of.

That is the shape of the problem. The one trial in the literature that measured before and after preparation, found the number moving, and reported that finding rather than burying it, could comply completely with the field's agreed reporting instrument and leave the content of the phase unrecorded. If somebody wanted to replicate that middle measurement, the published record tells them how many hours to buy and almost nothing about what to do with them.


What would have to change

Not much, and it is describable.

Items specific to the phase rather than to the study in general: what the participant was told and when, whether an intention was elicited and whether it was revisited, what was rehearsed rather than merely discussed, what was restricted and for how long, who was present and whether the participant chose them, what contact was permitted between sessions, what the participant was required to give up as a condition of entry.

A separation the literature currently collapses: what was done, why the people doing it say they did it, and what an analyst hypothesises it might alter. Those are three different claims and conflating them is how this subject goes wrong in both directions — by treating a stated rationale as a mechanism, or by substituting somebody's mechanism for their account of their own practice.

Expectancy and alliance reported as data alongside the protocol that produced them, or an explicit statement that neither was measured. Validated instruments exist. A single item from a credibility questionnaire, which is what the Ohio State trial had available, cannot answer the question in either direction, and "expectancy was measured and was not a factor" is exactly the sentence that detaches from a paper and travels.

And provenance of technique, where a preparatory practice has a documented origin: where it came from, from whom, with what involvement or permission. Not as an ethical preface, but as a fidelity variable, because a practice adopted from a named lineage with practitioner involvement and the same practice adopted from a secondary summary are not the same intervention and will not be delivered the same way.

I should declare an interest. This is the subject I am about to spend a year on, and readers should discount accordingly. But the argument does not depend on my doing it. It depends only on the observation that the field has started finding signals in the preparatory phase while holding an instrument that cannot describe what the phase contains — and that when the meta-analysts went looking, the only thing the record could give them was a number of hours.

The scanner in Copenhagen altered the experience it was built to observe. A reporting standard does something quieter and more consequential. It determines which questions can later be asked of the record. Preparation currently gets one item, and one item is not enough to find anything with.


Sources

Pronovost-Morgan, C., Greenway, K. T., Roseman, L., & The ReSPCT Experts. (2025). An international Delphi consensus for reporting of setting in psychedelic clinical trials. Nature Medicine, 31, 2186–2195. doi:10.1038/s41591-025-03685-9

Hultgren, J., Hafsteinsson, M. H., & Gruneau Brulin, J. (2025). A dose of therapy with psilocybin: a meta-analysis of the relationship between the amount of therapy hours and treatment outcomes in psychedelic-assisted therapy. General Hospital Psychiatry, 96, 234–243. doi:10.1016/j.genhosppsych.2025.07.020

Mertens, L. J., Koslowski, M., Betzler, F., et al. (2026). Efficacy and safety of psilocybin in treatment-resistant major depression: the EPIsoDE randomized clinical trial. JAMA Psychiatry, 83(5), 448–460. doi:10.1001/jamapsychiatry.2026.0132

Therapist manual for psilocybin-assisted therapy of treatment-resistant major depression (EPIsoDE), version 3.71.

Armstrong, S. B., Levin, A. W., Sepeda, N. D., et al. (2026). Safety, feasibility, and preliminary clinical outcomes of psilocybin-assisted therapy for veterans with severe, treatment-resistant PTSD: an open-label pilot clinical trial. Communications Medicine, 6, 411. doi:10.1038/s43856-026-01767-4

Levin, A. W., Lancelotta, R., Sepeda, N. D., et al. (2024). The therapeutic alliance between study participants and intervention facilitators is associated with acute effects and clinical outcomes in a psilocybin-assisted therapy trial for major depressive disorder. PLOS ONE, 19(3), e0300501.

Also referenced

Patch, K., & Smith, W. R. (2025). What are set and setting: reducing vagueness to improve research and clinical practice. Journal of Psychopharmacology, 39(9), 900–909.

Brennan, W., Kelman, A. R., & Belser, A. B. (2023). A systematic review of reporting practices in psychedelic clinical trials: psychological support, therapy, and psychosocial interventions. Psychedelic Medicine, 1(4), 218–229.


Previously on ARDMT: Before the First Dose, on the Ohio State veterans trial and the measurement in the middle; The Room and the Record, on what a Copenhagen PET study found and what it didn't.