# Psilocybin certainty and information analysis v0.1

**Framework:** The Shape of Harm v0.5  
**Estimand:** EST-AE-PSI-001  
**Outcome:** Investigator-attributed drug-related serious adverse event within seven days  
**Status:** Single-author provisional assessment. It is not an independently adjudicated GRADE assessment or a completed systematic review.

## The result is a bound, not a safety score

The two direct randomized MDD trials observed **0 events among 67 psilocybin administrations**. The two-sided 95% Clopper–Pearson interval is **0 to 53.6 events per 1,000 administrations**. The observation “zero occurred in these trials” is certain as a report of the extracted data; the inference that the underlying risk is very low is not.

A single missed or differently classified event would change the point estimate to **14.9 per 1,000** and widen the exact interval to **0.4–80.4 per 1,000**. That fragility is why the website must display the interval and evidence status next to the count.

## Provisional outcome-level GRADE assessment

Randomized trials begin at high certainty. This pilot rates the underlying risk estimate **very low certainty, provisionally**:

| Domain | Judgment | Reason |
|---|---|---|
| Risk of bias | Serious (−1) | Functional unblinding, subjective causality attribution, and outcome reporting that was not designed around the exact seven-day estimand. |
| Inconsistency | Not downgraded | Both direct trials observed zero drug-related SAEs. With two sparse studies, heterogeneity is not meaningfully estimable; this limitation is handled under imprecision. |
| Indirectness | Not serious for the narrow scenario | Both trials closely match pharmaceutical 25 mg psilocybin, adults with MDD, an active comparator, and a highly supported clinical setting. |
| Imprecision | Very serious (−2) | The upper confidence limit remains 53.6 per 1,000 and the information size is far below what is required to exclude uncommon serious harm. |
| Missing evidence / publication bias | Cannot assess | The complete database, registry, citation, and unpublished-evidence search has not been completed. |
| Overall | **Very low, provisional** | The true risk may be substantially different from the current point estimate. |

This follows GRADE's outcome-centric approach and its requirement to judge risk of bias, inconsistency, indirectness, imprecision, and missing evidence transparently. It does **not** turn GRADE into a numeric component of the harm score.

## Information-size targets

With zero observed events, the following total exposed sample sizes are required for the **two-sided 95% exact upper limit** to fall below each threshold:

| Upper bound sought | Required total n | Additional zero-event participants beyond 67 |
|---:|---:|---:|
| < 50 per 1,000 | 72 | 5 |
| < 25 per 1,000 | 146 | 79 |
| < 10 per 1,000 | 368 | 301 |
| < 5 per 1,000 | 736 | 669 |
| < 1 per 1,000 | 3,688 | 3,621 |

These are precision targets, not assertions that any threshold is socially or clinically acceptable. A later decision panel would have to justify any action threshold separately from the evidence synthesis.

## Claim rules now enforced

The public research page may say:

- no investigator-attributed drug-related SAE was observed among 67 participants in the two direct trials;
- the exact 95% interval extends to 53.6 per 1,000;
- certainty about the underlying risk is very low because the evidence is small and incompletely reviewed.

It may not say:

- the risk is zero;
- psilocybin is proven safe;
- psilocybin and niacin are equivalent in safety;
- the estimate applies to unscreened or recreational use;
- the estimate can be converted into a universal 0–100 harm score.

## Why most prespecified outcomes remain ungraded

GRADE evaluates a body of evidence for a defined outcome. All-cause serious events, severe psychiatric events in exactly seven days, and rescue or prolonged-observation outcomes are still **not estimable** because the studies report incompatible windows or units. A missing effect estimate should not be replaced with a certainty label.

## Next scientific gate

1. Peer review and translate the search strategy.
2. Complete all databases, registries, citation searches, and sponsor/author queries.
3. Have two reviewers independently screen, extract, and complete outcome-level RoB 2.
4. Adjudicate the GRADE profile without looking at the preferred public narrative.
5. Re-run the frozen sensitivity and information-size analysis.

## Method sources

- GRADE handbook and GRADE Book: https://gradepro.org/handbook/ and https://book.gradepro.org/
- RoB 2 official tool: https://www.riskofbias.info/welcome/rob-2-0-tool
- PRISMA 2020: https://www.prisma-statement.org/prisma-2020
- PRISMA-Harms: https://www.equator-network.org/reporting-guidelines/prisma-harms/
- Raison et al. 2023: https://jamanetwork.com/journals/jama/fullarticle/2808950
- Yngwe et al. 2026: https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2849099
