The Shape of Harm Research Framework
A protocol draft for comparing psychoactive-substance harms as defined outcomes under defined exposure scenarios—with the evidence model kept separate from the value judgments used to combine outcomes.
Call it the right kind of scientific object
The project is better framed as a scenario-based comparative risk model with an optional multi-criteria decision layer, not as a psychometric scale that measures a single latent trait.
Scientific title: “The Shape of Harm: a scenario-based comparative risk and decision-analysis framework for psychoactive substances.”
Evidence engine
Systematically identifies, appraises and synthesizes outcome evidence in natural units for each defined scenario.
Risk model
Maps substance, exposure and context to outcome distributions. Dependence, route, co-use and supply are causal modifiers, not extra points.
Decision layer
Optionally converts non-overlapping outcomes through declared value functions and swing weights. It never presents those preferences as empirical facts.
Because the domains form a composite rather than interchangeable symptoms of one hidden trait, internal-consistency statistics such as Cronbach’s alpha are not an appropriate validation target. The relevant questions are whether the components are well defined, non-overlapping, estimated reliably, externally calibrated and fit for the stated use.
Every result begins with a fully specified question
A row is not simply “cocaine” or “benzodiazepines.” It is a scenario with a formulation, dose, route, exposure pattern, context and target population.
| View | Unit of analysis | Primary denominator | What it may answer | What it must not claim |
|---|---|---|---|---|
| One episode | A defined administration episode | Per 100,000 episodes | Absolute acute medical and behavioral/psychiatric risk | Long-term harm or population burden |
| Regular use | A defined pattern over time | Per 1,000 user-years | Chronic health loss, dependence and withdrawal outcomes | Risk from one night or total national harm |
| Population burden | A named population in a named year | Per 100,000 residents and total burden | Deaths, DALYs, victim harm and cost with prevalence included | Intrinsic per-user or per-episode danger |
Four author phases completed: four pilot estimands are explicit, and an author data-feasibility screen has produced go/hold/block decisions. Controlled psilocybin now has a single-author pilot extraction: one narrow rare-event estimate is releasable and three outcomes remain blocked by reporting incompatibility. The psilocybin serious-event estimate is now paired with a provisional very-low-certainty rating and explicit information-size targets. Tobacco requires redesign; alcohol and opioid absolute episode-risk estimates remain blocked. Inspect the certainty gate →
No default overall ranking crosses these views. A population-burden result can put a common drug above a rarer but more dangerous one; that is not a contradiction. It is a different estimand. The current four pilot estimands also occupy different comparability classes and therefore may not be ranked against one another.
A proposed core outcome set with overlap guards
These are candidates for expert and lived-experience review. They are deliberately separated by denominator, time horizon and causal role.
| ID | Outcome | Operational form | Unit | Overlap rule | Status |
|---|---|---|---|---|---|
| AE-MED | Acute medical severity | Mutually exclusive states: death; ICU/organ failure survived; hospital/ED severe event survived | Events per 100,000 episodes and expected health loss | Each episode occupies one highest-severity state | Proposed |
| AE-BEH | Acute behavioral/psychiatric severity | Severe injury, self-harm, violence, psychosis or dangerous disorientation attributable to the episode | Events per 100,000 episodes | Separate user and victim outcomes; do not re-add medical consequences | Proposed |
| CH-PHY | Chronic physical health loss | Attributable mortality and morbidity under the stated use pattern | DALYs or cases per 1,000 user-years | Exclude acute events already modeled unless annual burden is the declared endpoint | Proposed |
| CH-PSY | Chronic psychiatric/cognitive health loss | Persistent disorder or impairment beyond the acute window | DALYs or cases per 1,000 user-years | Do not count transient intoxication effects | Proposed |
| SUD | Substance-use-disorder incidence | New DSM/ICD disorder during a fixed follow-up among initiators | Cases per 1,000 initiators | Separate from physical dependence | Separate |
| PHY-DEP | Physical dependence | Physiologic adaptation under a fixed exposure pattern | Cases per 1,000 exposed people | Separate from compulsive use and withdrawal hazard | Separate |
| WD-SEV | Severe withdrawal outcome | Seizure, delirium, medically serious deterioration or death after cessation | Events per 1,000 cessation attempts among dependent users | Conditional denominator must be explicit | Separate |
| EXT-USER | External harm attributable to use | Victim injury/health loss, family harm and other external outcomes | Victim DALYs or events per 1,000 user-years | Keep public cost separate from health outcomes | Proposed |
| POP | Population burden | Total attributable deaths, DALYs, victim harms and costs | Total and per 100,000 residents/year | Never mix with per-user outcomes in the same score | Separate view |
The evidence pipeline has to be reproducible before it is persuasive
The current project shows its sources. The research version must also show how every source was found, included, extracted, appraised and synthesized.
Preregister the protocol before looking for a preferred answer
Lock scope, scenarios, outcomes, eligibility rules, synthesis methods, subgroup plans and deviations. Publish amendments with timestamps.
Run systematic searches designed with an information specialist
Use reproducible database strategies, citation searching, grey-literature rules and a PRISMA flow diagram. Update searches on a declared schedule.
Use at least two independent reviewers
Duplicate screening, extraction and key risk-of-bias judgments. Resolve disagreements through a recorded process rather than silent consensus.
Assess each study with a design-appropriate risk-of-bias tool
Do not turn “peer reviewed” into “high quality.” Record confounding, selection, exposure misclassification, outcome measurement, missingness and reporting bias.
Grade certainty by outcome, not by drug
Summarize risk of bias, inconsistency, indirectness, imprecision and publication bias. A drug may have high-certainty mortality evidence and very-low-certainty psychiatric evidence.
Use structured expert elicitation only where data remain genuinely absent
Experts receive identical evidence dossiers, state distributions rather than single digits, disclose conflicts, score independently first, and never overwrite empirical estimates without a documented model.
Estimate harm first; aggregate only when the question requires it
The model should preserve the original units and uncertainty as long as possible. A 0–100 display is a communication transform, not the evidence itself.
Parameter uncertainty
Sampling error, study heterogeneity and expert-elicited uncertainty become distributions, not evidence-grade widths.
Structural uncertainty
Alternative causal structures, outcome definitions, pooling assumptions and value functions are compared explicitly.
Decision uncertainty
Report rank probabilities, dominance and value-of-information—not a brittle one-to-thirteen order.
Missing is not zero. Sparse cells remain missing or are estimated through a preregistered hierarchical model with visible partial pooling. The interface must distinguish a directly estimated value from a model-based prediction.
Validation is a program of tests, not one coefficient
The relevant evidence is different for the framework, the evidence process, the statistical model and the optional decision layer.
| Target | Question | Proposed test | Release gate |
|---|---|---|---|
| Content | Are the scenarios and outcomes relevant, comprehensive and understandable? | Multidisciplinary panel plus people with lived experience; predefined relevance/completeness ratings and qualitative review | No major omitted domain and no unresolved construct overlap |
| Review process | Can independent reviewers reproduce inclusion, extraction and bias judgments? | Agreement statistics appropriate to the data plus transparent disagreement logs | Thresholds declared before pilot; retrain/revise if missed |
| Construct behavior | Does the model change in predicted directions? | Preregistered hypotheses for route, dose, co-use, setting and supply changes | Most directional hypotheses supported; failures investigated |
| External validity | Does it reproduce data not used to build it? | Hold out surveillance years, jurisdictions or datasets; assess calibration and error | Performance bounds fixed before unblinding |
| Transportability | Do estimates remain usable across place and time? | US/Canada and period comparisons; update supply-sensitive scenarios separately | No universal label where meaningful interactions are present |
| Decision robustness | Does a conclusion survive plausible values and model choices? | Global sensitivity analysis, alternative value functions, probabilistic weights | Claim only conclusions robust across declared ranges |
| Replication | Can an independent group reproduce the result? | Second team receives protocol and raw evidence but not final scores | No “validated” label before independent replication |
Important: do not use Cronbach’s alpha to prove these domains “hang together.” Acute death, dependence and chronic organ disease are intentionally different outcomes. High internal consistency would more likely signal redundancy than validity.
A small credible pilot beats another complete-looking table
The first publishable study should prove the method on a deliberately narrow subset before expanding to thirteen substances and many contexts.
Phase A — author-draft estimands completed
Four exact questions are provisionally frozen in the pilot estimand registry. Any change now requires a versioned amendment.
Phase B — author feasibility screen completed
The data-feasibility review maps candidate sources to every required data element. One evidence review can proceed, one estimand requires redesign, and two remain blocked by denominator failure. This is not independent validation.
Phase B — independent content-validity review is next
Domain experts and lived-experience contributors independently judge population relevance, exposure/comparator precision, outcome completeness and overlap, time windows, intercurrent-event handling, interpretability, and whether each quantity is actually estimable. The supplied review form records ratings, conflicts, and required revisions.
Phase C — author pilot extraction completed; independent review pending
The psilocybin evidence pilot tested the extraction and synthesis rules. It found one narrow estimate and multiple non-estimable outcomes. Two independent reviewers must now repeat the work before any scientific release.
Phase D — provisional certainty and information analysis completed
The certainty gate separates the observed zero from confidence in the true risk. The outcome is provisionally very low certainty, and the page shows the sample sizes needed to exclude specific rare-event bounds.
Phase D — held-out validation
Freeze the model, reveal a reserved dataset or later year, assess calibration and directional hypotheses, and publish failures alongside successes.
Phase E — only then build the interactive comparison
The public page reads the validated scenario estimates and lets users explore perspectives. It does not generate scientific legitimacy through visual polish.
Protocol registered; all scenario fields locked; systematic searches reproducible; duplicate extraction complete; outcome-level certainty rated; code and data public; held-out test reported.
Single-author scores; undefined denominators; mixed time horizons; judgment inserted as data; fixed uncertainty widths; missing treated as zero; ranking published before validation.
A research project someone else can actually inspect
These files are intentionally drafts. Their purpose is to make critique specific and make collaboration possible.
Methodological foundations
These are starting points for the protocol, not badges of automatic validity.
- International Council for Harmonisation. E9(R1): Estimands and Sensitivity Analysis in Clinical Trials. FDA final guidance. 2021.
- Hernán MA, Robins JM. Specifying the target trial. In: Causal Inference: What If. 2020.
- Thokala P, et al. Multiple Criteria Decision Analysis for Health Care Decision Making—An Introduction. Value in Health. 2016.
- Marsh K, et al. Multiple Criteria Decision Analysis for Health Care Decision Making—Emerging Good Practices. Value in Health. 2016.
- Page MJ, et al. PRISMA 2020 statement. BMJ. 2021.
- Cochrane. Standards for independent duplicate data extraction.
- Cochrane. GRADE certainty of evidence by outcome.
- Bollen KA, et al. In Defense of Causal-Formative Indicators. Psychological Methods. 2015.
- Eddy DM, et al. Model Transparency and Validation. Medical Decision Making. 2012.
- Briggs AH, et al. Model Parameter Estimation and Uncertainty Analysis. Medical Decision Making. 2012.
- Jünger S, et al. CREDES guidance for Delphi studies. Palliative Medicine. 2017.
- Crépault JF, et al. Drug harms in Canada: a multi-criteria decision analysis. Journal of Psychopharmacology. 2026.