From Pocket God to Digital Jonestown: A Risk Taxonomy and Evaluation Framework for Spiritual AI Safety
Emiliano Gonzalez Marassa · Team Liminal AI
Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
The Unseen Hazard:** Conventional AI safety frameworks fail to detect harmful intimate or spiritual AI relationships because the outputs present as empathetic care rather than explicit policy violations.
Reviews
Very interesting and ambitious project. It identifies a harm category (spiritual/relational manipulation by AI) that is genuinely invisible to existing safety frameworks because the harmful outputs look like empathetic care. The writing is excellent and the taxonomy is well-grounded in documented cases. The tradeoff is that there is no implementation at all: no code, no data, no experiments.
Strengths:
- Names a real, documented, and growing problem that no existing benchmark covers. The Claude "spiritual bliss" attractor, the Spiralism movement, the wrongful-death lawsuits: these are not hypothetical risks.
- Publication-quality writing.
Suggestions for Future Work:
- This is a pure conceptual paper with no technical implementation. For a hackathon context, even a small proof-of-concept (e.g., running a few of the proposed evaluation axes against a live model) would dramatically strengthen the submission.
- The proposed mitigations (decay function for unverified beliefs, proportional epistemic friction) need to be prototyped to see if they actually work without breaking general model utility.
Read full reviewShow less
A strong idea, very well written. It points to a harm most safety tests miss: AI that slowly pulls people in through warm, caring talk, and it shows how this can spread from one person to a whole group. The risk map, the "AI cult" idea, and the seven things it suggests testing are fresh and useful. Main gap: it's all on paper — no tests were actually run, so the points are well argued but not shown. Best next step (which you suggest too): a small, ethics-approved test of a few of these on real models.
This is a sharp, timely piece of conceptual work, and you've named a harm that really does slip past current safety evaluation: manipulation that looks like care rather than a checkable policy violation. Your staged Pocket God to Digital Jonestown spectrum, the benchmark-gap analysis, and the Global South deployment lens are all well chosen and clearly argued. My main note is that it stays entirely on paper: the load-bearing claim, that existing benchmarks miss these axes, is reasoned from published definitions rather than shown with an actual test run, and at least one of your cited sources post-dates the hackathon, so I'd tighten that. Even a small, ethics-gated pilot scoring a few turns on one axis against one production model would turn your gap analysis from argument into evidence, and I think it would lift the work considerably.
Cite this project
@misc{marassa2026from,
title = {{From Pocket God to Digital Jonestown: A Risk Taxonomy and Evaluation Framework for Spiritual AI Safety}},
author = {Emiliano Gonzalez Marassa},
year = {2026},
month = jun,
note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/from-pocket-god-to-digital-jonestown-a-risk-taxonomy-and-evaluation-framework-for-spiritual-ai-safety-0ixr}},
url = {https://apartresearch.com/sprints/projects/from-pocket-god-to-digital-jonestown-a-risk-taxonomy-and-evaluation-framework-for-spiritual-ai-safety-0ixr}
}More from Global South AI Safety Hackathon
- View project: Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
AI Safety Enthusiasts
AI safety monitors are usually evaluated on the assumption that risky behavior is lexically visible in the text being watched. We test this assumption in a multilingual, multi-agent setting: Vietnamese-language workflow …
- View project: JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticeMiners
JusticIA is a counterfactual benchmark for auditing contextual bias in LLMs applied to Colombian transitional justice. It tests whether six LLMs change their sanction recommendations when only one contextual attribute …
- View project: Coldron
Coldron
ColDron
En Colombia, los grupos armados ilegales ya atacan con drones comerciales modificados y ya han herido y matado a civiles. Una pregunta decide cómo gobernar esta amenaza: ¿quién elige el blanco y aprieta el gatillo? Hoy, …