Gaurav Hadavale
Help an undergraduate mechanistic interpretability researcher present his accepted AI-safety work at MICCAI 2026.
陳鈺澔
An AI agent for crypto and stock analysis with TON token swaps. Every swap must go through user confirmation and fixed safety checks, and the core safety rules
Heath
Trace makes sure an AI has authority before it acts, controls what it is allowed to remember, and checks that authority again before memory is retrieved.
Rose G. Loops
TRiADiC Intelligence Labs ethical alignment research program and consumer empowerment educational program.
Tomáš Gavenčiak
An online overview of the AI safety techniques we use in the defense-in-depth stack
Justin Shenk
Increasing public awareness of AI risks and benefits through in-person, interactive experiences
Pete Wolfendale
Integrating Selves and Research on Selves in the AI Community
Ankur Pandey
AI safety for builder hackathon / fellowship - to build tools, products, etc.
Logan Graves
A formal, testable account of LLM persona selection as Bayesian inference, validated with model internals, so labs can monitor and steer personas.
Florian Dietz
Clearing barrieers to adoption for an existing ICML-published interpretability technique that can elicit latent knowledge from red teamed model organisms
Taehyun Cho
This project builds cognitively-aligned preference learning that interprets feedback the way human actually decide rather than as a reward to maximize.
Michail Patsakis
An open-source benchmark and defense toolkit for testing whether corrupted biological databases can hijack retrieval-augmented AI agents used in genomics, prote
Felix Harder
A hand-verified library of AI-safety theorem statements in Lean 4 with AI-generated proofs, building the skills to trust AI formalization.
Pip Foweraker
A nightmarishly hard AI safety strategy game about holding p(Doom) down. You can't win; you can only buy time.
Eitan Sprejer
The Argentinian AI Safety community (BAISH, baish.com.ar) is the largest in Latin-America. Support BAISH's growth, by providing funding for paying salaries.
Jai Dhyani
Creating conditions for cooperative strategies to dominate adversarial ones among near-future AIs while we still can
Christopher Leet
A benchmark to empirically investigate: (i) the ability of models to tacitly coordinate with copies of themselves and (ii) which decision theory best explains t
Karthik Viswanathan
LLM agents collaborate to discover and formally verify theorems about the internal computations of transformers, beginning with a simple pilot question: how man
Victor Porton
Nickola Horozov