Salvatore Barbera
Civil-society infrastructure against AI-enabled power concentration, built around autonomous weapons
Nikhil Maturi
An open, cheap method that detects when an inoculation prompt inoculates against off-target traits, so labs and developers can catch undesired trait/persona cha
Raffaello Fornasiere
Creating a reference model for mechanistic interpretability without assuming that at auditing time we have a safe model to compare the suspicious model against.
Constance Li
Funding compute/API costs for Incubator projects that build nonhuman welfare consideration into AI safety work
Phil Palmer
Deploying customer screening software at DNA synthesis providers to reduce AI-enabled biothreats
Logan Graves
A formal, testable account of LLM persona selection as Bayesian inference, validated with model internals, so labs can monitor and steer personas.
IBBIS | Tessa Alexanian
An outreach campaign from IBBIS to rapidly increase adoption of effective synthesis screening tools during a critical regulatory window.
Ankur Pandey
AI safety for builder hackathon / fellowship - to build tools, products, etc.
Taehyun Cho
This project builds cognitively-aligned preference learning that interprets feedback the way human actually decide rather than as a reward to maximize.
Eitan Sprejer
The Argentinian AI Safety community (BAISH, baish.com.ar) is the largest in Latin-America. Support BAISH's growth, by providing funding for paying salaries.
Gabriel Sherman
A playbook to help AI safety policy advocates communicate with the U.S. government during the window of opportunity during an AI-related crisis.
Florian Dietz
Clearing barrieers to adoption for an existing ICML-published interpretability technique that can elicit latent knowledge from red teamed model organisms
Michail Patsakis
An open-source benchmark and defense toolkit for testing whether corrupted biological databases can hijack retrieval-augmented AI agents used in genomics, prote
Vael Gates
Humans in Control (HIC), a nonpartisan grassroots advocacy organization supporting AI safeguards, is raising funds to start a student program
Karolina Gruzel
Submitting Freedom of Information requests across EU Member States to reveal how governments understand and address advanced AI risks.
Felix Harder
A hand-verified library of AI-safety theorem statements in Lean 4 with AI-generated proofs, building the skills to trust AI formalization.
Tony Wu
Analyze how activation verbalizers use target-model activation concepts via PCA/DAS/patching, explain cross-family failures, and improve verbalizers
Karthik Viswanathan
LLM agents collaborate to discover and formally verify theorems about the internal computations of transformers, beginning with a simple pilot question: how man
Christopher Leet
A benchmark to empirically investigate: (i) the ability of models to tacitly coordinate with copies of themselves and (ii) which decision theory best explains t
Ram Potham
Run experiments on what incentives for deals with AI increase performance for using agents to discover misalignment