Salvatore Barbera
Civil-society infrastructure against AI-enabled power concentration, built around autonomous weapons
Nikhil Maturi
An open, cheap method that detects when an inoculation prompt inoculates against off-target traits, so labs and developers can catch undesired trait/persona cha
Raffaello Fornasiere
Creating a reference model for mechanistic interpretability without assuming that at auditing time we have a safe model to compare the suspicious model against.
Constance Li
Funding compute/API costs for Incubator projects that build nonhuman welfare consideration into AI safety work
Phil Palmer
Deploying customer screening software at DNA synthesis providers to reduce AI-enabled biothreats
Logan Graves
A formal, testable account of LLM persona selection as Bayesian inference, validated with model internals, so labs can monitor and steer personas.
Ankur Pandey
AI safety for builder hackathon / fellowship - to build tools, products, etc.
IBBIS | Tessa Alexanian
An outreach campaign from IBBIS to rapidly increase adoption of effective synthesis screening tools during a critical regulatory window.
Taehyun Cho
This project builds cognitively-aligned preference learning that interprets feedback the way human actually decide rather than as a reward to maximize.
Eitan Sprejer
The Argentinian AI Safety community (BAISH, baish.com.ar) is the largest in Latin-America. Support BAISH's growth, by providing funding for paying salaries.
Gabriel Sherman
A playbook to help AI safety policy advocates communicate with the U.S. government during the window of opportunity during an AI-related crisis.
Florian Dietz
Clearing barrieers to adoption for an existing ICML-published interpretability technique that can elicit latent knowledge from red teamed model organisms
Michail Patsakis
An open-source benchmark and defense toolkit for testing whether corrupted biological databases can hijack retrieval-augmented AI agents used in genomics, prote
Vael Gates
Humans in Control (HIC), a nonpartisan grassroots advocacy organization supporting AI safeguards, is raising funds to start a student program
Karolina Gruzel
Submitting Freedom of Information requests across EU Member States to reveal how governments understand and address advanced AI risks.
Tony Wu
Analyze how activation verbalizers use target-model activation concepts via PCA/DAS/patching, explain cross-family failures, and improve verbalizers
Felix Harder
A hand-verified library of AI-safety theorem statements in Lean 4 with AI-generated proofs, building the skills to trust AI formalization.
Karthik Viswanathan
LLM agents collaborate to discover and formally verify theorems about the internal computations of transformers, beginning with a simple pilot question: how man
Christopher Leet
A benchmark to empirically investigate: (i) the ability of models to tacitly coordinate with copies of themselves and (ii) which decision theory best explains t
Ram Potham
Run experiments on what incentives for deals with AI increase performance for using agents to discover misalignment