Ankur Pandey
AI safety for builder hackathon / fellowship - to build tools, products, etc.
Tony Wu
Analyze how activation verbalizers use target-model activation concepts via PCA/DAS/patching, explain cross-family failures, and improve verbalizers
Gabriel Sherman
A playbook to help AI safety policy advocates communicate with the U.S. government during the window of opportunity during an AI-related crisis.
Karolina Gruzel
Submitting Freedom of Information requests across EU Member States to reveal how governments understand and address advanced AI risks.
Taehyun Cho
This project builds cognitively-aligned preference learning that interprets feedback the way human actually decide rather than as a reward to maximize.
Felix Harder
A hand-verified library of AI-safety theorem statements in Lean 4 with AI-generated proofs, building the skills to trust AI formalization.
Florian Dietz
Clearing barrieers to adoption for an existing ICML-published interpretability technique that can elicit latent knowledge from red teamed model organisms
Christopher Leet
A benchmark to empirically investigate: (i) the ability of models to tacitly coordinate with copies of themselves and (ii) which decision theory best explains t
IBBIS | Tessa Alexanian
An outreach campaign from IBBIS to rapidly increase adoption of effective synthesis screening tools during a critical regulatory window.
Vael Gates
Humans in Control (HIC), a nonpartisan grassroots advocacy organization supporting AI safeguards, is raising funds to start a student program
Michail Patsakis
An open-source benchmark and defense toolkit for testing whether corrupted biological databases can hijack retrieval-augmented AI agents used in genomics, prote
Eitan Sprejer
The Argentinian AI Safety community (BAISH, baish.com.ar) is the largest in Latin-America. Support BAISH's growth, by providing funding for paying salaries.
Karthik Viswanathan
LLM agents collaborate to discover and formally verify theorems about the internal computations of transformers, beginning with a simple pilot question: how man
Logan Graves
A formal, testable account of LLM persona selection as Bayesian inference, validated with model internals, so labs can monitor and steer personas.
Nickola Horozov
Ram Potham
Run experiments on what incentives for deals with AI increase performance for using agents to discover misalignment
lucas.irwin
A policy memo, co-authored with the Institute for Public Policy Research, resolving the open technical, economic, and legal questions blocking real-world implem
James Calder Knight
Jordan Arel
A scalable fellowship training researchers to develop interventions for achieving high-value long-term futures
Joshua Reiners