Jordyn Harland-Graham
Persistent memory in AI using LoRAs
Stewy Slocum
Nikhil Maturi
An open, cheap method that detects when an inoculation prompt inoculates against off-target traits, so labs and developers can catch undesired trait/persona cha
Abeer Sharma
Raffaello Fornasiere
Creating a reference model for mechanistic interpretability without assuming that at auditing time we have a safe model to compare the suspicious model against.
Constance Li
Funding compute/API costs for Incubator projects that build nonhuman welfare consideration into AI safety work
HEATH ERWIN PARISH
Testing whether AI can be governed at the moment it acts, then putting that protection to work for organizations that need it most.
Mekaoui Helmy
When an AI agent spawns sub-agents, its safety limits do not follow. I build and deploy the layer that makes them inherited and non-strippable.
陳鈺澔
An AI platform for crypto and stock analysis, news verification, scam detection and wallet safety, with an open-source Safety Kernel tested on TON.
Gary Welz
Agent Roles, a Constitution and Governance in the Research Workflow
Naufal Ridwan
Testing whether dynamic boundaries, history, feedback, and uncertainty-aware decisions can make AI behavior more interpretable and auditable.
Tomáš Gavenčiak
An online overview of the AI safety techniques we use in the defense-in-depth stack
Trey Anderson
Independent Behavioral Research on Open-Weight AI Models
Lidia
AISafety, AI and Science, AI for Human Reasoning
Pete Wolfendale
Integrating Selves and Research on Selves in the AI Community
Anthony Ozerov
Help me evaluate the safety, ethics, and values of the quantized and fine-tuned open-weight LLMs that individuals and enterprises are actually using.
Logan Graves
A formal, testable account of LLM persona selection as Bayesian inference, validated with model internals, so labs can monitor and steer personas.
Ankur Pandey
AI safety for builder hackathon / fellowship - to build tools, products, etc.
Gaurav Hadavale
Help an undergraduate mechanistic interpretability researcher present accepted AI-safety work at Mechanistic Interpretability workshop of Medical model 2026.
Florian Dietz
Clearing barrieers to adoption for an existing ICML-published interpretability technique that can elicit latent knowledge from red teamed model organisms