Nicholas Volta
Facilitating AI X-risk Education in India, Southwest Cameroon, and beyond!
Sydney AI Safety Space
An additional year funding for the co-working hub in Sydney, Australia, which offers free office space for people working in the field of AI safety.
NTU AI Safety
Habeeb Abdulfatah
Trent Maziarz
the public, evidence-cited record of what companies actually do with AI
Constance Li
Animal Welfare Midtraining Data Creation
Muhammad
An open-source study of AI safety, dialect accuracy, and hallucination in farming contexts.
Amritanshu Prasad
Detailed models of how delegating critical functions to AI agents could destabilise strategic stability: pathways, likelihoods, and defenses.
Stewy Slocum
Abeer Sharma
Salvatore Barbera
Civil-society infrastructure against AI-enabled power concentration, built around autonomous weapons
Nikhil Maturi
An open, cheap method that detects when an inoculation prompt inoculates against off-target traits, so labs and developers can catch undesired trait/persona cha
Raffaello Fornasiere
Creating a reference model for mechanistic interpretability without assuming that at auditing time we have a safe model to compare the suspicious model against.
Funding compute/API costs for Incubator projects that build nonhuman welfare consideration into AI safety work
Elliot Tower
"ICML published ML researcher — evaluation validity, mechanistic interpretability applied to biology
Justin Shenk
Increasing public awareness of AI risks and benefits through in-person, interactive experiences
Charan Gowda D
Testing when VLA world models (Robots) simulation doesn't matchs the reality before it becomes an incidents
陳鈺澔
An AI platform for crypto and stock analysis, news verification, scam detection and wallet safety, with an open-source Safety Kernel tested on TON.
Matteo Leonesi
Alignment, interpretability, evals, security, governance. Open to everyone.
Tomáš Gavenčiak
An online overview of the AI safety techniques we use in the defense-in-depth stack