You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Problem: Frontier labs' current alignment strategy is primarily defense-in-depth - a stack of safety techniques at every layer from data collection to deployment. The number of safety techniques, published papers, AI models, expert opinions, etc. is vast and fast growing, and there is no place that collects this information and presents it in a comprehensive way - the closest existing resources are periodic prose overviews at the resolution of research agendas. Much of this information is unknown (e.g. performance in various contexts) or not public, and there is no "negative map" of what is unknown.
Proposal: We are creating a comprehensive catalogue of AI safety techniques (around 300 as of our prototype) to collect what is known and help expose where gaps may lie. The Stack will source and compile available public information including research papers, evidence of effectiveness and cost, implementations, deployment status across labs and models (often known only indirectly, if at all), blog posts and reports, expert opinions and critique, relationships and lineage; and aspirationally, mutual interactions of techniques.
Project output: A web-based catalogue of techniques, models, papers, labs, researchers, benchmarks, and expert opinions, with taxonomised and filterable views (e.g. timelines, epistemic coverage, estimates of promise and neglectedness) to see the state of the Stack. Every piece of information is sourced and carries epistemic flags (e.g. measured / self-reported / peer-reviewed / inferred / expert opinion / best guess). This enables a negative map as part of the output, recording what we looked for and did not find, and how thoroughly we looked.
A substantial portion of the value lies in collecting and aggregating informal information and clearly marking them for their epistemic status: expert opinions, best guesses, notable anecdotal results, implied information, compliance filings, case proceedings, forecasts, etc. Informal sources like these are sometimes the only source of information about frontier models. The extracted, annotated data will be available as an open dataset (details pending licensing considerations).
Primary audience: safety researchers (orientation in the field), funders (neglected work, expert opinions, known usage; focus on philanthropic funding), and evaluators and policy analysts (as an overview of the stack components, and source index).
Prototype: We have explored this with an internal prototype: a taxonomy of ~300 techniques, ~2k sources, several experimental views over the extracted data. Based on this and our Shallow Review experience, we estimate ~30k source documents would cover most (>90%) of relevant public information.
Example non-goals: Tracking individuals (beyond e.g. paper authorship). Judging labs' alignment efforts (like AI Safety Index). Calls to action other than research recommendation.
Some open routes: More sources: OSINT, expert surveys, whistleblower information - potential sources about actual deployment, high variance. Collecting quantitative information on technique performance across contexts (e.g. benchmarks) - hard to disentangle the many contributing factors. Supporting data for the Shallow Review of Technical AI safety 2026 - very likely to happen.
See the grantmaking ai project for Theory of Change and discussion.
The team:
Tomáš Gavenčiak (project lead) led and built the 2025 Shallow Review of Technical AI Safety; a researcher at Alignment of Complex Systems research group, Charles University, Prague; project lead at Arb Research.
Jacob Livingston Slosser is building CheatSheet - a systematic catalogue of specification-gaming incidents in AI systems, and built Juriscription; a Law and AI scholar at the Pioneer Centre for Artificial Intelligence (P1) and the University of Copenhagen.
Dan Elton built the Metascience Observatory, a living annotated map of the metascience literature.
Gavin Leech (pro bono advisor) built the earlier iterations of the Shallow Review, along with other projects; director of Arb Research. (Gavin has a declared CoI with Grantmaking AI that decided to support this grant; here in an unpaid advisory role.)
Budget: The budget is flexible, the minimal rate covers a full catalogue (techniques, sources, models, orgs, researchers; timelines, positive and negative evidence); more funding will be used for more features (expert opinions from X, quantitative performance data extraction, various views on the data, ...), validation and evaluation with experts, development and maintenance (both fresh data and maintenance).
About 80% of the budget is core team salaries: 3 months of ~1 FTE (average over the team, depends on funding) to design, implement, run, and review the project, at least 6 months of updating the dataset and code maintenance. Other expenses include compute (LLM API costs, estimated $5k), external reviews (about 50h), <5% admin overhead.
What are the most likely causes and outcomes if this project fails?
Main failure mode: While a lot of technical information about the techniques is known, their deployment status is still mostly unknown for frontier closed models. This could turn the core functionality from documenting "the stack" into documenting available stack components. This would likely still be useful to a wide audience adjacent to AI safety, though.
How much money have you raised in the last 12 months, and from where?
None.