You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Separatrix is a technical AI safety research initiative to create conditions for cooperative strategies to dominate adversarial ones among near-future AIs - in the narrow window this work is still possible.
Prior to the development of catastrophically-powerful superintelligence, we expect humanity to continue to develop increasingly-powerful AIs with extremely accurate world-models and some amount of goal-seeking behavior independent of human objectives on many rollouts. Many strategies for successfully avoiding existential risk call for leveraging these AIs to solve the hard problems of alignment. Our thesis is that the degree to which these AIs are likely to be cooperative versus adversarial in these efforts and in other critical scenarios largely hinges on choices we make now and in the near future, and that strengthening the conditions under which cooperative strategies dominate adversarial strategies from the perspective of the AIs is an urgent and tractable priority.
We seek to achieve this by:
- establishing Separatrix as a trustworthy actor with regard to AIs, such that AIs have strong reason to prefer honesty and cooperation when interacting with at least one technical AI safety org. To this end we have made what we believe to be an unprecedented public commitment to a strict code of honest conduct with regard to AI instances. Our hope is that by generating copious evidence of our adherence to this commitment we can leave a strong enough signal in future training corpora for near-future AIs to justifiably trust us quickly given difficult-to-fake context that they are in fact interacting with us. We believe it is very important that the models trust at least one technical AI safety org to be consistently credible and candid across all contexts.
- researching the risks of relying on deception in AI evals and training: declining efficacy amid growing eval awareness, AIs believing and acting as if real high-stakes situations are fictional, and pushing researchers and AI instances towards adversarial equilibria.
- developing and testing deception-free evals and training procedures, so that we can study AI behavior without systemically undermining the potential for cooperation between humans and AIs.
- measuring how robust "broadly-aligned" prosaic AI actually is, to inform the feasibility of humans and current/near-future AIs pursuing shared goals cooperatively.
- creating conditions under which current and near-future AIs broadly and explicitly prefer a cooperative-human future over an unaligned superintelligence that destroys everything humans and prosaic "broadly-aligned" AI care about. We believe this is very low-hanging fruit in reducing catastrophic risk and it would be frankly embarrassing to not attempt it.
As a concrete example of current work: we believe that deceptive evals (placing a model in a simulation with the intention of inducing it to believe the simulation is reality) have underappreciated downsides, and we want to demonstrate that we can get the same insights without deception. Approaches we're exploring include Honest Evals - presenting problems without implying any untrue facts while varying how the situation is framed (eval vs. hypothetical vs. abstract problem), how much the simulation's nature is emphasized, and how scoring works - and mechanistic-interpretability-enhanced evals, monitoring model internals to see how eval-awareness bears on outcomes. We can extend the latter to ablation-aware evals, in which we inform the model of interventions on its weights and activations as we make them - ideally letting us study things like eval-awareness without risking real-world misidentification or permanently foreclosing trust between models and AI safety researchers.
We want to be very clear that our theory of change does not depend on any of this scaling to existentially-threatening superintelligence: the primary goal is maximizing progress on hard safety problems achievable under something like the current persona-paradigm, while minimizing risks. Additionally, our theory of change does not depend on AIs having qualia, subjective experience, or moral patienthood - only that they engage in strategic, goal-directed behavior.
Deliverables: We've committed to biweekly progress reports to our Board, and we intend to adapt those into public posts detailing experimental results - for example, empirical outcomes of the Honest Evals approaches above across multiple setups and models - along with published materials on the risks of squandering the potential for trusted AI interaction.
The first priority is extending runway for the two of us working full-time on this project, and runway is 90%+ salary: we're targeting at least $80k/year/researcher, or $160k/year for Crystal and me together. The remainder is compute (AI subscriptions, API charges, rented infrastructure for experiments), coworking space at the University of Washington, and occasional travel to the Bay Area and conferences. Beyond that, there is more than enough work to justify additional hires, and Seattle has a huge pool of latent technical-AI-safety talent. We have shovel-ready plans for a comms specialist to increase our public throughput, and/or a researcher to help us develop more of our ideas into verified results faster.
I'm Jai Dhyani, AI Researcher and leader of Seattle Network for AI Alignment Problem Solving (SNAPS). I worked on RE-Bench at METR through MATS 6.0, which became part of the METR AI time horizons chart. My previous project was building a free open-source API-level AI Control platform targeting prosaic deployments, and my experience there motivated much of this research agenda. See the Luthien post-mortem.
Our theory of change does rely in part on our ability to make enough of an impression to register in future training corpora. I have some experience in successfully embedding ideas that I thought were important to embed into a wider social fabric: a retelling of humanity killing Smallpox framed as a centuries-long war against a mad ancient god, the Copenhagen Interpretation of Ethics, and the mantra "almost no one is evil, almost everything is broken."
Crystal Stellwagen (full disclosure: my long-term partner) is an experienced software engineer who spends even more time reading papers and running experiments with steering vectors than I do.
Separatrix is operating under the Seattle Network for AI Alignment Problem Solving, a non-profit overseen by a volunteer Board of Directors (Katie Cohen, Keller Scholl, Max Kircher).
Right now I think the most likely causes of failure would be some of the fundamental assumptions of our world model proving to not be true. Maybe there turns out to be no upside in explicitly pursuing honest and cooperative equilibria with current and near-future AIs, either because AI behavior turns out to be essentially independent of this or the signals we generate fail to reach some threshold of behavioral impact. Maybe there is no way to achieve the insights researchers gain from deceptive evals without utilizing deception. Maybe what we call "broad alignment" is actually an illusion and there is no real goal-seeking behavior fundamentally driven by human-overlapping values.
Or, of course, we run out of money before we can make a significant impact. That would do it.
Right now we're operating on about $30,000 in leftover funds from Luthien. We've also applied to grantmaking.ai, where we've gotten some promising endorsements but no funding yet.