You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Most approaches to AI safety focus on producing a static set of weights at the end of training that define an "aligned model". We think this view is too narrow. Instead, we should focus on alignment techniques that preserve aligned behavior at all times, even as training progresses and model weights change. This is important for two reasons:
recent alignment failures (e.g. Hugging Face) involved systems that took misaligned actions during training, rather than after deployment
if per-user, weight-level continual learning becomes common, we will no longer be able to assume that all users are served the same static model with consistent propensities.
We propose a 4 month research pilot focused on two subjects.
Measuring and quantifying drift over time in alignment-relevant model characteristics in response to realistic post-training and continual-learning. This will focus on two sub-areas:
reward hacking during RLVR
alignment drift during online continual learning from user-provided feedback, specifically in the form of sycophancy and drift to non-assistant personas.
The development and implementation of countermeasures for the observed drift. Primarily, I plan to investigate gradient-routing based methods for the modularization of undesirable model propensities, such as reward hacking and sycophancy. These techniques are largely inspired by my team's previous research into new techniques for alignment pretraining.
The intended outcome will be a research report and scientific publication, targeting the ICML 2027 submission deadlines. We will also be in active contact with frontier lab researchers, with whom we intend to share results and recruit as collaborators on any publication that results from this grant. If the results end up less promising than we hope, we still intend to release a report publicly, which we expect will still be a useful artifact for the alignment community.
The proposed budget is approximately $70,000 for 4 months.
$40,204 for GPU rental
8xH100s GPUs, at $3.49 / GPU / hour, for 1440 hours.
$2,000 for data storage and research infrastructure (S3, coding agent usage, etc)
$24,796 to support researcher labor
$3,000 for travel and conference expenses.
As the expected expenses are mostly proportional to the duration of the research, a lower amount of funding would enable the same research for a proportionally smaller amount of time.
Ethan Roland is an independent AI researcher. For the past 1.5 years, he's led a small research team in collaboration with Anthropic, developing methods for mitigating dangerous capabilities in frontier models. His work was published at a top-3 AI conference (ICML) and recognized as a spotlight paper (top 2% of submissions). His work was also recognized by Dario Amodei as a "promising approach to improving the safety of open-weight models". He is also a co-author on a recent paper investigating alternatives to traditional AI architectures.
The most ambitious version of this project produces new techniques for the robust preservation of alignment-relevant characteristics in models undergoing weight changes, reducing concerns related to reward hacking and other emergent failure modes of continual learning. The mostly likely way the project does not meet this ambitious vision is either due to insufficiently positive research results or due to limited information about what research will be valuable in the future.
Specific failure modes include: (a) the versions of RLVR and continual learning the project considers do not faithfully reflect what is implemented / will be implemented within frontier labs. (b) we successfully measure and quantify drift in alignment-relevant characteristics, but are not able to successfully develop new techniques for the mitigation of this drift. The first can be mitigated via careful research into the techniques most likely to be in-use by frontier labs and by investigating the characteristics of a variety of RLVR and CL techniques. The second is a technique development question, which while we're optimistic for, we can't guarantee until having done the prerequisite research.
There are no bids on this project.