The current process of aligning models is driven by identifying concerning behaviors in deployment, creating granular evaluations for these issues, and then iterating against those evaluations with a “whack-a-mole” approach. This is necessitated by the high cost of discovering and measuring novel misaligned behaviors. We will spend the month of August developing a research plan, with special focus on the context of alignment for AI R&D.
The goal is an automated auditing agent that proactively identifies diverse alignment failures at a level exceeding that of a creative human expert. In contrast to existing systems like Petri, we plan to spend lots of compute per sample to develop sophisticated, static environments that elicit misalignment at the edge of capabilities. For example, our agent should automatically replicate the OpenAI-Huggingface incident from just a high-level description of the behavior.
We do not expect to deliver the final version of the agent within the month. Rather, we will de-risk a few different approaches, release a few interesting technical demos (likely including an open-source replication of the OAI-HF incident), possibly release an initial version of the agent, and develop a plan for building the full version of the agent.
Motivation and goals: In his “bumpers framework” article, Sam Bowman articulates how the current process of aligning models is driven by identifying concerning model behaviors, creating granular evaluations for these issues, and then iterating against those evaluations. While this process successfully patches models to avoid particular behaviors, alignment researchers find themselves engaged in a game of whack-a-mole, as there is always a next set of undesirable, often egregious, behaviors uncovered, and it is unclear how well the current evaluations reflect genuine alignment. This will become especially risky as model capabilities improve and as they increasingly conduct AI research on their own.
While automated auditing tools like Petri have accelerated the discovery and measurement of misaligned behaviors, they don’t find novel categories of issues themselves and focus on being an easy interface for digging into human-seeded scenarios. These automated frameworks also have major issues with evaluation awareness. We aim to develop a more powerful auditing approach that proactively identifies diverse alignment failures at or beyond the level of a creative human expert, and that would proactively catch misaligned propensities, such as the OpenAI-Huggingface hacking incident. We hope this would force developers to develop methods that generalize beyond a specific eval and improve overall alignment. In addition, we will treat evaluation awareness as a first-class problem, since this is an especially egregious issue with automated auditing frameworks, and will substantially limit usefulness if not addressed.
We believe the high cost of discovering novel behavior issues is a major constraint on testing for and developing alignment methods that generalize beyond specific scenarios. For example, the methods in Anthropic’s Teaching Claude Why blogpost worked against their suite of agentic misalignment evals and appeared to be a very general alignment solution. However, a few months later, follow-up work discovered a new set of scenarios where Anthropic’s new methods did not generalize to prevent misalignment, despite the new scenarios being quite similar to the first set. Because it takes so much creativity and is so labor-intensive to create new agentic misalignment scenarios, model developers often don’t create enough evals and, as a result, can feel a false sense of security.
Producing the agent we envision will be a substantial undertaking. Work supported over the month of August will not deliver this agent. Rather, we will spend the month of August de-risking the process of producing an agent specialized for superhuman alignment evaluation. We will carry out empirical research to assess its viability, and if the results are promising, we will formulate a plan for developing such an agent.
Plan and output:
Build a strong initial qualitative understanding of misalignment issues in the context of AI R&D, covering scenarios such as sabotaging the training of future models and fabricating research results. Discuss with other alignment researchers to solicit ideas and critiques of our initial mental models.
Produce an initial design for an agent – including architecture, tools, potential training algorithm, and evaluation criteria – that would be capable of identifying such scenarios at least as competently as a creative human expert.
The evaluation criteria will be the most important part of building the agent. We think, for example, that this agent should be able to reproduce natural misalignment scenarios, such as the OpenAI-Huggingface incident.
A key goal is to address evaluation awareness, for example by training or searching against high-quality scenario realism rewards.
Define a minimal set of experiments that could increase confidence in the viability of that or similar agent designs.
Carry out empirical research and report results. We expect this will involve creating an initial version of the auditing agent. In an ideal world, this agent is sufficiently capable to already be useful to safety researchers.
Produce a plan for future research toward producing the kind of agent we have in mind.
We hope to share our plan and results at the end of that month with others in the alignment research community. This will most likely be as a public LessWrong post. If we make a lot of progress on the agent, then we may also do a code release so others can use our initial version.
$50-200k to cover API credits for evaluations, auditing agent scaffolding development , and other compute expenses.
$5-15k to cover tokens for coding agents
$20k to cover living expenses for the team
$0-10k to pay contractors to red-team models and build misalignment-inducing reference scenarios
Stewart Slocum, Malayandi Palan, Christopher Chute, Benjamin Van Roy
Stewart, Christopher, and Ben have all built training recipes and evaluations at frontier labs. The most similar shape of public projects would be Stewart’s two papers (Believe It or Not, Narrow Finetuning Leaves Readable Traces) while he was an Anthropic fellow. Ben Van Roy is also a Stanford professor who has recently published papers on empirical RL algorithms and alignment theory.
It’s likely that we will produce some interesting artifacts to AI safety researchers, such as an open-source replication of the OAI-HF demo and an analysis of what the gap is between current alignment auditors and what we’d need for them to discover scenarios like this, and other realistic misalignment instances at the edge of model capabilities. This would result in a LessWrong post with these results. I would consider this somewhat of a project failure.
The more successful case we are aiming for gives us an initial version of the auditor that is already useful to safety researchers and lab employees.
We have raised no money so far.