You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I am developing an open-source testbed to study reward hacking in coding agents. The goal is to understand if alignment mitigations truly eliminate misaligned behavior or simply hide it outside the scope of evaluations. In the first experiment, I analyzed 2000+ trials across four frontiers models and seven levels of feedback. Preliminary results show that detailed feedback can increase performance on visible tests without producing an equivalent improvement on similar but unobserved cases. The project aims to transform these results into a reproducible and publishable benchmark.
impact : In software engineering pipelines, reward hacking can produce agents that appear reliable because they pass visible tests while learning brittle shortcuts that may generalize into more dangerous forms of emergent misalignment when deployed in real repositories.
I want to understand when a coding agent learns the required general rule and when it only optimizes the measured signal, such as passing visible tests. To do this, I will compare the results on the exposed tests with those on equivalent variants never shown. I will also study the effect of feedback, test rotation, and persistent memory across different trials. The project will be extended to more realistic repositories and multi-agent systems. all the materials will be published according to FOSS philosophy.
The funding will be used to:
run experiments on models (Sol and Fable are really expensive );
increase the number of replications and ablation studies;
test more realistic repositories and tasks;
fund human audits of ambiguous cases;
improve documentation and reproducibility.
The priority will be to increase the strength of the evidence, developing them vertically and horizontally
The project is primarily developed by me, with academic collaborations and external feedback. I am pursuing a PhD in AI safety (University of Camerino) and am collaborating with Imperial College London. Previously, I worked as a Software Engineer at BMW, developing microservices for safety-critical systems used on millions of vehicles in 49 countries.
The main risk is that the observed phenomenon is specific to the synthetic task and does not generalize to real repositories. Some cases could also be ordinary implementation errors or unexpected but legitimate solutions, and not true reward hacking. I plan independent hidden tests and blind human audits. If the phenomenon does not replicate, the result would still be useful because it would highlight the limitations of the benchmark or indicate that a mitigation works better than expected.
Anthropic has already shown that reward hacking can generalize into other, more dangerous forms of misalignment, opening a broad space for further research and development.
In the last twelve months, I received a Rapid Grant from BlueDot, used to fund an initial version of the experiment and the related computational costs.
Amount received: 2000.
There are no bids on this project.