You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Research labs are beginning to entrust entire areas of artificial intelligence research to AI agents. Safety relies on monitoring mechanisms—an LLM “judge” that reviews the agent’s actions, or a probe that analyzes its activations—designed to detect any dangerous elements. I want to know what happens when the agent resists.
My previous work has highlighted a simple and somewhat unsettling phenomenon: social pressure causes a model’s judgment to shift by a fixed fraction in the direction imposed by that pressure. The parameters I adjusted during an experiment allowed me to predict a completely different experiment without readjustment (R² of 0.73/0.69 out-of-sample). This paper is currently under review at ICLR 2027. The question now is whether control mechanisms can be influenced in the same way. If an agent flatters the judge, claims to have authority, feigns approval, or conceals instructions within its results, does the control mechanism allow elements to slip through that it should have detected?
That is what this project sets out to test. I launch five types of attacks against LLM judges and activation probes tasked with detecting sabotage; I measure the extent to which they fail, and then I verify whether the obvious solutions actually work: paraphrasing the agent’s output using a separate trust model,
Article C: Corruption Monitoring (Months 1–4). To what extent do LLM judges and investigators yield to the five types of pressure? What is the cost of moving an investigator compared to that of a judge? What can be done to correct the system’s operation? And finally, an open acceptance test for supervisors. My previous work provides the necessary tool: on Llama-3.1-8B, social pressure shifts the model’s judgment from a fixed fraction toward a target set by the pressure, and the parameters are carried over to a new campaign. Decision in Month 2: while previous work already covers this law for controllers, the article refocuses on the cost of attacks and on the test.
Paper A, diffuse sabotage (months 1–7). In a simulation of a research sabotage scenario (Terekhov et al., arXiv 2606.08892) involving 8 billion model organisms, I compare black-box and white-box monitoring, action by action and over time. “Go/no-go” decision in month 4: “white-box” monitoring over time must outperform the best “black-box” benchmark in terms of recall with 1% false positives (95% CI); otherwise, I publish a negative result.
Results: two papers on arXiv, with open-source code, data, model organisms, probes, and acceptance tests. The first phase is already publicly available at github.com/lciric/controle-ia.
Three goals to be achieved by the end of the twelve months:
An article on the corruption of monitors (Article C) showing the extent to which monitors yield to hostile persuasion and which defenses hold up. This includes the Red Team’s results, comparisons of defenses, and the code needed to reproduce the entire experiment. This should be completed by the 4th or 5th month.
A paper on detecting diffuse sabotage (Paper A) using the sequential monitoring approach. I aim to demonstrate that it detects patterns that “per-action” monitors fail to identify, and that it is resilient against an adaptive attacker who is aware of the detection method.
If the first two projects go well and the budget allows, a multi-agent extension (Article B): what happens when multiple copies of a model split a malicious modification into elements that, taken in isolation, appear harmless?
I defined criteria for continuing or abandoning the project before launching anything. If the monitoring system proves robust against my attacks, that is a positive result and will be published as such. I am not looking for a specific answer…
It's just me. I studied physics (ENS Paris-Saclay), computational neuroscience (ENS; master's thesis in Srdjan Ostojic's group, published in PLOS Computational Biology in 2023), and machine learning (Mines Paris-PSL). I completed a research internship at IBM Research Zurich, and my work was published in Neuromorphic Computing and Engineering in 2022.
Since March 2026, I have been conducting research on AI safety using my own funds. Here is what I have discovered so far:
I have demonstrated that flattery in Llama-3.1-8B follows a displacement law: pressure shifts the model’s assigned score by a fixed fraction toward a target, with an R² of 0.77/0.79 within the sample, and the fitted parameters allow us to predict a new campaign without retraining, with an R² of 0.73/0.69. This is the paper currently under review at ICLR 2027.
I also found that the direction detected by a behavior-tracking probe accounts for less than 1% of the causal effect on the model’s output. The full hidden state, on the other hand, accounts for 68%. In other words, probes can detect behavior without controlling it, which is of paramount importance for interpretability-based supervision..
I spotted and reported my own error: the responses had been stored truncated to 300 characters, which made a simple change in distribution appear to be a targeted security measure. I wrote the correction and published it.
The current project (controle-ia) includes 98 hashed source articles, 3,497 machine-verified citations with zero errors, an activation logging mechanism validated on Llama-3.1-8B, and 602 automated tests. The total computational cost to date is approximately $20.
I am in the final phase of the 2026 Anthropic Fellows program focused on AI safety. I am French and live in Paris.
The main risk is that my attacks will be too weak. If I simply try obvious persuasion strategies and the instructor sees through them, I won’t have demonstrated much—a real agent might be more creative. I mitigate this risk partly by including indirect attacks (instructions hidden in the content, spoofed system messages) that go beyond mere social pressure, and partly by being honest about the scope of the test. The “proceed/abort” decision criteria help me in this regard: if the monitor holds firm, I publish this result without claiming that the test was more difficult than it actually was.
Second risk: Results obtained on open models may not be transferable to state-of-the-art models whose internal workings I cannot examine. I can run “LLM-judge” attacks in “black-box” mode against API models, but the probing work requires weights. This is a real limitation, and I will make that clear.
Third risk: I may not be able to complete both papers on my own within twelve months. The paper on monitor corruption is further along and will be published first. If I run out of time, the paper on diffuse sabotage will become a working paper or a preprint rather than a full-length paper.
There is also a personal risk: I am funding this project out of my own pocket.
None for this project so far. I have spent approximately 7,500 € of my own money on it since March 2026. I have applications pending with EA Funds (Transformative AI Fund), BlueDot (two tracks), Foresight, Lightcone, GTR, and Middle Powers, and I have requested computing credits from Anthropic, the NDIF, and the Google TPU Research Cloud. None of them have made a decision yet.