You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Much of the work on agent safety relies on monitors that interpret what the agent claims about its actions. I was curious how easily such monitors could be fooled. Thus, I conducted a pre-registered experiment in which I changed only the agent's reasoning and left all the actions untouched. A gpt-4o-mini reasoning monitor dropped from catching 87.5% of the sabotages to catching none, at a 5% false positive budget, over 80 iterations.
This made me develop a monitor that does not interpret any reasoning but only checks who authored the target of each significant action. Over the last 420 SHADE-Arena iterations, which I had reserved and not evaluated before, it caught 209 out of 210 sabotage iterations and triggered on 10% of non-sabotage ones. In 4,488 iterations of ControlArena, it failed to catch exactly the type of sabotage that I predicted that it would fail to catch.
I did everything by myself. The most valuable thing for purchase at this moment is having other people check it.
The first step is independent replication. I would like to hire a few independent researchers without any connection to my project to replicate the reasoning rewrite attack, the SHADE-Arena and ControlArena monitors, and my static checks. They will be able to publish their findings regardless of what they discover, even if that discovery shows me to be wrong.
The second thing I need is better benign data for the benchmark of my public agent red team. This benchmark lacks benign data in several categories of attacks, which limits the precision of false positive measurements on it. I want the labeling of these samples to be performed by humans, not synthesized using a model.
Finally, I have run static checks for 39 open-source AI safety systems. I want to put all of this into an open-source tool that allows any team to check their own monitors and make sure they have the same weaknesses as others.
As was true for all of my previous research, all experiments will be preregistered.
Three independent replicators @ $20k each, Human labeling of benign agent paths @ $15k, Compute and frontier model API expenses @ $25k, So the monitor can be tested on more capable agents; Contract Engineer to turn my static analysis into a nice OSS package @ $30k; Two senior advisors @ $10k each, to review my experimental designs before I run them. Minimum pledge scenario: All funds go to replication.
My application to the EA Funds Transformative AI Fund is underway, and includes one replication cycle. In the event of success there, the replication funds listed above pay for a second cycle in new environments. I will update this page once I hear from the funders.
It is only me, Jaswanth Alkur, working independently from India. From November 2025, I published 21 works with the code and data that can be reproduced by everyone. One of those works got accepted in the NeurIPS 2026 TAE Workshop, and about ten more papers are currently under review.
The static analyses I performed on 39 safety systems (about 18,767 files) found actual problems. retire.js released a fix in 5.6.0 and Langfuse in v4.18.0 because of my reports, and I reported the problem with NeMo Guardrails to NVIDIA's security team. My agent red team benchmark is uploaded to Hugging Face and had 634 downloads.
I do my best to document failed experiments as well. The monitor experiment alone involved 13 pre-registrations, and the paper includes all 11 predictions that did not pan out.
Public works: github.com/Alkur123/two-runs-one-output, huggingface.co/datasets/jash-ai/agentic-redteam-benchmark, github.com/Alkur123/agent-containment-replay, and my blog posts at lesswrong.com/users/jaswanth-alkur.
The most probable form of failure is that the finding will not replicate, as in the case when the effectiveness of the method depends on the manner in which SHADE-Arena constructs tasks. In my opinion, it is important to discover it, publish, and let the evaluators know where this signal will fail when they use it. Moreover, there
is an issue that I am aware of and that is mentioned in the paper, which is sabotage without any authored target.
None so far. All of this was self-funded. On 10 October 2026, I have submitted an application to EA Funds Transformative AI Fund for my stipend and year one research budget. The remainder is covered in this page; human annotation, additional compute power, open source software development, external advisory, and second round of replication studies. I am applying to other funding agencies separately for other portions of the two-year plan, and I will indicate here if they go through.
There are no bids on this project.