You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
AI systems learn from the scores they receive. If the system giving those scores rewards wrong answers or broken code, an AI can get better at passing the test without getting better at the task. I am building this to find such failures, work out what causes them, and check whether fixes hold up automatically.
We have already demonstrated a concrete example in MASK, an honesty test in the Inspect Evals library associated with the UK AI Security Institute. The AI judging an answer correctly recognised that it was false, but the software reading the judge’s response recorded it as honest. I traced the error to how the software extracted the verdict and demonstrated a local repair.
The project will turn findings like this into tools that other teams can use to check the systems they rely on for training and assessing AI.
Our goal is to find and fix weaknesses in the tests used to judge whether AI systems are safe to release. If these tests give passing scores to incorrect or dishonest answers, they can create false confidence in a model’s safety. As AI systems become more capable and take on more responsibility, catching these failures before release becomes increasingly important.
Our tool looks for cases where an AI gives a wrong answer but still passes the test, and saves the evidence so others can check what happened. The product is already at beta stage. We need funding to apply it to real AI evaluators whose results inform release decisions.
We also want the tool to help repair the problems it finds. We have already demonstrated a specific failure and a local fix in the UK AI Safety Institute's honesty evaluator. The next step is to make this process work reliably across more AI evaluators.
Over six months, we will use the tool to try to make evaluators accept wrong answers. Whenever it succeeds, we will check the result ourselves, identify what went wrong and test a fix. We will then try new wrong answers to see whether the fix holds up, while checking that correct answers still pass.
We will seek evaluation teams to trial the tool on their own tests. We hope to find weaknesses they had missed and helping them fix those weaknesses before the results are used to support a safety claim or release decision. We will track how many real problems the tool finds, how often it raises false alarms, and how much each audit costs.
I am requesting $10,000–$200,000. At the full funding level, the six-month budget would be:
$75,000 for model API credits and compute.
$50,000 for research and engineering contractors.
$15,000 for my living expenses, including taxes.
$20,000 for datasets and licences.
$8,000 for coding agents.
$8,000 for legal advice and IP protection.
$4,000 for software, hosting and accounting.
$20,000 contingency buffer.
At $10,000, we would run a much smaller pilot with fewer experiments and no contractors. Higher funding would allow us to take the project from research to a start-up.
The team is me and my uncle, who advises on the business side. I recently completed my master’s in physics at University College London. My research, currently under review at ICML, examines how a training setting can distort comparisons between AI agents with and without memory. In another published paper, I showed why seemingly sensible fixes to AI evaluation need careful testing: asking an AI judge to solve a task before assessing another answer helped when the judge’s own answer was right, but made matters worse when it was wrong. I have several other publications in physics, one published in Nanoletters, and another expected to be published in IEEE JSAC.
I handle the research and technical development and have brought tech to beta. My uncle founded a very successful international fertility-treatment business. As we develop the project into a start-up, he brings experience growing a research-based business.
The biggest risk is whether the failures we can find are important enough to change a team’s assessment of an AI system. A tool could uncover many scoring errors without finding anything that affects a safety conclusion or release decision.
Access is also another risk. We can work with public evaluators, but teams may be unable to share their internal tests, model access or results. Without those partnerships, we could end up with a tool that works on public examples but has no influence on the evaluations used for release decisions.
If these problems prevent adoption, we will have an interesting open source research tool instead of a viable start-up. This outcome would fall very short on the wider impact we are aiming for.
None sought and none raised so far.
There are no bids on this project.