You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I am testing a fairly simple question: if an AI research agent reaches a wrong conclusion, how likely is another AI agent to catch it?
This sounds obvious, but there is an important complication. The verifier may use the same model, see the same reasoning, share the same assumptions, or simply repeat the same error. I want to measure how much genuinely independent verification helps.
I will build a set of research tasks where the proposer produces a conclusion that can be independently checked.
For each task I will run different verification conditions: same model vs different model, blind vs exposed verification, and different amounts of compute.
The main quantity I want to estimate is: P(detection | the proposed conclusion is materially wrong).
I also want to measure the cost of gaining additional detection probability.
I already have an experimental pipeline running and preliminary verification experiments. Funding would let me increase the number and variety of tasks, models and repetitions, and make the resulting dataset and methodology public.
At the minimum funding level, I would run a focused version of the experiment with several models and enough repetitions to estimate the main effects.
Additional funding would allow more task families, more independent replications, stronger models, larger sample sizes, and external review of the methodology.
The main costs are model/API compute, research and engineering time, data processing, replication, and publication of the results and code.
I am Ilya Nikitin, founder and director of The Null Institute. I have a Ph.D. in microelectronics and a background in research, engineering and entrepreneurship.
The Null Institute is a new independent research project built around reproducible analysis, adversarial testing and independent verification. I have already built the first version of the research-agent system and am using it for experimental research projects.
Project website: https://nullinstitute.org/
The main possibility is that independent verification adds much less than expected. Different models may still make correlated errors, or the effect may depend strongly on the type of research task.
I would consider that a useful negative result rather than a failed experiment. The purpose is to measure whether the proposed safety mechanism actually works, not to demonstrate that it does.
Another risk is that the task set is too artificial. I plan to address this by including several task types and by publishing the tasks and protocol so others can challenge or reproduce the result.
No external funding has been received for this project so far. The work to date has been self-funded.
I currently have funding applications pending with several other funders, including EA Funds / Transformative AI Fund, Iliad, Emergent Ventures and Lightcone. If overlapping work is funded elsewhere, I will reduce or redirect the Manifund request rather than fund the same work twice.
There are no bids on this project.