You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
A lot of current AI safety evaluations use LLMs as judges. What is less clear is how reliable these judges actually are when the same answer is presented in a slightly different way.
In our preliminary experiments, we found a surprisingly large change in agreement. Cohen’s κ went from 0.78 to 0.24 after making changes that should not have changed the meaning of the answer. These included simple typos, replacing words with synonyms, changing the order of sentences, and changing the writing style.
We want to test whether this is a small effect from our initial experiments or a broader problem with LLM-based evaluation.
For JudgeRobust v1.0, we plan to test 10 judge configurations across 500 prompts from three evaluation datasets: MT-Bench, WildBench, and SafetyBench. We will apply 20 different types of meaning-preserving changes and run roughly 100,000 evaluations in total.
The main result will be an open benchmark, together with a κ-stability metric and the code needed to reproduce the experiments. The benchmark will make it possible to test how much a judge's decision changes when the underlying meaning stays the same.
The funding will mainly allow us to run the full experiment instead of having to reduce the number of models, datasets, or perturbations because of inference costs. It will also cover the research time and the engineering work needed to release the experiment in a reproducible form.
The main question is straightforward: if two versions of an example mean the same thing, will an LLM judge evaluate them in the same way?
We will test this by taking examples from MT-Bench, WildBench, and SafetyBench and applying 20 different perturbations. These will include things such as typos, synonym changes, reordering, and changes in style.
We will then run the resulting examples through 10 judge configurations, giving us approximately 100,000 individual judgments.
For the analysis, we will calculate κ-stability for each judge, dataset, and perturbation type. We will also use bootstrap confidence intervals so that we can report the uncertainty around the results rather than relying only on point estimates.
Once the experiments are complete, we plan to release the benchmark, code, results, and evaluation pipeline publicly.
Total requested: $19,800
Compute / API calls — $7,500: Model inference for approximately 100,000 evaluations across the different judge configurations.
Researcher time — $8,000: Experimental design, implementation, running the experiments, statistical analysis, interpreting the results, and writing the final research output.
Engineering / reproducibility — $2,500: Docker, CI/CD, dependency pinning, experiment management, and the infrastructure needed to reproduce the evaluations.
Contingency — $1,800: Additional runs, API price changes, and unexpected research costs.
Most of the budget goes toward inference and researcher time. The number of evaluations matters here because we want to compare several judges across several datasets and perturbation types. Cutting the experiment down too much would make some of those comparisons difficult to interpret.
The engineering budget is there because we do not want the project to end as a set of results that are difficult for someone else to reproduce. The goal is to make it possible for another researcher to run the same experiments and later add new judges or perturbations.
PI / Research Lead — Nurlan Orucov
Nurlan is an AI safety researcher working on technical AI safety, AI evaluation, and technical research. He will lead the project, including the experimental design, implementation, analysis, and writing.
Researcher — Rovnag Nadirov
Rovnag is an AI safety researcher at LASR (London AI Safety Research) and has experience with AI safety research and technical R&D.
The team has also contributed to Anthropic R&D, giving us experience with practical AI research and evaluation in addition to our own research work.
The project comes directly from our preliminary experiments. Those experiments were self-funded and cost approximately $150 in API usage. They were enough to identify the initial effect, but not enough to tell us how general it is.
The main possibility is that the effect we saw in the preliminary experiments turns out to be smaller or less consistent when tested on a larger set of data.
It is also possible that the effect is present on one dataset but not on others, or that some judge configurations are much more stable than others.
This is why we are using three different datasets and multiple judge configurations rather than building the project around one evaluation setup.
If the larger experiment does not reproduce the initial result, we will still publish the measurements and the negative result. That would tell us something useful about where LLM judges are stable and where our initial observation does not hold.
There is also a practical cost risk. API usage will be capped and monitored, and the evaluation pipeline will be modular. If prices change or some experiments become too expensive, lower-priority runs can be removed while keeping the core experiment intact.
We have not received institutional funding for this project during the last 12 months.
The preliminary experiments were self-funded and used approximately $150 in API credits.
We are currently requesting $19,800 from Manifund to run the full evaluation and develop JudgeRobust v1.0.