You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I’m building a small independent research project to test whether widely used AI evaluations are actually as stable and trustworthy as their headline scores suggest.
I already built a public MMLU robustness audit that tests answer-order sensitivity using fixed-seed cyclic reordering and frozen results. I’ve also built deterministic evaluation items with machine-verifiable checkers and a reproducibility-first research workflow.
This 90-day pilot would turn that work into a reusable evaluation-auditing framework, apply it to two additional public AI evaluations, and publish the code, frozen artifacts, failure cases, and a short technical report.
The core question is simple: if an evaluation score changes meaningfully when we alter things that should not matter—answer order, formatting, seed, evaluator settings, or other non-semantic details—how much confidence should we place in that score?
Existing work:
https://github.com/GrobeStreet/mmlu-robustness-audit
https://github.com/GrobeStreet/ai-eval-work-sample
https://github.com/GrobeStreet/bobby-research-os
The 90-day pilot has four concrete goals:
Turn my current MMLU robustness audit into a reusable framework for perturbation and reproducibility testing.
Run full robustness audits on two additional public AI evaluations.
Implement at least one audit in a format compatible with Inspect AI or another widely used evaluation framework.
Publish a short Evaluation Reliability Report with code, frozen results, null findings, limitations, and reproducibility instructions.
Each audit will start with a frozen research question and perturbation plan before results are inspected. I’ll reproduce the reference evaluation first, then test controlled changes such as option order, formatting, random seeds, scoring configuration, calibration, and other implementation choices.
The goal is not to produce another leaderboard. It is to test the reliability of the measurement process itself.
Most of the funding would buy protected research and engineering time so I can work on this full-time for the pilot period.
At full funding, I expect roughly:
$24,000 — three months of research/engineering time
$7,000 — model API usage and compute
$4,000 — independent replication/statistical or technical review
$3,000 — research workstation / local compute
$2,000 — storage, software, hosting, and reproducibility infrastructure
I will use free credits, open-source tools, and local compute wherever practical rather than spend grant money unnecessarily.
The project is useful at lower funding levels too. With partial funding I would reduce the number of audits or external replication work rather than weaken the reproducibility standards.
I’m Robert “Bobby” Morong, an independent research engineer working on AI evaluation, statistical stress-testing, reproducibility, and scientific verification.
I currently work independently. My public work includes:
MMLU robustness audit:
https://github.com/GrobeStreet/mmlu-robustness-audit
Machine-verifiable AI evaluation work sample:
https://github.com/GrobeStreet/ai-eval-work-sample
Research/reproducibility operating system:
https://github.com/GrobeStreet/bobby-research-os
GitHub:
https://github.com/GrobeStreet
I use AI heavily for research orchestration, hypothesis generation, literature review, code assistance, and adversarial critique, but I try to keep empirical claims grounded in deterministic code, tests, frozen artifacts, and reproducible experiments.
There are a few ways this could fail.
The evaluations I test may turn out to be much more robust than expected. The failure modes I find may not generalize beyond a specific benchmark. Or I may build technically sound tooling that other evaluators find too cumbersome to adopt.
Those are still informative outcomes. I plan to publish null results and failed experiments rather than only successful findings.
The worst outcome would be producing a complicated framework that nobody uses. To reduce that risk, the pilot focuses on a small number of concrete audits and integration with existing evaluation tooling rather than building a large standalone platform first.
I have not raised dedicated funding for this project in the last 12 months. I have submitted a separate application to Lightcone Commons for a larger nine-month version of the Open Evaluation Robustness Lab. It is currently pending. Any overlapping funding would be disclosed and budgets/milestones adjusted to avoid double-funding.
There are no bids on this project.