You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I built RCV-Bench, a benchmark that checks whether an AI agent can tell a real research result from a broken or fabricated one. You hand the agent a claimed number and the code behind it, and it has to return a verdict (reproduced, deviation, fabricated, or robust-turned-fragile), the value it regenerated, where the problem is, and how confident it is. v0 is already public and running at github.com/GrobeStreet/rcv-bench, with a live leaderboard. It has 14 tasks built from 4 real, pinned public research repos, a scorer that behaves the same way every run, and CI that rebuilds the whole benchmark on every commit. I am asking for funding to take it from a proof-of-concept to something other people can evaluate their own agents against.
The point of the benchmark is a gap I can already show. An agent that just re-runs the code looks fine on the surface, 9 of 14, but it goes 0 for 3 on results that fall apart under a stated perturbation and 0 for 2 on results with no real backing artifact. Re-running catches a wrong number. It does not catch a fragile or fabricated claim, and verifying a claim and reproducing a number are different skills. Almost nothing measures the first one. With about three months of focused work I would take v0 from 14 tasks to roughly 45, each built from a real public repo with a documented, reproducible way of generating the gold answer. I would add a held-out, private split so agents can be scored without the answers leaking, and add real LLM-driven agents as baselines instead of the simple reference policies that are there now. Everything stays public and reproducible.
This funds my time as an independent researcher, plus compute. Researcher time, roughly 10 to 12 weeks: building and checking the new tasks, the private held-out split, and the live-agent baselines. Compute: sandboxed regeneration of the new tasks, which means model inference for the audit and MMLU-style tasks and re-running the source pipelines at their pinned commits. I set a minimum of $10,000 and a goal of $35,000. At the minimum I finish the core of v1; the full amount lets me cover more repositories and add stronger agent baselines.
It is just me, an independent researcher. My approach is reproduction-first and honesty-first: I pre-register predictions and publish the ones I get wrong. The strongest external check on my work is the FAIR Universe weak-lensing reproduction, a NeurIPS 2026 competition graded on Codabench, where I reproduced the official out-of-distribution baseline and found and fixed a scoring bug in it. My other public work (mmlu-robustness-audit, de-stress-lab, arc-agi-2-occam-baseline) is on github.com/GrobeStreet. RCV-Bench itself is already built, public, and running, so this is scaling something that works rather than starting from scratch.
The most likely failure is scope: verification tasks are slow to build well, so I could end up with fewer high-quality instances than planned. I would rather ship 30 solid, reproducible tasks than 45 shaky ones, and I would say so plainly. A second possibility is that strong agents solve the benchmark quickly, which would be a good outcome and would tell me where to make it harder. The smallest useful result is a public, honest benchmark that shows the reproduce-versus-verify gap on real repositories, and even the minimum funding gets that shipped.
Nothing so far for this project; it has been self-funded solo work. I just applied to BlueDot Impact's Rapid Grants for a smaller $10,000 version of the same work, and this Manifund proposal is the fuller three-month version. No other funding is committed.
There are no bids on this project.