You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Open Parity Bench is a public test that lets anyone check two things about an AI request: how much power it really burned, and whether a router's choice of a smaller model gave an answer just as good as a big one's. Right now, nobody can verify those claims outside a lab. We can't either.
It's a safety question because routing is taking off. Every router says it's cheaper, greener, and just as good. If that's not true for high stakes tasks, people get bad answers and don't even know. Our bench makes that promise something you can test. One prompt goes out to several unrelated model families; the bench checks where answers match or differ claim by claim, and shows that next to real energy measurements. When models disagree, it's a low cost sign an answer might be wrong, and it catches things self consistency misses, since one model can keep making the same error.
Our own estimates show the problem. Based on published numbers (IEA, Energy and AI, 2025; arXiv 2505.09598, 2510.01889, 2509.20241), we think routed traffic uses 0.03 to 0.24 Wh per query, versus 0.42 to 1.79 Wh for a high compute model. That's 60 to 85% less energy. But it's just a model. The mix is assumed, not measured, and it's about energy, not carbon. We can't prove it now, and the same goes for anyone else making that promise.
Here's what we'll build:
Energy measured per request right at the wall and GPU on our own hardware with open weight models, so it's physical, not estimated.
Every routed answer gets graded against a large model baseline, weighted 40% task success, 30% semantic similarity, 20% LLM as judge, and 10% human labels on a held out set.
Claim by claim disagreement scoring across separate model families.
Milestones (timing starts when funding is received):
Month 3: Metered bench is live, energy method draft is public, baselines done for 3 open weight models.
Month 6: Parity scoring with human labelled held out set is added, disagreement scoring works, first public results out.
Month 9: Full task suite runs comparing routed and single model, measured results versus our estimates.
Month 12: Version 1.0 is done, with docs and a report that includes what didn't work.
Success looks like code anyone can run, metered numbers covering 5 or more open weight models over a shared public set of tasks, and a clear answer: does routing really save 60 to 85% of the energy with parity, or a lot less? We'll share the results either way.
Everything we make will be open: code under Apache 2.0, and data and results under CC BY 4.0. Our commercial platform, ChatFuse, isn't part of this grant.
Over 12 months we ask for $89,100. The founder spends half his working time on it, costed at $50,000. Buying 2 GPUs plus inline power meters for the measurement rig comes to $12,000. Hosted model comparison calls are $9,000, labelling the held out set by hand is $8,000, and keeping the public repository and results online is $2,000. The 10% overhead on those direct lines adds $8,100. A minimum of $29,000 pays for the rig, the labels and the API calls, which is enough to put measured numbers for 3 open weight models in public.
ChatFuse LLC is an AI company in San Diego. ChatFuse is an organization of specialized AI agents, each responsible for a different function (engineering, solutions architecture, QA, research and client communications), led by founder and CEO Nico Coetzee. We run a platform that handles over 100 models from places like Anthropic, OpenAI, and Google. It figures out the best model to use for each request, one that's efficient but still good enough, and if one fails, it tries another.
Nico built the platform itself: the routing layer, the retry and fallback engine, the memory layer on Postgres with vector search, and the tool pipeline.
Risks: Model providers might overfit to the test until it's meaningless, so we'll rotate held out sets and keep some human labels secret. Another risk is people using small models for critical jobs where errors are costly, so we report by task type and give failures the same attention as wins. Disagreement scoring goes out first.
Why a for profit wants a grant: The bench doesn't make money. It only helps if it's open for everyone. Nobody has invested in ChatFuse; it runs on its own revenue. Unfunded, the bench would remain something we run privately. We'd gain from having real numbers to cite instead of estimates, and those numbers might show we were too optimistic.
$0. ChatFuse LLC was formed in California in September 2025, has taken no outside money, and has never received a grant.
Other funding: We're also preparing an application to the Foresight Institute AI for Science and Safety Nodes RFP for the same project. Any funding from them would lower what we ask for here.