You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
LLM judges — the scoring step in RLAIF, Constitutional AI, and most automated evals — decide which AI outputs are good enough to keep. I took 200 outputs from my own production pipeline, each mechanically verified to contain fabricated claims, and tested the judges on them.
I tested them across four models — gpt-oss-120b, Llama-3.3-70b, grok-3-mini, and claude-haiku-4.5 — for 2,000 judgments total (10 runs × 200 items). Every result is sealed on the Bitcoin blockchain (OpenTimestamps), and the decisive follow-up tests were pre-registered — interpretation thresholds anchored before a single call was made, so the methodology couldn't be tweaked later.
Judges approve a surprising amount of fake info: When given no source text to double-check, judge models approved between 43% and 100% of confirmed fabrications. Every single model failed — some just failed worse than others.
Standard fixes are unreliable: Giving the judge a basic rule like "Here is the source text; reject anything unsupported" gave wildly inconsistent results. One model blocked everything (0% leaked), while another let 65.5% of the fabrications slide right through. Swapping judge models under the hood can silently tank your application's reliability.
Making the judge "do the math" actually works: When I forced the judge (using gpt-oss) to break down the text line-by-line — citing the exact supporting quote or explicitly marking it missing — the failure rate dropped from 11% (that same judge's with-source baseline) down to 0.5%. And no single fabrication survived every defense we tested.
I didn't rely on another AI's opinion to label what was fake. Instead, every claim was broken down and checked against source material, then validated against a separate human-anchored dataset.
I stumbled across this problem while building autonomous AI systems and creating an open-source auditing tool for them (BIJOTEL). This project packages those findings into a citable research paper, complete with open, verifiable datasets and a practical test harness so anyone can benchmark their own AI judge before blindly trusting it.
We already know that LLM judges tend to reward fluent, plausible text over factual accuracy. Recent papers have tackled parts of this: "More Convincing, Not More Correct" (arXiv:2607.05904) demonstrates reward-hacking in self-play with reference-free judges, while "Faithful or Fabricated?" (arXiv:2605.23970) identifies rationalization bias on static summaries.
Where the existing literature falls short isn't in proving that judges are biased — we know they are. The real gap is that we don't know how unreliably standard grounding fixes perform across different LLM judges on actual fabricated outputs, nor do we have a proven protocol to fix it. We're addressing this directly by measuring cross-judge variance and introducing a counting protocol that actually closes the gap, backed by verifiable provenance.
Key Deliverables:
The Paper: A preprint targeted at a scalable-oversight workshop, featuring our cross-model matrix and a thorough related-work analysis.
Open Datasets: We're releasing both the 200-item cross-model matrix (2,000 judgments across 10 runs) and a 2,091-pair historical corpus under MIT/CC licenses. To keep everything fully verifiable, these come with OpenTimestamps/Bitcoin proofs and a fixed extraction rule (linking/hashing source texts where copyright limits redistribution).
The Testing Harness: An open tool you can point at any OpenAI-compatible judge endpoint to instantly benchmark its blind, grounded, and forced-counting baselines in a sealed, reproducible format.
Funding will support expanding our validation work — specifically running our counting protocol across more model families (we've tested it on gpt-oss so far) and scaling up the human-labeled held-out set behind our oracle.
A $30k grant covers roughly 6 months of my full-time work to turn this project into rigorous, reproducible research. That includes:
- Writing and publishing the paper.
- Cleaning and releasing both datasets alongside their provenance proofs.
- Building and documenting the test harness.
- Running cross-model validation for the counting protocol.
- Hardening and documenting BIJOTEL so other testbeds can adopt the same provable provenance standard.
Because running open-weight models is so cheap — the entire 2,000-call matrix barely cost anything — I don't need staff hires or new infrastructure; just a small budget for a second independent annotator on the expanded validation set.
Funding Tiers
At $10k (Minimum): I'll deliver the core public goods — the preprint and both provenance-verified datasets.
At $30k (Full Goal): I'll be able to complete the full scope, adding the harness, cross-model counting validation, and BIJOTEL hardening + documentation.
I'm a solo researcher and the founder of Aisophical SRL, an EU-based AI safety SME in Romania. A few receipts of my work:
- Research: Author of the arXiv preprint "Emergent Formal Verification" (arXiv:2603.21149).
- Open Source: Built BIJOTEL (~14k PyPI downloads) and substrate-guard.
- Production: Running a live autonomous system with over 40,000 cryptographically sealed model calls logged since May.
I take calibration seriously. Before asking for funding, I intentionally tried to disprove my own headlines — and succeeded twice. My initial finding that "with-source approval stays at 45%" collapsed when scaled up (n=20 to n=200), and another claim about an "11% floor" turned out to be mostly an artifact of the prompt setup. The second collapse happened under interpretation thresholds I had pre-registered and Bitcoin-anchored before execution — I couldn't move the goalposts after seeing the data. What you're looking at now is what survived.
ORCID: 0009-0007-1106-2644
Here's a realistic look at the risks. First, the counting protocol has only been tested on one model, so there's no guarantee it works across the board yet. That said, mapping out where it breaks is still a valuable, publishable result in its own right.
Content variety: All 200 fabrications come from one specific pipeline (tech and AI news). Other domains might behave differently, but that's exactly why the testing harness was built — so anyone can test their own data.
Small test set: The human-verified benchmark set is currently tiny (13 items, two independent review passes). Expanding this set is already budgeted and funded work.
Bandwidth: I'm running this solo, which naturally caps how fast things move.
The bottom line: Even in a worst-case scenario where the core hypothesis fails, nothing is wasted. The datasets, testing harness, and raw results — positive or negative — will be fully open-source. The claims might fail, but the tools will still be useful to the community.
We're 100% self-funded — no outside capital or cash grants so far. We were accepted into Cloudflare for Startups and NVIDIA Inception, which cover a big chunk of our compute costs through credits, but no actual cash has come in yet. This would be our first direct funding.
There are no bids on this project.