You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I worked 23 years in bank credit in Adana, Turkey. Last year, I started publishing open analysis of the complete US federal mortgage dataset (HMDA, covering 1,187,606 FHA credit decisions in 2025) on my site, financeratecalc.com.
Then I noticed a pattern. Whenever an AI model quoted my numbers, the digit itself was almost always right. But the sentence built around it was almost always wrong. "22.1% of decisioned FHA applications were denied in 2025" somehow mutated into "a quarter of mortgage applicants get rejected." The target population vanished, the loan program vanished, the time frame vanished. Standard accuracy benchmarks miss this entirely because the raw number is sitting right there.
So I built a system to track and measure this.
Every statistic I publish now ships with a machine-readable contract (defining what the number means, what it excludes, and how it can be rephrased) alongside a receipt: a cryptographic hash of the claim that lets any quote be verified, ensuring updated figures explicitly mark themselves as legacy. On top of that, I designed an evaluation setup (Inspect AI framework, 12 fixed benchmark questions across 4 environments: no tools, web search, my dedicated MCP server, and a receipt-aware client) plus an RL environment (verifiers setup) where rewards are calculated without relying on another LLM: Is the figure present? Is the receipt authentic? Has anything been hallucinated or forged? Did the output cross a defined boundary?
Testing a frontier model (12 questions x 3 runs) yielded clear results: the digit was accurate in 100% of the numerical answers. However, full contract compliance scored between just 1 and 11 out of 36 responses, depending on the tool output provided. The receipt survived 0 out of 36 times in plain prose, but reached 23 out of 36 when the client explicitly treated it as part of the data point. In 2 out of 36 cases, the model actually fabricated a receipt even when valid ones were supplied by every tool. The RL environment runs end-to-end (initial rollout: 0.75 mean reward, 9/9 receipts valid, 0 forged).
The entire system is public under CC BY 4.0: the specification, the verifier, the task suite, every evaluation run log with named graders, and my own errors. During development, I uncovered 13 distinct error types on my own platform and cataloged them using the exact same error codes applied to the models.
A snapshot of all claims is signed via Sigstore into the Rekor transparency log. All figures are recomputed directly from the official CFPB source file inside GitHub Actions with build attestations—so no one has to blindly trust my local machine, including me.
Expand the benchmark from 12 to roughly 200 questions. With 386 contracted claims backed by data-generated and self-checked keys, scaling the test suite is straightforward and will provide solid statistical power for the fidelity metrics.
Bring in a cross-provider secondary grader. Relying on a single model grader is the most obvious vulnerability in the current evaluation pipeline (a limitation explicitly acknowledged in the write-up). Running a second independent provider will eliminate single-judge bias and solidify the agreement scores.
Train a small open-weights model on the deterministic reward environment. Test whether the policy of carrying the receipt without forging it generalizes to unseen claims out-of-distribution. If the behavior fails to transfer, publish the negative finding with the exact same rigor as a success.
Deploy the contract specification to an external publisher. The success condition is hardcoded into the spec itself: at least one independent consumer publicly attesting to validating against a claims.json manifest within 90 days. If adoption sits at zero, report the null result cleanly and transparently.
Minimum ($6,000): Roughly $2,500 in API credits across two providers (200 questions x 3 runs x 4 environments x 2 model graders), around $1,500 in compute to fine-tune one small open model, and the remaining amount as a stipend covering six months of evening and weekend work. I hold a full-time job; this project is built after hours.
Full ($18,000): Expands the project to include a second external publisher, integrates a third independent grader, and funds a part-time helper to handle outreach for the 90-day adoption test.
Just me. Ziya Yetiş—bank branch manager with 23 years on the credit side. I've published four SSRN working papers analyzing this dataset (7156938, 7309319, 7341481, 7423798), released a benchmark on Hugging Face, and maintain an MCP server in the official public registry.
To date, this project has received zero funding, accepts no lender money, and sells no leads. My public corrections log is the piece of work I am most proud of.
If the second grader disagrees with the first on a majority of answers, my measurement reduces to grader noise, and I will publish that finding directly. If the fine-tuned model only carries receipts on claims present in its training data, the reward environment is teaching memorization rather than generalization, meaning the setup is weaker than assumed. If no consumer attests to checking against a manifest within 90 days, the format remains a specification that nobody enforces. All three outcomes will be documented on the null results page with a date stamp, in line with every other log on the project.
$0, from nowhere.
Links
Project & Open Data: financeratecalc.com
Hugging Face Benchmark: huggingface.co/datasets/ziyayetis/fha-denial-reconciliation-2025
SSRN Research Papers: ssrn.com/author=8328657 (Papers: 7156938, 7309319, 7341481, 7423798)
Model Context Protocol (MCP) Server: github.com/modelcontextprotocol/servers (Registry: frc-mcp)
Public Corrections Log: financeratecalc.com/corrections