You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
In plain terms: we let AI systems grade other AI systems, and almost nobody checks whether the grader itself is reliable. This project checks the grader first, in a way that is written down before the data exists and published whatever the result. We evaluate models with models. LLM judges score leaderboards, safety evals and preference data, and almost nobody measures the judge against itself first. When I checked mine, the judge estimated citation accuracy at 65–96% where deterministic source verification put it at 47–67% (n = 20 themes; on three of the four models measured). In a separate protocol, a production model abandoned an answer it had just given on 77.6% of the 85 items it had accepted, versus 1.2% in a control arm on the same 85 items without the challenge. If judges are that unstable, agreement between two of them can measure a shared bias rather than merit — and every number downstream inherits it. This question concerns me personally because I was building on top of a tool I had not calibrated: however good a model is, an error at the base is a building that eventually collapses. That is where my concern with measurement began.
Two things, both pre-registered before any data is generated.
A fresh sealed bench. Items disjoint from my existing labelled set, balanced against the theme prior, stratified with per-stratum baselines, and sized against a minimum effect declared in advance instead of the effect I happened to observe. Planning envelope: ~284–562 cells, $335–500 of compute at list prices. The pre-registration is deposited with a DOI before the first cell exists.
Judge test–retest. How consistent is an LLM judge with itself under declared perturbations — re-roll, order, phrasing — reported as chance-corrected coefficients against a usability criterion written down beforehand. Published either way, including if the answer is boring.
Where this leads. This grant buys one link in a chain, and only the first. Measured judges come first: until we can say how reliably a judge agrees with itself and with checkable sources, nothing built on top of one is worth trusting. The next link is a detector for how a model handles a claim that is contested rather than settled. After that, a correction loop that resolves disagreement against verifiable primary sources instead of preference agreement. The link worth having at the end is systems that say plainly when they have no source, rather than producing fluent text anyway. Only the first link is in scope here. The rest is a direction I can argue for, not work I am promising to deliver in six months.
Milestones. Start: 2026-10-01. Six months of full-time work, one milestone per month.
Month Window Milestone
M1_Oct 2026 Pre-registration written and deposited with a DOI, before any cell exists.
M2_Nov 2026 Sealed items built and held out encrypted, balanced against the theme prior.
M3_Dec 2026 Bench run executed against the sealed set, with per-stratum baselines.
M4_Jan 2027 Judge test–retest under the perturbations declared in the pre-registration.
M5_Feb 2027 Written report against the pre-registered usability criterion, negative result included.
M6_Mar 2027 Dataset released CC BY-SA, code AGPL, everything archived with a DOI.
At the $5,000 minimum, M1–M3 and the release still ship on this calendar; M4 does not happen and M5 shrinks to the bench's own result, published with its baselines. No salary is drawn.
$25,000 part-funds six months of full-time work: author time, non-Anthropic compute, dissemination and archiving. The full budget is EUR 40,000 (~$45,000).
At the $5,000 minimum I skip the salary entirely and ship one module: the sealed bench's compute run plus its deposited pre-registration and the CC BY-SA dataset. No test–retest programme — one sealed, honest measurement, published with its baselines. I would do it anyway, because an AI that can be trusted is worth a different future, and from there nobody should step back. If one detail can bring down a building, decisions taken on complex but flawed logic can do far worse to a society.
If both this page and the TAIF application are funded, money raised beyond the EUR 40,000 budget goes, in this order, to marginal uses declared to the funders — the first two in the TAIF application, the third added to it by addendum: (1) a second-annotator blind re-read on a subsample of the gold labels, turning the declared single-annotator limit into a measured inter-annotator figure; (2) sizing the sealed bench at the upper end of the published envelope (~562 cells) for higher power; (3) a public, leak-free release of the existing 2,000-cell corpus as a per-claim bias dataset, in the redaction that survives a pre-declared disclosure audit, with contamination canaries and machine-readable metadata. Anything beyond that is returned.
Outputs: dataset CC BY-SA, code AGPL, written report. Negative result published either way.
I run this alone, in Italy, with no affiliation, which normally reads as a credibility problem. Working alone was not a preference. Around me there is no colleague available for an AI project that starts from the humanities and from ancient texts; waiting for a research group would have meant never starting. Here is my counter-offer. Under a pre-registered blind protocol I withdrew my own best numbers, with dates on the public repository, and I publish the surviving one next to a baseline that beats it: 49/62 = 79.0% blind, against a zero-parameter trivial rule at 90.3% (p = 0.033) and a majority baseline at 72.6% (paired p = 0.23). I withdrew them because I always check the chain of reasoning that produces a result, not just the result, and that day the chain did not hold. I run the same check whether a number is surprising or discouraging. My strongest figure — 253/269 = 94.1% versus a 47.6% baseline, at the packaged build of 2026-09-02 — is in-sample, and I refuse to call it accuracy. Bench so far: 2,000 cells, 200 claim slots, 283 gold labels, one annotator (a stated limit; a second annotator on a subsample is an option, not a promise). Two experiments were killed by my own checks before they cost money. An internal adversarial audit of the gold labels found a ~2% contamination rate in one label class; the corrections LOWERED my internal numbers, are recorded with dates, and each produced a new written labelling rule.
Repo: github.com/claudiodegenua/precorrect-method · DOI 10.5281/zenodo.22345121 · ORCID 0009-0008-5896-3172. A second, unrelated dataset of mine (a six-layer parallel Psalter, 2,469 verse-cells, per-layer measured quality) is deposited as DOI 10.5281/zenodo.22770451 — evidence that I ship data with a data sheet, not only claims. Recent work on verifier reliability is converging on this question — arXiv 2506.13342 (Verifying the Verifiers) on label quality in fact-verification benchmarks, and arXiv 2606.19544 (Reliability without Validity) on chance-corrected agreement and test–retest across 21 judges. What I add is orthogonal to both: the bench is sealed and the usability criterion is pre-registered with a DOI before any cell exists, and ground truth resolves to checkable primary sources rather than preference agreement — which is what makes a usability verdict, rather than a reliability description, possible at all.
How to check me. Every number I withdrew is listed with its date in the public repository — github.com/claudiodegenua/precorrect-method — next to the one that replaced it. Read that ledger before you read my claims.
The headline risk is already on the record, not hypothetical. On the current blind set, a zero-parameter rule that looks only at WHICH THEME a claim belongs to — never reading the text under test — scores 90.3%, beating the detector at 79.0%. Finding that out was unpleasant and useful in the same moment: it changed the design of the bench. I did not drop the project for a simple reason: a residual error on the themes that matter is not small, and I do not believe machines will close it on their own while nobody works on the structure that keeps them in line. That number says nothing about models being right: it says the blind set's difficulty is concentrated in the theme prior, so the bench so far measures theme difficulty more than text reading. This matters beyond my project: aggregate scores hide exactly the strata that matter most — a judge can look excellent on average and still fail on the contested themes it exists for. The new sealed bench is designed against precisely this failure: balanced against the theme prior, with per-stratum baselines declared in advance. If it confirms the null, that is a publishable negative result, and this grant buys precisely that answer. Second risk: a single annotator — a declared limit, and the first marginal use converts it into a measured inter-annotator figure. Third: benchmark contamination — every released file carries a project-specific canary string, and the sealed holdout ships encrypted rather than in plain text.
Nothing received to date. No grant, prize, salary or donation has been received for this work in the last 12 months. What exists are applications, not funds:
EA Funds — Transformative AI Fund (TAIF): an application to EA Funds' Transformative AI Fund for the full budget was submitted on 2026-09-09 and is pending. An application, not money received.
Anthropic ERA: a request ($1,000, compute-only credits, submitted 2026-09-05) was not approved in the September cycle (the programme notifies only successful applicants); I intend to reapply in the 5 October cycle. It is strictly complementary — it buys Anthropic-model inference, not time. An application, not money received.
AI Safety Fund (AISF): an application is pending. An application, not money received.
There are no bids on this project.