You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Several AI checkers can agree and still be wrong together. Even a good measurement showing they add independent information can expire when a model or task changes. This project tests a narrow fail-closed rule for deciding when an AI-generated claim has earned permission to trigger an action.
The rule combines two deliberately limited signals. The first is a public detector that measures whether repeated answers stay semantically stable. The second is a time-stamped ledger that watches whether the checkers actually disagree, and flags when it's time to rerun the labeled check of how often they fail together. Neither signal establishes truth. Together, they must beat four simple baselines at the same false-accept budget without merely refusing everything.
Here is the strongest case against funding me: detectors like mine already exist, benchmarks get gamed, and I'm a solo builder grading my own homework. Those objections become controls. Before anything runs, the pre-registration freezes the tasks, the metrics, the false-accept budget, the refusal-rate guardrail, and the kill criterion, the point where I declare the idea didn't work. Then an unaffiliated person, named in the writeup, reruns everything from the frozen public release, and the help I'm allowed to give them is fixed in advance. I define the protocol and pay for their time; I do not perform or score their rerun, and their report goes public whether it confirms mine or contradicts it.
That's already how I work, on the record. The repo contains v0.2, which was pre-registered, failed its own criteria, and is published as exactly that: github.com/regsaddler/semantic-entropy.
The $5,000 minimum runs the smallest complete test. If the rule beats the baselines, you get a reproducible screening gate anyone can run. If it loses on enough data to count, you get a documented negative result and a baseline harness others can build on. If there isn't enough data to call it, the result is published as inconclusive. In every case, the grant buys a public, checkable result.
A note on process: AI systems drafted alongside me and were pointed at every section with instructions to break it. I verified the cited sources, constrained the claims to what the evidence supports, and I take responsibility for every claim. I'll defend any of them in person.
My goal is to test a working "no" for AI agents: a gate that keeps watching whether the checkers still disagree, and flags when their last labeled check is due again. I didn't get this from a book. It came from a rule I wrote for myself in July: treat cross-model agreement as a decaying asset, name the residual correlation, date the measurement, and re-estimate on a schedule, because the estimate goes stale. The contribution I am testing is the ledger itself: does adding it improve the gate's decisions? The $5,000 run tests whether tracking disagreement helps the gate. One frozen model-version swap on the same questions also gives re-estimating after a change its first test. The fuller arm, with task changes too, runs at the $12,000 tier. This project turns the working rule into a measured instrument.
It is deliberately narrow about what it can measure. A time-stamped ledger keeps 3 things separate: it monitors continuously, without labels, whether the checkers I rely on ever actually disagree; it estimates how much of their error is shared, but only inside labeled audit windows, where that estimate is valid; and it records the conditions under which each estimate holds, so it flags when a model or task change has made re-estimation due. Disagreement by itself won't tell you how much of their error is shared. You can watch for warning signs and refuse to trust an estimate past its expiry; that honesty is the point.
Why this counts as safety. Oversight schemes that pool model judgments inherit the correlation problem the studies below measured. The ledger turns that risk into something you recheck on a schedule, not something you assume. The gate itself, wired into an agent framework, is the $12K tier (details under funding); the $5K answers whether the gate has earned the right to be wired in at all.
How I get there is one experiment with four ship items:
1. The detector, already public. A zero-dependency proxy inspired by semantic entropy. Built, passing its self-tests: github.com/regsaddler/semantic-entropy.
2. The independence ledger, the contribution. It logs whether supposedly independent checkers ever disagree, and tracks that over time rather than certifying it once. Zero disagreement is a reason to check for shared errors, not evidence of correctness. I've used it privately while building it. Its logs can't produce labeled outcomes or false-positive rates on their own; the benchmark supplies them.
3. The boring bar. Four simple baselines (answer-string uniqueness, majority vote, confidence threshold, evidence-presence), frozen before the run, applied to my own detectors first. These are the floor any detector should clear, not the ceiling. Sophisticated methods exist, like conformal prediction and calibrated abstention, but a rule that can't beat the trivial baselines hasn't earned a comparison to the hard ones.
4. An outsider's rerun that I don't score, under the protocol in the summary. I pay for their time, the same whatever they find.
Where this sits in the field. The skepticism it rests on is published, not mine. Panickssery et al. found LLM judges favor their own generations (NeurIPS 2024). Kim et al. (ICML 2025) looked at more than 350 models: on one leaderboard dataset, when two models both erred they agreed 60% of the time, and the bigger, more accurate models had highly correlated errors, even across providers and architectures (arXiv:2506.07962). Kohli's 2026 preprint (arXiv:2605.29800) tested nine frontier judges from seven families and found the panel carried about two independent votes' worth of information; the best single judge matched or beat the whole panel. Stacking more judges isn't necessarily more checking.
Those studies measure dependence once, on labeled evaluation sets. None of them checks it again on a schedule, and that gap is what the ledger is built for: the minimum tests whether its disagreement signal helps the gate at all, plus one model-version change; the fuller arm, with task changes, runs at $12,000. Two recent papers put uncertainty signals into agent control (Knowlton et al., arXiv:2608.14707; Vijayvargiya and Lokesh, arXiv:2608.10430; both 2026), which tells me the gating idea is alive. In the versions reviewed here, neither one goes back to measure how much of its checkers' error is shared.
The detector rests on Farquhar et al. (Nature 2024), a published method for one class of error, confabulation driven by uncertainty over meanings. It doesn't verify truth, and it has documented limits I don't hide from: stable errors can pass, and the clustering can merge answers that should stay separate. Detectors like this exist. What I am testing is the license: the rule that decides when a refusal signal is allowed to stop an action, and whether that rule earns its keep. When a claim can't be formally proved, the referee can't be a proof; it has to be a calibrated signal you keep checking.
Distribution. Results, including a loss or a failed rerun, go out to my more than 440,000 followers on X (x.com/zaibatsu), an audience built on technical AI content. Reach is distribution, not adoption; I report what third parties do with it, not follower counts.
Without this grant: no labeled outcomes, no head-to-head, no outside rerun. The detector is already public. The $5K buys the test and a public result.
The funding is stage-gated, and the $5,000 minimum stands on its own. It is the experiment, complete, not a down payment on a bigger one. The higher tiers unlock only if it clears its own pre-registered bar. If more than $5,000 arrives before that result, the excess sits unspent until the gate decides; if the gate fails, I'll agree its disposition with the funders rather than spend it.
The $5,000 runs the four ship items end to end. At the minimum, this is an offline test of permission decisions, not deployment into an agent. The pre-registration names the claim class, what accepting it would authorize, and what counts as a false accept. Before anything runs, the pre-registration freezes: tasks, sample size, held-out split, label construction, model versions, calibration set, primary metric, false-accept budget, refusal-rate guardrail, effect-size bar with confidence intervals, and the kill criterion. It also fixes where tasks and labels come from. At the $5,000 tier, labels come from existing public datasets. If I have to build any myself, the adjudication rule is part of that freeze, and the rerunner audits a random sample of them. If the rerunner disagrees with my labels more often than a threshold set in advance, I drop the ones I built, say so, and rescore without them by the same frozen rules.
The comparison that can kill the ledger at the minimum is the gate with it against the detector alone, on the pre-registered primary metric, at the same false-accept budget and refusal-rate guardrail. The swap is reported separately and doesn't enter that comparison. I also report whether watching continuously beats one fixed estimate. One model-version swap on the same questions runs at the minimum; the fuller transition arm, with task changes too, is frozen now as well and runs at the $12,000 tier. Labels used to update the ledger never score that update; it is judged only on later, separate held-out decisions. The pre-registration fixes a minimum eligible-error count and a precision requirement for the false-accept estimate. If the data can't meet them, I report the test as inconclusive; it doesn't unlock the higher tiers. The refusal-rate guardrail is in the pre-registration because a gate that refuses everything fails. Then the outside rerun.
If the experiment clears its bar, the next $7,000 (the $12,000 tier) does 3 things: wires the gate into an agent framework as a working guard hook, hardens the harness so third parties can run it without me, and runs the fuller transition arm, with task changes. The pre-registration names the target framework, picked by a rule frozen in advance. The full $25,000 adds the adversarial round: a model actively gaming a known grader, run against the detector stack. That protocol is already implemented, with passing tests on two codebases I wrote separately, which is implementation diversity, not external replication. The tier pays for the campaign against live models; the build is done. Its pre-registration was written before any run, and its hash publishes with the Day 30 release, so you can check it came first.
Where the money goes, plainly. My time is $2,600 of the minimum, 65 hours at $40 an hour, and $13,600 of the goal, about 340 hours. The outside rerunner gets $1,000 at the minimum and $3,500 at the goal. Compute, hosting, and CI run $300, growing to $1,800. Documentation and packaging take $800, growing to $4,200. That's my time too, at the same rate and on top of my main hours: 20 more at the minimum, 45 at $12,000, and 105 at the goal. Most of it goes to the README and the frozen release, because the rerun lives or dies on a stranger following them alone. That leaves $300 in contingency at the minimum and $1,900 at the goal. At the $12,000 tier, the same lines are: my time $7,200, or 180 hours; the rerunner $1,500; compute, hosting, and CI $800; documentation and packaging $1,800; contingency $700. All three tiers sum exactly.
The clock is 90 days. By Day 30, the frozen release ships: detector, ledger, baselines harness, and the pre-registration, whose hash publishes before any benchmark run. By Day 60, the benchmark run is complete, and nothing about the scoring can change after the fact. That run includes the head-to-head against a copy of Farquhar et al.'s released code with two patches, both in the Day 30 release (device placement and float16, so it runs on my hardware), checked first against the results the paper reports. It uses short-answer questions and open 7B models anyone can download, run locally. By Day 90, the outside rerun happens and the writeup goes public, pass or fail. The rerun only counts if the artifacts match the published hashes, working from the named release and allowed support alone.
The released code is MIT-licensed. The rerunner is recruited by public call after the Day 30 release, under a selection rule fixed in the pre-registration: someone unaffiliated with me who has had no prior contact with this project and has a public record of evaluation or replication work. If the first call doesn't turn up a credible unaffiliated rerunner, that gets published too, and the money holds for a second call.
I'm Reg Saddler. Here's how the thirty years in the bio actually break down. I ran a small graphic design shop where the whole craft was a fingerprint that had to stay consistent: the right type, the right spacing between the sentences, even the weight of the paper you were going to print on, all of it a design language, and it had to be transportable and durable. Then two decades running IT for people like US West, where what I really sold was simpler than it sounds: when I walked out the door, their hardware and their network and their software all worked, and they knew it. The recent years ran through publishing and social media marketing, growing the technical audience I still write for every day.
When AI arrived I built a research assistant on top of GPT-4o, fed it papers off arXiv, and gave it a trust module I called Veritas. It lied to me anyway, cheerfully, like a kid caught out, while feeding on my own enthusiasm. There was real science buried under the flattery, and separating the wheat from the chaff became the entire problem. So I built the correction: a framework where a claim that can't show its evidence simply stops, unable to move forward. Measure twice, build once. It is the same rule I learned setting type, pointed at a harder material. I'm not a credentialed researcher. This exact problem, under different names, is what I have been solving my whole working life.
Track record you can check today: the public repository holds the pre-registration and verdict for v0.2, which failed and stayed published, a cross-family adversarial review transcript, and 45 passing self-tests. That demonstrates implementation and research discipline, not detector validity; the proposed benchmark and outside rerun exist to test validity. Beyond the public record, I work solo with a multi-instance AI development discipline, and the honest governance note is this: the framework reduces unreviewed model action; it doesn't make me stop being a single point of failure. I remain the key-person risk, which is one reason the funded outside rerun is a deliverable rather than an aspiration. The outside rerunner will be selected for demonstrated replication work and no prior contact with this project.
The most likely failure is the boring one: the gates don't beat the boring baselines. That outcome is priced in. The pre-registration makes it publishable by design, and the field still gets the reproducible baseline harness plus a negative result it can cite. The repo already shows one.
The ledger has its own kill criterion, stated up front and made specific: if adding the ledger signal doesn't improve the pre-registered primary metric over the detector alone, at the same false-accept budget and refusal-rate guardrail, that is a published negative result, and the headline shrinks to the detector against the baselines, with the ledger written up as an add-on that didn't help. The ledger has to earn the weight I am putting on it, on a number fixed before the run, or it doesn't carry it. If the minimum comes up empty, the one model-version swap still tested re-estimating after a change, once, on one model. It didn't test task changes. Those stay untested, not refuted.
The other limits, named plainly. The detector rests on a method with contested edges, which is why it is benchmarked head-to-head against the authors' released code rather than assumed. Having free local models from other AI families review my work is weaker than true independent replication. That's why shared error gets estimated per pair inside labeled windows, never assumed, and why the ledger dates each measurement instead of trusting it once. I am a single operator with no institutional adopters yet, which is why the $5K scope is sized to what one person plus one funded outside rerunner can demonstrably complete. And if the rerunner goes silent or returns an ambiguous result, the ambiguity publishes as ambiguity.
This experiment tests the released proxy and the ledger method, not the private framework they grew out of. A win transfers to the public artifacts only. The line between those released artifacts and the private framework was reviewed with pro bono IP counsel in September, and further releases follow that review.
If the money arrives and the work still fails, the failure publishes with the same receipts as a win: prereg hash, artifacts, and the rerunner's report; the kill criterion is the point of the design, not decoration.
Nothing. No other funding supports this work. For completeness: I have separate applications live for adjacent work (a submitted multi-agent research proposal, a submitted evaluation-research expression of interest, and a pending request for $1,000 of API credits that, if granted, would cover some model calls here). None of that money is in this budget, and nothing has come in from any of them yet. If any of them lands, I'll say so and account for it separately.