You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I built this because coding agents were sometimes giving me summaries I could not fully trust. The code was often fine, but the report could contain an old number, a skipped test, or a claim that no longer matched the run.
So I added a small mechanical layer. A marked claim must trace back to a recorded run. The check returns match, mismatch or unresolved, and unresolved fails.
The public version is here:
https://github.com/hilberspace-dev/evidence-admission
It uses two Git hooks and a small resolver. It has no dependencies. The public extract has only been tested inside the project so far.
I want to test whether this layer adds anything over a second AI review or the same rules written as instructions.
I will compare three setups: second AI review, prose rules only, and the mechanical gate plus the same second review.
The tasks will include five planted problems: stale citation, number drift, missing artefact, edit after review and ambiguous acceptance.
The main measure is how many unsupported claims reach the final report without being flagged. I will also track false blocks and human checking time.
The study will be preregistered. An independent person will prepare the held-out tasks and answer key. I will use at least two model families and publish the tasks, transcripts and results, including a null result.
I will also try to break the gate myself.
I am asking for $20,000.
12,000 for my time over six months
4,000 for API costs
2,500 for an independent evaluator
1,500 for taxes, transfer costs and contingency
The minimum useful amount is 5,000 dollars.
I am Serhat Atılgan. I am in my final year of physiotherapy in Kayseri, Türkiye, and I also work independently as a software developer.
Since September 2026 I have used coding agents in my own computational research workflow with preregistration, run logs and separate review sessions. Seven preregistered experiments were run that way in September.
I also keep failures on record. One guard silently passed through a folder alias; I later found the same issue in more places and added regression tests.
Outside this project, I reproduced a deterministic defect in ERC-4337 EntryPoint v0.8 in two separate environments, including a negative control.
I design the system and decide what should be tested. AI agents write much of the code.
The most likely failure is a null result: the gate may catch nothing that a second AI review would miss.
If that happens, I will publish it. It would mean the simpler approach is enough for this class of problem.
Other risks are that the planted defects are too artificial, that unmarked numbers escape the resolver, or that Git hooks are bypassed. These are part of the evaluation.
Nothing so far.
I also applied to the Digital Science Catalyst Grant, the GTR AI Safety Fund and a BlueDot Rapid Grant on 4 October 2026. I am waiting for the results.
If more than one application is funded, I will not charge the same costs twice.
There are no bids on this project.