I will start with the project summary, then. I have been using agents in my research for a long time, and I wanted to assess and verify which sources they used, what inferences they drew, and whether the information I would be using in the experiment was accurate. The biggest problem was that some of the things the agent said could be incorrect, and these errors could accumulate to cause a massive problem later on; to solve this, I developed a monitoring system. Therefore, rather than testing the accuracy of the incoming information semantically, I needed to test it numerically, and I developed the ‘match’, ‘mismatch’ and ‘unresolved’ system. To give an example : for instance, if a measurement of 109.56 ms was recorded in an experiment and this matches the recorded source exactly, then this is a ‘match’; if it does not match, it is a ‘mismatch’; and if it cannot be definitively verified whether it matches or not, it is marked as ‘unresolved’ . In this way, I am able to measure source-experiment discrepancies in numerical experiments; there is also a ‘guar’ to prevent ‘push’. My main aim is to record the agent’s claims in a documented evidence file for verification and to verify this numerically and mechanically as well.
https://github.com/hilberspace-dev/evidence-admission
What I really want to investigate is whether mechanical checking actually works better than a second agent review. In other words, does it catch more errors and provide better checking? I will test this by comparing three approaches: ‘second agent review only’, ‘rules as prose only’, and ‘mechanical gate combined with the same second agent review only’. The aspects I will be looking at in the results are: how many incorrect or unsupported claims make it into the final report? I will also measure false positives and the time taken for human review. I will publish the results once I have obtained them.
How will this funding be used?
I am looking for a total of 20000 dollars. The breakdown of the funds is as follows: my own working time (6 months) approximately 12,000 dollars, API costs approximately 4,000 dollars (running at least two agent types), an independent evaluator approximately 2,500 dollars, and taxes approximately 1,500 dollars. The minimum meaningful budget for me is 5000 dollars.
The team consists of just one person – it’s just me. My name is Serhat; I live in Turkey, I’m a final-year undergraduate student, and I’ve been carrying out independent research for about six months. I design and manage the system, have agents carry out the implementation, and monitor the results. I’ve also previously worked on various security vulnerability studies.
The reviewing party is usually associated with a model agent, typically Claude code. For this reason, it still relies on human oversight, and git hooks can also be bypassed. In addition, if the resolver is not marked as a number, it may not displa
I haven’t received any money; I’ve simply applied to funds such as Digital Science Catalyst