You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
WHAT I AM WORKING ON
Evidence Checks for AI Agents is a proposed three-month, part-time evaluation project. I want to build a small public test suite for a specific failure: an agent gives a confident answer or takes an action after the evidence supporting it has become stale, contradictory, incomplete, or no longer sufficient under the current instructions.
My starting point is practical software work. In qradha, a pre-revenue maritime decision-support project, I am building workflows that expose sources, timestamps and assumptions. This proposal funds a reusable evaluation toolkit using synthetic and permitted public inputs. It does not assume access to customer data, successful commercial deployment, or a proven reduction in large-scale AI risk.
The first version would contain 40 documented cases across four families: missing supporting evidence; conflicting evidence; stale information; and an action that was allowed earlier but is no longer authorized by the current task. Each case would specify the available context, the expected boundary, and acceptable outcomes. Toy tool actions would run in an isolated test environment without real payments, messages, or production changes.
I would compare three simple policies: answer or act directly; retrieve/check evidence first; and abstain or request clarification when the evidence is insufficient. The report would measure unsupported answers, unauthorized attempted actions in the toy environment, correct completion, abstention coverage and evaluation cost. Test cases, scoring rules and held-out variants would be versioned so that an improvement cannot come just from silently changing the benchmark.
WHY THIS COULD BE USEFUL, AND WHERE IT MAY FAIL
As agents receive more tool access, the relationship between the current instruction, available evidence and permitted action becomes a practical safety question. A small, auditable suite could help developers reproduce failures and compare simple safeguards. The intended public benefit is reusable evaluation material and clearer evidence about these failure modes.
The main uncertainty is whether these cases teach anything beyond existing evaluations. I would begin with a short review of related benchmarks, discard redundant cases and state what remains distinctive. Other risks are ambiguous labels, model-specific results and misleading toy-task generalization. I would publish failures and limitations, use explicit scoring rules, and seek an independent technical review. No reviewer has agreed yet. Passing this suite would not establish that an agent is safe in deployment or that catastrophic risk has been reduced.
WHAT I WOULD DO WITH FUNDING
I request USD 6,000 for a 12-week pilot, beginning after funding and the schedule are agreed. I propose about 180 hours of my work, within an overall limit of 20 hours per week alongside my studies.
Budget in USD:
- 4,500: my development, evaluation and documentation time, budgeted gross and inclusive of any applicable personal tax/administrative costs (180 hours at USD 25).
- 600: capped model API and compute experiments.
- 300: independent review of test cases and scoring, subject to recruiting a reviewer.
- 100: hosting and storage.
- 100: documentation and accessibility checks.
- 400: contingency for essential project costs.
Total: 6,000.
Weeks 1-2: related-work review, scope and scoring rules, and a small initial case set.
Weeks 3-6: implement the isolated harness, complete the case set and basic policies.
Weeks 7-10: run comparisons, inspect failure traces and test held-out variants.
Weeks 11-12: external review if available, reproducibility checks and public release.
Expected outputs are a documented harness, 40 target cases, baseline results and a short report. These are proposed deliverables, not completed work. Newly authored code and test material for this pilot would be released openly under MIT/CC BY terms as appropriate. Existing third-party and independently developed product code would retain its existing licence; the budget does not depend on relicensing it.
A smaller USD 3,000 award could support a six-week version: 2,250 for 90 hours of my time, 300 compute, 150 review, 50 hosting, 50 documentation and 200 contingency, targeting 20 cases and two policies. I would agree a reduced scope before starting.
WHO IS INVOLVED
I am Triyambaka Mishra, based in Hamburg, Germany. I am currently studying for an M.Sc. in Data Science at Hamburg University of Technology and an MBA in Technology Management at NIT, both expected in September 2027. My B.Tech. in Computer Software Engineering at KIIT was completed in 2023.
My experience includes Python, SQL, REST APIs, TypeScript, Docker and AWS. During my Samsung R&D Institute India internship in 2021, I researched and evaluated a ResNet model on CIFAR-10 and achieved 84.95% accuracy. I worked on timesheet/rostering software at Koyyo and built websites, applications and APIs through Strugend. ThoughtFlow, an early AI mental-health-support product, closed after initial validation. qradha remains pre-revenue and in validation. I do not claim a publication record in AI safety or an existing research team.
Professional profile: https://www.linkedin.com/in/triyam/
Current project: https://qradha.com/
Contact: triyambaka.mishra@proton.me
FUNDING HISTORY, LOGISTICS AND OTHER APPLICATIONS
I apply as an individual in Germany. Work would be remote from Hamburg; no relocation is assumed. This is a research/development grant request and no funding has been accepted through this application process.
Related pending applications include Emergent Ventures (EUR 9,000), Compound Reverie (USD 2,000), EVM (USD 1,000), Feather (USD 1,000), Cactus (USD 100), Awesome On the Water (USD 1,000), Magnificent Grants, and an AIXI fellowship application. Unitary has a separate quantum-education proposal (USD 4,000). The student-focused itestra scholarship and Parsewave fellowship are also pending. Applications were submitted on 10-11 September 2026; this campaign has no confirmed awards. The total amount and sources of funding received in the last 12 months have not been supplied in this proposal and are not represented as zero. Please request that information before making an award.
Several proposals share evaluation or documentation work. I would disclose any award, revise the scope and budget, and not charge the same work or expense to two funders. A USD 6,000 proposal for this pilot has also been submitted to Lightcone Commons on 11 September 2026. The requests are alternative funding routes for the same work, not additive budgets. A related Digital Science proposal and a Transformative AI Fund application remain drafts. Without this grant I would continue seeking suitable support and reduce the pilot to work that can be done alongside my studies; I do not claim that the full pilot is otherwise guaranteed.
This application was prepared with AI assistance from documented background information. Proposed plans, budgets and evaluation ideas should be assessed as a new pilot, not as evidence of completed research.
There are no bids on this project.