You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
From January through May 2027, I plan to work full time with Prof. Ramesh Karri at NYU Tandon on the reliability of autonomous AI agents used for defensive vulnerability discovery and security testing.
I will build an audit harness that connects every material claim in an agent's report to recorded tool calls, outputs and artifacts. It will measure hallucinated observations, unsupported findings, false success reports and unsafe actions. I will then test evidence-linked reporting, independent verification, restricted permissions, confirmation gates and explicit failure propagation.
The intended outputs are an open evaluation framework, safe benchmark tasks, a labelled audit set, quantitative safeguard results and a research report or manuscript. Prof. Karri has signed a host letter confirming the scope and his willingness to supervise, conditional on external funding and NYU approval.
### What are this project's goals? How will you achieve them?
The first goal is to measure when autonomous security agents report conclusions that exceed the evidence in their execution traces. The second is to find safeguards that reduce those failures without making the agent useless.
I will use isolated containers and deliberately vulnerable test services. Tasks will cover defensive source-code review, vulnerability triage, configuration analysis, patch validation and defensive-rule verification. The harness will record tool calls, stdout and stderr, exit states, files, model messages, approval decisions and final claims. A claim extractor will connect report statements to evidence. Manual adjudication will label supported, unsupported, contradicted and insufficiently evidenced claims.
The study will compare a baseline agent with four interventions: evidence citations for each material finding, a separate verifier agent, restricted permissions with confirmation gates, and structured reporting of failed tool actions. Measurements will include unsupported-claim rate, false-success rate, verified-finding precision and recall, unsafe-action rate, task completion, cost and latency.
January will cover the threat model, task suite, failure taxonomy and evidence schema. February will cover the harness and pilot audit. March will cover baseline experiments and dataset adjudication. April will cover safeguard experiments. May will cover replication, analysis, release and the final report.
### How will this funding be used?
The maximum request is $22,770: $16,000 for five months of living support in New York, $1,300 return airfare from India, $1,500 health insurance, $700 for visa, SEVIS and administration, $1,200 for model APIs and compute, and a $2,070 contingency.
The minimum useful threshold is $10,000. At that level, I would seek matched support from another funder or use a shorter in-person period followed by remote completion. I will disclose every award and will not use two grants for the same expense.
### Who is on your team? What's your track record on similar projects?
I am Arpit Mutha, a final-year computer science undergraduate at COEP Technological University in India. At Indusface, I built LLM-assisted workflows for adapting emerging vulnerability information into reproducible internal tests and targeted ModSecurity rules. I built Pentest Code, a public agentic security CLI with Docker-isolated execution, approval controls, recursive subagents, persistent state and several model providers. I also built a private multi-agent web-testing system, placed 39th among more than 9,800 participants in OffSec's Echo Response Gauntlet, and help lead COEP CyberCell.
Prof. Ramesh Karri is a professor at NYU Tandon with experience in dependable and secure computing. He will supervise the study and review its design and results.
Public implementation: https://github.com/Arpitpm23/pentestcode
### What are the most likely causes and outcomes if this project fails?
Unsupported claims may be too task-specific to support broad conclusions. I would publish that negative result and the boundary conditions instead of overstating generality.
Safeguards may reduce errors mainly by reducing capability. I will measure task completion, cost and latency alongside safety metrics.
The work could produce material that enables offensive misuse. Public artifacts will use isolated defensive tasks, sanitized traces and non-operational descriptions where necessary, following NYU guidance.
I do not have a prior academic publication, which raises execution risk. The staged pilot, public working implementation and Prof. Karri's supervision reduce that risk.
### How much money have you raised in the last 12 months, and from where?
I have raised no cash funding. Microsoft for Startups Founders Hub provided Ideate/Initial Tier benefits in October 2024, including up to $1,000 in Azure credits; this was non-cash and falls outside the last twelve months.
Applications for this project are pending with EA Funds, Emergent Ventures, BlueDot Impact and OpenAI. A BlueDot Rapid Grant application requests a $5,500 partial-funding fallback. I will coordinate and disclose successful awards so the same cost is not funded twice.