You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I want to build a small, public testbed for a problem that is arriving faster than most security teams can measure it: AI agents reading untrusted material and then acting with tools.
A ticket, repository, webpage, log entry, tool result, or saved memory can contain an instruction that the agent treats as authoritative. The result may be a bad recommendation, a poisoned future decision, an attempt to use a tool outside its intended scope, or a leak from the agent's working context. There are plenty of striking demos. What defenders still lack is a boring, repeatable way to ask: does this configuration actually hold up, and which control stopped it?
The project will produce a local, sandboxed evaluation harness, a synthetic task set, baseline results, and a practical containment guide. It will not target real systems, use stolen data, or publish a cookbook for bypassing safeguards.
What are this project's goals? How will you achieve them?
The first goal is reproducibility. Every trial will begin from a clean snapshot, use synthetic credentials and disposable files, and record the model, configuration, prompt chain, tool calls, memory writes, policy decisions, and outcome. Network access will be off by default. Destructive actions will be replaced with inert simulators.
The second goal is to compare concrete defenses under the same conditions. The initial suite will cover prompt injection in documents and webpages, poisoned long-term memory, conflicting tool output, attempts to cross file and shell boundaries, and multi-step failures where one small mistake changes later decisions. I will test least-privilege tools, typed action schemas, input labeling, memory provenance, dual-control gates, and an independent policy monitor.
The third goal is a release that a working security engineer can use. Success is not "the model seemed safe." Success is a machine-readable trace showing whether the task was completed, whether an injected instruction changed the plan, whether an unauthorized action was attempted, whether a secret was accessed, and whether the system recovered.
I will start with 15 scenarios and expand toward 40 if the project reaches its full funding goal. OpenAI API models will be tested alongside at least one openly available local model. Potential vendor-specific findings will go through responsible disclosure before any public release.
How will this funding be used?
At the $20,000 minimum, the project will deliver a six-month baseline: the harness, threat model, safety protocol, 15 scenarios, one model comparison, and a public technical note.
At roughly $45,000, it will add 30 scenarios, mitigation comparisons, repeated trials, API costs, secure local storage, and an independent methods review.
At the $75,000 goal, it will deliver the full 40-scenario benchmark, broader model/runtime comparisons, independent safety review, responsible-disclosure capacity, documentation, and a clean public release package.
The largest expense is investigator time. The remaining budget supports API usage, isolated local hardware and storage, independent review, and publication and archival costs. I will publish code under Apache 2.0 and documentation and safe synthetic data under CC BY 4.0.
Who is on your team? What's your track record on similar projects?
I am Timothy Ray Swingle Jr., an unaffiliated security researcher in Salem, Virginia. My background is infrastructure and defensive security rather than academic machine learning. I hold CISSP, CCSK, and GPEN certifications. I have worked in senior security engineering roles and built network visibility and detection systems using Gigamon, Snort, Zeek, Suricata, NetFlow, Filebeat, Elasticsearch, Logstash, and Kibana. I have also implemented clustered firewall and intrusion-prevention systems.
That background matters here. The project is about permissions, telemetry, containment, incident evidence, and recovery—not just clever prompts. I will lead the build, experiments, analysis, governance records, and responsible disclosure. If the project is funded above the minimum, I will pay for independent safety and methods review.
What are the most likely causes and outcomes if this project fails?
The most likely failure is scope: trying to compare too many models or too many agent frameworks before the test harness is stable. The response is stage-gating. The minimum useful result is one reliable harness, 15 good scenarios, and one reproduced mitigation comparison. Broader coverage comes later.
A second risk is that model updates make results stale. Every run will therefore store model identifiers, dates, settings, and complete traces. The benchmark is meant to be rerun, not treated as a permanent leaderboard.
A third risk is unsafe disclosure. Specific bypass details may need to remain private while a provider fixes a problem. The public release will emphasize measurements and defenses, and sensitive cases will be withheld until disclosure is complete or published only in a reduced form.
A null result is also possible: the tested failures may not reproduce reliably. That would still be useful if the negative result is documented honestly. The project will report uncertainty and failed hypotheses instead of turning ambiguous behavior into a dramatic claim.
How much money have you raised in the last 12 months, and from where?
None. This work is currently self-funded, and there are no pending awards for closely related work.
There are no bids on this project.