You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
This is the security arm of my company Small Mind. I red-team agentic AI products launching all over the place that you put data into and they act on it. I also test whether the benchmarks we use to call these products secure actually catch the real attacks. I find the holes in real agent products through prompt injection. Then I prove that the leaderboards vendors point to when they say "enterprise-grade security" are actually evaling the wrong thing.
Two parts. I'm working on a red-team harness aimed at direct and indirect prompt injection in agentic systems mostly from the confused deputy class. This is where an agent with tool access and standing permissions reads attacker controlled text from emails, document uploads, pictures, etc and then acts on it. This is the hole in every "AI agent" product shipping right now, most of them don't have an answer for it. They just take it at face value the models work and the benchmark stamp holds up. I run the attacks against open weight models locally and against commercial agent APIs that are deployed. Record it all. Then I publish the findings. That way providers can patch the findings while independents can expand on the work. This is already what I'm doing by hand at volume and I want to extend and expand the eval integrity work I've already shipped. I took the JailbreakBench leaderboard and every attack they run and changed only the judge model they use to score them. Four of five alternate judges reordered the leaderboard. Minimum Kendall tau was 0.115. The ranking is almost a coin flip depending on which judge you pick. A parser defect alone in a reference judge flipped 534 of 1637 rows on one model. The whole paper was pre-registered and committed at github.com/Threadborne/eval-sensitivity. The funding lets me run the 70B reference judge locally which is the pushback against this kind of work. The claim is that because I used "smaller" models as judges, the work isn't "apples to apples". Running the 70B models would put it to rest. I achieve these 2 goals with unlimited local runs on owned hardware, rented burst compute for the fontier scale open models and metered API spend to test the deployed commercial agents. I will continue to use my pre-registered methods on the eval side so the results can't just be dismissed.
Down to the penny. Hardware and compute only:
Local runs:
2x Lenovo ThinkStation PGX workstations (NVIDIA GB10 Grace Blackwell, 128gb unified memory, 4tb Gen5 self-encrypting NVME, DGX OS) @ $7,989.00 each = $15,978.00 total
UPS + metered PDU (multi day evals will survive brownouts) = $329.98
Rented compute & model access:
GPU burst for frontier open weight models (vast/runpod) = $2,000
Commercial agent API credits (the deployed agents I'm testing) = $3,000
Total: $21,307.98
The 2 ThinkStations are the backbone. I need thousands of adversarial generations per attack class they have to be uncensored by provider side guardrails and not metered. This is in the millions of tokens. The self encrypting disks are also a must, I'm running client adjacent data and these encrypt at rest. The rented side covers the models too big to run locally, period. And the commercial agents that are only available behind API.
Just me. Small Mind LLC
I was raised in forums and the piracy scene moving into ethical hacking, red teaming and pen testing. I've had a computer since 2nd grade and I've been online through every platform generation watching how they've been broken and fixed since yo-yos were hype. I red team models on Gray Swan Arena, HTB and put 6 months into OSCP. Before any of the AI work I was in the 101st Airborne, company RTO and heavy weapons team leader. Deployed in OEF 10 and held secret clearances.
Most likely causes of it failing are the vendors just straight up don't want to see the results. This is all still new and even though everyone talks about security with AI its usually about them breaking out of sandboxes. Every day I see companies and vendors just straight up ignoring security where it involves people's sensitive data. And the second part of that is the eval-sensitivity results being dismissed. For the same reason. All the benchmarks are using AI to grade and judge AI. Even preregistered they want to nitpick results because theyre so heavily invested in the leaderboards.
Even if the project fails it doesn't fail hard. The results will be published, the hardware this grant would provide doesn't get shut off with a pat and a "we tried". I keep going and keep publishing and creating open tooling. If they want to keep arguing the eval work and dismissing it. Thats the point. The conversation drives the change. The worst case is the money provides local research and a mountain of working attacks anyone can run themselves, out in the open. There is no version of this where the computer sits idle. Breaking things and pushing the security industry is what I do. It doesn't matter if anyone pays me for it or not.
Nothing yet. Entirely out of my own pocket. I have one other fundraiser live on Manifund right now for the research arm of Small Mind. "Persistent maybe conscious AI and whether it chooses for itself" ($15k goal, closes in mid nov) which as of writing this is unfunded. This would be the first outside money my work will have received.