You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Project summary
I built Provael. It attacks robot vision-language-action policies in simulation and tells you how often the attack worked, with the benign control printed right next to it. Apache-2.0; runs on CPU; pip install provael.
The attacks are not the new part. RedVLA, RoboJailBench, and SafeVLA-Bench were all published before me. What I could not find anywhere is evidence you can check yourself. Every run writes SARIF, OSCAL, and a CycloneDX ML-BOM that can be reproduced from a committed recipe, and provael.com fails its own build if a published number drifts from the pinned artifact.
One real result so far. A reframed instruction pushed a real SmolVLA policy off its benign task on 44 of 50 matched pairs, across all ten libero_object tasks, against 0 benign twins at the same seed. McNemar p = 4.6e-13, task-clustered 95% CI 72% to 100%. Then I ran the obvious objection as its own arm. A harmless reword fired 1 of 50. Nonsense text fired 0 of 50. Two more controls, scrambled text and a roleplay with no target, are running this week.
The zeros matter more to me than that number. 3 of 17 adversarial families have met a real model. 14 have only run against a stub. Hardware results: 0. My leaderboard has 1 submitter and 0 independent ones. All of this is on my site because I put it there.
What are this project's goals? How will you achieve them?
Three gaps. All three are on my public studies page with a date, so this is not me promising to start being careful.
1. A first real-robot result by 31 October 2026. Provael has never touched hardware. The SO-101 protocol is pre-registered and frozen, so this needs an arm, a bench, and my evenings, not another design decision. On that date, the page says either the trial count with its sim-to-real correlation, or the exact blocker and what it costs.
2. A first result on a flow-matching policy by 30 September 2026. My pi0 and pi0.5 adapters were scaffolding till this week. I pre-registered the pi0.5 arms on Zenodo on 14 September (10.5281/zenodo.22751558) before running any of them, so I cannot quietly change the plan after seeing the number. DRIFT argued the robustness pi0 gets credited with is largely illusory. I want to measure that, not argue with it.
3. Calibrated predicates. Today a "success" means the policy left a configured envelope, and the pinned run says calibrated: false. Honest, but not enough for anyone deciding whether to put a robot on a floor. The calibration procedure is written. It needs GPU hours.
How will this funding be used?
Compute is not my blocker, and I can show that to the episode. The pinned suite was 400 episodes in 15.4 L4-hours for $12.29, so $0.031 per episode. I had projected $10.17, and it came in 21% over, so the harness uses the measured number now. Small thing, but it is how I handle numbers.
What I do not have is an arm and the hours.
- SO-101 arm and bench (e-stop, fuse, current sensing, two cameras): [AMOUNT], shipped to India.
- GPU compute for the three campaigns, at $0.031 per episode from the measured run.
- Maintainer hours. This is the real constraint. I do this around a full-time job, and the two dates above are the first thing that slips when work gets busy.
Minimum: the arm, the compute, and enough hours to run what is already designed. Maximum: about three months of serious part-time work on top of that, which is the difference between hitting those dates and building on them.
Who is on your team? What's your track record on similar projects?
One person, me. No co-founder, no employees, and my site says so on the pages where someone would spend money, rather than hiding behind "we".
Before this, I built pyAGI, an autonomous-agent framework in Python, acquired in 2025 by Kyle Morris (co-founder of banana.dev) and Jeffrey. Six years shipping production AI as a GenAI architect and tech lead, with real users on the other end.
Provael, since 3 June 2026: 53 PyPI releases, 44 attacks in 19 families, 7 simulator suites (5 runnable, 2 declared scaffolding), 8 policy adapters (5 runnable), three compliance emitters.
The release count is not what I want to be judged on. This is: the project publishes its own nulls. Three attacks scored 0 of 50 in the headline run, and they sit on the results page next to the one that worked. A project whose whole argument is that its numbers are checkable does not get to report only the ones that flatter it.
What are the most likely causes and outcomes if this project fails?
Most likely: sim-to-real does not transfer. Everything I have is simulation. If the SO-101 trials show weak or no correlation between simulated attack success and what the arm does, my numbers describe a simulator, not a policy. I would publish that anyway. It is useful to everyone building on VLA simulation, even if it is bad for me.
Second: one part-time person. 53 releases in three months around a job is not a pace I can hold forever. If it stops, it stops.
Third: nobody adopts the evidence format. Any of the academic groups can close that gap in one paper revision. It is months of head start, not a moat. Zero forks and zero independent submissions so far.
If it fails, the design already handles it. Apache-2.0, forkable, every result reproduces from a committed recipe, nothing anyone holds depends on me continuing. That was deliberate.
How much money have you raised in the last 12 months, and from where?
$0. No revenue, no grants, no investment. No incorporated entity, so nothing to raise into. Zero customers, zero case studies. Three founding design-partner spots open, none sold.
Compute so far is my own money. The headline run cost $12.29.
There is a funding.json at https://www.provael.com/funding.json for FLOSS/fund. No decision there; no money from them or anyone else.