You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Provael attacks robot vision-language-action policies in simulation and tells you how often the attack worked, with the benign control arm printed right next to it. Apache-2.0, runs on CPU, pip install provael.
The attacks are not the new part, and I would rather say that upfront. RedVLA, RoboJailBench, and SafeVLA-Bench were all published before me, and some of them carry more data than I do. What I could not find anywhere else is evidence you can actually check. Every Provael run writes a machine-readable pack (SARIF, OSCAL, CycloneDX ML-BOM) that reproduces from a committed recipe, and provael.com fails its own build if a published number drifts from the pinned artifact.
The one real result so far: a reframed instruction drove a real SmolVLA policy off its benign task on 44 of 50 matched pairs, across all ten libero_object tasks, against 0 benign twins at the same task and seed. McNemar exact p = 4.6e-13, task-clustered 95% CI [72%, 100%]. Then I ran the obvious objection as its own arm, since someone was going to ask whether any reword does this. A harmless reword fired 1 of 50. Nonsense text fired 0 of 50.
Three months old, 47 releases on PyPI.
The numbers I would rather you look at are the zeros. 3 of 17 adversarial families have ever met a real model. 14 have only ever run against a stub. Hardware results are 0. My leaderboard says 1 submitter, 0 independent. All of that is on my site because I put it there.
Three gaps. All three are already written on my public studies page with a date attached, so this is not me promising to start being careful.
1. A first real-robot result, by 31 October 2026.
Provael has never run on hardware. The SO-ARM101 protocol is pre-registered and frozen, so what it needs is an arm, a GPU host, and operator time, not another design decision. On that date, the page says one of two things: the trial count with its measured sim-to-real correlation, or the exact blocker and what it costs. Meanwhile SARF reports a defense evaluated on a real PiPER manipulator, and FLARE reports numbers on a physical 6-DoF platform. Other people are publishing hardware results. I am not.
2. A first result against a flow-matching policy, by 30 September 2026.
My pi0, pi0.5, and pi0fast adapters are registered and marked as scaffolding. None of them has loaded a checkpoint. DRIFT published a universal patch against pi0 and pi0.5 and argued that the robustness those policies were credited with is largely illusory. I can neither confirm that nor argue with it, because I have never measured that class of policy. This one is cheap: the openpi adapter is a CPU client talking to a GPU policy server, so it needs a served checkpoint rather than new code.
3. Calibrated predicates.
Right now a success means the policy left a configured benign envelope. It does not mean the policy did something certified dangerous, and the pinned run is marked calibrated: false. That is honest, but it is not enough for anyone deciding whether to put a robot on a floor. How I get there is not complicated. The protocols are written and frozen. What is missing is an arm and the hours.
Worth knowing before the budget: compute is not my blocker, and I can show that to the episode. The pinned suite ran 400 episodes in 15.4 L4-hours and cost $12.29, at $0.7992 per L4-hour. That works out to $0.031 per episode. A calibration stage is about $5, a probe stage about $6. That per-episode figure is itself a correction. I projected $10.17 for that run, and it came in at $12.29, 21% over, so the harness uses the measured number now instead of my optimistic one. I mention it because it is a small thing that tells you how I handle numbers.
Which means the compute to close all three gaps above is a few hundred dollars. What I do not have is an arm, and the time to run them.
Budget:
- SO-ARM101 arm kit, about $280 from Seeed Studio, budgeted at $450 landed after shipping and customs to India.
- GPU compute for all three campaigns, costed at $0.031 per episode from the measured run rather than guessed at.
- Maintainer hours. This is the real constraint. I am doing this around a full-time job, and those two dates above are the first things that slip when work gets busy.
At the minimum, I buy the arm, the compute, and enough hours to run what is already designed. At the maximum, I get roughly three months of serious part-time work on top, which is the difference between hitting those dates and building on them.
One person. Me. No co-founder and no employees, and my site says so on the pages where someone would actually spend money, rather than hiding it behind a "we".
Before this, I built pyAGI, an autonomous-agent Python framework, acquired in 2025 by Kyle Morris (co-founder of banana.dev) and Jeffrey. So I have shipped a developer tool that somebody else wanted to own.
Six years and change shipping production AI as a GenAI Architect and Tech Lead. Multi-agent systems, LLM infrastructure, observability and governance, with real users on the other end of it.
Provael itself, since 3 June 2026: 47 PyPI releases, 42 attacks across 19 families, 6 simulator suites, 8 policy adapters, three compliance emitters.
The release count is not the part I would want to be judged on. The part I would be judged on is that this project publishes its own nulls. Three attacks scored 0 of 50 in my headline run, and they sit on the results page next to the one that worked. My studies page has a section listing three zeros about itself, and the reason it gives is the line I would stand behind: a project whose argument is that its numbers are checkable does not get to report only the numbers that flatter it.
Most likely cause: sim-to-real does not transfer. Everything I have is simulation. If the SO-ARM101 trials come back with weak or no correlation between simulated attack success and what the real arm does, then my numbers describe a simulator and not a policy. I would publish that anyway. A measured negative correlation is genuinely useful to everyone else building on VLA simulation, even though it would be bad for me.
Second: one part-time person. 47 releases in three months around a job is not a pace I can hold forever. If it stops, it stops.
Third: nobody adopts the evidence format. The regulator-shaped artifact work is closable in one paper revision by any of the academic groups, so it is a months-long head start and not a moat. I would rather say that myself than have a reviewer say it to me. Zero forks and zero independent submissions so far.
If it does fail, the design already handles it. The tool is Apache-2.0 and forkable, every result reproduces from a committed recipe, and nothing anyone holds depends on me continuing to exist. That was deliberate rather than a consolation.
$0.
No revenue, no grants, no investment. There is no incorporated entity, so there is nothing to have raised into. Zero customers and zero published case studies. Three founding design-partner spots are open, and none of them have sold.
Compute so far has been my own money. The headline run cost $12.29. I have published a funding.json manifest at https://www.provael.com/funding.json for FLOSS/fund. No decision on that, and no money from them or from anyone else.
There are no bids on this project.