You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I am building Ruhusa (https://github.com/Claire56/ruhusa) , an open source authorization framework for AI Agents and multi- agent systems.
In situations where a model makes a bad decision, follows malicious instructions or even replans after denial , Ruhusa can stop the action from turning into a real world action
I have four main goals.
To build a public benchmark for AI-agent authorization attacks, for example replanning after denial, delegation-based privilege escalation, replay, revoked authority, cross-task permission reuse, and tool substitution.
I want to run controlled experiments with model-driven agents and measure how often the agents attempt unauthorized actions and how often those attempts succeed.
Compare systems that rely mainly on model instructions with systems that enforce authorization independently at the tool-execution boundary.
I want an opportunity to , publish the benchmark, evaluation code, attack scenarios, results, and documented failure modes so other researchers can reproduce and extend the work.
I’ll use Ruhusa as the authorization layer and run repeated experiments across different prompts, task variations, and agent behaviors.
Model/API evaluation runs
Cloud execution environments
Benchmark packaging and reproducibility infrastructure
Tracing, storage, and evaluation tooling
CI and automated regression testing
Documentation and publishing infrastructure
I am currently the primary researcher and maintainer of Ruhusa.
I’m a PhD student in Artificial Intelligence studying authorization and security in multi-agent AI systems. I also have professional experience building production agentic AI systems and AI-platform infrastructure.
Ruhusa is already publicly available on GitHub and PyPI. I have built the authorization framework, threat model, attack benchmarks, and automated security tests covering delegation, replanning, revocation, replay, tool identity, confused-deputy risks, and execution-time authorization.
This funding would support the experimental phase rather than starting a new project from scratch.
One possibility is that external authorization works well against simple attacks but provides less protection in more complex agent workflows.
Another risk is that the benchmark scenarios are too synthetic and do not represent realistic agent failures.
There is also a usability tradeoff, an authorization system that blocks too much could look safe while making agents ineffective.
I’ll address these risks by testing both allowed and denied actions, measuring false denials as well as successful attack prevention, using multiple attack families, and publishing limitations alongside the results.
Even a negative result would be useful. If certain authorization approaches fail, that would help identify where model-level or other system-level safeguards are still needed.
As of now i haven't raised any money yet.
There are no bids on this project.