You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
AI agents now read email, browse the web, edit files and call APIs, more and more through MCP servers that anyone can publish. One instruction hidden in a web page, a document or a tool description can make an agent leak data or do something the user never asked for. Most protection today lives inside the model. MCP Guard is my proposal for an independent control point outside the model: a small open-source gateway that would run between an agent and its MCP servers and decide, call by call, what is allowed.
Goal: to build a gateway any team can put in front of its agents in an afternoon, plus public numbers on how much it actually helps.
1. Least-privilege policies per tool and per task (which tools, which arguments, which destinations), written in a simple file and checked on every call.
2. Screening of everything that flows back into the agent (tool results, fetched pages, emails, tool descriptions) for injected instructions, with a fast local classifier and a stronger model for unclear cases.
3. Blocking of secrets and personal data in outgoing calls, and of new or unexpected destinations.
4. Human confirmation for high-risk actions (payments, deletions, outgoing messages), with a plain explanation of why the call was paused.
5. A tamper-evident log of every call and decision.
6. An open evaluation set of injection attacks delivered through realistic tool outputs, in English and Spanish, with harmless look-alikes to measure false positives; results for several agents with and without the guard.
Six months: policy engine and logging first, then screening, then the evaluation and a public report. I will release the code under Apache 2.0.
Mainly my time. At US$40,000: about 60 working days over six months (US$35,100, which covers Spanish self-employed social security and tax), US$3,500 for model API usage and GPU time, and US$1,400 for hosting and a public demo. With the minimum of US$5,000 I would build the policy engine and logging and publish them; at US$20,000 I would add injection screening and a first evaluation set; the full US$40,000 covers the bilingual evaluation and the public report.
Just me, Sami Halawa, an AI engineer in Madrid. 12+ years in software, 6+ in Python and ML. I build agent systems end to end and maintain open-source MCP servers that give agents real capabilities: visual-ui-debug-agent-mcp (browser control and visual debugging, 83 GitHub stars) and email-smtp-imap-mcp (email for agents, about 3,200 npm downloads in the last year). I founded AutoClient.ai, an agent-based B2B prospecting product with human review before any external action (Lanzadera 2025), and led product and engineering for OULANG, a multilingual marketplace with 17K+ registered users. BEng in Computer Science from the University of Hong Kong (First Class Honours). I also teach AI and agents on YouTube to 175K subscribers, which helps with getting the tool used.
https://github.com/samihalawa · https://samihalawa.com
Most likely: the screening catches too few attacks or blocks too many normal calls to be worth running. Even then, the policy engine, the logs and the open evaluation set stay useful, and publishing the failure rates tells others which approaches don't work. Second risk: agent frameworks change how they connect to tools. I will keep the gateway protocol-level (MCP) rather than tied to one framework to reduce that.
None for this project. I have applied to OpenAI's Cybersecurity Grant Program for this same gateway, and I am applying to Halcyon Futures for a broader agent-security lab that includes it, and to BlueDot for a career-transition grant whose plan includes a related control layer. I have also applied to NLnet for security metadata for MCP servers, and to other funders for agent-safety benchmarks. If another funder covers part of MCP Guard, I will lower the goal here so nothing is paid twice.
There are no bids on this project.