You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Computer-use and coding agents now run with a shell, a browser and real logins on real machines. That machine usually holds the API keys, cloud credentials, SSH keys and browser sessions the agent needs, and nothing independent records what the agent did. A misaligned model or a single prompt injection can read those credentials, exfiltrate data or act quietly, and there is no reliable record afterwards.
I'm building an open (Apache-2.0) containment layer from two tools I already maintain:
herds' sandbox (https://github.com/teddyoweh/herds, about 650 PyPI downloads a month): each agent run gets its own HOME and TMPDIR and a macOS Seatbelt profile that rolls writes back and denies reads of credential stores (~/.ssh, ~/.aws, ~/.config/gh, Keychains, browser cookies).
proxyagent (https://github.com/teddyoweh/proxyagent): the agent's machine holds only a scoped, revocable token. The real model and tool keys stay on a proxy that swaps them in upstream and logs every call.
Put together: the agent operates a machine without ever holding a real credential, all of its traffic passes one choke point with a kill switch, and every action lands in a tamper-evident log stored off the machine.
Credential broker. Extend proxyagent from model and tool keys to every secret an agent uses (git and cloud tokens, website sessions), with per-task scope, TTL, spend limits and one-command revocation that stops the agent immediately.
Sandbox and egress control. Harden the macOS sandbox, add a Linux profile, and route all network egress through the proxy with an allow-list, so the only way out is logged and can be cut.
Tamper-evident audit log. Every shell command, file write, browser step (with screenshot), model call and tool call is appended to a hash-chained log off the agent's machine, with a replay viewer and JSONL export for monitoring research.
Containment evaluation. An open suite of adversarial tasks where the agent, or content injected into a web page or file, tries to read credentials, exfiltrate data, persist, or disable logging. Run on at least 3 frontier models with the layer off and on; publish credential-exposure rate, blocked attempts and detection rate.
Release. 1.0 with integrations for Claude Code, Codex, goose and mantis-agent-sdk, plus a public write-up.
Success looks like: zero real credentials reachable from the agent's machine in the eval suite with the layer on, every escape or exfiltration attempt either blocked or recorded, and a layer anyone can add to the major agent harnesses in a few commands.
$15,000 over 6 months:
$8,000: part of my time building the broker, sandbox, egress control and audit log
$4,000: model API usage for the containment evaluation (3+ frontier models, layer on and off, repeated runs)
$3,000: external red-team of the sandbox and credential broker
At the $6,000 minimum: the credential broker, audit log and a smaller evaluation on 2 models.
Me, Teddy Oweh (https://teddyoweh.net, https://github.com/teddyoweh).
Google, software engineering intern (2026): LLM agent tooling and visual grounding from screenshots.
Apple (Apple Home, two internships): ML and software. LinkedIn: software engineering and ML/relevance.
USC ISI: built a "Git for models" registry that cut checkpoint storage 73%, saves 12 TB a month, adopted by 8 teams.
Maintainer of herds, proxyagent, mantis-agent-sdk (Apache-2.0 cross-provider agent SDK with conformance tests across Claude, OpenAI, Gemini, Grok and open-weight models) and iphone-mcp.
Founder of Spawn Labs, whose product Universe (https://universe.works) runs agents on the user's own Mac and records every step (per-step browser screenshots, full transcripts, files) so the work is auditable.
Senior, Computer Science and Mathematics, Morgan State University.
The core pieces already exist and ship, so the realistic range of outcomes is about scope: the credential broker and audit log are the first milestones and land regardless, and the evaluation grows or shrinks with funding. Everything produced along the way is released under Apache-2.0.
None for this work; it has been self-funded. Pending applications: EA Funds Transformative AI Fund ($40,000, this same work, applied 28 Sep 2026), GitHub Secure Open Source Fund ($10,000, security hardening of herds and proxyagent), goose Grant Program ($100,000, a goose runtime on herds). Manifund funds would extend the evaluation to more models and the red-team scope.
There are no bids on this project.