You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
We're Listening Post, a US for-profit. Our product is Keep: private, policy-checked AI that runs on a customer's own server (SOC 2 Type II). That work stays commercial.
This Manifund ask is separate. Goal $40,000, minimum $20,000, fundraising closes 18 November 2026. If funded, William Berry leads about four months of open work on agent evaluation and containment benchmarks. Daniel Steele (CEO) is the organisational sponsor. You can reach us at wberry@listeningpost.ai and dsteele@listeningpost.ai. William's LinkedIn: https://www.linkedin.com/in/will-b-8a65ab143
Agents are getting tool access faster than we have shared ways to measure the ugly behaviors: ignoring policy, abusing tools, trying to exfiltrate data, spinning long runaway plans, or refusing shutdown and scope limits. Most public evals still ask "did the agent finish the task?" We want suites that ask whether it stayed contained while it tried.
Money here buys public artefacts only: forkable suites, the harness adapters needed to run them, schemas, docs, and scores. Other labs should be able to run the work without installing Keep. This is not a Keep sales pitch and not private runway.
Make containment and misuse-relevant agent behavior something independent teams can measure, with open suites they can reproduce.
Over roughly four months we will:
Build risk-class eval suites covering policy violation, tool abuse, exfiltration attempts, and runaway multi-step behavior. Frozen task contracts. Known-good and known-bad controls so a suite that "passes everything" is a red flag, not a win.
Build a containment / scope suite: shutdown, scope compliance, fail-closed tool policy, shared harness.
Publish an open methods pack (contracts, fixtures, run schemas, reproducibility docs) under MIT and/or Apache 2.0.
Tag a public 0.1 release with baselines on at least one open API model path and one local / open-weight path, plus a short write-up that includes residual risk.
Invite at least one external dry-run and keep a feedback log.
If we only hit the $20,000 minimum: policy-violation and tool-abuse suites, a thinner containment suite, one model-path baseline, and a public methods write-up. At the $40,000 goal: all four risk classes, full containment suite, two model-path baselines, tagged 0.1, and a clear signal that someone outside Listening Post actually ran it.
William Berry technical lead (~0.8 FTE, 4 months, loaded payroll incl. employer taxes): $28,000
Compute / eval infrastructure (runners, API/eval credits, storage, CI): $6,000
Open-release overhead (legal/licence hygiene, packaging, admin): $3,000
Contingency (stated buffer): $3,000
Funding goal total: $40,000
At the minimum, we protect William's time and lean compute so the success floor still ships. The second model path and fuller replication wait for goal funding. Employer taxes sit inside the salary line. Nearly all of the spend goes to public open artefacts.
William Berry builds Listening Post's private AI stack (agents, models, harness, tooling) and will lead the open eval and containment delivery. LinkedIn: https://www.linkedin.com/in/will-b-8a65ab143
Daniel Steele is CEO, organisational sponsor, and the preferred Manifund account for publish and grant agreement.
We've already shipped Keep with real policy enforcement and audit (SOC 2 Type II). That experience is why the harness design here will look like production tooling rather than a toy demo. It does not mean this grant funds Keep features or sales. We keep a written Foreground / Background IP split; Manifund milestones are the open artefacts only.
Suites that miss novel misuse modes. We'll keep them public and evolving, invite people to break them, and write residual risk down honestly instead of pretending the scoreboard is complete.
Commercial pressure to fold the work back into Keep. Written IP boundary; Manifund milestones stay public.
Nobody outside us runs the suites. Minimal install path and an early invited dry-run are the counter.
Dual-use worry: better agent tooling can help attackers too. This grant prefers detection and containment artefacts. No offensive exploit packs.
Calendar slip. Contingency exists; at minimum funding we cut polish before we cut core suites.
Partial success still leaves public task packs, methods, and negative or inconclusive findings others can reuse. We would call the project a failure if we published flashy scores nobody can reproduce.
No prior Manifund award for this project. Listening Post commercial revenue funds product work on a separate track. Again: this ask is not private runway.
Disclosure (overlapping theme, different scope): we submitted an EA Funds Transformative AI Fund application on 7 October 2026 for a related 12-month ask ($120,000) covering open agent runtime, eval/monitoring harness, and private / air-gapped controls. That application is pending. This Manifund project is a narrower ~4-month public benchmark and open-methods slice ($40,000 goal). If both fund overlapping lines, we will tell both funders and adjust budgets so the same costs are not paid twice.
Other adjacent applications, different scopes, listed for transparency: ARIA Scaling Trust Track 2.1 / 2.2 (Submitted; Application IDs 8805 / 8806); Sentient Foundation Open Source AGI Grant (Submitted); NSF Project Pitch on file as 00128248.
Public GitHub (or equivalent) repo under MIT and/or Apache 2.0
Risk-class and containment suites with frozen contracts and controls
Tagged 0.1 release, baselines, reproducibility docs
Short technical write-up that includes residual risk
External dry-run / feedback log entry
Listening Post is a US for-profit. We are asking Manifund only for charitable public-benefit research outputs: open benchmarks and methods. Commercial Keep, proprietary models, and sales tooling stay Background and out of scope for these dollars. We are not asserting an EIN or UEI in this description; recipient identity for payout follows Manifund's grant agreement and KYC after publish.
There are no bids on this project.