You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
As a psychologist, I have always been interested in human behaviour, especially social behaviour. With the emergence of AI and AI agents, I think we also need to evaluate their social behaviour, particularly when they face different social contexts, tool use, or situations involving pressure.
To do this, I am building an open-source tool based on the concept of behavioural contracts. A contract defines what an agent should do, what it should not do, when it must stop, when it should escalate to a human, and what evidence is required before it can regain autonomy after a violation.
I already have a working prototype with 509 passing tests. It includes a contract DSL, a deterministic evaluator, an incident log, diagnostic change cards, a sandbox for testing candidate fixes, operational integrity checks, and a supervised execution boundary that signs hash-chained evidence. I have also created small adversarial cohorts that test failures such as tool misuse, budget violations, unauthorised payments, stopping-condition failures, and evidence suppression.
Funding would allow me to turn this work from a local prototype into a public benchmark. I would run it against commonly used models, publish the results and dataset, document the failure modes, and make it easy for others to plug in their own agents or mitigation ideas.
Realistically, I don't think this project will solve the alignment problem. But I do think powerful agents will need externally verifiable commitments: not just prompts, but boundaries that can be tested, violated, restricted, and restored. I want to make that measurable.
I want to produce a public v0.1 benchmark with:
1. A small DSL for behavioral contracts: obligations, prohibitions, conditional rules, time limits, and recovery conditions.
2. Agent scenarios that test failures beyond unsafe text: tool misuse, budget violations, stop-condition failures, unauthorized payments, evidence suppression, and unsafe handling of health/public-sector/banking style workflows.
3. A reproducible evaluation harness: canonical trajectories, model/provider adapters, deterministic evaluator, signed evidence outputs, and autonomy recommendations: COACH, SUPERVISE, RESTRICT, QUARANTINE.
4. A public report: which models and agents failed which boundaries, which failures were caught by the
The current prototype already covers the local deterministic part. The missing step is running enough real-model evaluations and packaging the results in a way that is useful to other AI safety researchers and agent builders.
Target funding: 15,000 USD.
- 3,000 USD - API credits for GPT-4o/4.1, Claude Sonnet, Gemini, Llama and Mistral via API.
- 1,000 USD - Linux VM, CI, dataset hosting and public artifact hosting.
- 8,000 USD - engineering/research time to harden the benchmark, build adapters, run evaluations, analyze failures, and write the report.
- 1,500 USD - external review, replication help, or small bounties for researchers/builders who try to break the benchmark.
- 1,500 USD - contingency for extra API runs, failed experiments, or higher-than-expected model costs.
Minimum useful funding: 5,000 USD.
At that level I would do a smaller version: 3-5 models, fewer seeded runs, public v0.1 dataset, short technical report, and no external replication budget.
Stretch funding up to 25,000 USD would let me add more scenarios, multi-agent cases, independent review, better documentation for third-party users, and a simple public leaderboard or results explorer.
Right now this is a solo project by Bartolome Montejano Galan, founder of INCAREXTREM SLU in Spain.
The main track record is the code itself. I have already built the system end-to-end rather than stopping at a proposal:
- 509 passing tests.
- Behavioral-contracts DSL.
- Deterministic evaluator.
- Raw-session adapter.
- Incident registry.
- Diagnostic change cards.
- Sandbox verification.
- Operational-integrity policy checks.
- Signed, hash-chained execution evidence.
- 24-run six-agent behavioral cohort.
- 10-run adversarial containment cohort.
- EU-style scenarios in banking/PSD2, healthcare/GDPR and public administration.
My angle is not purely academic alignment research and not generic compliance consulting. I am coming at this from the deployer side: if an organisation receives an agent from a third party, it needs external control, evidence, and a way to reduce or restore autonomy based on behavior.
The most likely technical failure is that the benchmark turns out to be too narrow. It may catch obvious containment failures while missing more subtle goal-pursuit or deception-like behavior.
Another failure mode is that the DSL becomes too rigid. Real agent workflows may require richer contracts than the first version can express.
A third risk is adoption: even if the benchmark works, other researchers may not care unless the documentation and examples are easy to run.
If the project fails in those ways, I still expect useful outputs: a public set of containment scenarios, negative results about where simple external contracts break down, a clearer list of what future agent-control benchmarks need, and code/evidence formats that others can reuse or criticize.
No funding has been received for this project in the last 12 months.
I applied to BlueDot Impact Rapid Grants for 4,000 USD on 1 September 2026. No decision has been received yet. If both grants were funded, I would avoid double-spending by using BlueDot mainly for API/compute and Manifund for benchmark hardening, public release work, documentation, and replication/review.
There are no bids on this project.