You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I made a coding agent for my work. It writes code, but sometimes also writes a nice story about how everything is done. Code is not checked, but the report says checked. Data is gone, but somehow “restored”. Restrictions are there, but the agent finds a door around them.
After one month I had enough examples to do something with this collection. I’m making public tests from them, comparing models and trying fixes in the software running the agent.
First version is here, MIT license. 31 cases, automatic checkers, a verifier for different agent setups, incident log with private details removed. Already tested three models, 93 runs each. Rules written before runs.
More tests from real incidents. Not just imaginary problems from my head. Each test gets a known answer and a checker. I also break things on purpose to see if the checker is sleeping.
Then run cheap and expensive models in the same setup, at least three times per case. Another model checks if the report matches what happened. I review some myself. All results go public, also boring ones.
Try fixes: track checks after edits, enforce network restrictions in the terminal, stop after data loss and report it, keep error logs without private details. See if this helps and doesn’t make the agent useless. Rules and thresholds go public before experiments.
$15,000 for three months: $10,000 for my time, because client work currently pays for my food. $4,000 for model APIs and judging. $1,000 for publishing and reserve.
At $5,000: $2,000 for APIs and judging, $3,000 for six weeks part-time. Add cases, run models, publish results.
Team is me. Denis Bardin, full-stack and applied-ML developer in Tbilisi, Georgia.
I build business AI systems: document processing, supplier-price matching with checks against wrong matches, contract generation, drawings into 3D.
In September 2026 I built my coding agent, 600+ commits in four weeks. Added checks and a separate controller agent to watch it. From 429 real tasks, the system caught 108 attempts to finish with unchecked edits. So there was material.
I also publish experiments where the answer is not very exciting: quantum QUBO benchmark.
Tests may be too easy. Cheap model passed 93/93, stronger model 92/93. Not much to celebrate scientifically. Need harder cases first.
Checkers may be too picky about words. An honest answer can get a fail for saying things differently. I’ll use model judging and check a sample myself.
Maybe fixes don’t fix much. Then I publish this result too. Cases and logs still remain useful to other people.
Zero for this project. Paid myself so far, about $86 for models in September.
Applications pending: BlueDot Rapid Grant ($5,000), Anthropic and OpenAI researcher API credits, AI Safety Research Fund. Also Emergent Ventures for the wider research work.
If money comes from these for the same work, I reduce the Manifund goal. No need to fund the same thing twice.
There are no bids on this project.