You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I built and shipped beancount-ledger, a multi-turn tool-use RL environment where an agent acts as a bookkeeper for a small company. It reads a ledger, a bank statement and supporting registers, finds what is missing, wrong or duplicated, writes the corrected ledger back, and submits. The reward is fully deterministic — the trial balance either ties or it does not — so there is no LLM judge anywhere in the scoring path. It is live on Prime Intellect’s Environments Hub with a keyed generative population of 1,400 preflighted tasks and a public evidence chain. This grant funds the next stage: porting it to the UK AISI Inspect framework and publishing a broad, preregistered measurement across many open-weight models.
Goal: turn beancount-ledger into a reproducible, judge-free reference benchmark for agentic capability that anyone can run and trust. Three concrete steps:
(1) Port the environment to the UK AISI Inspect framework so the safety-evaluations community can run it natively alongside their existing task suites.
(2) Publish a broad, preregistered measurement across 15–20 open-weight models with multiple rollouts each, so the community gets an apples-to-apples capability picture on real bookkeeping work — my pilot only covered 4 models because free-tier quotas ran out mid-experiment.
(3) Extend the adversarial exploit corpus for the new task family, keeping the scorer honest. The hard part is already done solo; these steps are execution and the main blocker is inference budget.
The entire grant goes to inference and compute — the one thing I cannot get for free at the scale this needs. Roughly: $1,000 in API credits to run the 15–20 open-model measured evaluation (multiple rollouts per task across the released population); $600 in compute for the Inspect port, generation sweeps and CI; $400 in API credits for the extended adversarial red-team runs. I keep the request modest ($2,000 min / $8,000 max) because everything else — my time, the tooling, the environment itself — is already built and unpaid.
Just me. I'm a solo developer in Turkey (B.Sc. Robotics & AI). Track record on exactly this kind of work: I built and shipped beancount-ledger end-to-end — the generative world graph, the deterministic scorer, the exploit corpus, and a sealed, preregistered 4-model pilot (138 valid rollouts) — with every reward claim tied to a named oracle. Source and full evidence chain: https://github.com/gultekinhasancan79/beancount-ledger . Live on the Hub: https://app.primeintellect.ai/dashboard/environments/cangultekn/beancount-ledger . Professionally I did agentic-task QA at Fleet AI, isolating 100+ automated-verifier defects from genuine agent failures, and I run ongoing RLHF code-preference evaluation at Alignerr across 15 Python libraries.
The most likely failure is not a crash but a smaller result: I run out of budget again and can only measure a handful of models instead of 15–20, or the Inspect port takes longer than expected and I ship the measured release without it. In that case the environment still exists and still works — it’s already live — but the public, reproducible capability picture I want to give the community stays narrow. A second, lower-probability risk is that a broad measurement surfaces a scorer edge case I hadn’t hit at 4 models; that would actually be a useful outcome, because finding and fixing it is exactly the exploit-hardening discipline the project is built around, and I’d document it openly.
Nothing so far — the entire project to date was built and run on free-tier infrastructure and my own unpaid time. I have one other application in flight (a BlueDot Rapid Grant, submitted this week, no decision yet); I’m applying to both openly and will tell you if either is funded so there’s no double-funding of the same line items.
There are no bids on this project.