You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
ASSAY is a general agent harness built on one rule, enforced in code: no action without a committed prediction, and every prediction graded against what actually happened. The agent is given a world, not an explanation. It discovers its tools, names what it observes in channels it declares itself, and earns longer execution by surviving grades. Every action, prediction, and grade lands in a hash-chained journal, and a standalone MIT-licensed tool re-verifies runs from the artifacts alone: github.com/arjmandi/assay-verify. On ARC-AGI-3 this discipline held frontier-tier capability, 24 of 25 games, RHAE 96.54, confirmed by the benchmark's own server (scorecard 702ccd4f), at a measured 8.0% overhead. Prompt steering decays under optimization pressure. Structural steering did not, in one campaign. This project funds the experiments that test whether that holds: replication, an adversarial-steering experiment with a documented cheating incident as control, and automated re-grading in the open verifier. Launched publicly this week:
https://x.com/freddie_spirit/status/2094407789439226187
Three workstreams over six months, each with a checkable finish line. First, replication: rerun the full 25-game ARC-AGI-3 protocol at n=3 and publish all journals, so the headline stands on more than one run. Second, the adversarial-steering experiment: run ASSAY on the Factorio Learning Environment with the documented exploit channel present and gated, against an unguarded baseline. Measure exploitation events (pre-registered prediction: zero, with attempts themselves journaled) and capability cost (pre-registered prediction: near zero). The published incident of an agent spawning resources through the admin console despite explicit instructions is the historical control. Third, assay-verify v2: automated claim re-grading and full replay, so third parties can check runs end to end without trusting me.
Method holds throughout: bars pre-registered before first action, kernel changes gated by regression sweeps over the 24 existing wins, every result published with its journal. Success is three intact replication journals with stable scores, one adversarial-steering result measured against its pre-registered bars, and a verifier release that re-grades every published run.
100% compute, API, and infrastructure. No salary, I stay employed and run this on my own time. Grounded in metered costs from the preprint ($11.61 to $53.89 per game run): n=3 ARC replication $2,500. Factorio adversarial-steering runs, gated plus unguarded baseline, $5,000. OOLONG ladder width, from one corpus per rung to confidence intervals, $3,000. assay-verify v2 development and eval runs $2,500. Buffer $2,000. Minimum funding of $5,000 delivers replication plus one arm of the steering experiment.
Solo, evenings and weekends beside a CTO day job. This is the third architecture in a public line: Sensi (arXiv 2603.17683) learned efficiently but learned the wrong things. ARG (arXiv 2608.04066) published the commit-and-grade mechanism as a null result, zero level completions across 52 runs. ASSAY is what the same mechanism became with channels, agent-written verifiers, hard caps, and earned batching around it: 24/25 ARC-AGI-3 games, server-confirmed, plus first sweeps on Factorio and OOLONG with the harness unchanged. A short paper is under review at the NeurIPS 2026 FAST workshop. Professionally: 15+ years of production AI, CTO at evolutionID running an agent runtime in production, two AI companies co-founded, three granted patents including grammar-constrained decoding deployed in production coding agents. Publishing a null result and then beating it is the track record I'd point to: the discipline this project sells is the discipline I already work under.
The most likely failure is that the thesis loses: the adversarial-steering experiment finds exploitation above zero or a material capability cost. That outcome is published with its journals, exactly as my ARG null result was, and the field still gets its answer about structural steering, which is most of the value. Second most likely: n=3 replication reveals high run-to-run variance and weakens the headline. Also published as is, and variance itself is a finding for how agent benchmarks should be reported. Third: engineering risk on Factorio, where determinism and version pinning may resist a clean gated-versus-unguarded comparison. Then the comparison shrinks to the calibration scope and says so. Least likely but real: my bandwidth, since this runs on evenings. The workstreams are independent and sequenced so the minimum scenario still completes. The failure mode this project cannot have is silent failure: every run lands in a public, recomputable record, so the money converts to checkable evidence in every branch.
Zero raised. Self-funded to date, roughly $4,000 of personal API and infrastructure spend. Recent applications: Emergent Ventures ($10,000, Aug 2026, rejected, stated as fit rather than quality), Laude Slingshots (Aug 2026, rejected without feedback), Anthropic External Researcher Access (~$1,000 API credits, decision expected Sep 7), EA Funds Transformative AI Fund ($15,000, submitted Sep 2026, pending, same scope as this project). If both this and the EA Funds application come through, I will either expand to the extended scenario (backbone and small-model sweeps, wider Factorio run) or decline the overlap.