You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Autonomous agents drift from their goals — and an agent watching only itself provably can't catch it, because each step looks fine next to the last while it wanders far from where it began. laserbrain is a fixed outside reference an agent checks in one MCP line; we proved a fixed, unchangeable reference is necessary and sufficient to detect this displacement, and that no self-monitoring agent can be. The detection is a theorem, the grammar is free and open, and the preregistered studies are public — including the null: we tested whether returning the agent then helps, three ways, and could not show it, so we don't claim it. This grant funds the open research layer — extending the detector to multi-agent and long-horizon oversight, and funding the human evaluation needed to finally settle the benefit question that LLM judges can't (inter-rater κ ≈ 0.1 on open-ended work).
The core is already built, deployed, and proven, so this funds the next layer, not a bet on whether it works.
1 — Extend the detector to multi-agent and long-horizon oversight. The theorem holds for one agent against a fixed goal; the same fixed-reference argument extends to a dialogue of agents, where drift becomes echo/agreement spirals, topic-drift, and deliberation stall — collective grounding losing its reference. We'll formalize that case, ship it as an open MCP tool beside the existing one, and measure coverage on real multi-agent runs. Deliverable: an open multi-agent drift check + a short write-up.
2 — Settle the benefit question honestly, with human evaluation. Our preregistered studies show the detection works but could not establish that returning the agent helps — open-ended agent quality has no ground truth, and LLM judges can't resolve it (inter-rater κ ≈ 0.1). We'll fund human raters on the fixed rubric and run one powered, preregistered study to finally answer it, and publish whatever it shows — including, if it comes to it, another null. Deliverable: the powered, human-evaluated benefit result, public.
3 — Keep it a public good. The grammar, theorem, and MCP detector stay free and open, and every study is published (already live at phronesis.world/laserbrain/research). Deliverable: a documented oversight primitive any agent framework attaches to in one line.
Human evaluation — ~30% (~$12k). Paid human raters on the fixed rubric for the one powered benefit study. This is the most decision-relevant line: LLM judges can't resolve open-ended agent quality (inter-rater κ ≈ 0.1), so human raters are the only way to actually settle whether returning the agent helps. Funds enough raters × pairs to clear a real agreement threshold.
Compute & model API — ~25% (~$10k). The funded model keys for the powered agent runs and the multi-agent study — the exact bottleneck that has kept the benefit question open until now (every experiment has been gated on a funded key).
Development — ~40% (~$16k). Formalizing the fixed-reference theorem for the multi-agent / long-horizon case and shipping it as an open MCP tool, plus keeping the open detector documented and maintained. Solo-founder time, or one contract collaborator.
Infrastructure & misc — ~5% (~$2k). Hosting the free, open detector (Cloudflare Workers/Pages) and small tooling.
The team consists solely of the developer. Track record is successful at designing projects to completion.
1 — The benefit stays null, even with human raters. Our own studies already show the detector works but that returning the agent isn't established as helpful. The deepest risk is that human evaluation confirms the same — or that open-ended agent quality resists reliable measurement even by humans, because it may genuinely have no ground truth. Outcome: we publish another null. That damages the product story, but it is not waste: "the harness reliably detects drift, and stopping the agent does not demonstrably help" is a result the field should have, and reporting it is exactly what our discipline is for.
2 — The multi-agent extension doesn't hold as cleanly as the single-agent proof. The theorem is airtight for one agent against a fixed goal; a dialogue of agents may lack a well-defined fixed reference for collective grounding, so the extension could be partial or need a different formalism. Outcome: a weaker or negative theoretical result — which we'd publish, and which still sharpens where the argument applies.
3 — A correct primitive nobody uses. Even proven and free, the drift it catches may not be the drift that bites in real deployments, and its domain (open-ended, criterion-absent work) is narrower than "all agents." Outcome: an open, correct, unadopted tool.
4 — Human evaluation is harder than budgeted. Qualified raters for nuanced judgments are scarce, and agreement may stay low even among people. Outcome: the money is spent and the benefit question is still open.
$0