You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
AI control protocols assume a sandbox that contains a caught model. Real agent deployments leak authority through legitimate channels: hidden state left for the next worker, mid-task credential requests, self-reported success. This project builds and attacks the layer that closes that gap. Authority is fixed before execution and never raised; each step runs in a disposable Xen Cell that inherits no state or credentials; only canonicalized artifacts cross boundaries; evidence is produced outside the Cell; a passing check can only activate a pre-committed successor, never mint privilege.
Answer, with evidence, a question control protocols currently leave to the sandbox: even if a compromised agent wins inside its step, what can it carry into the next one? The substrate is already measured in a closed lab (rump host surface 46–55 hypercalls, reproducible builds 190/190, Xen Cells with zero residue over 30 create/destroy cycles, boundary observable via xentrace). The funded work: (1) build the authority plane — Root Job Contract, deterministic transition gate, verified promotion boundary, hypervisor-side observer; (2) run a real coding agent inside two-Cell and fan-out workflows; (3) attack it — successor contamination, escalation, evidence replay, validator fuzzing, covert channels through valid outputs, Byzantine liveness — and report every result, including negative ones; (4) release code, harness, adversarial corpora, TLA+ model and report under open licences.
USD 78,000 for 6 months: researcher stipend 60,000 (replaces the professional income currently funding this out of pocket); model API usage for worker, reviewer and adversarial runs 6,000; lab additions 5,000; targeted specialist review (formal methods, Xen internals, crypto protocol) 5,000; publication and reproducibility 2,000. At the 39,000 minimum: 3 months, ending at a two-Cell evidence-gated transition with root ceiling, external observer, deterministic gate, promoted artifact and no state inheritance, which already proves or falsifies the core invariant.
Solo, by design at this stage. Twenty-seven years of hands-on systems security: former Cyberdefense Director for Latin America at Itaú, the largest bank in the region, detecting and eradicating real intrusions including APT activity; member of Trustix in Norway, a pioneering security-focused Linux distribution; independent vulnerability research (disclosed CVE, RSA errata); CSC, IBM; currently leading global DFIR work. I built the lab and measured the substrate this proposal depends on; artefacts (scripts, SHA256-summed logs, xl configs, xentrace dumps) available to any reviewer. Chile-based, Patagonia. No publications in the AI safety community; my track record is operational.
Most likely technical cause: the rumpuser backend for a Xen PVH guest (Rumprun is historical) takes longer than planned; mitigation is to run the authority plane on a minimal-Linux Cell first, since the invariant does not depend on rump. Most likely scientific outcome: the gate's deterministic predicates govern fewer real coding tasks than hoped, or covert capacity through valid artifacts stays high. Both are published as results; they tell defenders exactly where sandbox-plus-policy leaks authority, which is actionable today.
None. Self-funded to date. Applications pending at the OpenAI Cybersecurity Grant Program (small research grant + credits, submitted 30 Sep 2026) and EA Funds' Transformative AI Fund (USD 78,000, submitted this week). Accepted into Claude for Startups; credits never materialized.
There are no bids on this project.