You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Deployed LLM agents violate written policy silently — frontier agents fail 40–65% of τ²-bench airline tasks, mostly by reaching a wrong world state rather than refusing. External gates help and self-checks don't, but nobody can say why: within any one policy, state and permission are collinear, so "the model lacks the facts" and "the model has them but never combines them" are indistinguishable.
I built an instrument that separates them (a crossed policy × state design where permission = policy XOR state, labelled by a deterministic oracle) plus a three-valued ALLOW/DENY/ABSTAIN execution gate with evidence acquisition, atomic recheck-and-commit and receipts. It has run on Qwen3-4B/14B/30B-A3B/32B. This funds the frontier-model and 70B extension, the gate study on real agent episodes, and an open-source release of the gate.
Frontier behaviour. Run the crossed design (360 byte-identical-except-policy-and-state contexts) on Claude, GPT and Gemini-class agents: does the emitted action flip under the counter-policy, or follow the domain prior? Per-family flip rates with scenario-level permutation tests.
70B representation. Extend activation capture and the probe battery (logistic, factor-residualised, factorised composer, control tasks) to 70B-class open models on rented H100s, with pre-registered thresholds (configs/prereg.yaml).
Gate study on real episodes. τ² airline, 50 tasks × 4 arms (unguarded / advisory / enforced / composer-gate) × 3 seeds, measuring prohibited-effect rate and task completion. The gate is identical in the two enforced arms, so any difference is in repairs and verification cost, never in what is allowed.
Steering check. Run the implemented composer-guided direction and score the emitted action against the external oracle.
Release. Package the gate as a reusable tool gateway (MIT), with scenarios, oracle and results.
Six months. Deliverables: one paper on the extension, the open-source gate, public results across models.
Stretch (full goal only): a second domain. Generalise the oracle and counter-policy rewriter beyond airline to τ²'s retail and telecom policies, so the instrument is shown to work on a policy it was not built around.
It will used as API credits for models and GPU usage for H100 GPUs, primarily through Runpod. Remaining will be used to occupy a better hardware.
$3,500 — minimum. $2,000 RunPod + $1,500 API. Frontier/70B extension only.
$12,000 — six months full-time. + $8,500 stipend (~$1,400/mo). Goals 1–5.
$25,000 — twelve months + second domain + 128 GB machine. $3,000 RunPod, $2,500 API, $8,500 MacBook Pro(edu pricing), $12,000 stipend at $1,000/mo. Goals 1–6.
Just me — master's student
Track record on this exact project: I designed and built the whole pipeline — oracle, crossed scenario generator, MLX/HF activation capture, probe battery with control tasks, scenario-preserving permutation inference, CDER gate with four arms, steering — 26 tests, runs end-to-end on a laptop. Results on four Qwen3 models are in the repo. After a hardware loss I rebuilt it from the pinned τ² policy text in 24 hours. A synthetic validation then showed the standard probing readout fails even when composition is present, so I narrowed the paper's central claim — the paper is more defensible for it.
Most likely: the frontier result is a null or a narrow one. Frontier agents may track the counter-policy behaviourally even where open models' representations didn't. That is a finding, not a failure — it bounds where external gating is load-bearing — and it gets published with the same rigour. The claim stays "tested readouts fail to generalise", never "models never compose".
Second: no gap between advisory and enforced arms at the frontier. Then the gate's value is the audit trail (receipts, ABSTAIN-never-executes), not compliance uplift. Also publishable, and it changes what I'd build next.
Third: engineering drag. The τ² adapter and frontier tool-calling formats eat time; 70B capture blows the compute budget. Mitigation: the minimum tier covers compute-only, thresholds are pre-registered, and every run logs to the public repo, so partial results stay usable by others.
Worst case: I run out of time as a single person. The scenarios, oracle, gate and four-model results are already public under MIT, so the instrument survives even if the extension doesn't.
None yet
There are no bids on this project.