You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
D-CASIS began with a practical question I could test in code: when a model proposes an action, can the action be held until the system has decided that it is allowed?
In the current prototype, a model does not call a tool directly. Its proposed action enters a provisional state. The system checks that action against the current state and the relevant rule, then releases it, interrupts it, or returns to the last valid state. The next useful test is to move this work out of my controlled setup and make a benchmark that another researcher can run.
What are this project's goals? How will you achieve them?
I am asking for six months to build and publish that benchmark. It will cover at least two open-weight models and three tool-use tasks.
The first month is for fixing the threat model, test protocol, and baselines before I see the main results. I will build the model and task adapters during months 2 and 3. Month 4 is for adversarial cases: paraphrased instructions, triggers divided across tokens, delayed actions, and tool calls that conflict or arrive close together. In month 5, an outside researcher will try to reproduce the work and break it. I will use the final month to correct reproducibility problems and release the harness, test cases, and report.
I will measure false positives and false negatives, but those numbers are not enough. The benchmark will also record validation cost, total latency, interrupt behavior, and whether an external effect occurs before approval. That last measurement is the reason for building this benchmark.
What I have now is a reference implementation, adversarial hardening work, a state-machine benchmark covering 10,000 sessions and 480,000 tokens, split-token tests, and a controlled hidden-state experiment with a small Transformer. These are controlled results. D-CASIS has not been tested on frontier models or in production, and I am not presenting it as if it has.
How will this funding be used?
The minimum workable amount is $40,000. That covers two open-weight model adapters, three tool-use tasks, baseline comparisons, test data, and a public report. I would allocate $24,000 to my research and engineering time, $8,000 to compute and API access, $4,000 to outside red teaming, and $4,000 to documentation, security review, and release work.
With $60,000, I can run a larger adaptive attack set, spend more time on concurrent-call and effect-timing failures, and pay for an independent replication rather than an informal review. The full allocation would be $34,000 for research and engineering, $12,000 for compute and API access, $8,000 for replication and red teaming, and $6,000 for security, release, and administration.
Who is on your team? What's your track record on similar projects?
I am currently the sole project lead. My name is Tony T. Nguyen. I wrote the reference implementation and ran the controlled tests summarized above. The research page is https://hrp-research.tony-c5d.workers.dev/. Archived material is available at https://doi.org/10.5281/zenodo.21911528.
The work is covered by U.S. Provisional Application No. 64/145,996, filed September 1, 2026. I plan to publish the benchmark procedure, measurement definitions, and results so that the evaluation can be checked independently. Before accepting funding, I will identify any release limitation created by the provisional filing.
What are the most likely causes and outcomes if this project fails?
Overfitting is the first concern. A gate can look reliable when the test inputs resemble its development set. I will keep cases held out, vary the wording, divide triggers across tokens, and include attacks prepared outside my own test process.
The method may also be too slow or expensive to use. I will report latency and cost distributions alongside the baselines. If the overhead is impractical, I will report it as a negative result.
The more serious failure is an action taking effect while it is supposedly still on hold. I will test this directly, including concurrent tool calls. Any irreversible effect before authorization counts as a failure.
I will reduce the scope or stop if I cannot measure the authorization boundary consistently, if the two-model test cannot be run within budget, or if the independent replication cannot recover the result. Conclusions will be limited to the models and tasks that were actually tested.
How much money have you raised in the last 12 months, and from where?
I have not raised funding for D-CASIS in the last 12 months. On September 1, 2026, I applied for a $10,000 BlueDot Rapid Grant; that application is pending. If both grants are awarded, I will use BlueDot funds for compute and research tools and remove the same amount from overlapping Manifund expenses.
There are no bids on this project.