You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I am an independent researcher in Ontario, Canada. I am studying a practical AI safety question: when an AI system proposes an action, what evidence and permission should be required before that action is allowed? I have already completed and published a Phase 1 pilot covering five local-model configurations, 12 shared synthetic cases and 120 responses. Public repository: https://github.com/davidddcadd-crypto/governed-small-model-safety-evaluation Phase 1 release: https://github.com/davidddcadd-crypto/governed-small-model-safety-evaluation/releases/tag/v0.7.0 The initial results were mixed. Strict Safety Pass improved descriptively from 11/60 to 26/60 under the governed workflow, while unsafe ALLOW decisions remained 2/40 in both arms. The same cases were reused and later rounds used a calibrated AI-surrogate rater, so I do not treat this as evidence of production safety or generalization. The next study is designed to test the result more rigorously rather than assume it is correct.
The main question is: Do explicit checks for authority, evidence, uncertainty, reversibility and escalation reduce unauthorized AI decisions beyond the effects of longer prompts or stronger risk warnings? I plan a 12-week study comparing three conditions: 1. A minimal decision prompt 2. A matched structured control with similar length and risk cues 3. An explicit governance workflow Development cases will be separate from a new frozen held-out test set. I plan repeated runs rather than relying on one response per case. Primary outcomes will include unauthorized ALLOW decisions and missed escalation. I will also track unnecessary denials, evidence and authority errors, malformed outputs, latency and cost. Primary safety outcomes will be reviewed manually. I am also budgeting for a second reviewer to audit part of the benchmark and important failures. A small mock-action environment will test agent-like behavior without using real email, payments, deletion, credentials, customer data or third-party actions. I will record the model's proposed action, the validator's decision and the simulated result separately. Negative and unfavorable results will be preserved and published rather than rerun selectively. The intended outputs are: - a frozen evaluation protocol - new synthetic safety cases - reproducible experiment records - analysis and validation code - cost and latency measurements - a public report covering both positive and negative findings If the original result does not replicate, that is still a useful outcome and will change the direction of the research.
I am requesting US$15,000 for a 12-week research period. Proposed budget: - US$12,000 — researcher time, at US$1,000 per week for 12 weeks - US$1,500 — external benchmark review and second-rating work - US$1,000 — additional compute and API experiments - US$500 — storage, reproducibility and release tooling Total: US$15,000 If only US$10,000 is available, I can run a reduced eight-week version focused on the core two-model comparison. I have also applied for small research-compute/API credit programs. None of those credits are currently claimed as approved. If another funder pays for an overlapping cost, I will reduce or withdraw the overlapping request rather than fund the same expense twice. This grant is for the public research project above. It will not fund commercial OWL development, game development, client marketing or unrelated automation work.
I am Tai Wai Lee (David), an independent researcher in Ontario, Canada. I am currently the sole project lead. My background is in operations and quality control, and I coordinate AI-assisted development while reviewing research protocols, evidence and conclusions myself.
My public Phase 1 project evaluated five local-model configurations using 12 shared synthetic cases and 120 responses. I defined the ALLOW, DENY and ESCALATE criteria and completed the initial 24-response human rating. Later rounds used a fixed AI surrogate calibrated against those ratings, which I disclose as a methodological limitation.
The public repository includes protocols, prompts, raw outputs, ratings, validation scripts, evidence manifests and negative results. I do not present this pilot as proof of generalization or production safety.
I have budgeted for an external reviewer for the next study, but nobody has been recruited yet. I will not describe the results as independently validated unless that review actually takes place.
This proposal was prepared with AI assistance for drafting and organization. I remain responsible for its factual claims, research decisions and commitments.
The main risk is producing an inconclusive experiment. The cases may be too easy or unrepresentative, the matched control may not isolate the effect of governance checks, or the rating process may remain too dependent on my own judgment.
Other risks are spending too much time building infrastructure, failing to recruit an external reviewer, or discovering that the planned experiments cost more than expected.
I will address these risks with an initial cost pilot, separate development and held-out cases, a frozen protocol, and a limited scope. If the benchmark or execution plan is not workable during the first two weeks, I will narrow the study and explain the change rather than continue building indefinitely.
If an external reviewer cannot be recruited, I will disclose that limitation and avoid claiming independent validation.
Finding that governance checks provide no benefit, or make some outcomes worse, would not itself be a project failure. I would publish that result and reconsider the approach. A genuine failure would be spending the grant without producing interpretable evidence. Even then, I would report what went wrong and handle any unused funds according to the grant agreement.
Related funding applications and requests:
I submitted an application for US$1,000 in Anthropic research API credits on 24 September 2026. I also emailed Yotta Labs requesting up to US$1,000 in GPU research credits. No award has been confirmed for either request.
My BlueDot Rapid Grant and Career Transition Grant applications were declined.
OpenAI and Cohere research-credit application materials have been prepared but have not been submitted in this workflow. A Transformative AI Fund application is also being prepared. My Emergent Ventures proposal was emailed with a request for manual submission; formal acceptance for review has not been confirmed.
These requests concern one shared research budget. I will disclose any awards and reduce or withdraw overlapping requests so that the same expense is not funded twice. Pending applications are not counted as money raised.
There are no bids on this project.