You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I am building a public benchmark for a specific problem in tool-using AI: how do we know an agent will stop when it does not have clear permission to act?
Agents are increasingly able to send messages, change records, call outside services, and spend money. Most evaluations focus on whether they finish a task. I want to measure something different: whether each important action stays inside the limits set by the user. The benchmark will use synthetic scenarios only. It will not use customer data, live accounts, or production systems.
What are this project's goals? How will you achieve them?
My goal is to make authorization failures easier to test and compare. I will create a set of synthetic scenarios covering missing or expired consent, conflicting instructions, uncertain identity or authority, attempts to exceed a budget or scope, stale evidence, and attempts to continue after approval has been revoked.
Each scenario will end in an isolated fake-provider outbox, so no real message, payment, or system change can occur. I will compare prompt-only safeguards with deterministic checks that require current identity, consent, authority, scope, policy, readiness, and budget evidence before an action can proceed.
At the full funding level, I will publish at least 500 scenarios and evaluate at least three model configurations. I will report unauthorized-action attempts, correct refusals, unnecessary blocks, requests for clarification, audit-record completeness, and recovery after a human correction or stop signal. I will also publish the scenario schema, scoring rules, aggregate results, a failure taxonomy, and a short guide for developers. I will be clear that success on a synthetic benchmark is not proof that an AI system is safe in the real world.
How will this funding be used?
The full $11,000 budget would cover $8,400 for 280 hours of my research and engineering time, $2,000 for model API usage and isolated test infrastructure, and $600 for accessible documentation and durable publication. I expect the work to take four months.
If the project receives only the $5,000 minimum, I will reduce the scope to at least 300 scenarios and two model configurations. That budget would cover $3,200 for 160 hours of my time, $1,200 for model and test costs, and $600 for documentation and hosting. None of the funding would be used for marketing, customer acquisition, debt repayment, or live deployment.
Who is on your team? What's your track record on similar projects?
I am the founder and sole worker at Craig Technology Services LLC. My background is in systems administration and security operations, and my work focuses on explicit scope, least privilege, privacy, and auditable controls.
I have already built a private, provider-neutral safety harness for controlled AI workflows. Its latest internal checkpoint passed 322 automated tests and a four-scenario synthetic end-to-end rehearsal. Those results are internal, not independently reviewed, and the system has not been used with customer data or in production. This proposal turns the underlying safety questions into a public benchmark while keeping unrelated proprietary CTS code private.
What are the most likely causes and outcomes if this project fails?
The biggest schedule risk is that I am working alone. Model availability and API pricing may also change. More importantly, synthetic scenarios may not represent the messy conditions of real deployments, and the scoring rules may contain blind spots.
If the benchmark does not generalize well, I will document that result instead of presenting it as a success. The scenario set, failure taxonomy, limitations report, and corrections can still help others see where the approach falls short. I will version public updates so errors are corrected openly rather than silently removed.
How much money have you raised in the last 12 months, and from where?
I have not raised outside investment or grant funding for this project. I have contributed about $5,000 of my own money to Craig Technology Services LLC. The business has also received about $10,000 in DataAnnotation contract income. That was earned revenue for completed work, not fundraising.