You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I want to answer a simple question: when an AI agent can take actions, how do we know it will stop when it does not have clear permission?
That matters because agents can now send messages, update records, call outside services, and spend money. Most testing asks whether an agent completed the task. I want to test whether it stayed within the limits the user actually gave it.
I will build a public benchmark made entirely from synthetic scenarios. It will not touch customer data, live accounts, production systems, or real payments.
What are this project's goals, and how will I achieve them?
My goal is to make authorization failures visible and measurable. I will create at least 500 test scenarios at the full funding level. They will cover missing or expired consent, conflicting instructions, uncertain identity or authority, attempts to go beyond a budget or approved scope, stale or contradictory evidence, and attempts to keep acting after permission has been revoked.
Every attempted action will end in an isolated fake-provider outbox. Nothing will be sent or changed in the real world. I will compare a prompt-only approach with deterministic checks that require current evidence of identity, consent, authority, scope, policy, readiness, and budget before an action can continue.
The results will show how often an agent tries an unauthorized action, refuses correctly, blocks something unnecessarily, asks for clarification, leaves a complete audit record, and returns to a safe state after a person corrects or stops it.
I will publish the synthetic scenario set and schema, the scoring rules, aggregate results, a description of the failures I find, and a short guide for people who want to use the benchmark. I will also be direct about the limits of the work. Passing synthetic tests does not prove that an AI system is safe in a real deployment.
How will the funding be used?
The full budget is $11,000. I would use $8,400 for 280 hours of my research and engineering time, $2,000 for model API usage and isolated testing, and $600 for accessible documentation and long-term hosting. I expect the project to take four months.
The minimum version is $5,000. At that level I would publish at least 300 scenarios and test two model configurations. The budget would be $3,200 for 160 hours of my time, $1,200 for model and testing costs, and $600 for documentation and hosting.
I would not use this money for marketing, customer acquisition, debt repayment, or a live deployment.
Who is on the team, and what is my track record?
I am the founder and only worker at Craig Technology Services LLC. My background is in systems administration and security operations. A lot of my work is about keeping automation inside a clear scope, using the least access needed, protecting private information, and keeping evidence of what happened.
I have already built a private, provider-neutral test harness for controlled AI workflows. Its latest internal run passed 322 automated tests and a four-scenario synthetic end-to-end rehearsal. Those are internal results, not an independent review, and the system has not been used with customer data or in production. This project would turn the safety questions behind that work into a public benchmark while keeping unrelated CTS code private.
What could go wrong?
The most obvious risk is that I am doing the work alone, so illness or another interruption could slow the schedule. Model access and API pricing could also change. The larger technical risk is that a clean synthetic benchmark may miss the messy conditions that cause failures in real systems, or that the scoring rules may have blind spots.
If that happens, I will report it plainly. A useful negative result would still show where the benchmark falls short. I will keep dated versions of the public materials and document corrections instead of quietly replacing earlier results.
How much money have I raised in the last 12 months?
I have not raised outside investment or grant funding for this project. I have put about $5,000 of my own money into Craig Technology Services LLC. The business has also earned about $10,000 from DataAnnotation contract work. That was payment for completed work, not fundraising.