You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Platforms now hand most of their moderation decisions to AI classifiers, and almost nobody outside the vendors can measure how well those classifiers follow a written policy. I've run that purchasing process from the inside, across several vendors and billions of items a month, and the honest answer is that we bought on sales demos and spot checks. I want to build the public test that should have existed: policy-scored case sets, an open runner that works on open models and commercial APIs alike, and published results anyone can reproduce.
Three deliverables over six months. First, five case sets of 300 to 500 items each, covering child safety and grooming, harassment, self-harm, extremism and spam, each tied to a plain-language policy of a few hundred words, with labels agreed by at least two reviewers. The cases are written and synthetic where the content would be illegal or harmful to hold, and I'll publish the generation method so others can check it. Second, a runner that scores any model against a policy and case set. I already have a working version for OpenAI's open gpt-oss-safeguard model in my T&S Workbench, and I'll extend it to the major commercial moderation APIs and to any model served through Ollama. Third, a public results page with precision, recall and over-enforcement rates per harm area, rerun quarterly, plus a short write-up of what the numbers mean for a team choosing a vendor.
Everything ships as open content under a permissive license in the repos I already maintain. The Workbench is ten free Trust & Safety tools used by product and safety teams, and this becomes its eleventh.
Mostly my time. I'd take this on part time for six months, outside my job and without any employer data. The rest is API spend to run the commercial vendors, a second reviewer paid hourly to label cases so no label rests on one person, and a small amount for hosting and a design pass on the results page. If funding comes in above the minimum, the extra goes to a third harm area per quarter and to audio and image cases, which is where the vendors differ most and where almost no public testing exists.
It's me. I've spent ten years in Trust & Safety, the last several leading the function and global compliance at a gaming UGC platform. I built that program from zero, set up its NCMEC partnership, led the migration to AI-assisted scanning across several vendors with human review, and shipped a user reputation system projected to save over $1M a year. Evaluating classifiers against policy is a thing I've done for a living, with real budgets attached.
In 2026 I built and published the T&S Workbench on my own: ten tools, including an abuse pre-mortem with 58 risks and 104 safeguards, an incident tabletop with 43 scenarios, a metrics framework with 36 metrics and SQL, and a classifier eval that already runs on an open model. All of it is open content across eight GitHub repos. I also contribute to ROOST's open safety tooling community. This project is the natural next step of work I've already shown I finish.
The most likely failure is that the synthetic cases don't look enough like real abuse, so the scores flatter the models. I'll reduce that by writing cases from patterns I've seen operationally and by inviting other practitioners to add and challenge cases. Second, vendors could refuse API access or object to being scored. If so, the open models still get tested and the vendor column says "declined," which is itself useful information. Third, I could run short on time while employed full time. The work is staged so each quarter leaves a usable release even if the next one slips. In the worst case, funders get a public case set and runner for two harm areas rather than five, and the method for anyone to finish the rest.
None. The Workbench and the related repos have been self-funded and built on my own time.
There are no bids on this project.