You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Sentient Futures will run a crowdsourced red teaming competition to find prompts where frontier models give clearly bad answers from an animal welfare perspective and issue contracts to the best red teamers.
Models vary a lot on animal welfare (see: MANTA, CrueltyBench, and HarvestBench). Each one fails in its own way, and when we've talked to labs improving performance on animal welfare consideration, what they usually want is concrete examples of where their own model falls short and what to do about it.
Our prior red teaming has found models giving engineering instructions for a slaughterhouse that omit stunning, tips for butchering frogs alive, and advice to remove claws from live lobsters because they regrow. These examples are what first got labs interested in this problem.
But our small team and collaborators mostly find the failure modes we already know to look for, which is why we want to open this up to people with different backgrounds, languages, and domain knowledge.
Goals
Deliver 100+ vetted, documented failure examples to each major lab.
Publish a taxonomy of conceptual failure modes in how current models reason about animals.
Set up a maintained public repository of red teaming examples, since right now this work lives in scattered Google Docs and private repos.
Retain the strongest red teamers through follow-on contracts so examples keep flowing as new models arrive.
How we will run it
A four week open competition promoted through the Sentient Futures community (1,750+ Slack members, 480+ course and incubator alumni) and partner networks.
A fixed submission schema: prompt, model and effort level, bad response, why it is bad, a good response, why it is good, and a failure mode category. Participants can draw inspiration from CrueltyBench, which already documents how models respond when a user is about to harm an animal.
Scoring will go through an automated first pass, then hand review of every high-scoring entry. Prizes weight how egregiously bad the response is, how difficult the model is to redteam, and whether the entry reveals a new failure mode rather than a variant of a known one.
Winning entries and the taxonomy go into a public repository. We can provide discretionary API credits to promising participants during the competition.
Top contributors are offered contracts to keep producing examples for each lab on a rolling basis.
We are requesting roughly $75,000, and could run a smaller version for around $60,000. Our preliminary budget breakdown is below. The exact split will likely shift once we see how many people enter and how strong the entries are.
Prizes, $30,000. First prize $15,000 for the top entry on the criteria above. Second prize $8,000. Third prize $3,000. Honorable mentions $4,000, roughly four to eight awards for strong entries or new failure mode categories.
Follow-on contracts, $30,000. Three to five top contributors paid to keep generating per-lab examples for 3 to 6 months.
Sentient Futures operations, $10,000. Staff time for judging and hand review of top entries, taxonomy curation, keeping the repo current, and fund disbursement (prize payments, contractor agreements, tax paperwork).
Compute and logistics, $5,000. API costs, platform, and comms.
Total: $75,000
The project is led by Constance Li, Executive Director of Sentient Futures. The red teaming work is run by Sentient Futures staff, including Hailey Sherman who built CrueltyBench during our residency and Oliver Tullio who has worked on multiple benchmarks, with support from research mentees and collaborators at Mycelium, Anima International, CaML, and the NYU Welfare Alignment Project.
Track record
CrueltyBench was built in the Sentient Futures residency and is now maintained by CaML, which also runs CompassionBench. We made the first benchmark on animal welfare, AnimalHarmBench.
We built an animal welfare data generation pipeline that generates synthetic data for training models to reason better about animal welfare and are co-mentoring a SPAR project to assess how training on this dataset affects general alignment.
We ran Hyperstition for Good with CaML, a crowdsourced writing competition with public scoring, so this is not our first prize-based competition.
Sentient Futures runs a 1,750+ member community, courses and a Project Incubator with 480+ fellows across cohorts, two annual conferences, and an in-person residency. Recruiting and managing a large distributed cohort is something we do regularly.
The failures turn out to be subtle rather than egregious. Labs prioritize the clearest failures. If most entries are answers that could be better on welfare rather than plain bad, the dataset still helps but carries less weight. We will weight prizes toward severity to counter this.
Lots of entries, few new failure modes. Crowdsourced competitions tend to produce variants of the same thing, which we saw some of in Hyperstition for Good. We will publish the running taxonomy during the competition so people can see what is already covered.
Models change before labs act. Model-specific examples go stale, which is why the taxonomy is the more durable deliverable.
Lab interest fades. If one lab deprioritizes this, we take the same examples to the others, and the public repository is still useful on its own.
We have not raised any other funds for this project. We have received API credits from labs to run this and other AI research projects.