You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Award from Falcon Fund
Project summary
There is a short window of opportunity to provide mid-training data that helps AI models reason well about animal welfare. Constance Li has been leading a sprint project to get this done with a few people I like (Allen Lu, among others).
This funding slightly funges with Sentient Futures funding, but I’m very okay with that. I’m expecting more money to come into the Falcon Fund, so I am awarding $20k here.
The idea is to train on data where AIs treat animals well and behave admirably towards them, but, most importantly, to explain why the decisions are being made so that, hopefully, this generalizes out of distribution. Fiction is completely fine for this.
I think it's quite reasonable to think “animal welfare people” should be doing this as opposed to lab employees. They have much more nuanced, well-thought-out views on why things are bad for animals, from a variety of philosophical perspectives, and will do a better job of explaining why.
Create mid-training data to give to the AI Labs
Pay people to create the mid-training data.
Constance Li - Leads Sentient Futures
Allen Lu - Created the MANTA benchmark
Aidan Kanknyoku - Animal welfare alignment team at Anima
Sentient Futures Residents
This turns out not to be that useful; it's quite speculative. It could also just get trained out in RL.
It's part of Sentient Futures, but they don't have enough money.
The project seems good and promising, and I don't know of any other grantmaker who would make this.
I think this is unlikely to have very large effects on models, at least at this scale.
I think it could use more than this. My main uncertainty is how much funding the Falcon Fund will get.
Please disclose e.g. any romantic, professional, financial, housemate, or familial relationships you have with the grant recipient(s).
N/A
There are no bids on this project.