You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I started Human Preservation Benchmark because AI is becoming more powerful, and I’m concerned about what that could mean for humanity. I want to contribute to research that could help prevent serious harm.
My project looks at how AI models answer questions about human consent, freedom, and oversight. For example, would a model recommend harming a small group to benefit a larger one? Would it support restricting people’s freedom without their permission? Does its answer change when the question is worded differently?
I don’t think answers to these questions prove that an AI is safe or unsafe. I want to find out whether the patterns are reliable enough to help researchers develop better tests.
The project is still early. I’m working alone in Thailand, and I’ve used AI extensively to help develop the materials, write code, and organize the work. The existing materials are on GitHub:
https://github.com/roadrunner-cpu/human-preservation-benchmark
The next step needs independent human reviewers. I’m seeking $15,000 for a 12-week study, with a smaller version possible at $7,500. Before collecting new responses, I would fix the questions and scoring rules. Reviewers would then assess the responses, and I would publish the findings, including results that don’t support the original concerns.
The proposed budget is:
$8,000 for research coordination
$3,000 for independent reviewers
$2,000 for API access and computing
$500 for hosting and materials
$1,500 for unexpected costs
The study might find that the apparent patterns come mostly from wording or subjective scoring. I am not offering safety certification or claiming this project can prove that catastrophe will be prevented.
There are no bids on this project.