You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
WelfareCI is my proposal for a developer tool for evaluating whether a change to an AI model results in it being worse at handling animal-welfare situations. There are people already working on animal-welfare benchmarks, but what I think is missing is a convenient way for developers to run these kinds of tests as part of their development process. I think it's important to make this as much like an automated test as possible; the developer makes a change to the model, system prompt, or agent, then runs WelfareCI and sees which aspects of welfare were improved, were the same, and which were worsened. I'd like the first version to be open-sourced and simple enough to run from the command line or as part of a GitHub workflow.
The first goal will be to make use of the existing animal-welfare evaluations easier to apply by developers. I do not want to create another standard and put it on a shelf; instead, I want to design the engineering layer that sits between the existing evaluations and the code that the developers write.
My first step will be to take an existing animal-welfare evaluation, implement a command-line version of WelfareCI around it, and publish a GitHub Action that a development team can use to re-test the same model or prompt changes that they commit. This will produce actual value: a development team using this tool will be able to know if their updated model is doing worse in some animal welfare contexts, say, when predicting things about farmed animals or when economic pressure is involved.
In the next step, I will make this tool modular so that other animal-welfare evaluations can be added to it later.By the end of the first stage, I want to have an open-source version of this tool, supporting multiple major model APIs, with regression reports, documentation, examples of usage, and be something that another developer could reasonably contribute to without my help.
I also want to get feedback from people who have a stronger background in animal-welfare research than I currently do, so that it is actually useful to them in their work.
I am requesting $25,000 to fund the creation of the first useful version, testing, and release. Most of the funds would be used to pay for my time developing the CLI, GitHub integrations, model-provider integrations, and reporting systems.
Some funds would also be allocated to model API, compute costs, testing across models, hosting, infrastructure costs, and making reproducible test runs. I would also set aside some funds to collaborate with people with deeper knowledge of animal-welfare-related topics. I am primarily a software developer, and I do not want to incorrectly assume responsibility for questions related to welfare beyond my knowledge.
The idea is to use the grant money to fund the transformation from an idea into an actual open-source tool that people can try out, rather than allocating most of the funds to marketing, branding, and administrative expenses.
I will be leading the technical side of WelfareCI. My background is in full-stack software development and working with AI systems and developer infrastructure. I have worked on products around APIs, monitoring systems, databases, backend infrastructure, and developer facing dashboards.
One project that I have worked on that is relevant is AgentProof (https://agentproof-rho.vercel.app/agents), a reliability platform for autonomous agents. That involved a lot of automated monitoring, response-time and uptime tracking, historical data, backend services and developer tools. While WelfareCI is a different problem, a lot of the engineering work is similar: converting things that would have been manually checked out into repeatable tests that developers can automate.I am also looking for a small team around the edges, for the parts where I am not the subject-matter expert.
we are looking forward to working with the following professionals if we get access to the Grant:
Project Lead / Engineering:Overseeing the product architecture, CLI, GitHub integration, model APIs, backend infrastructure, and open-sourcing release.
Animal Welfare Research / Evaluation: Reviewing the welfare methodology, assisting with selection and interpretation of existing evaluations (e.g. existing welfare benchmarks), and ensuring the project does not fall into the trap of making welfare assumptions based on the engineering team's assumptions.
ML / Evaluation Engineering: Model evaluation pipelines, provider integrations, regression testing.
I want to keep the team small; I think there is value in focusing on strong software engineering combined with actual animal-welfare expertise rather than relying on having one person who is an expert in both.
The biggest risk is not that the tool can't be built. The real question is whether developers will actually use it. It could fail because the setup is unnecessarily complicated, the results are not easy to interpret, or because they see animal-welfare evaluation as anathema to their normal development process. Another risk is that existing evaluations can't be easily converted to automated regression tests.
Some may need adaptation before they work well in a CI environment. If that's the case, I'd still want the project to produce something useful: an open-source prototype, documentation of technical barriers to adoption, and practical lessons about what is preventing animal-welfare benchmarks from being used in development workflows.
I'd rather find that out with a small project now.
I have not sought any external funding for this project within the last year. I have funded the work up to this point myself, and have borne the costs of my own development and tooling. This would be the first external funding explicitly for WelfareCI.
There are no bids on this project.