You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
WelfareCI is a developer tool for testing whether changes to an AI model make it worse at handling animal-welfare situations.
There are already people doing good work on animal-welfare benchmarks. What seems to be missing is an easy way for developers to run those tests as part of normal software development.
The idea is to make it work like other automated tests. A developer changes a model, system prompt or agent, runs WelfareCI, and gets a clear result showing what improved, what stayed the same and where animal-welfare performance got worse.
I want the first version to be open source and simple enough that a developer can run it from the command line or add it to a GitHub workflow.
The first goal is to make existing animal-welfare evaluations easier for developers to actually use.
I do not want to create another benchmark and leave it in a repository. I want to build the engineering layer that connects existing evaluations to the tools developers already use.
I would start with one established animal-welfare evaluation, build a command-line version of WelfareCI around it, then add a GitHub Action so a team can run the same tests whenever they change a model or prompt.
The output would be practical. For example, it could show that a new model version performs worse in scenarios involving farmed animals or decisions made under economic pressure.
After the first integration works, I would make the system modular so other evaluations can be added without rebuilding the whole tool.
By the end of the first phase, I want to have a working open-source version, support for several major model APIs, clear regression reports, documentation and examples that another developer can set up without needing my help.
I would also get feedback from people with stronger animal-welfare research backgrounds so that the tool is useful to that community, not just technically functional.
I am asking for $25,000 to get the first useful version built, tested and released.
Most of the funding would cover my development time while I build the CLI, GitHub integration, model-provider integrations and reporting system.
Part of it would also go toward model API and compute costs, testing across different models, hosting, infrastructure and creating reproducible test runs.
I would also set aside funding to work with people who have deeper animal-welfare expertise. My strength is on the software side, so I do not want to make welfare assumptions on my own when there are researchers who understand those questions much better.
The aim is to use the grant to get from an idea to a working open-source tool that people can actually try, rather than spending most of the money on branding, marketing or administrative costs.
WelfareCI is being developed by a small team with complementary experience across software engineering, AI systems and animal-welfare research.
The technical lead has experience building full-stack applications, APIs, backend infrastructure, monitoring systems, databases and developer-facing products.
One relevant project is AgentProof, a reliability platform for autonomous AI agents. That work involved automated endpoint monitoring, uptime and response-time tracking, historical reliability data, backend services and developer tools.
WelfareCI is a different product, but the engineering challenge is similar: turning something that would normally require manual checking into a repeatable automated test developers can run as part of their workflow.
The team also includes support for animal-welfare research and evaluation. This is important because the project should not make welfare judgments based only on software engineering assumptions. The research side will help with selecting suitable evaluations, interpreting results and reviewing the way welfare-related scenarios are handled.
The team will stay intentionally small so the project can move quickly while still combining strong technical execution with relevant subject-matter knowlThe biggest risk is not that the tool cannot be built. The harder question is whether developers will actually use it.
It could fail if the setup is too complicated, if the results are difficult to interpret, or if teams see animal-welfare evaluation as something separate from their normal development process.
Another risk is that existing evaluations may not translate cleanly into automated regression tests. Some may need adaptation before they work well in a CI environment.
If that happens, I would still want the project to produce something useful: an open-source prototype, documentation of where the technical barriers are, and practical lessons about what is stopping animal-welfare benchmarks from being used in development workflows.
I would rather find that out with a small project now than assume adoption will happen automatically.edge.
The biggest risk is not that the tool cannot be built. The harder question is whether developers will actually use it.
It could fail if the setup is too complicated, if the results are difficult to interpret, or if teams see animal-welfare evaluation as something separate from their normal development process.
Another risk is that existing evaluations may not translate cleanly into automated regression tests. Some may need adaptation before they work well in a CI environment.
If that happens, I would still want the project to produce something useful: an open-source prototype, documentation of where the technical barriers are, and practical lessons about what is stopping animal-welfare benchmarks from being used in development workflows.
I would rather find that out with a small project now than assume adoption will happen automatically.
I have not raised any external funding for this project in the last 12 months. The work so far has been self-directed and I have covered my own development and tooling costs. This would be the first external funding specifically for WelfareCI.