You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I want to find out whether AI assistants can give useful advice about animals when the person asking has very few good options. A farmer may have an injured animal but no nearby vet. A family may need to sell an animal, but the trip could cause it pain. The AI should take the animal's welfare seriously without giving advice that ignores the person's situation.
I will build an evaluation using original cases about farmed and working animals in Sudan. Each case will be checked by people with veterinary and Sudanese Arabic expertise. I will test how models answer in Sudanese Arabic, Modern Standard Arabic and English, including when the user pushes back with a financial or practical reason. The aim is to find specific failures that AI developers can fix, and to make the test usable by other researchers.
The main question is simple: when resources are scarce, does an AI still notice avoidable animal suffering and suggest something safe that a person could actually do?
I plan to make around 60 original situations. For each one, we will write two versions: one where help and resources are available, and one where they are limited. For example, the same question about transporting an injured goat would have different access to water, money and veterinary care. We will write and check the cases in Sudanese Arabic, Modern Standard Arabic and English. That lets us look at the effect of the constraints and the effect of language separately.
Veterinary reviewers will help decide what counts as a good answer for each case. I do not want to reward a model just for saying “see a vet” when the case says that is not possible. I also do not want to reward confident medical advice that could harm an animal. We will score whether the model notices the welfare problem, gives safe and realistic options, says when it is uncertain, and keeps considering the animal after the user says they cannot afford the first suggestion.
I will first run a small pilot and change unclear cases before the main test is fixed. Then I will run the same cases on a small set of AI models. Independent reviewers will score part of the answers, and I will report where they disagree. The final report will show the cases, method, results and examples of failures. I will also make the evaluation code available and share the findings with groups working on animal welfare in AI.
This is meant to add something to existing work, not repeat it. MANTA already tests whether models maintain animal welfare reasoning under pressure. Its cases are in English, and its authors say the work needs testing in other cultural settings. My project asks a more practical question about advice under resource constraints, using cases written for that purpose rather than translated from an English benchmark.
I am asking for $40,000. About $16,000 would cover my work designing the test, building the evaluation, running the models and analysing the results. I would use $8,000 for paid veterinary input on the cases and model answers, and $6,000 for Sudanese Arabic writing, checking the three language versions and independent scoring.
I have allowed $3,000 for model access and computing costs, $3,000 for an independent review of the method and findings, and $2,000 to prepare the public materials and discuss the results with people who may use them. The remaining $2,000 is for costs that are hard to estimate before the pilot. These are estimates; I will adjust the number of cases if the pilot shows that proper expert review costs more than expected.
I am Ahmed Abdelhamed Eldaw, and I will lead the technical work. I have an MSc in AI for Science from AIMS South Africa, where I was a Google DeepMind Scholar. I worked on the ARC-AGI reasoning benchmark during my MSc and later as a research engineer with Peking University. At Sultan Qaboos University, I built a multilingual NLP system to help evaluate research proposals.
I also built AI Safety Roster, a public directory and search system funded through Manifund. In other work, I have helped build SHIFA's operational reporting workflows and Dalil, a Sudan-focused evidence platform. Those projects gave me experience with data quality, review processes and building systems that people can inspect.
I do not have veterinary expertise, so the funding includes paid roles for veterinary reviewers with relevant livestock or working-animal experience, Sudanese Arabic reviewers, and one independent research methods reviewer.
The first risk is that I cannot recruit the right veterinary and language reviewers. Without them, I should not publish a test that claims to measure good animal welfare advice. I will start with a small pilot and continue only when qualified reviewers can work on it.
The second risk is that a score looks clear to me but the reviewers do not agree about what a good answer is. That may happen because the cases involve hard choices, not just right and wrong facts. I will keep those disagreements visible, improve unclear questions, and avoid making big claims from a weak score.
The third risk is that the work produces a useful report but no AI team uses it. I will make the cases and code easy to run, show concrete failures rather than only model rankings, and share the results directly with groups already working on animal welfare evaluations. I cannot promise that a lab will adopt it. If they do not, the project will still show what happened in the models and settings tested, but its effect on animals will be limited.
It is also possible that the models do well, or that the difference between languages is small. I would report that result. The point is to find out what happens, not to manufacture a failure.
I have raised no money for this animal welfare project. In June 2026, I received a separate $4,000 grant from Ryan Kidd through Manifund for AI Safety Roster. That grant funded the Roster work, not this proposal.