You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Our research investigates how to make AI systems safer and more reliable when providing agricultural information in high-stakes, real-world settings. We are developing an initial agricultural question-and-answer dataset and evaluation framework to test how well language models understand agricultural problems, provide appropriate recommendations, express uncertainty, and avoid harmful or misleading advice.
The first phase will establish baseline performance across several models, followed by targeted fine-tuning and systematic evaluation. We will compare model behavior before and after training and identify failure patterns that could lead to unsafe agricultural decisions.
Our goal is to produce a reproducible research benchmark and experimental results that contribute to the broader study of AI safety, reliability, and human-centered evaluation in specialized domains. In the longer term, this work could inform the development of agricultural AI systems that are better grounded in expert knowledge and real-world evidence.
Project goals;
Firstly we will establish a reliable baseline for how current language models perform on agricultural questions, including accuracy, safety, reasoning, and uncertainty.
And we will Identify failure modes where AI produces incorrect, misleading, overconfident, or potentially harmful agricultural recommendations.
To test whether targeted training improves reliability by fine-tuning models on expert-reviewed agricultural data and comparing them against their original versions.
And develop a reproducible evaluation benchmark that can measure agricultural AI safety across different models and scenarios.
Lastly we will Study how AI safety and reliability transfer across languages and contexts, particularly where training data and expert resources are limited
The outcome of this project will be a reproducible benchmark, experimental results, and research paper that can contribute to safer AI development for high-stakes domais.
We will first expand our agricultural dataset with human and domain expert input, we will then evaluate several open weight language models using a held out test set, establish baseline results, and fine-tune selected models using parameter-efficient methods. Finally, we will compare pre and post training performance using standardized safety, accuracy, hallucination, and uncertainty metrics.
$30,000. Dataset creation & expert validation expand the agricultural dataset, human annotation, quality control, and review by agricultural experts.
$30,000. Compute & model training Fine tuning and evaluating multiple language models across different experimental configurations.
$15,000. Safety benchmark & evaluation develop rigorous tests for accuracy, hallucination, harmful advice, uncertainty, and model failure modes.
$10,000. Research personnel support for 3 months.
$5,000. Field validation Collect real world agricultural scenarios and validate model responses against practical conditions.
$5,000. Infrastructure & tools: Research software, data storage and AI assistant subscription
$5,000. Publication & dissemination Research documentation, open source release, benchmark publication, and preparation of the research paper.
This is me Muhammad Zakari – Founder & Lead Researcher software engineering student & Self taught AI safety researcher Built the agricultural safety dataset, evaluation framework, and QLoRA fine tuning pipeline.
Jin Wang – Research Advisor PhD candidate in Economics at the University of Arizona. Advises on evaluation methodology, dataset design, and farmer trial protocols.
Muhammad Khalilullah Uthman – Policy Advisor
Technical Assistant to the Katsina State Governor on Engineering and Automation. Provides policy guidance and helps connect the research to public sector adoption.
· Abdulrahman Uthman
Professional software developer with 6 years experience and open-source contributor with over 4,000 GitHub commits. Supports code quality and development workflows.
Track Record on Similar Projects have already built and openly released the core tools for this research a 1,000-entry multilingual agricultural safety dataset, evaluation scripts that score models on safety and dialect accuracy, and a QLoRA fine-tuning pipeline. We also trained a proof of concept adapter and published it on Hugging Face, demonstrating that our pipeline works end to end.
In addition, build & tested a earlier agromind.chat with more than 100 farmers, documented the failures of frontier models, and used that evidence to shape the current benchmark. The work has been recognised through an $5k Tinker research grant, selection for the Africa Impact Challenge 2026 cohort, and selection for the United Nations Institute for Training and Research (UNITAR) Empowering Youth and Women programme in 2026 (ongoing) and an invitation to interview at The Engine (MIT).
Website: https://agroguardai.com
Chatbot: https://agromind.chat
GitHub: https://github.com/agroguardaiOSS
Hugging Face: https://huggingface.co/AgroguardAI
The most likely causes of failure are limited access to cloud compute and budget shortfalls. The project depends on running multiple fine‑tuning experiments on open‑source models, and without consistent GPU access, the main technical milestones cannot be completed on time. Dataset quality also remains a risk, especially around dialect accuracy and removing duplicate or code switched entries. If these issues are not corrected before training, the models may learn the wrong patterns and weaken the safety evaluation.
If the project fails, the most probable outcome is a narrower, slower version of the original plan the open source dataset, evaluation scripts, and fine tuning pipeline would still remain public on GitHub and Hugging Face. The documented evidence that frontier models give unsafe agricultural advice in local languages would also survive. In the worst case, the work becomes a foundation for another researcher or team to continue, and the safety problem it exposed remains visible to the wider AI community. So the failure would be incomplete execution, not a total loss of value.
I have raised $6,000 in the last 12 months this includes a $5,000 research grant from Thinking Machines and a $1,000 grant from an individual researcher secured through a cold email there's no institutional funding or venture capital has been raised so far