You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
AfriSafe-Eval will develop an open benchmark for evaluating the safety, robustness, and alignment of AI models in low-resource African languages, beginning with Hausa. Current AI safety evaluations are largely designed and tested in high-resource languages, creating a potential blind spot where models may behave differently or become less safe when interacting in African languages. The project will develop a curated Hausa-English evaluation dataset covering harmful requests, jailbreaks, misinformation, code-switching, transliteration, and other language-specific safety challenges. We will evaluate selected AI models, compare their safety performance across languages, identify systematic failure modes, and develop a reproducible scoring framework and open-source evaluation toolkit. The project will produce a public benchmark, technical report, model evaluation results, and safety failure taxonomy that can support researchers and AI developers in building safer multilingual systems. By starting with Hausa and designing the methodology for future expansion to other African languages, AfriSafe-Eval aims to address an underexplored gap in global AI safety research and improve our ability to measure whether AI safety generalizes beyond English.
Goal 1: Develop an African-language AI safety benchmark, beginning with Hausa.
We will create a structured benchmark for evaluating the safety, robustness, and alignment of AI models in Hausa and comparable low-resource settings. The benchmark will cover harmful requests, jailbreak attempts, misinformation, code-switching, transliteration, ambiguous prompts, and other language-specific safety challenges.
Goal 2: Measure whether AI safety generalizes across languages.
We will evaluate selected AI models using matched English and Hausa prompts to identify differences in refusal behavior, harmful-content generation, hallucination, instruction following, and other safety-relevant outcomes. This will help determine whether models that appear safe under English-language evaluations remain safe in Hausa.
Goal 3: Identify and characterize previously under-measured safety failures.
We will conduct systematic adversarial testing involving code-switching, Hausa-English mixing, transliteration, linguistic variation, and culturally contextualized scenarios. We will categorize observed failures into a reproducible taxonomy that can be used by future researchers.
Goal 4: Build an open and reproducible evaluation toolkit.
We will develop an evaluation pipeline and scoring framework that researchers can use to test AI models consistently. Where safety and licensing considerations permit, we will release the benchmark data, evaluation code, methodology, and documentation openly.
Goal 5: Generate evidence that can improve multilingual AI safety.
We will publish a technical report documenting the methodology, model results, and key failure modes, and disseminate the findings to AI safety researchers, model developers, and African technology and policy communities. The initial Hausa benchmark will also provide a foundation for extending the methodology to additional African languages.
The project will be implemented over six months through five stages: benchmark design, dataset development and expert annotation, model evaluation and adversarial testing, toolkit development, and publication/dissemination. The work will combine machine-learning expertise, Hausa-language knowledge, human review, and reproducible experimental methods. We will prioritize responsible research practices by avoiding unnecessary publication of harmful prompts and applying appropriate safeguards when constructing and releasing evaluation materials.
The requested funding will support a six-month research and development project to build, validate, and release the AfriSafe-Eval benchmark.
The funding will primarily be used for:
Research personnel: Support researchers and research assistants responsible for benchmark design, literature review, dataset development, experimentation, analysis, and documentation.
Dataset development and expert annotation: Develop Hausa-English evaluation cases, conduct human review, and ensure the benchmark accurately captures multilingual and culturally relevant safety challenges.
Computing and model evaluation: Cover cloud computing, GPU resources, and API costs required to evaluate selected AI models and conduct systematic safety and robustness experiments.
Adversarial testing: Support structured red-teaming to test jailbreaks, code-switching, transliteration, prompt manipulation, and other potential safety failures.
Benchmark and software development: Build the open-source evaluation toolkit, scoring framework, documentation, and reproducible evaluation pipeline.
Expert review: Engage AI safety, NLP, and Hausa-language experts to review the methodology, benchmark quality, and interpretation of results.
Publication and dissemination: Produce the technical report, research outputs, documentation, and public dissemination materials.
We will prioritize direct research outputs over administrative expenditure. The goal is for the funding to produce a reusable public-good research asset: an African-language AI safety benchmark, evaluation methodology, documented safety failure taxonomy, and open-source tools that other researchers can build upon.
Rufai Bello — Project Lead & Technical Researcher
Rufai Bello is a Data Scientist and AI researcher with 8+ years of experience applying machine learning, data science, NLP, and AI to real-world problems. He holds a B.Sc. in Computer Science from Bayero University Kano and an M.Sc. in Computer Science from Kaduna State University. His current doctoral research focuses on multimodal explainable fake-news detection for Hausa using cross-lingual contrastive learning and few-shot adaptation, providing directly relevant experience in low-resource language NLP, multilingual learning, multimodal AI, and misinformation detection. He has also led data science and technology projects involving government and community stakeholders.
Dr. Mustapha Lawal — Co-Founder & Research/Technical Advisor
Dr. Mustapha Lawal holds a Ph.D. in Computer Science with a specialization in Artificial Intelligence from Abubakar Tafawa Balewa University. He provides expertise in AI research, technical methodology, and research oversight, supporting the design and validation of the project's experimental approach.
Dr. Abdulmatalib Abdullahi — Co-Founder & Operations Lead
Dr. Abdulmatalib Abdullahi supports organizational coordination, partnerships, project implementation, and operational management, ensuring that the research team can execute the project efficiently and responsibly.
Organizational Track Record
Coded Foundation for Technology Development (CFTD) is a Nigerian technology nonprofit established to use technology, education, and innovation to address social and development challenges. The organization has implemented technology and digital-skills initiatives across Northern Nigeria, including coding education, ICT access, data and digital-literacy programmes, and technology-focused community projects.
While AfriSafe-Eval is a new research project rather than a continuation of an existing AI-safety programme, the team brings relevant prior experience in machine learning, Hausa NLP, multimodal AI, data science, research, software development, and technology project management. This combination gives us a strong foundation for executing the proposed benchmark and building the technical infrastructure required to make its results useful to the wider AI-safety community.
The most likely causes of failure are insufficient high-quality Hausa-language data, difficulties recruiting qualified language annotators and AI-safety reviewers, higher-than-expected model/API and computing costs, and challenges developing evaluation criteria that are both technically rigorous and culturally appropriate. There is also a risk that some models or providers may change access policies during the project, limiting reproducibility or the number of systems we can evaluate.
If these challenges are not adequately managed, the project may produce a smaller benchmark than planned, evaluate fewer models, or require additional time to achieve reliable results. In the worst case, we may be unable to establish sufficiently robust evidence about multilingual safety differences.
We will reduce these risks through staged development, expert review, pilot testing, careful dataset quality control, use of multiple model providers and open-weight models where possible, and prioritization of the core benchmark before secondary features. We will also document methodological limitations transparently rather than overstating findings.
Even if the full project does not succeed, the work should still produce useful outputs such as a documented research methodology, pilot evaluation data, identified safety failure modes, and lessons for future African-language AI-safety research. The primary failure we want to avoid is producing a benchmark that appears comprehensive but lacks sufficient validity or reproducibility; therefore, research quality and transparency will take priority over the number of languages, models, or examples covered.
Over the last 12 months, Coded Foundation for Technology Development has raised a limited amount of external funding. Our primary focus during this period has been developing and implementing projects while pursuing competitive grants and partnerships.
We have submitted applications and concept proposals to a range of international foundations, development organizations, and technology-focused funding programmes. However, we do not count unsuccessful applications, pending applications, or proposed grants as funds raised.
Our limited fundraising history is also why we are seeking Manifund support for AfriSafe-Eval. This project is designed as a focused, technically rigorous research initiative that can establish a stronger track record in AI safety research and generate open research infrastructure for the wider community.