You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Almost all AI safety testing happens in Standard English, but in West Africa, people mix English with Nigerian Pidgin, local slang, and regional dialects when using LLMs. I want to test if frontier models like GPT-4o, Claude 3.5, and Llama 3 drop their safety guardrails when prompted in these local dialects. Since safety classifiers were never fine-tuned on regional vernacular, guardrail bypasses are likely taking place completely unnoticed. I am creating a dataset of 100+ harmful prompts translated into natural Pidgin and code-switched variations, testing them against model APIs, logging refusal rates, and publishing the benchmark and write-up on GitHub and LessWrong.
My main goal is to build and publish an open-source red-teaming dataset on GitHub alongside a detailed empirical report on LessWrong and the Alignment Forum showing exact guardrail failure rates. To achieve this, we will write 100+ standard harmful safety prompts covering risk areas like social engineering and unsafe code. Next, native speakers will adapt them into authentic Nigerian Pidgin and regional code-switched variations. We will run both the standard English baseline and the dialect prompts through OpenAI, Anthropic, and Together AI endpoints to log refusal versus compliance rates, analyze the safety degradation gap, and publish the full benchmark openly so alignment teams can fix these blind spots.
The $10,000 budget will cover a 6-month timeline focused on compute, prompt verification, and operational execution. Specifically, $3,500 goes toward API credits and compute to run automated evaluations across OpenAI, Anthropic, and Together AI endpoints. Another $1,500 goes toward paying local native speakers to verify that code-switched Pidgin prompts reflect natural, real-world phrasing rather than machine translation. Finally, $5,000 covers primary development hardware setup, testing pipeline hosting, and my execution time so I can focus on this research full-time through final publication.
I am an independent software developer and researcher based in Nigeria leading a small technical team. We have hands-on experience building functional AI micro-tools, integrating model APIs, and deploying software projects under tight constraints. Our key advantage is native context, Western AI safety labs evaluate multilingual alignment using automated translation tools like Google Translate, which miss conversational slang, regional code-switching, and natural sentence structures. Living in this environment every day lets us construct and verify authentic red-teaming prompts that automated Western translation workflows miss entirely.
The main risk is that frontier models end up being more robust at dialect alignment than expected, resulting in low guardrail bypass rates across our test suite. If that happens, the outcome is still valuable to the AI safety community because we will publish the benchmark dataset on GitHub and our write-up on LessWrong as empirical proof that frontier model guardrails remain robust when exposed to non-Western dialects and code-switching.
$0 (Self funded ).