You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Most LLM safety filters are red teamed in standard Western English. When users prompt models using code switched vernacular or regional dialects(Nigerian pidgin), or mixed languages, models frequently degrade or fail completely. This project builds an open-source evaluation suite to benchmark and log these dialect based safety bypass across models like Llama 3, Claude 3.5, and GPT-4o.
My goal is to publish an open source red teaming dataset on Github and an empirical findings report on LessWrong/Alignment Forum detailing exact guardrail failure rates. I will curate 100+ standard harmful test prompts, convert them into authentic code switched dialect variations, run them systematically through model APIs, and document refusal vs compliance rates.
The $10,000 will be deployed directly into testing infrastructure, data annotation, and research execution over a 6 months period.
$3500: API & Compute Infrastructure - to run automated, high volume red teaming evaluations across OpenAI (GPT-4o), Anthropic (Claude 3.5), Together AI (Llama 3 70B/405B), and local open source inference endpoints.
$1500: Data Annotation & Translation - to verify authentic dialect, code switched (Nigerian Pidgin), and regional nuance accuracy across the 100+ prompt test suite.
$5000: Technical Hardware & Execution Operations - to get a durable primary development machine setup for local model evaluations, testing workspace operations, data pipeline hosting, and direct research execution (publication).
I am an independent tech founder and student developer based in Nigeria leading a small technical team. We have experience building functional AI tools, working with model APIs, and deploying software projects under tight constraints. Our native linguistic context gives us an advantage in creating authentic code switched test suites that western AI safety labs overlook.
The primary risk is that frontier models have already improved dialect alignment better than expected, resulting in low guardrail bypass rates. If so, the outcome is still valuable in the sense that, we will publish the benchmark dataset proving model robustness in regional dialects.
$0 (Self funded ).