Hey everyone, I’m Norbert, a Full Stack AI Engineer based in Uganda working across React, Next.js, and Tailwind CSS on the frontend, alongside Python, PyTorch, FastAPI, and Docker on the backend. I’m jumping into Common Voice to help advance voice AI for East African language specifically Kiswahili and Kinyarwanda by fine-tuning speech models (like Whisper) on local accents, building automated audio processing pipelines, and packaging low-latency inference APIs for low-resource environments. I’m eager to collaborate with the community, contribute to open datasets, and start taking on open GitHub issues immediately!
Hi @Norbert_Ndungutse, welcome to the Common Voice Community! The main way to contribute to the platform is through adding sentences, recording, or validating recordings. It looks like there are some tasks open for the Kinyarwanda dataset, which I can link to below.
Kiswahili is open for contribution on both Scripted Speech and Spontaneous Speech! Kinyarwanda is available for contribution on Scripted Speech, but if you’re a speaker, we would be delighted to have you open an issue here to get it added to Spontaneous Speech too!
As far as accent-specific tuning, it looks like dialects are specified for Swahili, so contributors can set their dialect in their profile. Kinyarwanda doesn’t have this feature enabled yet, but you can open a GitHub issue here to allow for selection of dialect.
If you’re interested in helping build community and organize events for speakers, you should apply to the Common Voice Volunteer Support Program here.
If you have any questions, feel free to post on this forum, email commonvoice@mozilla.com, or reach out to us on Discourse, Matrix, or Discord!