# Building an Arabic dataset for common voice

**URL:** <https://discourse.mozilla.org/t/building-an-arabic-dataset-for-common-voice/29728>\
**Category:** Common Voice\
**Tags:** sentence-collection\
**Created:** [June 29, 2018, 11:24am UTC](https://discourse.mozilla.org/t/building-an-arabic-dataset-for-common-voice/29728 "2018-06-29T11:24:37Z")\
**Posts on this page:** 1\
**Showing post:** 5

<div class="post-metadata">

**Author:** ![tinok](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/tinok/32/25881_2.png) [@tinok](https://discourse.mozilla.org/u/tinok)\
**Post date:** [January 30, 2019, 4:24pm UTC](https://discourse.mozilla.org/t/building-an-arabic-dataset-for-common-voice/29728/5 "2019-01-30T16:24:16Z")

</div>

Hi all, we are looking for collaborators to build out the various Arabic datasets, but we need to address some major questions first. So far we have [92%](https://pontoon.mozilla.org/ar/common-voice/) of interface localized and exactly 0 sentences submitted to [Sentence Collector](https://common-voice.github.io/sentence-collector/).

If I’m not mistaken, we can now start the process of adding sentences even while the UI localization is not complete (though only 33 strings are missing).

But, a _major question_ needs to be decided on at this point: There is only one “Arabic” in Common Voice right now. Does this refer to Modern Standard Arabic (MSA)? If the goal of this platform is to “help teach machines how real people speak” then we need the different colloquial Arabic variants separate from MSA, i.e. ar-SA, ar-LB, etc. Google’s STT engine [claims to support](https://cloud.google.com/speech-to-text/docs/languages) 15 different types of Arabic. As anyone with some knowledge of Arabic knows, [both vocabulary and pronunciation vary greatly between different countries](https://en.wikipedia.org/wiki/Varieties_of_Arabic) (and yes, sometimes within them). These different types of locally spoken Arabic are often considered as separate languages and should not be seen as equivalent to Irish, American, or Australian English.

So, I suggest renaming “Arabic” in the sentence collector to “Arabic (MSA)” and adding relevant locally spoken Arabic as separate languages.

For the purpose of language collection, it may make sense to start with MSA and then “translate” sentences to reflect local vocabulary, as needed.

But whereas it may be tempting to focus on MSA at first, it is not a language spoken naturally between most Arabic speakers, so for the purpose of creating a useful STT engine, MSA may not have much value.

I’d love other Arabic speakers, especially people with linguistics and translation backgrounds to weigh in.

---

_[View the full topic](https://discourse.mozilla.org/t/building-an-arabic-dataset-for-common-voice/29728)._
