# Building an Arabic dataset for common voice

**URL:** https://discourse.mozilla.org/t/building-an-arabic-dataset-for-common-voice/29728
**Category:** Common Voice
**Tags:** sentence-collection
**Created:** [June 29, 2018, 11:24am UTC](https://discourse.mozilla.org/t/building-an-arabic-dataset-for-common-voice/29728 "2018-06-29T11:24:37Z")
**Posts on this page:** 1
**Showing post:** 11

<div class="post-metadata">

### Author: ![tinok](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/tinok/32/25881_2.png) [@tinok](https://discourse.mozilla.org/u/tinok)
#### Post date: [February 23, 2019, 7:29pm UTC](https://discourse.mozilla.org/t/building-an-arabic-dataset-for-common-voice/29728/11 "2019-02-23T19:29:12Z")

</div>

A quick update: We just uploaded 170 new Arabic (MSA) sentences to the sentence collector to be verified. These sentences were machine translated from the verified English corpus and verified for accuracy by a native speaker. So far 76% of the translations were accurate.

Please [help review them here](https://common-voice.github.io/sentence-collector/#/review/ar).

We have another 3000 sentences ready to go but need more volunteers: If you can, please open [this spreadsheet](https://docs.google.com/spreadsheets/d/1cS3-FwuW9hLQpNQq6WufOrpL_g-tq9ERKVcrMawMdF4/edit?usp=sharing) and mark any sentence that is correct as ‘1’. We will upload verified sentences every few days.

I hope with this method we can get to 5,000 more quickly and start recording audio.

We should still collect sentences from other sources, especially colloquial / conversational speech, and phrases with non-MSA Arabic words.

---

_[View the full topic](https://discourse.mozilla.org/t/building-an-arabic-dataset-for-common-voice/29728)._
