Building an Arabic dataset for common voice

A quick update: We just uploaded 170 new Arabic (MSA) sentences to the sentence collector to be verified. These sentences were machine translated from the verified English corpus and verified for accuracy by a native speaker. So far 76% of the translations were accurate.

Please help review them here.

We have another 3000 sentences ready to go but need more volunteers: If you can, please open this spreadsheet and mark any sentence that is correct as ‘1’. We will upload verified sentences every few days.

I hope with this method we can get to 5,000 more quickly and start recording audio.

We should still collect sentences from other sources, especially colloquial / conversational speech, and phrases with non-MSA Arabic words.