# Russian speech

**URL:** <https://discourse.mozilla.org/t/russian-speech/18572>\
**Category:** Common Voice\
**Tags:** sentence-collection\
**Created:** [August 30, 2017, 9:28am UTC](https://discourse.mozilla.org/t/russian-speech/18572 "2017-08-30T09:28:42Z")\
**Posts on this page:** 1\
**Showing post:** 27

<div class="post-metadata">

**Author:** ![pdsfjd](https://avatars.discourse-cdn.com/v4/letter/p/a9adbd/32.png) [@pdsfjd](https://discourse.mozilla.org/u/pdsfjd)\
**Post date:** [March 4, 2019, 8:52am UTC](https://discourse.mozilla.org/t/russian-speech/18572/27 "2019-03-04T08:52:00Z")

</div>

And I found another big source for Russian sentences. United Nations documents published under PD, and there is already tagged corpus available [https://cms.unov.org/UNCorpus/](https://cms.unov.org/UNCorpus/) I wrote a script that extract only proces-verbaux (transcripts) records from corpus and validate it using same method sentence-collector use, and got more than 300k unique sentences. Not all sentences are good, so they need to be validated by human additionally. Should I upload it to sentence-collector fully?

It also can be useful for other United Nations official languages (Arabic, English, Spanish, French, Chinese).

---

_[View the full topic](https://discourse.mozilla.org/t/russian-speech/18572)._
