# Russian speech

**URL:** <https://discourse.mozilla.org/t/russian-speech/18572>\
**Category:** Common Voice\
**Tags:** sentence-collection\
**Created:** [August 30, 2017, 9:28am UTC](https://discourse.mozilla.org/t/russian-speech/18572 "2017-08-30T09:28:42Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Roman\_Frantov](https://avatars.discourse-cdn.com/v4/letter/r/df705f/32.png) [@Roman\_Frantov](https://discourse.mozilla.org/u/Roman_Frantov)\
**Post date:** [August 30, 2017, 9:28am UTC](https://discourse.mozilla.org/t/russian-speech/18572/1 "2017-08-30T09:28:42Z")

</div>

Hi there!  
Let’s start contributing Russian speech to the project!  
We need projects like this due to complexity of language. How can I be of help?

---

<div class="post-metadata">

**Author:** ![mhenretty](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/mhenretty/32/21321_2.png) [@mhenretty](https://discourse.mozilla.org/u/mhenretty)\
**Post date:** [September 25, 2017, 1:40pm UTC](https://discourse.mozilla.org/t/russian-speech/18572/2 "2017-09-25T13:40:21Z")

</div>

Hi Roman,

Thanks for your interest in this project!

Right now, beyond copying the code and setting up your own version of the Common Voice website with all the sentences translated, there is not much we have in place to help you. On the bright side, we plan on doing a big push for localization, and enabling more communities soon. Sadly, this won’t happen for another couple of months.

In the meantime, what you can do is look for a large collection of sentences in Russian that are part of the public domain. Once we have that, collecting voice samples is the easier part.

---

<div class="post-metadata">

**Author:** ![Roman\_Frantov](https://avatars.discourse-cdn.com/v4/letter/r/df705f/32.png) [@Roman\_Frantov](https://discourse.mozilla.org/u/Roman_Frantov)\
**Post date:** [September 26, 2017, 1:57pm UTC](https://discourse.mozilla.org/t/russian-speech/18572/3 "2017-09-26T13:57:10Z")

</div>

Hi Michael!  
Thank you for your reply!

I will be grateful for any help! I want to do a local instance of Common Voice and ready to invest in translating sentences.

How can we start?

BR

Roman

---

<div class="post-metadata">

**Author:** ![mhenretty](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/mhenretty/32/21321_2.png) [@mhenretty](https://discourse.mozilla.org/u/mhenretty)\
**Post date:** [September 26, 2017, 2:41pm UTC](https://discourse.mozilla.org/t/russian-speech/18572/4 "2017-09-26T14:41:14Z")

</div>

Localizing and spinning up a new instance of the Common Voice website is the easy part. The hard part is finding a large collection of public domain Russian sentences for your visitors to read. I’d start by looking for this.

Translating the English sentences we use isn’t a good approach because we don’t have enough sentences yet, and we want typical Russian way of talking (not translated Russian way of talking).

Once we find a large collection of these Russian sentences we can use, I can guide you through getting a copy of Common Voice up and running.

---

<div class="post-metadata">

**Author:** ![Roman\_Frantov](https://avatars.discourse-cdn.com/v4/letter/r/df705f/32.png) [@Roman\_Frantov](https://discourse.mozilla.org/u/Roman_Frantov)\
**Post date:** [September 26, 2017, 2:55pm UTC](https://discourse.mozilla.org/t/russian-speech/18572/5 "2017-09-26T14:55:18Z")

</div>

You’re right. The languages are so different that we can’t sometimes directly translate a sentence - we need to alternate the meaning a bit.  
Where did you get your sentences?

I believe forums and news are not the best idea, right?

Books? Call center calls?

I have an access to a quality monitoring recording of one Outsourcing contact center. But the speech there is quite boring 🙂

---

<div class="post-metadata">

**Author:** ![mhenretty](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/mhenretty/32/21321_2.png) [@mhenretty](https://discourse.mozilla.org/u/mhenretty)\
**Post date:** [September 27, 2017, 9:16pm UTC](https://discourse.mozilla.org/t/russian-speech/18572/7 "2017-09-27T21:16:14Z")

</div>

> [@Roman\_Frantov](#):
>
> Where did you get your sentences?

We have been collecting them from users lately, but in the past I have used public domain books and movies (War of the Worlds, It’s a Wonderful Life). The problem is that those texts are old and the language is outdated, so we are trying to get people to donate stuff they would say. It hasn’t been a very scaleable approach yet, but I think it’s a good direction.  
[https://github.com/mozilla/voice-web/issues/341](https://github.com/mozilla/voice-web/issues/341)

> [@](#):
>
> I believe forums and news are not the best idea, right?

It really depends. Either can be pretty technical and use a lot of proper nouns, which is less useful I think. But the right forum might have some good stuff, you just gotta find it.

> [@](#):
>
> Books? Call center calls?

Books also have the “not very conversational” problem. Movies scripts are better. Call center calls are really good, that’s how google trained their original engine.

> [@](#):
>
> I have an access to a quality monitoring recording of one Outsourcing contact center. But the speech there is quite boring 🙂

Is it public domain? License free? If so, this would be an amazing source granted we could use it.

---

<div class="post-metadata">

**Author:** ![Roman\_Frantov](https://avatars.discourse-cdn.com/v4/letter/r/df705f/32.png) [@Roman\_Frantov](https://discourse.mozilla.org/u/Roman_Frantov)\
**Post date:** [September 28, 2017, 2:43pm UTC](https://discourse.mozilla.org/t/russian-speech/18572/8 "2017-09-28T14:43:50Z")

</div>

I found great public domain collection of Russian sentences.  
[https://tatoeba.org/eng/](https://tatoeba.org/eng/)  
The whole DB can be downloaded here

[https://tatoeba.org/rus/terms\_of\_use](https://tatoeba.org/rus/terms_of_use)  
[https://tatoeba.org/rus/downloads](https://tatoeba.org/rus/downloads)

---

<div class="post-metadata">

**Author:** ![mhenretty](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/mhenretty/32/21321_2.png) [@mhenretty](https://discourse.mozilla.org/u/mhenretty)\
**Post date:** [September 29, 2017, 11:05am UTC](https://discourse.mozilla.org/t/russian-speech/18572/9 "2017-09-29T11:05:01Z")

</div>

That’s a great find @Roman_Frantov, and thanks for looking into this.

We have investigated Tatoeba in the past, and the problem is that their license CC-2, is not compatible with our CC-0 license. That project is amazing though, so I will reach out to them and see if we can work something out.

---

<div class="post-metadata">

**Author:** ![Roman\_Frantov](https://avatars.discourse-cdn.com/v4/letter/r/df705f/32.png) [@Roman\_Frantov](https://discourse.mozilla.org/u/Roman_Frantov)\
**Post date:** [October 20, 2017, 2:40pm UTC](https://discourse.mozilla.org/t/russian-speech/18572/10 "2017-10-20T14:40:12Z")

</div>

Sorry, can’t really tell the difference between cc-2 and cc-0 in this case. Can I be of help communicating with them or we shall find another corpus?

---

<div class="post-metadata">

**Author:** ![mhenretty](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/mhenretty/32/21321_2.png) [@mhenretty](https://discourse.mozilla.org/u/mhenretty)\
**Post date:** [October 23, 2017, 2:17pm UTC](https://discourse.mozilla.org/t/russian-speech/18572/11 "2017-10-23T14:17:54Z")

</div>

Thank you @Roman_Frantov for the offer to help! Are you part of the Tatoeba community?

We are already talking to Trang, the founder of Tatoeba, about collaboration. The conversations so far are sounding very positive, and we may start some sort of collaboration in early 2018.

---

<div class="post-metadata">

**Author:** ![Roman\_Frantov](https://avatars.discourse-cdn.com/v4/letter/r/df705f/32.png) [@Roman\_Frantov](https://discourse.mozilla.org/u/Roman_Frantov)\
**Post date:** [October 26, 2017, 10:14am UTC](https://discourse.mozilla.org/t/russian-speech/18572/12 "2017-10-26T10:14:32Z")

</div>

@mhenretty I joined the community recently. Good to hear that you’re moving on with them!

---

<div class="post-metadata">

**Author:** ![RLarissa](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/rlarissa/32/17882_2.png) [@RLarissa](https://discourse.mozilla.org/u/RLarissa)\
**Post date:** [December 5, 2017, 10:15pm UTC](https://discourse.mozilla.org/t/russian-speech/18572/13 "2017-12-05T22:15:54Z")

</div>

Hi! This is an interesting part. How to manage “kinda”, " dunno"? Another good case is " ain’t", different by nature though. Anyway, this is the easiest part. Voice music is more challenging.

---

<div class="post-metadata">

**Author:** ![RLarissa](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/rlarissa/32/17882_2.png) [@RLarissa](https://discourse.mozilla.org/u/RLarissa)\
**Post date:** [December 6, 2017, 4:03pm UTC](https://discourse.mozilla.org/t/russian-speech/18572/14 "2017-12-06T16:03:55Z")

</div>

My guessing is that professional assistance is required at least for cases ( nominative, dative etc.) of numerals. They are indeed complicated, some are being distorted in colloquial language. Personally, I wouldn’t decide in this capacity.

---

<div class="post-metadata">

**Author:** ![nikita.mekh](https://avatars.discourse-cdn.com/v4/letter/n/eada6e/32.png) [@nikita.mekh](https://discourse.mozilla.org/u/nikita.mekh)\
**Post date:** [June 11, 2018, 9:33am UTC](https://discourse.mozilla.org/t/russian-speech/18572/15 "2018-06-11T09:33:41Z")

</div>

Any news about russian speech? How can we help to start Russian localization?

---

<div class="post-metadata">

**Author:** ![luc.salommez](https://avatars.discourse-cdn.com/v4/letter/l/898d66/32.png) [@luc.salommez](https://discourse.mozilla.org/u/luc.salommez)\
**Post date:** [June 11, 2018, 8:24pm UTC](https://discourse.mozilla.org/t/russian-speech/18572/16 "2018-06-11T20:24:13Z")

</div>

Hi Nikita, before a language can be released it is needed to gather written sentences people will be able to speak.

You can contribute written sentences on this website : [https://voice-sprint.mozilla.community/](https://voice-sprint.mozilla.community/)  
Even though it is written “10-11th May”, the website still works and you can contribute there.

Once enough sentences will be gathered, the Russian language should be available for speech contribution.

Be aware though that the sentences you contribute must be free of copyright.

Thanks you for contributing to this project :).

---

<div class="post-metadata">

**Author:** ![nikita.mekh](https://avatars.discourse-cdn.com/v4/letter/n/eada6e/32.png) [@nikita.mekh](https://discourse.mozilla.org/u/nikita.mekh)\
**Post date:** [June 15, 2018, 1:36pm UTC](https://discourse.mozilla.org/t/russian-speech/18572/17 "2018-06-15T13:36:01Z")

</div>

Thanks. I will try to contribute as much as possible.  
How can I check current work process? How many sentences are already collected for russian? What is the goal? This information can be very useful for contributors

---

<div class="post-metadata">

**Author:** ![luc.salommez](https://avatars.discourse-cdn.com/v4/letter/l/898d66/32.png) [@luc.salommez](https://discourse.mozilla.org/u/luc.salommez)\
**Post date:** [June 15, 2018, 2:43pm UTC](https://discourse.mozilla.org/t/russian-speech/18572/18 "2018-06-15T14:43:19Z")

</div>

As a contributor myself I do not have access to those informations.  
In my opinion, as Russian is spoken by a lot of people it shouldn’t be too long before it is released.

Note though that you have to be on the russian version of the Common Voice website to access the russian contribution section for voice.

---

<div class="post-metadata">

**Author:** ![odinho](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/odinho/32/23800_2.png) [@odinho](https://discourse.mozilla.org/u/odinho)\
**Post date:** [June 23, 2018, 9:43pm UTC](https://discourse.mozilla.org/t/russian-speech/18572/19 "2018-06-23T21:43:26Z")

</div>

Elsewhere it was said that more sentences is better, but these things will allow a language to be initially released for starting to gather some recordings:

- the website fully translated on [http://pontoon.mozilla.org/](http://pontoon.mozilla.org/)
- at least 2000 sentences (and more in pipeline for review)

---

<div class="post-metadata">

**Author:** ![Orless](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/orless/32/30660_2.png) [@Orless](https://discourse.mozilla.org/u/Orless)\
**Post date:** [August 6, 2018, 10:18am UTC](https://discourse.mozilla.org/t/russian-speech/18572/20 "2018-08-06T10:18:39Z")

</div>

I also want to push Russian forward.

It’s a few week since the last reply, could we please recap what do we have to do to get started?

What I read so far:

- Translate the website [https://voice.mozilla.org](https://voice.mozilla.org) using [https://pontoon.mozilla.org/](https://pontoon.mozilla.org/). Not quite sure what is mean here, [https://voice.mozilla.org/ru](https://voice.mozilla.org/ru) seems to be translated.
- Deliver 2000+ Russian sentences (public domain texts only) - in which form, how exactly?

Is this correct? Could please someone from Mozille confirm this is the way to go and clarify how could we contribute texts or sentences.

There are many texts of classical Russian literature which are now in Public Domain. The language there may be not the most modert, but I think this is a good way to start with something.

---

<div class="post-metadata">

**Author:** ![mhenretty](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/mhenretty/32/21321_2.png) [@mhenretty](https://discourse.mozilla.org/u/mhenretty)\
**Post date:** [August 6, 2018, 4:14pm UTC](https://discourse.mozilla.org/t/russian-speech/18572/21 "2018-08-06T16:14:15Z")

</div>

The best way to contribute right now would be to find and review (or write) sentences in the public domain, and submit at PR to the main repo here: [common-voice/server/data at master · common-voice/common-voice · GitHub](https://github.com/mozilla/voice-web/tree/master/server/data)

Soon, though, we hope to have some better tools for reviewing and filtering bad sentence. See here for a discussion around that:

> [@We want your feedback: Improving the sentence collection](https://discourse.mozilla.org/t/we-want-your-feedback-improving-the-sentence-collection/30358):
>
> Hello everyone, I’m [Rubén](https://mozillians.org/u/nukeador/) and I’m working at the Mozilla Open Innovation Team. During this quarter I’ll be investing more of my time to work with @mhenretty to help the Common Voice project, specially around Community strategy. As we know, collecting sentences in different languages is an important step to advance the project, without valid sentences we won’t be able to offer something to read to people who want to donate their voice. In the past we have taken different approaches to solve t…

[Next page](https://discourse.mozilla.org/t/russian-speech/18572.md?page=2)
