I started to check audio clips in Russian Scripted Speech tab. And about 40-50% of them from 100-200 I checked were just synthesized voices reading these sentences. Is it health practice and should be accepted?
Hi @Libra. Thank you for bringing this to our attention. Synthesized voices are definitely not welcome in the dataset. If you come across one, please vote no on the recording. I’ll speak to the team about some approaches we can try for preventing contributions of synthesized voices.
If you have some ideas for how to address it, we always welcome suggestions, either here or especially through opening an issue on the Common Voice GitHub. You can find issue templates here.