# Common voice sentences are the opposite of "common"

**URL:** <https://discourse.mozilla.org/t/common-voice-sentences-are-the-opposite-of-common/52345>\
**Category:** Common Voice\
**Tags:** participation, sentence-collection, feedback, issue\
**Created:** [January 20, 2020, 1:41pm UTC](https://discourse.mozilla.org/t/common-voice-sentences-are-the-opposite-of-common/52345 "2020-01-20T13:41:19Z")\
**Posts on this page:** 8\
**Page:** 2

<div class="post-metadata">

**Author:** ![nukeador](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/nukeador/32/20533_2.png) [@nukeador](https://discourse.mozilla.org/u/nukeador)\
**Post date:** [February 3, 2020, 9:56pm UTC](https://discourse.mozilla.org/t/common-voice-sentences-are-the-opposite-of-common/52345/21 "2020-02-03T21:56:11Z")

</div>

Children voices is currently a legal limitation. About the rest, I agree, there are currently some of this open questions we should evaluate when checking the quality of our dataset as well as how models trained with it perform in the real world.

/cc @rosana because of the insightful comments about quality and diversity.

---

<div class="post-metadata">

**Author:** ![david-song](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/david-song/32/33437_2.png) [@david-song](https://discourse.mozilla.org/u/david-song)\
**Post date:** [February 3, 2020, 10:05pm UTC](https://discourse.mozilla.org/t/common-voice-sentences-are-the-opposite-of-common/52345/22 "2020-02-03T22:05:30Z")

</div>

What’s the legal issue with using children’s voices? Is it that you need permission from a third party (legal guardian)? Would it not be possible to get schools onboard, get them to use a separate site to submit clips? Or can the recognition models be composed, so the actual voice data isn’t shared, only the histograms that form the training data? That would remove GDPR/privacy liabilities, if not copyright ones.

---

<div class="post-metadata">

**Author:** ![dabinat](https://avatars.discourse-cdn.com/v4/letter/d/8dc957/32.png) [@dabinat](https://discourse.mozilla.org/u/dabinat)\
**Post date:** [February 4, 2020, 9:39am UTC](https://discourse.mozilla.org/t/common-voice-sentences-are-the-opposite-of-common/52345/23 "2020-02-04T09:39:47Z")

</div>

In many countries the data of minors has more legal restrictions than those of adults.

By the way, I have an ongoing project to clean up the English wiki sentences, focusing mainly on removing difficult-to-pronounce foreign and scientific terms. Although I’ve really only just scratched the surface, it is slowly improving things.

---

<div class="post-metadata">

**Author:** ![nukeador](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/nukeador/32/20533_2.png) [@nukeador](https://discourse.mozilla.org/u/nukeador)\
**Post date:** [February 4, 2020, 10:15am UTC](https://discourse.mozilla.org/t/common-voice-sentences-are-the-opposite-of-common/52345/24 "2020-02-04T10:15:23Z")

</div>

See this topic for reference

> [@Building a training data-set of kids voices](https://discourse.mozilla.org/t/building-a-training-data-set-of-kids-voices/28028/):
>
> We are building an educational platform for economically disadvantaged kids aged 4 - 6 and are planning on incorporating Common Voice into it in order to help kids improve their English reading skills. Initially, we plan to build a game where the child reads out individual words of a story, and our game gives feedback on whether or not the child has pronounced them correctly. Before we do that, we obviously need a training data-set for kids voices. We have the ability to collect kids voices,…

---

<div class="post-metadata">

**Author:** ![Filippo\_Davalli](https://avatars.discourse-cdn.com/v4/letter/f/f4b2a3/32.png) [@Filippo\_Davalli](https://discourse.mozilla.org/u/Filippo_Davalli)\
**Post date:** [February 12, 2020, 1:50pm UTC](https://discourse.mozilla.org/t/common-voice-sentences-are-the-opposite-of-common/52345/25 "2020-02-12T13:50:39Z")

</div>

Here a true idealistic proposal for easy collecting common voice sentences.  
Think about sharing Alexa voice history.  
Ask Amazon through [change.org](http://change.org) to add a button “I want to donate my voice history to CommonVoice” or at least “download the archive of my voice history”.  
Now you can actually only delete or play the history.  
Is it possible to realize a browser addon able to download voice data from alexa account?

---

<div class="post-metadata">

**Author:** ![cjbaker](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/cjbaker/32/72432_2.png) [@cjbaker](https://discourse.mozilla.org/u/cjbaker)\
**Post date:** [February 12, 2020, 8:16pm UTC](https://discourse.mozilla.org/t/common-voice-sentences-are-the-opposite-of-common/52345/26 "2020-02-12T20:16:19Z")

</div>

Our current data is what’s called a “read speech corpus”, because each sentence is read from a prompt. It can also be very useful to have a “spontaneous speech corpus”. In this case, the speech is produced spontaneously, and the transcription is created later on by listening to it.

I have long thought that that the creation of a spontaneous corpus would also be an ideal application for crowd-sourcing, not sure if it could one day be included in the scope of this project. You would initially contribute a recording of your voice, then other users would verify and transcribe it. Your Alexa voice history might be good for that, or also just recording your side of (phone) conversations, etc.

---

<div class="post-metadata">

**Author:** ![twinfrosty](https://avatars.discourse-cdn.com/v4/letter/t/ee7513/32.png) [@twinfrosty](https://discourse.mozilla.org/u/twinfrosty)\
**Post date:** [September 7, 2024, 1:59am UTC](https://discourse.mozilla.org/t/common-voice-sentences-are-the-opposite-of-common/52345/28 "2024-09-07T01:59:15Z")

</div>

You can get lots of casual english sentences from fandom wiki, which uses CC-SA. They’re mostly on media/entertainment topics.  
[https://www.fandom.com/licensing](https://www.fandom.com/licensing)

---

<div class="post-metadata">

**Author:** ![bozden](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/bozden/32/51381_2.png) [@bozden](https://discourse.mozilla.org/u/bozden)\
**Post date:** [September 7, 2024, 9:19am UTC](https://discourse.mozilla.org/t/common-voice-sentences-are-the-opposite-of-common/52345/29 "2024-09-07T09:19:39Z")

</div>

> [@twinfrosty](#):
>
> which uses CC-SA

Hey @twinfrosty, welcome. In Common Voice, only CC-0 sentences are allowed.

But a new Spontaneous Speech dataset creation application is under development, where people give spontaneous answers to prompts and then transcribe it.

[Previous page](https://discourse.mozilla.org/t/common-voice-sentences-are-the-opposite-of-common/52345.md?page=1)
