Building a training data-set of kids voices

Hi,

I see that is not easy to sketch a suitable process that integrate under 19 years old contribute, compliant with legal terms (see also Bad words words list for your languages ), but I don’t see the “children’s voice” as a so specific / different realm, if our goal is to achieve high quality, including “diversities” in common spoken “language model” definition.

Under 19 are people as adults… and excluding their contribution will build a biased dataset. That’s bad, immo.