Common Voice v1 corpus design problems, overlapping train/test/dev sentences

(Copying my response from the github issue)

Hi @bmilde,

First of all, thank you for reporting this bug. It is indeed very critical. One of the reasons we wanted to release this data so quickly was to get this kind of feedback from people like you, so bravo!

I have spoken about this split with our machine learning group (which is different than the Common Voice team that I am a part of), and there are a couple of solutions we are investigating.

One, we can redo the split between dev/train/test to make sure there are no overlapping speakers or sentences. The problem with this is that the dev and test datasets will probably need to become a lot smaller due to our limited sentence corpus.

Another approach is to modify our Common Voice server (ie. this repo), to have special sentences that are in quarantined to the test/dev sets, and then make certain users only get those sentences for reading. This is a better approach in the long term, since it means we could grow the test/dev sets larger, and wouldn’t have to worry about throwing out any training data (again due to our small corpus size).

We will be investigating the above approaches in the coming weeks, and will definitely fix this in the next release of the data (v2). In the meantime, any advice or information you (or anyone reading this) has in regards to this problem would be very welcomed.