OK, I downloaded CV English just to check it. There are 8396 audio clips and 2800 of them are reported (>2100 because of incorrect language). Great! High quality dataset /s
I’m not ready to waste my time for English dataset anymore, because it is abnormal situation that I need to report 2/3 audios. I even counted it… It is just wasting of my or of someone’s else time.
Even random 5 examples of questions and answers have 1 question and answer in a language other than English.
I guess you downloaded the “Spontaneous Speech” dataset. That’s a newer approach which has only recently been introduced, and it currently has some UI problems that caused contributors to submit recordings under the wrong locale.
You may want to try the “Scripted Speech” dataset instead. It contains around 1 million clips, including accent metadata:
Hello @Libra. We appreciate your feedback. If you are looking for a cleaner English dataset, I would echo Irvin’s suggestion to use the Scripted Speech dataset. The current Spontaneous Speech interface, wherein English is the default page, has resulted in frequent non-English contributions or otherwise problematic contributions to the English dataset. Additionally, some contributors seem to have confused the Spontaneous Speech platform with Scripted Speech, so you occasionally find recordings where they are reading the question instead of answering, resulting in the problems you noted in your post.
We are currently working through a series of methods to clean the English dataset on our end, but as there are many recordings, the process is slow. Previous version of Spontaneous Speech have not released the English dataset due to these issues.
However, holding back the dataset means that any useful contributions are also held back. While we work on a better solution, we have decided to release this version with quality tags in column S of the accompanying tsv. The quality tags allow for filtering out likely problematic recordings without requiring us to withhold or delete recordings without review. Further, using the prompt_upvotes column and prompt_reports columns should get you to a cleaner dataset.
As for the datasheet, sample sentences are randomly selected, so the inclusion of non-English samples is reflective of the state of the dataset and unavoidable when randomly sampling a dataset that uses quality tags instead of excluding recordings.
Use of English Spontaneous Speech certainly requires significant filtering/dataset cleaning, which is not a task all users will be up to. We hope future releases will produce a clean, quality dataset.