# Spanish dataset

**URL:** <https://discourse.mozilla.org/t/spanish-dataset/34292>\
**Category:** Common Voice\
**Tags:** sentence-collection\
**Created:** [December 14, 2018, 6:20pm UTC](https://discourse.mozilla.org/t/spanish-dataset/34292 "2018-12-14T18:20:03Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![mar\_martinez](https://avatars.discourse-cdn.com/v4/letter/m/db5fbb/32.png) [@mar\_martinez](https://discourse.mozilla.org/u/mar_martinez)\
**Post date:** [December 14, 2018, 6:20pm UTC](https://discourse.mozilla.org/t/spanish-dataset/34292/1 "2018-12-14T18:20:03Z")

</div>

Hi,

I would like to contribute to start collecting spanish voices, how can I do it?.  
How could I contribute, for example, with sentences (0/5000)?.

Best regards,  
Mar

---

<div class="post-metadata">

**Author:** ![carlfm01](https://avatars.discourse-cdn.com/v4/letter/c/ec9cab/32.png) [@carlfm01](https://discourse.mozilla.org/u/carlfm01)\
**Post date:** [December 14, 2018, 9:47pm UTC](https://discourse.mozilla.org/t/spanish-dataset/34292/2 "2018-12-14T21:47:21Z")

</div>

I was about to ask the same, how do we enable Spanish in the portal, I can can contribute too 🙂

---

<div class="post-metadata">

**Author:** ![carlfm01](https://avatars.discourse-cdn.com/v4/letter/c/ec9cab/32.png) [@carlfm01](https://discourse.mozilla.org/u/carlfm01)\
**Post date:** [December 15, 2018, 7:23am UTC](https://discourse.mozilla.org/t/spanish-dataset/34292/3 "2018-12-15T07:23:44Z")

</div>

> [@open\_book Readme: How to see my language on Common Voice](https://discourse.mozilla.org/t/readme-how-to-see-my-language-on-common-voice/31530):
>
> triangular_flag_on_post This information is also now available on the [About Pages](https://commonvoice.mozilla.org/about) on Common Voice Website. Please help us to localise this by joining [Pontoon](https://pontoon.mozilla.org/projects/common-voice/)open_book [Mozilla Voice Community Playbook](https://common-voice.github.io/community-playbook/): The source of truth for setting up and maintain self-sustainable communities. Hello everyone, I would like to open this topic to summarize some of the most asked question we are getting: How do I get my language in Common Voice. There are three steps to have your language ready: …

---

<div class="post-metadata">

**Author:** ![mar\_martinez](https://avatars.discourse-cdn.com/v4/letter/m/db5fbb/32.png) [@mar\_martinez](https://discourse.mozilla.org/u/mar_martinez)\
**Post date:** [December 15, 2018, 9:35am UTC](https://discourse.mozilla.org/t/spanish-dataset/34292/4 "2018-12-15T09:35:02Z")

</div>

Yes, thanks Carlos.  
The step 3, the one I am concerned about, is… still blocked?

How did the other languages in use managed? are there alternatives?

---

<div class="post-metadata">

**Author:** ![nukeador](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/nukeador/32/20533_2.png) [@nukeador](https://discourse.mozilla.org/u/nukeador)\
**Post date:** [December 17, 2018, 5:13pm UTC](https://discourse.mozilla.org/t/spanish-dataset/34292/5 "2018-12-17T17:13:21Z")

</div>

Hi,

We are still finishing the sentence collection tool:

> [@Sentence collection tool development topic](https://discourse.mozilla.org/t/sentence-collection-tool-development-topic/33390):
>
> Continuing the discussion from [We want your feedback: Improving the sentence collection](https://discourse.mozilla.org/t/we-want-your-feedback-improving-the-sentence-collection/30358/36): Hi all, This topic is aimed just to developers who would like to help (react and kinto skills required) [Source code](https://github.com/Common-Voice/sentence-collector)[Live site](https://common-voice.github.io/sentence-collector/) What do we need? Fork the project and test that you can run the environment locally following the instructions. Is everything working as expected? If not, submit [a new issue](https://github.com/Common-Voice/sentence-collector/issues/new). Review the pending issues on the next [milestone](https://github.com/Common-Voice/sentence-collector/milestones). Create a [new PR](https://github.com/Common-Voice/sentence-collector/compare) to fix any of the existing issues in t…

Ideally we would have a beta version to test before the end of the year, if the QA of that beta is satisfactory we can start using it to collect and review sentences.

And I have to say I understand it’s frustrating but please, keep collecting sentences so we can submit them through the tool as soon as it’s ready.

Gracias por vuestras paciencia 😉

---

<div class="post-metadata">

**Author:** ![fatimaig](https://avatars.discourse-cdn.com/v4/letter/f/3ab097/32.png) [@fatimaig](https://discourse.mozilla.org/u/fatimaig)\
**Post date:** [December 26, 2018, 5:40pm UTC](https://discourse.mozilla.org/t/spanish-dataset/34292/6 "2018-12-26T17:40:36Z")

</div>

And where do we contribute with spanish sentences?  
Me and some colleges from the University would like to colaborate in the Spanish part of this project.  
Ty 😉

---

<div class="post-metadata">

**Author:** ![nukeador](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/nukeador/32/20533_2.png) [@nukeador](https://discourse.mozilla.org/u/nukeador)\
**Post date:** [December 26, 2018, 5:45pm UTC](https://discourse.mozilla.org/t/spanish-dataset/34292/7 "2018-12-26T17:45:02Z")

</div>

> [@fatimaig](#):
>
> And where do we contribute with spanish sentences?  
> Me and some colleges from the University would like to colaborate in the Spanish part of this project.

You can start collecting sentences with public domain license anytime and use any form to store them in the meantime. As soon as the tool is ready you will be able to submit them for peer-review and approval.

Please, note that we want the sentence collection tool to [enforce some hard requirements](https://github.com/Common-Voice/sentence-collector/issues/33) that are necessary for sentences to be useful for the machine learning algorithm:

- Numbers. There should be no digits in the source text because they can cause problems when read aloud. The way a number is read depends on context and might introduce confusion in the dataset. For example, the number “2409” could be accurately read as both “twenty-four zero nine” and “two thousand four hundred nine”.
- Abbreviations and Acronyms. Abbreviations and acronyms like “USA” or “ICE” should be avoided in the source text because they may be read in a way that does not coincide with their spelling. Additionally, there may be multiple accurate readings for a single abbreviation. For example, the acronym “ICE” could be pronounced “I-C-E” or as a single word.
- Punctuation. Special symbols and punctuation should only be included when absolutely necessary. For example, an apostrophe is included in English words like “don’t” and “we’re” and should be included in the source text, but it’s unlikely you’ll ever need a special symbol like “@” or “#.”
- Foreign letters. Letters must be valid in the language being spoken. For example, “ж” is a letter in the Russian alphabet but is never used in English and so should never appear in any English source text.

---

<div class="post-metadata">

**Author:** ![mar\_martinez](https://avatars.discourse-cdn.com/v4/letter/m/db5fbb/32.png) [@mar\_martinez](https://discourse.mozilla.org/u/mar_martinez)\
**Post date:** [January 25, 2019, 12:40pm UTC](https://discourse.mozilla.org/t/spanish-dataset/34292/8 "2019-01-25T12:40:05Z")

</div>

Hi,

Any update about the sentence collection tool availability for Spanish?

Thanks,  
Mar

---

<div class="post-metadata">

**Author:** ![nukeador](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/nukeador/32/20533_2.png) [@nukeador](https://discourse.mozilla.org/u/nukeador)\
**Post date:** [January 25, 2019, 12:46pm UTC](https://discourse.mozilla.org/t/spanish-dataset/34292/9 "2019-01-25T12:46:02Z")

</div>

We plan to launch the beta version of the tool next week.

You can follow the development in [Sentence collection tool development topic](https://discourse.mozilla.org/t/sentence-collection-tool-development-topic/33390)

---

<div class="post-metadata">

**Author:** ![mar\_martinez](https://avatars.discourse-cdn.com/v4/letter/m/db5fbb/32.png) [@mar\_martinez](https://discourse.mozilla.org/u/mar_martinez)\
**Post date:** [January 30, 2019, 6:18pm UTC](https://discourse.mozilla.org/t/spanish-dataset/34292/10 "2019-01-30T18:18:28Z")

</div>

Hi,

The sentence collection tool is ready and now I am adding and reviewing sentences in Spanish. Great.

But the Common Voice site localization in Spanish is never ending, but indeed it is getting worse (from 95% completion in December to 75% now). Where can I fix this?. Directly with a github pull request or any other tool?.

Tanks,  
Mar

---

<div class="post-metadata">

**Author:** ![nukeador](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/nukeador/32/20533_2.png) [@nukeador](https://discourse.mozilla.org/u/nukeador)\
**Post date:** [January 30, 2019, 6:31pm UTC](https://discourse.mozilla.org/t/spanish-dataset/34292/11 "2019-01-30T18:31:00Z")

</div>

> [@mar\_martinez](#):
>
> But the Common Voice site localization in Spanish is never ending, but indeed it is getting worse (from 95% completion in December to 75% now). Where can I fix this?. Directly with a github pull request or any other tool?.

Website localization is handled via pontoon:

> **[Common Voice · Spanish (es)](https://pontoon.mozilla.org/es/common-voice/)**
>
> Mozilla’s Localization Platform

You can ping people with reviewer rights in [this telegram channel](https://t.me/mozilla_l10n_es) or [this forum](https://foro.mozilla-hispano.org/c/traduccion).

---

<div class="post-metadata">

**Author:** ![mar\_martinez](https://avatars.discourse-cdn.com/v4/letter/m/db5fbb/32.png) [@mar\_martinez](https://discourse.mozilla.org/u/mar_martinez)\
**Post date:** [January 30, 2019, 7:25pm UTC](https://discourse.mozilla.org/t/spanish-dataset/34292/12 "2019-01-30T19:25:05Z")

</div>

Thanks a lot,

On the other hand, to fulfill the accent list, I assume that is required a github pull request, the required list will be roughly by countries (e.g. Español de México, Español de Guatemala, Español de España…) or with local accents inside each country?.

Regards,  
Mar

---

<div class="post-metadata">

**Author:** ![nukeador](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/nukeador/32/20533_2.png) [@nukeador](https://discourse.mozilla.org/u/nukeador)\
**Post date:** [January 30, 2019, 8:34pm UTC](https://discourse.mozilla.org/t/spanish-dataset/34292/13 "2019-01-30T20:34:22Z")

</div>

> [@mar\_martinez](#):
>
> On the other hand, to fulfill the accent list, I assume that is required a github pull request, the required list will be roughly by countries (e.g. Español de México, Español de Guatemala, Español de España…) or with local accents inside each country?.

We should probably follow and official list of accents. @josh_meyer how did we get the English one?

---

<div class="post-metadata">

**Author:** ![josh\_meyer](https://avatars.discourse-cdn.com/v4/letter/j/bc8723/32.png) [@josh\_meyer](https://discourse.mozilla.org/u/josh_meyer)\
**Post date:** [January 30, 2019, 10:29pm UTC](https://discourse.mozilla.org/t/spanish-dataset/34292/14 "2019-01-30T22:29:23Z")

</div>

I dont know how we got English accents, but I’ll ask around.

---

<div class="post-metadata">

**Author:** ![nukeador](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/nukeador/32/20533_2.png) [@nukeador](https://discourse.mozilla.org/u/nukeador)\
**Post date:** [February 1, 2019, 1:02am UTC](https://discourse.mozilla.org/t/spanish-dataset/34292/15 "2019-02-01T01:02:16Z")

</div>

@mar_martinez @carlfm01 @fatimaig I see a lot of activity for Spanish today 😃

Just noticed an old book with weird old language and some subtitles that have incomplete sentences, I hope community vote negative but we should probably warn people before uploading thousands of sentences without even checking them.

I’ve added a few tips on how to get a lot of valid sentences reusing Catalan previous work here

> **[Nueva herramienta para mandar frases a Common Voice](https://foro.mozilla-hispano.org/t/nueva-herramienta-para-mandar-frases-a-common-voice/24312/2)**
>
> Una cosa que me he dado cuenta es que en Catalán ya tienen miles y miles de frases en Common Voice. Como se que las herramientas de traducción hacen un buen trabajo español→catalán he probado una cosa. He tomado unos de los archivos ya validados...

---

<div class="post-metadata">

**Author:** ![daniel.cruzado](https://avatars.discourse-cdn.com/v4/letter/d/9de0a6/32.png) [@daniel.cruzado](https://discourse.mozilla.org/u/daniel.cruzado)\
**Post date:** [April 3, 2019, 7:43am UTC](https://discourse.mozilla.org/t/spanish-dataset/34292/16 "2019-04-03T07:43:43Z")

</div>

Hi, I have seen that Spanish already has 14 hours, that is far more than for example than Breton or Irish, but Spanish dataset is not available for download.

Do we know when will it be ready?

---

<div class="post-metadata">

**Author:** ![nukeador](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/nukeador/32/20533_2.png) [@nukeador](https://discourse.mozilla.org/u/nukeador)\
**Post date:** [April 3, 2019, 10:19am UTC](https://discourse.mozilla.org/t/spanish-dataset/34292/17 "2019-04-03T10:19:11Z")

</div>

Please read this explanation about the current dataset release process:

> [@Add Basque to the dataset page](https://discourse.mozilla.org/t/add-basque-to-the-dataset-page/37609/2):
>
> The languages currently available for download are the ones that last year were already launched and had validated hours before we started the dataset creation later last year (Basque or Spanish for example were not even launched by that time) Dataset releases are really time consuming for the team and we are still trying to figure out how to be able to release them more often (I’ve added this as a todo on my list to discuss with the team). You can say with confidence that Basque is fully laun…

---

<div class="post-metadata">

**Author:** ![daniel.cruzado](https://avatars.discourse-cdn.com/v4/letter/d/9de0a6/32.png) [@daniel.cruzado](https://discourse.mozilla.org/u/daniel.cruzado)\
**Post date:** [April 3, 2019, 10:56am UTC](https://discourse.mozilla.org/t/spanish-dataset/34292/18 "2019-04-03T10:56:34Z")

</div>

Ok, thanks a lot for your answer and for all of your work!!
