# Problems finding public domain sentences

**URL:** <https://discourse.mozilla.org/t/problems-finding-public-domain-sentences/34790>\
**Category:** Common Voice\
**Tags:** sentence-collection\
**Created:** [January 3, 2019, 7:26pm UTC](https://discourse.mozilla.org/t/problems-finding-public-domain-sentences/34790 "2019-01-03T19:26:23Z")\
**Posts on this page:** 1\
**Showing post:** 13

<div class="post-metadata">

**Author:** ![tinok](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/tinok/32/25881_2.png) [@tinok](https://discourse.mozilla.org/u/tinok)\
**Post date:** [February 1, 2019, 7:34pm UTC](https://discourse.mozilla.org/t/problems-finding-public-domain-sentences/34790/13 "2019-02-01T19:34:38Z")

</div>

Hi all, I understand the requirement for CC0 so that no licensing restrictions may limit future outcomes of the DeepSpeech models. I do wonder though if this isn’t unncessarily holding back more rapid progress of this great project.

FB’s LASER project has just reached a major milestone for rapidly translating phrases from 93 languages (see this [detailed post](https://code.fb.com/ai-research/laser-multilingual-sentence-embeddings/) and [this full paper](https://arxiv.org/abs/1812.10464)). All code is released on [https://github.com/facebookresearch/LASER](https://github.com/facebookresearch/LASER) under `CC-BY-NC`. This leverages the huge [Tatoeba corpus](https://tatoeba.org/eng/), some of which is CC0 but most of it is under `CC-BY`. FB’s advances wouldn’t have been possible without this great resource.

Given that `CC-BY` places no restrictions other than needing to attribute that some source material may have come from [tatoeba.org](http://tatoeba.org), what are the arguments against allowing this kind of source/license? It would certainly allow jump-starting some languages where collected sentences are still at 0 (e.g., [Arabic](https://discourse.mozilla.org/t/building-an-arabic-dataset-for-common-voice/29728/5)). For example, Tatoeba has 31,481 sentences available in Arabic, but none of them are `CC0`, all are `CC-BY`. I ran a quick summary on the `sentences.csv` that can be [downloaded](https://tatoeba.org/eng/downloads); [here are the number of sentences by language](https://0bin.net/paste/5H8bG0CSIN0fcLOz#0TkRLgPZeVyrddDvLb9v5oSbyhGLDlCZtuVP2om1TeO).

I’d love to understand better what the hard arguments are for excluding `CC-BY` given such great resources.

cc @nukeador @mhenretty

---

_[View the full topic](https://discourse.mozilla.org/t/problems-finding-public-domain-sentences/34790)._
