# Problems finding public domain sentences

**URL:** <https://discourse.mozilla.org/t/problems-finding-public-domain-sentences/34790>\
**Category:** Common Voice\
**Tags:** sentence-collection\
**Created:** [January 3, 2019, 7:26pm UTC](https://discourse.mozilla.org/t/problems-finding-public-domain-sentences/34790 "2019-01-03T19:26:23Z")\
**Posts on this page:** 7\
**Page:** 2

<div class="post-metadata">

**Author:** ![Geor](https://avatars.discourse-cdn.com/v4/letter/g/eb8c5e/32.png) [@Geor](https://discourse.mozilla.org/u/Geor)\
**Post date:** [February 22, 2019, 9:52pm UTC](https://discourse.mozilla.org/t/problems-finding-public-domain-sentences/34790/21 "2019-02-22T21:52:31Z")

</div>

Thank you for the answer. Yeah, it will be interesting to see the summary.

---

<div class="post-metadata">

**Author:** ![tinok](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/tinok/32/25881_2.png) [@tinok](https://discourse.mozilla.org/u/tinok)\
**Post date:** [February 23, 2019, 7:32pm UTC](https://discourse.mozilla.org/t/problems-finding-public-domain-sentences/34790/22 "2019-02-23T19:32:02Z")

</div>

Maybe a useful approach to others: We [started using translations from the English CommonVoice corpus as a source for adding sentences in Arabic.](https://discourse.mozilla.org/t/building-an-arabic-dataset-for-common-voice/29728/11?u=tinok)

This requires native speakers to confirm the accuracy of the translations (because we wouldn’t want blindly translated phrases to go to the sentence collector). But it’s a starting point that may work for other languages as well that struggle finding enough CC0 phrases.

---

<div class="post-metadata">

**Author:** ![tinok](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/tinok/32/25881_2.png) [@tinok](https://discourse.mozilla.org/u/tinok)\
**Post date:** [April 18, 2019, 2:16pm UTC](https://discourse.mozilla.org/t/problems-finding-public-domain-sentences/34790/23 "2019-04-18T14:16:08Z")

</div>

@nukeador Is there any new guidance on using sentences from Tatoeba? The sentences we are considering for Arabic are licensed at `CC-BY`. But [as I wrote above](https://discourse.mozilla.org/t/problems-finding-public-domain-sentences/34790/18?u=tinok), I think the license applies to the entire database, not each individual sentence (all of which have been used in many places previously, including other copyrighted material).

We have analyzed the Tatoeba database and [found 31,806 Arabic sentences](https://raw.githubusercontent.com/dighr/audio_analysis/master/transcription_performance_test/ar/190412/arabic_sentences.csv). They are all good quality. Would randomly selecting 5,000 or any other number for inclusion in the sentence collector be a violation of the CC license?

It would be great to have definitive guidance since I’m sure many other people are finding and collecting sentences from other sources but are unsure about the legal questions (or simply go ahead regardless).

Maybe reaching out to Tatoeba would be possible so that including random subsets (rather than the entire database) would get an explicit exemption?

---

<div class="post-metadata">

**Author:** ![nukeador](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/nukeador/32/20533_2.png) [@nukeador](https://discourse.mozilla.org/u/nukeador)\
**Post date:** [April 24, 2019, 1:11pm UTC](https://discourse.mozilla.org/t/problems-finding-public-domain-sentences/34790/24 "2019-04-24T13:11:34Z")

</div>

Later this week I’ll be posting an update on sentence collection from the team that I hope will help in this matter.

---

<div class="post-metadata">

**Author:** ![ktaa](https://avatars.discourse-cdn.com/v4/letter/k/839c29/32.png) [@ktaa](https://discourse.mozilla.org/u/ktaa)\
**Post date:** [May 28, 2019, 2:38pm UTC](https://discourse.mozilla.org/t/problems-finding-public-domain-sentences/34790/25 "2019-05-28T14:38:03Z")

</div>

I uploaded around 14,000 modern standard Arabic sentences in the sentence collector that need verification.

---

<div class="post-metadata">

**Author:** ![davidak](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/davidak/32/25339_2.png) [@davidak](https://discourse.mozilla.org/u/davidak)\
**Post date:** [June 8, 2019, 3:30pm UTC](https://discourse.mozilla.org/t/problems-finding-public-domain-sentences/34790/26 "2019-06-08T15:30:24Z")

</div>

Any update on the usage of CC-BY?

---

<div class="post-metadata">

**Author:** ![nukeador](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/nukeador/32/20533_2.png) [@nukeador](https://discourse.mozilla.org/u/nukeador)\
**Post date:** [June 10, 2019, 1:40pm UTC](https://discourse.mozilla.org/t/problems-finding-public-domain-sentences/34790/27 "2019-06-10T13:40:35Z")

</div>

For now we should we stick with public domain. The update I posted was about the “fair-use” of some large sources of text and the work we started doing with wikipedia

> [@Extending our sentence collection capabilities](https://discourse.mozilla.org/t/extending-our-sentence-collection-capabilities/38783):
>
> Hello everyone, One of the most important components to build a strong dataset of voices is always being able to provide people with enough sentences to read in their language. Without this, voice collection is not possible, and as a team we have been putting in a lot of work to emphasize sentence collection since the launch of Common Voice. Some background To make the Common Voice dataset as useful as possible we have decided to only allow source text that is available under a [Creative Common…](https://creativecommons.org/share-your-work/public-domain/cc0/)

[Previous page](https://discourse.mozilla.org/t/problems-finding-public-domain-sentences/34790.md?page=1)
