# Експорт речень з української Вікіпедії

**URL:** <https://discourse.mozilla.org/t/topic/83480>\
**Category:** Українська (uk)\
**Created:** [July 21, 2021, 7:31pm UTC](https://discourse.mozilla.org/t/topic/83480 "2021-07-21T19:31:39Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![mytmpaccount2015](https://avatars.discourse-cdn.com/v4/letter/m/41988e/32.png) [@mytmpaccount2015](https://discourse.mozilla.org/u/mytmpaccount2015)\
**Post date:** [July 21, 2021, 7:31pm UTC](https://discourse.mozilla.org/t/topic/83480/1 "2021-07-21T19:31:40Z")

</div>

Шановні колеги,  
хочу повідомити, що я підготував експорт з української Вікіпедії для додання в Common Voice:

> <https://github.com/common-voice/cv-sentence-extractor/pull/152>
>
> This is an adaptation of the \[Belarusian rules\](https://github.com/common-voice/…cv-sentence-extractor/pull/118) to Ukrainian.
> 
> \> - How many sentences did you get at the end?
> 
> 654224 sentences.
> 
> \> - How did you create the blocklist file?
> 
> I took a grammatical dictionary of Ukrainian \[here\](https://github.com/brown-uk/dict\_uk/releases/download/v5.3.1/dict\_corp\_vis.txt.bz2), split the Wikipedia export into tokens and kept only those tokens in the blocklist that don't occur in the dictionary, no matter what is their frequency. Note that I didn't use the full export with \`--no-check\`, as it would bring many irrelevant tokens (non-Cyrillic spellings; words that only occur in the sentences which are filtered out anyway, etc.). Instead, I temporarily set \`max\_sentences\_per\_text\` to \`std::usize::MAX\`, in order to consider tokens only in those sentences that pass the rules.
> 
> \> - Get at least 3 different native speakers (ideally linguists) to review a random sample of 100-500 sentences and estimate the average error ratio and comment (or link their comment) in the PR.
> 
> Spreadsheet \[here\](https://docs.google.com/spreadsheets/d/1EY1ni2ahA8LLlphuZpUD-NtE-eNo3mD6/edit), not yet reviewed. As I'm not a competent speaker of Ukrainian myself, I'm going to contact the Common Voice Ukrainian community and update this PR once the review is complete.

Це дозволить додати ~600 тисяч нових речень для озвучення. Щоб pull request був прийнятий, треба перевірити достатньо велику випадкову вибірку нових речень і переконатися, що в них не більше як 5…7% помилок (докладніше написано [тут](https://discourse.mozilla.org/t/using-the-europarl-dataset-with-sentences-from-speeches-from-the-european-parliament/50184)). Випадкові 4000 речень з української Вікіпедії доступні в таблиці:

> **[ukwiki.v1.sample4000.xlsx](https://docs.google.com/spreadsheets/d/1EY1ni2ahA8LLlphuZpUD-NtE-eNo3mD6/edit)**
>
> This Sheet is private

Буду вдячний, якщо у вас буде можливість перевірити цю вибірку, відзначити і прокоментувати помилки. Я сам, не будучи грамотним носієм української мови, на жаль, не можу взяти участь в перевірці.

---

<div class="post-metadata">

**Author:** ![mytmpaccount2015](https://avatars.discourse-cdn.com/v4/letter/m/41988e/32.png) [@mytmpaccount2015](https://discourse.mozilla.org/u/mytmpaccount2015)\
**Post date:** [August 26, 2021, 3:59pm UTC](https://discourse.mozilla.org/t/topic/83480/2 "2021-08-26T15:59:03Z")

</div>

Після перевірки декількох сотень речень носіями мови (див. також [оновлену вибірку](https://docs.google.com/spreadsheets/d/1k_0oM9k7oQWjroSXJY6v3B8ZA5YwgFxV/edit) після удосконалення правил експорту речень) з’ясувалося, що відсоток помилок вище за допустимий максимум. Робота над pull request’ом зараз припинена, тому що більшість недоліків, які залишилися, складно відфільтрувати автоматично.
