No data to crawl? Monolingual corpus creation from PDF files of truly low-resource languages in Peru

Bustamante, G.; Oncevay, A.; Zariquiey, R.

No data to crawl? Monolingual corpus creation from PDF files of truly low-resource languages in Peru

Date

2020

Authors

Bustamante, G.

Oncevay, A.

Zariquiey, R.

Publisher

European Language Resources Association (ELRA)

URI

http://hdl.handle.net/20.500.14657/206833

Acceso al texto completo solo para la Comunidad PUCP

https://aclanthology.org/2020.lrec-1.356/

Abstract

We introduce new monolingual corpora for four indigenous and endangered languages from Peru: Shipibo-konibo, Ashaninka, Yanesha and Yine. Given the total absence of these languages in the web, the extraction and processing of texts from PDF files is relevant in a truly low-resource language scenario. Our procedure for monolingual corpus creation considers language-specific and language-agnostic steps, and focuses on educational PDF files with multilingual sentences, noisy pages and low-structured content. Through an evaluation based on language modelling and character-level perplexity on a subset of manually extracted sentences, we determine that our method allows the creation of clean corpora for the four languages, a key resource for natural language processing tasks nowadays.

Keywords

Yine, Ashaninka, Corpus creation, Endangered languages, Indigenous languages, Low-resource languages, Monolingual corpus, Pdf processing, Shipibo-Konibo, Yanesha

Collections

Artículos (DFI)

Full item page

No data to crawl? Monolingual corpus creation from PDF files of truly low-resource languages in Peru

Date

Authors

Journal Title

Journal ISSN

Volume Title

Publisher

URI

DOI

Acceso al texto completo solo para la Comunidad PUCP

Abstract

Description

Keywords

Citation

Collections

Endorsement

Review

Supplemented By

Referenced By