1. Building an Endangered Language Resource in the Classroom: Universal Dependencies for Kakataibo
- Author
-
Zariquiey, Roberto, Alvarado, Claudia, Echevarria, Ximena, Gomez, Luisa, Gonzales, Rosa, Illescas, Mariana, Oporto, Sabina, Blum, Frederic, Oncevay, Arturo, and Vera, Javier
- Subjects
Computer Science - Computation and Language - Abstract
In this paper, we launch a new Universal Dependencies treebank for an endangered language from Amazonia: Kakataibo, a Panoan language spoken in Peru. We first discuss the collaborative methodology implemented, which proved effective to create a treebank in the context of a Computational Linguistic course for undergraduates. Then, we describe the general details of the treebank and the language-specific considerations implemented for the proposed annotation. We finally conduct some experiments on part-of-speech tagging and syntactic dependency parsing. We focus on monolingual and transfer learning settings, where we study the impact of a Shipibo-Konibo treebank, another Panoan language resource., Comment: Accepted to LREC 2022
- Published
- 2022