Back to Search Start Over

RIDGES Herbology: designing a diachronic multi-layer corpus.

Authors :
Odebrecht, Carolin
Belz, Malte
Zeldes, Amir
Lüdeling, Anke
Krause, Thomas
Source :
Language Resources & Evaluation. Sep2017, Vol. 51 Issue 3, p695-725. 31p.
Publication Year :
2017

Abstract

This paper introduces a multi-layer corpus architecture with multiple tokenizations using the open source historical, diachronic corpus of German called Register in Diachronic German Science. The corpus contains herbal texts printed between the fifteenth and nineteenth centuries and is concerned with the development of a German scientific register, independent of Latin. We will discuss difficulties of transcribing, normalizing and annotating historical texts and will thereby argue for the advantages of multiple layers and multiple tokenizations. A virtually infinite number of annotations can be added to the corpus, without the need for deciding between or discarding interpretations. Thus, this flexible architecture enables multiple normalizations and types of annotation and is open to a wide range of research questions in the humanities. We provide case studies concerning the exploitation of our different normalizations as well as structural, register-specific and linguistic annotations. The corpus architecture allows for its reuse as a resource for corpus-based research approaches. [ABSTRACT FROM AUTHOR]

Details

Language :
English
ISSN :
1574020X
Volume :
51
Issue :
3
Database :
Academic Search Index
Journal :
Language Resources & Evaluation
Publication Type :
Academic Journal
Accession number :
124484774
Full Text :
https://doi.org/10.1007/s10579-016-9374-3