Abstract
This paper addresses the challenge of divergent lemmatization and part-of-speech (PoS) tagging practices for Latin participles in annotated corpora. We propose a solution through the LiLa Knowledge Base, a Linked Open Data framework designed to unify lexical and textual data for Latin. Using lemmas as the point of connection between distributed textual and lexical resources, LiLa introduces hypolemmas — secondary citation forms belonging to a word’s inflectional paradigm — as a means of reconciling divergent annotations for participles. Rather than advocating a single uniform annotation scheme, LiLa preserves each resource’s native guidelines while ensuring that users can retrieve and analyze participial data seamlessly. Via empirical assessments of multiple Latin corpora, we show how the LiLa’s integration of lemmas and hypolemmas enables consistent retrieval of participle forms regardless of whether they are categorized as verbal or adjectival.
| Lingua originale | Inglese |
|---|---|
| Titolo della pubblicazione ospite | Proceedings of the 19th Linguistic Annotation Workshop (LAW-XIX-2025) |
| Editore | Association for Computational Linguistics |
| Pagine | 103-114 |
| Numero di pagine | 12 |
| ISBN (stampa) | 979-8-89176-262-6 |
| Stato di pubblicazione | Pubblicato - 2025 |
Keywords
- Latin
- Lemmatization
- Linguistic Linked Data
- Participles
Fingerprint
Entra nei temi di ricerca di 'Harmonizing Divergent Lemmatization and Part-of-Speech Tagging Practices for Latin Participles through the LiLa Knowledge Base'. Insieme formano una fingerprint unica.Cita questo
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver