At the Crossroad of Cuneiform and NLP: Challenges for Fine-grained Part-of-Speech Tagging

Publication type: U
Publication status: In press
Authors: Smidt, G., De Graef, K., & Lefever, E.
Publisher: European Language Resources Association (Turin, Italy)
Conference: The 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (Turin, Italy)
Download
View in Biblio

Abstract

The study of ancient Middle Eastern cultures is dominated by the vast number of cuneiform texts. Multiple languages and language families were expressed in cuneiform. The most dominant language written in cuneiform is the Semitic Akkadian, which is the focus of this paper. We are specifically focusing on letters written in the dialect used in modern-day Baghdad and south towards the Persian Gulf during the Old Babylonian period (c. 2000-1600 B.C.E.). The Akkadian language was rediscovered in the 19th century and is now being scrutinised by Natural Language Processing (NLP) methods. However, existing Akkadian text publications are not always suitable for digital editions. We therefore risk applying NLP methods onto renderings of Akkadian unfit for the purpose. In this paper we want to investigate the input material and try to initiate a discussion about best-practices in the crossroad where NLP meets cuneiform studies. Specifically, we want to question the use of pre-trained embeddings, sentence segmentation and the type of cuneiform input used to fine-tune language models for the task of fine-grained Part-of-Speech tagging. We examine the issues by theoretical and practical approaches in a way that we hope spurs discussions that are relevant for automatic processing of other ancient languages.

April 8, 2024	Vacancy post-doctoral assistant at LT3
March 27, 2024	LT3 members involved in the organization of various shared tasks and workshops
Jan. 20, 2024	Veronique appointed as Francqui chair 2023-2024 at ULB
Nov. 7, 2023	Gilles-Maurice shows how ChatGPT can compile excellent dictionaries (for English)
Oct. 25, 2023	Meet the expert: Prof. Lynne Bowker