The GW/LT3 VarDial 2016 shared task system for dialects and similar languages detection

Publication type
C1
Publication status
Published
Authors
Zirikly, A., Desmet, B., & Diab, M.
Editor
Preslav Nakov, Marcos Zampieri, Liling Tan, Nikola Ljubešić, Jørg Tiedemann and Shervin Malmasi
Series
Proceedings of the Third Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial3)
Pagination
33-41
Publisher
The COLING 2016 Organizing Committee (Osaka, Japan)
Conference
COLING (Osaka, Japan)
Download
(.pdf)
View in Biblio
(externe link)

Abstract

This paper describes the GW/LT3 contribution to the 2016 VarDial shared task on the identification of similar languages (task 1) and Arabic dialects (task 2). For both tasks, we experimented with Logistic Regression and Neural Network classifiers in isolation. Additionally, we implemented a cascaded classifier that consists of coarse and fine-grained classifiers (task 1) and a classifier ensemble with majority voting for task 2. The submitted systems obtained state-of-the-art performance and ranked first for the evaluation on social media data (test sets B1 and B2 for task 1), with a maximum weighted F1 score of 91.94%.