A Comparative Study on Data Augmentation Techniques for Arabic Text Classification
Document Type
Conference Proceeding
Source of Publication
Lecture Notes of the Institute for Computer Sciences Social Informatics and Telecommunications Engineering Lnicst
Publication Date
5-1-2026
Abstract
Deep learning models for text classification require large, annotated datasets, which are often unavailable for many languages, including Arabic. This scarcity, combined with the linguistic complexity of Arabic, presents significant challenges, especially in sentiment and emotion analysis tasks involving short and long texts. In this paper, we evaluate several data augmentation strategies aimed at enhancing Arabic text classification performance, with a particular focus on a novel text generation approach based on fine-tuned transformer models. Our method demonstrates notable improvements in classifier accuracy for both short and long texts, achieving accuracy gains of up to 92.8% in low-resource scenarios. This comparative study highlights the effectiveness of text generation techniques over traditional augmentation methods and provides practical insights into optimizing data augmentation for Arabic NLP applications.
DOI Link
ISBN
[9783032166371]
ISSN
Publisher
Springer Nature Switzerland
Volume
677 LNICST
First Page
126
Last Page
139
Disciplines
Computer Sciences
Keywords
Arabic language, Data augmentation, Text classifier, Text generation
Scopus ID
Recommended Citation
Alkhatib, Manar and Belqasmi, Fatna, "A Comparative Study on Data Augmentation Techniques for Arabic Text Classification" (2026). All Works. 8017.
https://zuscholars.zu.ac.ae/works/8017
Indexed in Scopus
yes
Open Access
no