A Comparative Study on Data Augmentation Techniques for Arabic Text Classification

Document Type

Conference Proceeding

Source of Publication

Lecture Notes of the Institute for Computer Sciences Social Informatics and Telecommunications Engineering Lnicst

Publication Date

5-1-2026

Abstract

Deep learning models for text classification require large, annotated datasets, which are often unavailable for many languages, including Arabic. This scarcity, combined with the linguistic complexity of Arabic, presents significant challenges, especially in sentiment and emotion analysis tasks involving short and long texts. In this paper, we evaluate several data augmentation strategies aimed at enhancing Arabic text classification performance, with a particular focus on a novel text generation approach based on fine-tuned transformer models. Our method demonstrates notable improvements in classifier accuracy for both short and long texts, achieving accuracy gains of up to 92.8% in low-resource scenarios. This comparative study highlights the effectiveness of text generation techniques over traditional augmentation methods and provides practical insights into optimizing data augmentation for Arabic NLP applications.

ISBN

[9783032166371]

ISSN

1867-8211

Publisher

Springer Nature Switzerland

Volume

677 LNICST

First Page

126

Last Page

139

Disciplines

Computer Sciences

Keywords

Arabic language, Data augmentation, Text classifier, Text generation

Scopus ID

105041112131

Indexed in Scopus

yes

Open Access

no

Share

COinS