All Works

Authorship identification using ensemble learning

Ahmed Abbasi, Air University Islamabad
Abdul Rehman Javed, Air University Islamabad
Farkhund Iqbal, Zayed University
Zunera Jalil, Air University Islamabad
Thippa Reddy Gadekallu, Vellore Institute of Technology
Natalia Kryvinska, Univerzita Komenského v Bratislave

Document Type

Article

Source of Publication

Scientific reports

Publication Date

6-9-2022

Abstract

With time, textual data is proliferating, primarily through the publications of articles. With this rapid increase in textual data, anonymous content is also increasing. Researchers are searching for alternative strategies to identify the author of an unknown text. There is a need to develop a system to identify the actual author of unknown texts based on a given set of writing samples. This study presents a novel approach based on ensemble learning, DistilBERT, and conventional machine learning techniques for authorship identification. The proposed approach extracts the valuable characteristics of the author using a count vectorizer and bi-gram Term frequency-inverse document frequency (TF-IDF). An extensive and detailed dataset, "All the news" is used in this study for experimentation. The dataset is divided into three subsets (article1, article2, and article3). We limit the scope of the dataset and selected ten authors in the first scope and 20 authors in the second scope for experimentation. The experimental results of proposed ensemble learning and DistilBERT provide better performance for all the three subsets of the "All the news" dataset. In the first scope, the experimental results prove that the proposed ensemble learning approach from 10 authors provides a better accuracy gain of 3.14% and from DistilBERT 2.44% from the article1 dataset. Similarly, in the second scope from 20 authors, the proposed ensemble learning approach provides a better accuracy gain of 5.25% and from DistilBERT 7.17% from the article1 dataset, which is better than previous state-of-the-art studies.

DOI Link

10.1038/s41598-022-13690-4

ISSN

2045-2322

Volume

Issue

First Page

9537

Disciplines

Computer Sciences

Scopus ID

85131708338

Recommended Citation

Abbasi, Ahmed; Javed, Abdul Rehman; Iqbal, Farkhund; Jalil, Zunera; Gadekallu, Thippa Reddy; and Kryvinska, Natalia, "Authorship identification using ensemble learning" (2022). All Works. 5176.
https://zuscholars.zu.ac.ae/works/5176

Indexed in Scopus

yes

Open Access

Link to Full Text

COinS

All Works

Authorship identification using ensemble learning

Document Type

Source of Publication

Publication Date

Abstract

DOI Link

ISSN

Volume

Issue

First Page

Disciplines

Scopus ID

Recommended Citation

Indexed in Scopus

Open Access

Search

Browse

Contribute

Content Type

All Works

Authorship identification using ensemble learning

Author First name, Last name, Institution

Document Type

Source of Publication

Publication Date

Abstract

DOI Link

ISSN

Volume

Issue

First Page

Disciplines

Scopus ID

Recommended Citation

Indexed in Scopus

Open Access

Share

Search

Browse

Contribute

Content Type