Universiti Teknologi Malaysia Institutional Repository

Hate speech and offensive language detection: a new feature set with filter-embedded combining feature selection

Abdul Aziz, N. A. and Maarof, M. A. and Zainal, A. (2021) Hate speech and offensive language detection: a new feature set with filter-embedded combining feature selection. In: 3rd International Cyber Resilience Conference, CRC 2021, 29 January 2021 - 31 January 2021C, Virtual, Langkawi Island.

[img]
Preview
PDF
964kB

Official URL: http://dx.doi.org/10.1109/CRC50527.2021.9392486

Abstract

Social media has changed the world and play an important role in people lives. Social media platforms like Twitter, Facebook and YouTube create a new dimension of communication by providing channels to express and exchange ideas freely. Although the evolution brings numerous benefits, the dynamic environment and the allowable of anonymous posts could expose the uglier side of humanity. Irresponsible people would abuse the freedom of speech by aggressively express opinion or idea that incites hatred. This study performs hate speech and offensive language detection. The problem of this task is there is no clear boundary between hate speech and offensive language. In this study, a selected new features set is proposed for detecting hate speech and offensive language. Using Twitter dataset, the experiments are performed by considering the combination of word n-gram and enhanced syntactic n-gram. To reduce the feature set, filter-embedded combining feature selection is used. The experimental results indicate that the combination of word n-gram and enhanced syntactic n-gram with feature selection to classify the data into three classes: hate speech, offensive language or neither could give good performance. The result reaches 91% for accuracy and the averages of precision, recall and F1.

Item Type:Conference or Workshop Item (Paper)
Uncontrolled Keywords:feature selection, hate speech, machine learning
Subjects:Q Science > QA Mathematics > QA75 Electronic computers. Computer science
Divisions:Computing
ID Code:96045
Deposited By: Narimah Nawil
Deposited On:03 Jul 2022 04:31
Last Modified:03 Jul 2022 04:31

Repository Staff Only: item control page