Hate Content Detection via Novel Pre-Processing Sequencing and Ensemble Methods

Chhabra, Anusha, Vishwakarma, Dinesh Kumar

Sep-8-2024–arXiv.org Artificial Intelligence

Social media, particularly Twitter, has seen a significant increase in incidents like trolling and hate speech. Thus, identifying hate speech is the need of the hour. This paper introduces a computational framework to curb the hate content on the web. Specifically, this study presents an exhaustive study of pre-processing approaches by studying the impact of changing the sequence of text pre-processing operations for the identification of hate content. The best-performing pre-processing sequence, when implemented with popular classification approaches like Support Vector Machine, Random Forest, Decision Tree, Logistic Regression and K-Neighbor provides a considerable boost in performance. Additionally, the best pre-processing sequence is used in conjunction with different ensemble methods, such as bagging, boosting and stacking to improve the performance further. Three publicly available benchmark datasets (WZ-LS, DT, and FOUNTA), were used to evaluate the proposed approach for hate speech identification. The proposed approach achieves a maximum accuracy of 95.14% highlighting the effectiveness of the unique pre-processing approach along with an ensemble classifier.

classifier, dataset, detection, (15 more...)

arXiv.org Artificial Intelligence

Sep-8-2024

arXiv.org PDF

Add feedback

Country:
- Europe > France (0.14)
- North America
  - United States (0.04)
  - Canada (0.04)
- Asia
  - India (0.04)
  - China > Shanxi Province
    - Taiyuan (0.04)

Genre:
- Research Report > New Finding (0.67)

Industry:
- Information Technology > Services (0.88)

Technology:
- Information Technology > Artificial Intelligence > Machine Learning > Statistical Learning > Support Vector Machines (0.54)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found