Scinovex
article Open Access

A computational analysis of Hinglish cyberbullying detection using machine learning techniques

Abstract

The prevalence of cyberbullying across social media platforms has highlighted the need for effective detection systems, especially in multilingual regions like India, where code-mixed language use is common. Hinglish, a blend of Hindi and English, is widely used in online communication yet remains underrepresented in existing datasets for cyberbullying detection. This study addresses this gap by creating a Hinglish cyberbullying dataset of 22,523 annotated tweets, specifically designed to support machine learning approaches. Various machine learning models were trained and evaluated on this dataset, with Random Forest achieving the highest accuracy among the tested algorithms. Our findings emphasize the importance of targeted language resources for cyberbullying detection in multilingual contexts and demonstrate the potential of ensemble methods like Random Forest for classification tasks in code-mixed settings.

Hate Speech and Cyberbullying DetectionComputer scienceMachine learningArtificial intelligence
Citations
0
FWCI
0.00
field-weighted impact
References
0
Percentile
0%
vs. same field & year
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.

A computational analysis of Hinglish cyberbullying detection using machine learning techniques · Scinovex