A Machine Learning Framework for Real-Time Cyberbullying Detection, Monitoring, and Explainable Analytics on Social Media Platforms
DOI:
Keywords:
Index Terms—Cyberbullying detection, machine learning, NLP, TF-IDF, logistic regression, text classification, Explainable AI, social media analytics
Abstract
The emergence of social media has not only simplified communication but has brought about more cyberbullying cases. It is not easy to manually detect any cyberbullying case due to the vast numbers of posts generated each day in social media platforms. The current study aims at developing a machine learning framework that identifies cyberbullying cases and monitors the content of the social media through a web-based application. Text data will be processed using TF-IDF vectorization with unigram and bigram features while Logistic Regression, Multinomial Naïve Bayes and Linear SVM classifiers will be used to classify the data. Classification will be done by training and testing 14,602 cleaned and manually labeled social media posts. From the different classifiers considered, Logistic Regression classifier performed the best in terms of predicting the bullying class with an accuracy of 74.70% and an F1-score of 0.4703. Other than classification, the proposed framework offers other functionalities such as predictions explanations, severity level, prediction history, user feedback, administrative analytics and real time Reddit comment analysis.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.


