Detecting hate speech in roman Urdu using a convolutional-BiLSTM-based deep hybrid neural network
PeerJ Computer Science, vol.11, 2025 (SCI-Expanded, Scopus)
- Publication Type: Article / Article
- Volume: 11
- Publication Date: 2025
- Doi Number: 10.7717/peerj-cs.3342
- Journal Name: PeerJ Computer Science
- Journal Indexes: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Compendex, Directory of Open Access Journals
- Keywords: Convolutional layers, Data Mining and Machine Learning, Deep learning, Hybrid neural network, Natural Language and Speech, Neural Networks, Roman Urdu, Short-term memory, Text data, Text Mining
- Yıldız Technical University Affiliated: Yes
Abstract
The detection of hate speech on social media has become a pressing challenge, particularly in multilingual and low-resource language settings such as Roman Urdu, where informal grammar, code-switching, and inconsistent orthography hinder accurate classification. Despite progress in hate speech detection for high-resource languages, limited research exists for Roman Urdu content. This study addresses this gap by proposing a computationally efficient deep learning framework based on a hybrid convolutional neural network and bidirectional long short-term memory (CNN-BiLSTM) architecture. The model leverages FastText pre-trained embeddings to capture subword-level semantics and combines convolutional layers for local feature extraction with BiLSTM for global context modeling. We evaluate our approach on a labeled Roman Urdu dataset and compare it with traditional machine learning models and deep learning baselines. Our proposed CNN-BiLSTM model achieves the highest performance with an accuracy of 80.67% and an F1-score of 81.47%, outperforming competitive baselines. These findings demonstrate the effectiveness and practicality of our lightweight architecture in detecting hate speech in Roman Urdu, offering a scalable solution for multilingual and resource-constrained environments.