Divyansh Sharma, Yash Vashist, Aniket Malik, Astha Tripathi, Poonam Rani | International Journal of Digital Communication and Analog Signals | Vol 12, Issue 02 | ISSN: 2455-0329
Abstract
Identifying human emotional states from speech signals is the main goal of Speech Emotion Recognition (SER), which has grown to be a significant area of study in affective computing. However, class imbalance and efficient feature extraction from speech data pose problems for many of the current methods. This study suggests a hybrid framework that combines a Random Forest (RF) classifier with a multi-kernel Convolutional Neural Network (CNN) to address these problems. To extract various emotional patterns from speech spectrograms, the multi-kernel CNN uses parallel convolutional layers with various kernel sizes. The retrieved deep features are then used to train a Random Forest classifier for the purpose of classifying emotions. The RAVDESS and SUBESCO datasets, which include imbalanced emotion classes, are used in the experiments. The suggested method outperforms several current techniques for speech emotion recognition, achieving an accuracy of 87% on RAVDESS and 95% on SUBESCO. Comprehensive preprocessing methods, such as noise reduction, signal normalization, and spectrogram augmentation, are used prior to feature extraction to further increase the resilience of the model. The suggested methodology improves classification performance across a variety of speech samples by efficiently capturing both local and global emotional traits. When compared to traditional deep learning methods, extensive experimental evaluation shows improved accuracy, recall, and F1-score. The hybrid design is appropriate for real-time speech emotion recognition applications because it minimizes the effects of class imbalance while preserving computing performance. These results demonstrate the possibility for creating dependable, scalable, and intelligent emotion-aware human–computer interaction systems by combining deep feature extraction with ensemble machine learning approaches.
Keywords- Affective computing, Speech Emotion Recognition, Multi-kernel CNN, Random Forest, Deep learning
🔒 This is a subscription article
Full text is available to subscribers and institutional members. Please choose an option below to access it.
SubscribePurchase this articleInstitutional / Login accessReferences
- Tripathi A, Rani P. Multilingual speech emotion recognition using IGRFXG–Ensemble feature selection approach. Applied Acoustics. 2025 Dec 5;240:110905.
- Tripathi A, Rani P. An improved MSER using grid search based PCA and ensemble voting technique. Multimedia Tools and Applications. 2024 Oct;83(34):80497-522.
- Rani P, Tripathi A, Shoaib M, Yadav S, Yadav M. Multilingual Emotion Analysis from Speech. InInternational Conference on Innovative Computing and Communications: Proceedings of ICICC 2022, Volume 3 2022 Nov 8 (pp. 443-456). Singapore: Springer Nature Singapore.
- Rathi T, Tripathy M. Analyzing the influence of different speech data corpora and speech features on speech emotion recognition: A review. Speech Communication. 2024 Jul 1;162:103102.
- Jahan MS, Oussalah M. A systematic review of hate speech automatic detection using natural language processing. arXiv preprint arXiv:2106.00742. 2021 May 22.
- He L, Chan JC, Wang Z. Automatic depression recognition using CNN with attention mechanism from videos. Neurocomputing. 2021 Jan 21;422:165-75.
- Xu M, Zhang F, Zhang W. Head fusion: Improving the accuracy and robustness of speech emotion recognition on the IEMOCAP and RAVDESS dataset. IEEE Access. 2021 Mar 19;9:74539-49.
- Andayani F, Theng LB, Tsun MT, Chua C. Hybrid LSTM- transformer model for emotion recognition from speech audio files. IEEE Access. 2022 Mar 31;10:36018-27.
- Issa D, Demirci MF, Yazici A. Speech emotion recognition with deep convolutional neural networks. Biomedical Signal Processing and Control. 2020 May 1;59:101894.
- Bhavan A, Chauhan P, Shah RR. Bagged support vector machines for emotion recognition from speech. Knowledge-Based Systems. 2019 Nov 15;184:104886.
- Ancilin J, Milton A. Improved speech emotion recognition with Mel frequency magnitude coefficient. Applied Acoustics. 2021 Aug 1;179:108046.
- Koduru A, Valiveti HB, Budati AK. Feature extraction algorithms to improve the speech emotion recognition rate. International Journal of Speech Technology. 2020 Mar;23(1):45- 55.
- Belgiu M, Drăguţ L. Random forest in remote sensing: A review of applications and future directions. ISPRS journal of photogrammetry and remote sensing. 2016 Apr 1;114:24-31.
- Mencattini A, Martinelli E, Costantini G, Todisco M, Basile B, Bozzali M, Di Natale C. Speech emotion recognition using amplitude modulation parameters and a combined feature selection procedure. Knowledge-Based Systems. 2014 Jun 1;63:68-81.
- Yildirim S, Kaya Y, Kılıç F. A modified feature selection method based on metaheuristic algorithms for speech emotion recognition. Applied Acoustics. 2021 Feb 1;173:107721..
How to cite this article
@article{SharmaD2026,
author = {Divyansh Sharma and Yash Vashist and Aniket Malik and Astha Tripathi and Poonam Rani},
title = {Speech Emotion Recognition using Multi- Kernel CNN for Features extraction and Random Forest Classifier},
journal = {International Journal of Digital Communication and Analog Signals},
year = {2026},
volume = {12},
number = {02},
issn = {2455-0329},
url = {https://journalspub.com/publication/ijdcas/article=27476}
}