Standard

Fused YOLO and Traditional Features for Emotion Recognition From Facial Images of Tamil and Russian Speaking Children: A Cross-Cultural Study. / Mekala, A. Mary; Varalakshmi , M.; GOWDA, C. P. ACHYUTHA; KUMAR, LETI MANISH; Ляксо, Елена Евгеньевна; Фролова, Ольга Владимировна; Ruban, Nersisson.

в: IEEE Access, Том 13, 23.05.2025, стр. 86828-86840.

Результаты исследований: Научные публикации в периодических изданияхстатьяРецензирование

Harvard

APA

Vancouver

Author

BibTeX

@article{8f1bda82d16f4eaf90ff06df47e818d6,
title = "Fused YOLO and Traditional Features for Emotion Recognition From Facial Images of Tamil and Russian Speaking Children: A Cross-Cultural Study",
abstract = "Cross-cultural study that avoids bias in the formulation of emotion recognition models is indispensable to address the challenges in facial emotion classification for children, for it being a relatively difficult task. With evidences from literature for the hybrid feature extraction approaches to improve the image classification accuracy, this research work focuses on developing a hybrid framework for emotion recognition from facial images of Tamil and Russian children. The dataset is audio video recording of 28 Tamil speaking and 64 Russian speaking children. The data is collected in a controlled environment and labelled by experts. Traditional features like Grey Level Cooccurrence Matrix (GLCM) and facial landmark are extracted and are fused with You Look Only Once (YOLO V5) features. While facial landmarks and GLCM provide useful information about the facial expressions and texture of the image, YOLO V5 being a single-stage object detector makes the hybrid model super-fast and achieve high accuracy in detecting small objects and in low-light settings. The different classifiers that includes KNN, SVM, Random Forest, XGBoost, and Multilayer Perceptron, is employed yielding the accuracy result of 93%, 86%, 89%, 88%,and 90%. The use of majority voting ensemble of heterogeneous classifiers for the final prediction strengthens the model further, yielding an accuracy as high as 96% for the custom cross-cultural dataset, consisting of facial images of Russian and Indian children. This is consistent with the results obtained for Indian and Russian datasets. Further, the ablation study unveils the effect of feature fusion in boosting the performance and the dominance of YOLO V5 features over the other two.",
keywords = "Ensemble Classification, Facial Emotion Recognition, Facial Landmarks, Feature Fusion, Grey Level Co-occurrence Matrix (GLCM), YOLO V5",
author = "Mekala, {A. Mary} and M. Varalakshmi and GOWDA, {C. P. ACHYUTHA} and KUMAR, {LETI MANISH} and Ляксо, {Елена Евгеньевна} and Фролова, {Ольга Владимировна} and Nersisson Ruban",
year = "2025",
month = may,
day = "23",
doi = "10.1109/ACCESS.2025.3569771",
language = "English",
volume = "13",
pages = "86828--86840",
journal = "IEEE Access",
issn = "2169-3536",
publisher = "Institute of Electrical and Electronics Engineers Inc.",

}

RIS

TY - JOUR

T1 - Fused YOLO and Traditional Features for Emotion Recognition From Facial Images of Tamil and Russian Speaking Children: A Cross-Cultural Study

AU - Mekala, A. Mary

AU - Varalakshmi , M.

AU - GOWDA, C. P. ACHYUTHA

AU - KUMAR, LETI MANISH

AU - Ляксо, Елена Евгеньевна

AU - Фролова, Ольга Владимировна

AU - Ruban, Nersisson

PY - 2025/5/23

Y1 - 2025/5/23

N2 - Cross-cultural study that avoids bias in the formulation of emotion recognition models is indispensable to address the challenges in facial emotion classification for children, for it being a relatively difficult task. With evidences from literature for the hybrid feature extraction approaches to improve the image classification accuracy, this research work focuses on developing a hybrid framework for emotion recognition from facial images of Tamil and Russian children. The dataset is audio video recording of 28 Tamil speaking and 64 Russian speaking children. The data is collected in a controlled environment and labelled by experts. Traditional features like Grey Level Cooccurrence Matrix (GLCM) and facial landmark are extracted and are fused with You Look Only Once (YOLO V5) features. While facial landmarks and GLCM provide useful information about the facial expressions and texture of the image, YOLO V5 being a single-stage object detector makes the hybrid model super-fast and achieve high accuracy in detecting small objects and in low-light settings. The different classifiers that includes KNN, SVM, Random Forest, XGBoost, and Multilayer Perceptron, is employed yielding the accuracy result of 93%, 86%, 89%, 88%,and 90%. The use of majority voting ensemble of heterogeneous classifiers for the final prediction strengthens the model further, yielding an accuracy as high as 96% for the custom cross-cultural dataset, consisting of facial images of Russian and Indian children. This is consistent with the results obtained for Indian and Russian datasets. Further, the ablation study unveils the effect of feature fusion in boosting the performance and the dominance of YOLO V5 features over the other two.

AB - Cross-cultural study that avoids bias in the formulation of emotion recognition models is indispensable to address the challenges in facial emotion classification for children, for it being a relatively difficult task. With evidences from literature for the hybrid feature extraction approaches to improve the image classification accuracy, this research work focuses on developing a hybrid framework for emotion recognition from facial images of Tamil and Russian children. The dataset is audio video recording of 28 Tamil speaking and 64 Russian speaking children. The data is collected in a controlled environment and labelled by experts. Traditional features like Grey Level Cooccurrence Matrix (GLCM) and facial landmark are extracted and are fused with You Look Only Once (YOLO V5) features. While facial landmarks and GLCM provide useful information about the facial expressions and texture of the image, YOLO V5 being a single-stage object detector makes the hybrid model super-fast and achieve high accuracy in detecting small objects and in low-light settings. The different classifiers that includes KNN, SVM, Random Forest, XGBoost, and Multilayer Perceptron, is employed yielding the accuracy result of 93%, 86%, 89%, 88%,and 90%. The use of majority voting ensemble of heterogeneous classifiers for the final prediction strengthens the model further, yielding an accuracy as high as 96% for the custom cross-cultural dataset, consisting of facial images of Russian and Indian children. This is consistent with the results obtained for Indian and Russian datasets. Further, the ablation study unveils the effect of feature fusion in boosting the performance and the dominance of YOLO V5 features over the other two.

KW - Ensemble Classification

KW - Facial Emotion Recognition

KW - Facial Landmarks

KW - Feature Fusion

KW - Grey Level Co-occurrence Matrix (GLCM)

KW - YOLO V5

UR - https://www.mendeley.com/catalogue/89a2a81a-4f1e-32f0-b3eb-b41f493d2802/

U2 - 10.1109/ACCESS.2025.3569771

DO - 10.1109/ACCESS.2025.3569771

M3 - Article

VL - 13

SP - 86828

EP - 86840

JO - IEEE Access

JF - IEEE Access

SN - 2169-3536

ER -

ID: 135956206