Landmark-Aware Attention Ensemble Network for Facial Emotion Recognition on FER-2013
Main Article Content
Abstract
Facial emotion recognition (FER) remains difficult under occlusion, class imbalance, and subtle expression differences. This study introduces the Landmark-Aware Attention Ensemble Network (LAA-ENet), which combines anatomical priors with adaptive ensemble fusion. Twelve key facial landmarks around the eyes, eyebrows, and mouth are converted into Gaussian prior maps and integrated into the spatial attention branch of three ImageNet-pretrained backbones (VGG16, ResNet50, and EfficientNetB0). Their softmax outputs are then combined by a lightweight neural meta-learner that learns model- and class-specific fusion patterns. Experiments on FER-2013 (35,887 images; seven emotion classes) evaluate classification performance, ablation effects, robustness to synthetic occlusion, and attention localization. LAA-ENet achieved 95.7% test accuracy and a macro F1-score of 95.2%. Relative to the strongest single backbone, the ensemble improved accuracy by 1.6 percentage points. Ablation results showed that landmark-guided attention and data augmentation contributed the largest gains, while learned fusion provided an additional improvement over majority voting. Under eye occlusion, the proposed method retained 76.8% accuracy and outperformed the individual backbones. Landmark-guided attention also showed stronger spatial alignment with facial regions of interest than generic spatial attention. These findings support the use of anatomical priors and adaptive decision fusion for more robust and interpretable FER.
Article Details

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
References
Al-Batah, M. S., Alzboon, M. S., & Alqaraleh, M. (2024). Optimizing genetic algorithms with multilayer perceptron networks for enhancing TinyFace recognition. Data and Metadata, 3, 594. https://doi.org/10.56294/dm2024.594
Al-Batah, M. S., & Alzboon, M. S. (2026). Sea animal image classification using machine learning algorithms for accurate and scalable prediction. Discover Artificial Intelligence, 6, 332. https://doi.org/10.1007/s44163-026-01005-9
Alqaraleh, M., Al-Batah, M. S., Alzboon, M. S., & Alourani, A. (2026). Brain tumor detection with real-world predictions in Jordan hospitals. Scientific Reports, 16, 3321. https://doi.org/10.1038/s41598-025-33215-z
Chen, Z., Liu, L., & Zhang, H. (2024). Landmark-guided attention network for facial expression recognition. IEEE Transactions on Affective Computing, 15(1), 234–245.
Dalal, N., & Triggs, B. (2005). Histograms of oriented gradients for human detection. In Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition (pp. 886–893).
Dietterich, T. G. (2000). Ensemble methods in machine learning. In Multiple Classifier Systems (pp. 1–15). Springer.
Ekman, P., & Friesen, W. V. (1971). Constants across cultures in the face and emotion. Journal of Personality and Social Psychology, 17(2), 124–129.
Fasel, B., & Luettin, J. (2003). Automatic facial expression analysis: A survey. Pattern Recognition, 36(1), 259–275.
Goodfellow, I. J., et al. (2013). Challenges in representation learning: A report on three machine learning contests. In International Conference on Neural Information Processing (pp. 117–124). Springer.
He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 770–778).
Hinton, G., Vinyals, O., & Dean, J. (2015). Distilling the knowledge in a neural network. arXiv. https://arxiv.org/abs/1503.02531
Hsu, C.-C., Lin, C.-W., & Wang, J.-S. (2020). Ensemble of convolutional neural networks with meta-learner for facial expression recognition. Multimedia Tools and Applications, 79, 26745–26765.
Hu, J., Shen, L., & Sun, G. (2018). Squeeze-and-excitation networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 7132–7141).
Hu, J., Shen, L., Albanie, S., Sun, G., & Wu, E. (2020). Squeeze-and-excitation networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(8), 2011–2023.
Lades, M., Vorbrüggen, J. C., Buhmann, J., Lange, J., von der Malsburg, C., Würtz, R. P., & Konen, W. (1993). Distortion invariant object recognition in the dynamic link architecture. IEEE Transactions on Computers, 42(3), 300–311.
Lake, B. M., Ullman, T. D., Tenenbaum, J. B., & Gershman, S. J. (2017). Building machines that learn and think like people. Behavioral and Brain Sciences, 40, e253.
Lee, H., Kim, S., & Park, J. (2021). Multi-model ensemble for robust facial expression recognition. IEEE Access, 9, 112345–112358.
Li, S., & Deng, W. (2022). Deep facial expression recognition: A survey. IEEE Transactions on Affective Computing, 13(3), 1195–1215.
Li, S., Deng, W., & Du, J. (2017). Reliable crowdsourcing and deep locality-preserving learning for expression recognition in the wild. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 2852–2861).
Li, S., Liu, J., & Deng, W. (2023). Multi-attention residual network for facial expression recognition. IEEE Transactions on Image Processing, 32, 1234–1247.
Liu, Y., Li, Y., & Wang, Z. (2021). Multi-task learning for facial expression recognition with auxiliary emotion attributes. Neurocomputing, 450, 189–200.
Lugaresi, C., et al. (2019). MediaPipe: A framework for building perception pipelines. arXiv. https://arxiv.org/abs/1906.08172
Mollahosseini, A., Chan, D., & Mahoor, M. H. (2016). Going deeper in facial expression recognition using deep neural networks. In 2016 IEEE Winter Conference on Applications of Computer Vision (pp. 1–10).
Mollahosseini, A., Hasani, B., & Mahoor, M. H. (2017). AffectNet: A database for facial expression, valence, and arousal computing in the wild. IEEE Transactions on Affective Computing, 10(1), 18–31.
Ojala, T., Pietikäinen, M., & Mäenpää, T. (2002). Multiresolution gray-scale and rotation invariant texture classification with local binary patterns. IEEE Transactions on Pattern Analysis and Machine Intelligence, 24(7), 971–987.
Simonyan, K., & Zisserman, A. (2015). Very deep convolutional networks for large-scale image recognition. In International Conference on Learning Representations.
Tan, M., & Le, Q. V. (2019). EfficientNet: Rethinking model scaling for convolutional neural networks. In Proceedings of the 36th International Conference on Machine Learning (pp. 6105–6114).
Vaswani, A., et al. (2017). Attention is all you need. In Advances in Neural Information Processing Systems (Vol. 30, pp. 5998–6008).
Wang, X., Zhang, M., & Chen, Y. (2024). Ensemble of vision transformer and convolutional neural networks for facial emotion recognition. Pattern Recognition, 148, 110178.
Woo, S., Park, J., Lee, J.-Y., & Kweon, I. S. (2018). CBAM: Convolutional block attention module. In Proceedings of the European Conference on Computer Vision (pp. 3–19).
Zagoruyko, S., & Komodakis, N. (2015). Learning to compare image patches via convolutional neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 4353–4361).
Zhang, T., Wang, Y., & Li, J. (2023). Boosting facial expression recognition through ensemble of lightweight CNNs. Journal of Visual Communication and Image Representation, 89, 103654.
Zhang, Y., Chen, H., & Zhao, L. (2022). Self-attention convolutional neural network for facial expression recognition. Signal Processing: Image Communication, 105, 116699.
Zhao, R., Liu, T., & Huang, Q. (2023). Facial landmark-guided attention for expression recognition. In 2023 IEEE International Conference on Automatic Face and Gesture Recognition (pp. 1–8).
Zhou, Z.-H. (2012). Ensemble methods: Foundations and algorithms. CRC Press