Real-Time Sign Language Recognition Using Multi-Modal Landmark Extraction with MediaPipe Holistic

Authors

  • Muhammad Anas MJ CODERS, MOBIL1 PLAZA, Multan, 60000, Pakistan
  • Muhammad Azeem Umar MJ CODERS, MOBIL1 PLAZA, Multan, 60000, Pakistan

DOI:

https://doi.org/10.66108/mna.v5i01.90

Keywords:

Sign Language Recognition (SLR), MediaPipe Holistic Framework, Multimodal Feature Extraction, Real-Time Gesture Recognition, RGB-Based Computer Vision

Abstract

Sign language is a major mode of communication to millions of deaf and hard-of-hearing people, but unfamiliarity with the access of real-time interpretation technologies still restricts effective interaction and access to information. To overcome this, this paper proposes a real-time Sign Language Recognition (SLR) framework that utilizes a Sign Language Recognition system on top of Google MediaPipe Holistic framework that learns detailed 3D face, hand and whole-body landmarks with the help of standard RGB video input only. The system proposed has stable performance with an approximate real-time frame rate and feature vectors of the 1,662 dimensions being processed on consumer-grade hardware and operating in real time. Compared to current technologies, which utilize specialized sensors or only track hand gestures, the technique provides abundant spatial and temporal data of several body parts at once using only a single webcam. The system pipeline consists of data acquisition, landmark localization (468 facial landmarks, 21 key points per hand and 33 landmarks of the body pose per frame), multimodal feature vector creation and performance analysis using visualization features. Initial experiments show that it is not impossible to identify a small set of basic signs (e.g. greetings, like hello and thanks), having a demonstration of concept, not a trained large-vocabulary classifier. The framework offers an excellent base to integrate supervised temporal models in the future, including recurrent neural networks and transformer-based architectures, and includes low-cost, scalable, and inclusive SLR solutions to real-world problems in education, healthcare, and assistive technologies

Downloads

Download data is not yet available.

References

Bragg, D., Koller, O., Bellard, M., Berke, L., Boudreault, P., Braffort, A., Caselli, N., Huenerfauth, M., Kacorri, H., Verhoef, T., Vogler, C., & Ringel Morris, M. (2019). Sign Language Recognition, Generation, and Translation. Proceedings of the 21st International ACM SIGACCESS Conference on Computers and Accessibility, 16–31. https://doi.org/10.1145/3308561.3353774

Kadous, M. W. (1996, October). Machine recognition of Auslan signs using PowerGloves: Towards large-lexicon recognition of sign language. In Proceedings of the Workshop on the Integration of Gesture in Language and Speech (Vol. 165, pp. 165-174). Wilmington: DE.

Zafrulla, Z., Brashear, H., Starner, T., Hamilton, H., & Presti, P. (2011). American sign language recognition with the kinect. Proceedings of the 13th International Conference on Multimodal Interfaces, 279–286. https://doi.org/10.1145/2070481.2070532

Camgoz, N. C., Hadfield, S., Koller, O., Ney, H., & Bowden, R. (2018). Neural Sign Language Translation. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 7784–7793. https://doi.org/10.1109/cvpr.2018.00812

Lugaresi, C., Tang, J., Nash, H., McClanahan, C., Uboweja, E., Hays, M., ... & Grundmann, M. (2019). Mediapipe: A framework for building perception pipelines. arXiv preprint arXiv:1906.08172. https://doi.org/10.48550/arXiv.1906.08172

Pigou, L., Dieleman, S., Kindermans, P.-J., & Schrauwen, B. (2015). Sign Language Recognition Using Convolutional Neural Networks. Computer Vision - ECCV 2014 Workshops, 572–578. https://doi.org/10.1007/978-3-319-16178-5_40

Koller, O., Zargaran, S., Ney, H., & Bowden, R. (2016). Deep Sign: Hybrid CNN-HMM for Continuous Sign Language Recognition. Procedings of the British Machine Vision Conference 2016, 136.1-136.12. https://doi.org/10.5244/c.30.136

Huang, J., Zhou, W., Zhang, Q., Li, H., & Li, W. (2018). Video-Based Sign Language Recognition Without Temporal Segmentation. Proceedings of the AAAI Conference on Artificial Intelligence, 32(1). https://doi.org/10.1609/aaai.v32i1.11903

Prakash, K. B. (2020). Accurate Hand Gesture Recognition using CNN and RNN Approaches. International Journal of Advanced Trends in Computer Science and Engineering, 9(3), 3216–3222. https://doi.org/10.30534/ijatcse/2020/114932020

Qiao, S., Wang, Y., & Li, J. (2017). Real-time human gesture grading based on OpenPose. 2017 10th International Congress on Image and Signal Processing, BioMedical Engineering and Informatics (CISP-BMEI), 1–6. https://doi.org/10.1109/cisp-bmei.2017.8301910

Koller, O., Forster, J., & Ney, H. (2015). Continuous sign language recognition: Towards large vocabulary statistical recognition systems handling multiple signers. Computer Vision and Image Understanding, 141, 108–125. https://doi.org/10.1016/j.cviu.2015.09.013

Jolliffe, I. T. (1986). Principal Component Analysis and Factor Analysis. Principal Component Analysis, 115–128. https://doi.org/10.1007/978-1-4757-1904-8_7

Bradski, G. (2000). The opencv library. Dr. Dobb's Journal: Software Tools for the Professional Programmer, 25(11), 120-123.

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., N.Gomez, A., Kaiser, L., & Polosukhin, I. (2025). Attention Is All You Need. https://doi.org/10.65215/r5bs2d54

Additional Files

Published

2026-03-16

How to Cite

Muhammad Anas, & Muhammad Azeem Umar. (2026). Real-Time Sign Language Recognition Using Multi-Modal Landmark Extraction with MediaPipe Holistic. Machines and Algorithms, 5(01), 22–32. https://doi.org/10.66108/mna.v5i01.90

Issue

Section

Articles

Categories