MUMBAI, India, July 7 -- Intellectual Property India has published a patent application (202641081088 A) filed by Mlr Institute Of Technology on July 01, 2026, for 3d Gesture Recognition For Indian Sign Language Using Mediapipe Holistic And Temporal Graph Convolutions.
Inventors include Dr. P. Radhika; Ms. Sukrithi Bhattacharya; Ms. Nandini Pashwaan; and Ms. Bezawada Sohana.
The application for the patent was published on July 03, 2026, under issue no. 27/2026.
Abstract: In this invention, “3D Gesture Recognition for Indian Sign Language Using MediaPipe Holistic and Temporal Graph Convolutions” is disclosed as an intelligent assistive communication system designed to recognize and interpret Indian Sign Language (ISL) gestures in real time using advanced computer vision and deep learning techniques. The invention acquires gesture data through live webcam streams or uploaded video recordings and processes the captured frames using the MediaPipe Holistic framework. The MediaPipe Holistic module extracts three-dimensional landmarks corresponding to hand joints, body posture, facial expressions, and head movements, thereby generating a comprehensive skeletal representation of the signer’s gestures. The extracted landmark coordinates are subsequently normalized to eliminate variations caused by scale, position, orientation, and environmental conditions. The normalized landmark sequences are transformed into graph-based structures in which landmark points are represented as nodes and their anatomical relationships are represented as graph edges. A Temporal Graph Convolution Network (T-GCN) or Spatio-Temporal Graph Convolution Network (ST-GCN) analyzes the spatial and temporal dependencies among the landmark nodes across consecutive video frames to accurately classify dynamic sign language gestures. The system further incorporates a confidence threshold evaluation mechanism that validates classification outputs and improves recognition reliability by filtering low-confidence predictions. Recognized gestures are converted into meaningful outputs including text, speech, symbols, or multilingual communication formats, enabling effective interaction between hearing-impaired individuals and non-sign-language users. The invention supports continuous gesture recognition, adaptive learning, vocabulary expansion, and real-time deployment across educational institutions, healthcare facilities, workplaces, public service centers, and assistive communication platforms. By integrating MediaPipe Holistic landmark extraction, graph-based gesture modeling, temporal deep learning, confidence-based validation, and communication assistance within a unified framework, the invention significantly improves sign language recognition accuracy, enhances accessibility, reduces communication barriers, promotes social inclusion, and facilitates seamless human-computer interaction for individuals with hearing and speech impairments.
Disclaimer: Curated by HT Syndication.