MUMBAI, India, Aug. 12 -- Intellectual Property India has published a patent application (202641092034 A) filed by Madhankumar C; Dr. G. Apparao Naidu - Vignan'S Institute Of Management And Technology For Women; Mrs. P. Rupa - Vignan'S Institute Of Management And Technology For; Mrs. D. Geetha - Vignan'S Institute Of Management And Technology For Women; Mrs. B. Mamatha - Vignan'S Institute Of Management And Technology For Women; and Mr. K. Bharath Reddy - Vignan'S Institute Of Management And on July 29, 2026, for Multi-Modal Llm-Guided Surgical Skill Assessment From Video, Tool Telemetry And Eye-Gaze.
Inventors include Dr. G. Apparao Naidu - Vignan'S Institute Of Management And; Mrs. P. Rupa - Vignan'S Institute Of Management And Technology For; Mrs. D. Geetha - Vignan'S Institute Of Management And Technology For; Mrs. B. Mamatha - Vignan'S Institute Of Management And Technology For Women; and Mr. K. Bharath Reddy - Vignan'S Institute Of Management And Technology For Women.
The application for the patent was published on August 07, 2026, under issue no. 32/2026.
Abstract: Multi-Modal LLM-Guided Surgical Skill Assessment From Video, Tool Telemetry And Eye-Gaze Abstract The increasing adoption of robotic and minimally invasive surgery has created a demand for intelligent systems capable of providing objective, accurate, and scalable surgical skill assessment. Conventional evaluation methods primarily depend on expert observation and standardized scoring systems, which are time consuming, subjective, and difficult to scale across diverse clinical environments. This work proposes a Multi-Modal Large Language Model (LLM)-Guided Surgical Skill Assessment Framework that integrates surgical video streams, surgical tool telemetry, and surgeon eye-gaze data to deliver comprehensive and explainable performance evaluation. The proposed framework employs computer vision models to extract spatiotemporal surgical actions from operative videos, while tool telemetry data—including instrument position, motion trajectories, force measurements, velocity, and interaction events—are processed to quantify surgical dexterity and precision. Simultaneously, eye-gaze tracking captures visual attention patterns, fixation duration, scan paths, and cognitive workload indicators to assess decision-making behavior and situational awareness. These heterogeneous data sources are synchronized through a multimodal fusion architecture that combines temporal transformers with cross-modal attention mechanisms to learn robust representations of surgical performance. An integrated Large Language Model acts as an intelligent reasoning engine that interprets multimodal features, generates explainable feedback, identifies procedural errors, recommends corrective actions, and produces personalized training reports in natural language. The framework supports automated skill scoring based on internationally recognized surgical assessment metrics while also detecting deviations from optimal procedural workflows. Machine learning models continuously adapt to surgeon-specific learning patterns, enabling personalized skill progression and competency tracking. Experimental evaluation demonstrates that the proposed multimodal framework significantly improves assessment accuracy, consistency, and interpretability compared to traditional single-modal evaluation techniques. By combining visual perception, instrument behavior, and surgeon attention analysis with LLM-based reasoning, the system provides real-time, objective, and explainable surgical skill assessment. The proposed approach has strong potential for surgical education, simulator-based training, robotic surgery certification, and continuous professional development, ultimately contributing to safer surgical procedures and improved patient outcomes.
Disclaimer: Curated by HT Syndication.