MUMBAI, India, Oct. 5 -- Intellectual Property India has published a patent application (202641114743 A) filed by Vellore Institute Of Technology on September 24, 2026, for System And Method For Activation-Space Jailbreak Detection And Bounded Residual-Stream Steering In Transformer Language Models.

Inventors include Senthil Kumar T; Chirag S Das; Sujaa Shri S; Thanusha E; Elakiya R; and Ramya.

The application for the patent was published on October 02, 2026, under issue no. 40/2026.

Abstract: The present disclosure proposes a system (100) for activation-space jailbreak detection and bounded residual-stream steering in a transformer-based language model. The system (100) comprises a computing device (102), a frozen transformer-based language model (108), and inference-safety modules (110). An activation acquisition module (112) obtains harmful and benign residual-stream activations, while a refusal-vector generation module (114) generates a unit-norm refusal vector using difference-of-means, singular value decomposition, and null-space projection. A safe-reference generation module (116) determines an activation mean, covariance matrix, and detection threshold. During runtime inference, a forward-pass hook (120) intercepts a residual-stream activation, and a drift detection module (122) determines Mahalanobis distance. A bounded steering module (124) conditionally applies an L2-bounded refusal-vector perturbation verified by system-level assertion, and an activation reinjection module (126) propagates the steered activation through transformer layers without retraining or forward passes.

Disclaimer: Curated by HT Syndication.