MUMBAI, India, Oct. 5 -- Intellectual Property India has published a patent application (202647114764 A) filed by Google LLC on September 25, 2026, for Streaming, Array-Agnostic, Full- And Sub-Band Modeling Front-End For Robust Automatic Speech Recognition.

Inventors include Heitkaemper, Jens; Caroselli, Jr, Joseph Peter; Narayanan, Arun; and Howard, Nathan David.

The application for the patent was published on October 02, 2026, under issue no. 40/2026.

Abstract: A speech enhancement model (200) includes a first pre-processor block (210a), a second pre-processor block (210b), self-attention blocks (500), a first masking layer (250a), a second masking layer (250b), and a phrase extraction layer (260). The first pre-processor receives short-time Fourier transform (STFT) coefficients (212a) for a cleaned input signal (340) and generates a maximum value (454a) of an embedding dimension of the cleaned input signal. The second pre-processor receives STFT coefficients (212b) for a noisy input signal (206a) and generates a maximum value (454b) of an embedding dimension of the noisy input signal. The self-attention blocks receive a stacked input (232) of the embedding dimensions of the cleaned input signal and the noisy input signal and generates an unmasked output (580). The phrase extraction layer receives the unmasked output, a masked cleaned input signal (340M), and a masked noisy input signal (206M), and generates enhanced input speech features (262).

Disclaimer: Curated by HT Syndication.