MUMBAI, India, Aug. 10 -- Intellectual Property India has published a patent application (202617074646 A) filed by Google LLC on June 16, 2026, for Speaker Diarization Post-Processing With Large Language Models.

Inventors include Wang, Quan; Huang, Yiling; Liao, Hank; Zhao, Guanlong; Xia, Wei; and Clark, Evan.

The application for the patent was published on July 31, 2026, under issue no. 31/2026.

Abstract: A method (700) includes receiving audio data (108) including a plurality of spoken terms spoken by one or more speakers (10) during a conversation. The method includes generating diarization results (155) based on the plurality of spoken terms spoken by the one or more speakers during the conversation. The diarization results include a speech recognition result (120) including a series of predicted terms and a series of identity-agnostic speaker tokens (165). The method also includes processing the diarization results conditioned on a diarization prompt (116) to predict, as output from an LLM (170), updated diarization results (175). The updated diarization results include the speech recognition result including the series of predicted terms and a series of identity-specific speaker tokens (172).

Disclaimer: Curated by HT Syndication.