MUMBAI, India, Sept. 28 -- Intellectual Property India has published a patent application (202647113864 A) filed by Google LLC on September 23, 2026, for Processing Multimodal Prompts Using Text-Only Large Language Models.

Inventor includes Wang, Quan.

The application for the patent was published on September 25, 2026, under issue no. 39/2026.

Abstract: A multimodal system (150) includes a multimodal encoder (160), a vector quantization model (170), and a text-only large language model (LLM) (180). The multimodal encoder is configured to receive, as input, a multimodal input (106) for a prompt (184), and generate, based on the multimodal input, a sequence of embeddings (162). The vector quantization model is configured to receive, as input, the sequence of embeddings generated by the multimodal encoder, and generate, based on the sequence of embeddings, a sequence of textual tokens (172). The text-only LLM is configured to receive, as input text, a natural language command (104) for the prompt and the sequence of textual tokens generated by the vector quantization model, and generate, based on the natural language command and the sequence of textual tokens, a corresponding textual output (182).

Disclaimer: Curated by HT Syndication.