MUMBAI, India, Aug. 10 -- Intellectual Property India has published a patent application (202617071570 A) filed by Google LLC on June 09, 2026, for Annealed Policies For Finetuning Sequence Processing Models.
Inventors include Ferret, Johan; Bachem, Olivier Frdric; Hussenot, Lonard; Dadashi-Tazehozi, Robert; Geist, Matthieu Florent; Pietquin, Olivier Claude; Leurent, Edouard; Sessa, Pier Giuseppe; Vieillard, Nino Jean; Cideron, Geoffrey Virgil; Jacq, Alexis David; Momchev, Nikola Momchev; Stanczyk, Piotr Michal; Girgin, Sertan; Ramos Garea, Sabela; Sinopalnikov, Danila; and Hliou, Amlie Catherine.
The application for the patent was published on July 31, 2026, under issue no. 31/2026.
Abstract: Provided are techniques for training sequence processing models such as, for example, so-called large language models, with improved computational efficiency. In particular, the present disclosure provides a number of training frameworks which can be referred to as annealed language policies (ALP). ALP can operate to maximize a total reward while not deviating too much from a reference model. In particular, a hyperparameter, referred to as alpha, can control a trade-off between a reward term and a regularization term included in a loss function. The alpha hyperparameter can be "annealed" throughout the training process.
Disclaimer: Curated by HT Syndication.