MUMBAI, India, Aug. 12 -- Intellectual Property India has published a patent application (202641085540 A) filed by Dr. Sowmya B J on July 13, 2026, for Development Of Ai-Based System And Method For Correcting Ocr-Generated Kannada Text Using Linguistic Validation.
Inventors include Dr. J. Geetha; Mrs. Veena Gode Swamy Rao; Dr. Chandrika Prasad; Dr. Sowmya B J; Ms. Shraddha C. S Atreya; Ms. Y Sameeksha; Dr. Supreeth S; and Dr. Shruthi G.
The application for the patent was published on August 07, 2026, under issue no. 32/2026.
Abstract: 7. ABSTRACT OF THE INVENTION The Patent disclosure covers Development of Al-Based System and Method for Correcting OCR-Generated Kannada Text Using Linguistic Validation. System and Method for Al-Based Post-OCR Correction of Kannada Text Using Morphological, Syntactic, Semantic and Entity Consistency Analysis discloses an intelligent text restoration framework for automatically correcting errors present in Optical Character Recognition (OCR) generated Kannada text. Conventional OCR systems often produce noisy outputs containing character substitutions, missing symbols, invalid words, grammatical inconsistencies, conjunction errors, syntactic distortions, and semantic inaccuracies, particularly in morphologically rich Indic languages. The disclosed invention provides a multi-layer post-processing architecture "that receives'OCR fextasinputand-performsUnicodenormalization, tokenization, OCR erro_L _____ analysis, candidate generation, morphological parsing, grammatical validation, syntactic correction, conjunction correction, semantic consistency evaluation, and entity memory based contextual correction. The system further preserves gender, number, person, tense, pronoun references, and inter-sentence consistency across paragraphs. The invention employs a hybrid correction engine integrating rule-based finite state transducers, linguistic analyzers, dictionary validation modules, neural language models, and adaptive ranking mechanisms to select optimal corrected outputs. Unlike conventional spell-checking approaches, the disclosed framework performs comprehensive linguistic restoration by jointly addressing orthographic, morphological, grammatical, and contextual errors. The invention improves readability, searchability, archival quality, and downstream machine processing of digitized Kannada documents. The framework is scalable to printed and handwritten OCR outputs and can be extended to other Indic and morphologically rich languages
Disclaimer: Curated by HT Syndication.