Skip to content
Home India Ministry of Education Notifications Parliament Question: Mother Tongue Based AI Learni... (Official PDF)
Date: 12th August 2026 Category: Press Release Jurisdiction: India, Central Government

Parliament Question: Mother Tongue Based AI Learning System for Tribal and Multilingual Area

Issued by Ministry of Education

Read or download the official PDF of this gazette notification issued by the Ministry of Education on 12th August 2026. Classified under Press Release.

Executive Summary & Key Takeaways

Executive Summary This report outlines the Ministry of Education’s initiatives to implement mother tongue-based AI learning systems for tribal and multilingual areas in alignment with NEP 2020. The document, based on a parliamentary reply dated August 12, 2026, details the development of Large Language Models (LLMs) and digital translation tools to support 22 scheduled languages and various tribal dialects. Key action items include the public release of linguistic datasets and the launch of specialized language courses on digital platforms.

Key Points / Main Content

AI and LLM Development Initiatives

  • BharatGen Project: A multimodal Large Language Model (LLM) initiative anchored at IIT Bombay and supported by a consortium (including IIT Kanpur, IIT Madras, and others) to build inclusive AI infrastructure for all 22 scheduled languages.
  • BHASHINI Platform: An initiative under the Digital India Programme that hosts over 360 AI-based models and provides 22 specialized services such as Automatic Speech Recognition (ASR) and Text-to-Speech (TTS).
  • ANUVADINI: An AI-based translation tool developed by the All-India Council of Technical Education (AICTE) to facilitate content translation into Indian languages.

Tribal Language Support and Preservation

  • Targeted Tribal Languages: Specific digital and linguistic support is provided for tribal languages including Santhali, Bhili, Mundari, Kharia, Ho-Hindi, Kurukh, Korwa, Bhumij, and Malto.
  • Linguistic Datasets: The Linguistic Data Consortium for Indian Languages (LDC-IL) has developed a 50-hour Santhali speech dataset and parallel text corpora of 5,200 sentences each for Santhali and Mundari.
  • Foundational Primers: CIIL and NCERT have created foundational primers and linguistic datasets to promote mother tongue-based education in tribal regions.

Educational Resources and Digital Access

  • SWAYAM Platform: A 12-week Santhali language course has been launched to promote digital learning.
  • Public Repositories: Datasets and AI models are publicly accessible via the BHASHINI platform, AIKosh, and the Medha Bhashika website.
  • Bharatavani Portal: Hosts diverse resources including Bhashakosha (learning repository), Jnanakosha (encyclopedia), and Pathyapustakakosha (textbooks).

Impact Analysis

Academic and Technical Institutions Impact: Institutions like IIT Bombay and its consortium members are the primary drivers of AI research and infrastructure development for Indian languages. Action Required: Continue the core development of the BharatGen project and collaborate with BHASHINI's 70+ research partner institutes to refine AI models.

Tribal and Multilingual Students Impact: These communities gain access to educational content and technical tools in their mother tongues, reducing language barriers in education. Action Required: Engage with the 12-week Santhali course on SWAYAM and utilize the primers and textbooks available on the Bharatavani portal.

Developers and Researchers Impact: They receive public access to massive datasets, including 246 million parallel sentence pairs and 3.7 million monolingual text entries. Action Required: Access the AIKosh and Medha Bhashika platforms to utilize open-source models and datasets for building language-enabled digital services.

Government Bodies (AICTE, CIIL, NCERT) Impact: These organizations are responsible for creating the foundational content and tools that bridge the gap between traditional education and AI technology. Action Required: Maintain and update digital repositories and continue the digitization of text and speech data across all scheduled and tribal languages.

Key Entities Referenced

National Education Policy (NEP) 2020: The overarching policy framework that prioritizes multilingualism and the promotion of Indian languages through technology and research. BHASHINI: An initiative under the Digital India Programme that hosts AI-based language models and translation services for 22 scheduled and several tribal languages. BharatGen: A multimodal Large Language Model (LLM) project aimed at creating inclusive AI solutions and digital infrastructure for all scheduled Indian languages. Central Institute of Indian Languages (CIIL): A specialized body under the Ministry of Education that develops foundational primers, linguistic datasets, and digital resources for tribal languages like Santhali and Mundari. ANUVADINI: An AI-based translation tool developed by the All-India Council of Technical Education (AICTE) to facilitate the translation of educational content into Indian languages.
Official Gazette PDF Record Download Official PDF (Parliament Question: Mother Tongue Based...) →

Official Gazette Notification PDF Viewer

See Full Document Text & PDF Transcript
Ministry of Education Parliament Question: Mother Tongue Based AI Learning System for Tribal and Multilingual Area प्रव तथ: 12 AUG 2026 4:49PM by PIB Delhi The National Education Policy (NEP) 2020 highlights the importance of multilingualism and places strong emphasis on the promotion of all Indian languages. In alignment with the objectives of NEP 2020, the Government of India has undertaken several initiatives promoting education, preservation and research on Indian languages. These efforts have further been augmented through the adoption of Artificial Intelligence and Machine Learning (AI/ML) technologies, including the development of Large Language Models, translation tools and language-enabled digital services in Indian languages. Government of India has initiated the BharatGen project, which is a multimodal Large Language Model (LLM) focused on developing efficient and inclusive AI solutions that support all 22 scheduled languages and enable the creation of a robust digital AI infrastructure. BharatGen is anchored at IIT Bombay, where the core development of language technologies is undertaken. The initiative is implemented through a consortium of leading academic institutions including IIT Kanpur, IIT Madras, IIT Hyderabad, IIIT Hyderabad, IIM Indore, and IIT Mandi. Further, to facilitate the translation of content into Indian languages, technological advancements include the development of AI-based translation tools such as ANUVADINI by the All-India Council of Technical Education (AICTE) and BHASHINI, an initiative under the Digital India Programme. BHASHINI has developed state-of-the-art AI models for Indian languages through a collaboration of over 70 research partner institutes. BHASHINI platform hosts a repository of over 360 AI-based language models and provides more than 22 specialized language services. These services include Automatic Speech Recognition (ASR), Machine Translation (MT), Text-to-Speech (TTS), Optical Character Recognition (OCR) and Transliteration. The dataset corpus includes 246 million parallel sentence pairs and 3.7 million monolingual text entries. All datasets and models are publicly accessible via BHASHINI platform or through the Digital India BHASHINI Division account on the AIKosh platform. Further, the National Language Translation Mission through the BHASHINI platform is digitising large volumes of text and speech data across all 22 scheduled languages. The platform also supports tribal languages such as Bhili and Santhali. In addition, the Central Institute of Indian Languages (CIIL), a subordinate office under the Ministry of Education in collaboration with NCERT, has developed foundational primers and linguistic datasets in various tribal languages including Santhali. Furthermore, it has launched a 12-week Santhali language course on the SWAYAM platform to promote mother tongue-based education. Under the Linguistic Data Consortium for Indian Languages (LDC-IL), a project of CIIL, linguistic resources have been developed for tribal languages, including Santhali and Mundari. The resources available for Santhali include a speech dataset comprising 50 hours of audio recordings from 60 speakers. In addition, parallel text corpora covering linguistic features and structures have been developed, comprising 5,200 sentences each in Santhali and Mundari. LDC-IL has also developed web-based applications and tools to support research and development in Indian languages. Several of these tools, hosted at CIIL data centres, are accessible through the Medha Bhashika website https://medha.ciil. org CIIL has also prepared foundational primers in several languages namely Mundari, Kharia, Ho-Hindi, Kurukh, Korwa (Jharkhand and Chhattisgarh), Bhumij, Santhali-Hindi and Malto. Further, CIIL’s Bharatavani portal hosts diverse Santhali resources such as Bhashakosha (language learning repository), Jnanakosha (encyclopaedia), Pathyapustakakosha (textbooks) and Shabdakosha (dictionaries). This information was given by the Minister of State for Education, Dr Sukanta Majumdar in a written reply in the Rajya Sabha today. ***** RRTN (रलीज़ आईडी: 2298335) आगंतुक पटल : 497 इस वज्ञ को इन भाषाओ ंम पढ़: Urdu , ही , Bengali

Continue your research