Read or download the official PDF of this gazette notification issued by the Ministry of Education on 12th August 2026. Classified under Press Release.
Executive Summary
This report outlines the Ministry of Education’s initiatives to implement mother tongue-based AI learning systems for tribal and multilingual areas in alignment with NEP 2020. The document, based on a parliamentary reply dated August 12, 2026, details the development of Large Language Models (LLMs) and digital translation tools to support 22 scheduled languages and various tribal dialects. Key action items include the public release of linguistic datasets and the launch of specialized language courses on digital platforms.
Key Points / Main Content
AI and LLM Development Initiatives
BharatGen Project: A multimodal Large Language Model (LLM) initiative anchored at IIT Bombay and supported by a consortium (including IIT Kanpur, IIT Madras, and others) to build inclusive AI infrastructure for all 22 scheduled languages.
BHASHINI Platform: An initiative under the Digital India Programme that hosts over 360 AI-based models and provides 22 specialized services such as Automatic Speech Recognition (ASR) and Text-to-Speech (TTS).
ANUVADINI: An AI-based translation tool developed by the All-India Council of Technical Education (AICTE) to facilitate content translation into Indian languages.
Tribal Language Support and Preservation
Targeted Tribal Languages: Specific digital and linguistic support is provided for tribal languages including Santhali, Bhili, Mundari, Kharia, Ho-Hindi, Kurukh, Korwa, Bhumij, and Malto.
Linguistic Datasets: The Linguistic Data Consortium for Indian Languages (LDC-IL) has developed a 50-hour Santhali speech dataset and parallel text corpora of 5,200 sentences each for Santhali and Mundari.
Foundational Primers: CIIL and NCERT have created foundational primers and linguistic datasets to promote mother tongue-based education in tribal regions.
Educational Resources and Digital Access
SWAYAM Platform: A 12-week Santhali language course has been launched to promote digital learning.
Public Repositories: Datasets and AI models are publicly accessible via the BHASHINI platform, AIKosh, and the Medha Bhashika website.
Bharatavani Portal: Hosts diverse resources including Bhashakosha (learning repository), Jnanakosha (encyclopedia), and Pathyapustakakosha (textbooks).
Impact Analysis
Academic and Technical InstitutionsImpact: Institutions like IIT Bombay and its consortium members are the primary drivers of AI research and infrastructure development for Indian languages.
Action Required: Continue the core development of the BharatGen project and collaborate with BHASHINI's 70+ research partner institutes to refine AI models.
Tribal and Multilingual StudentsImpact: These communities gain access to educational content and technical tools in their mother tongues, reducing language barriers in education.
Action Required: Engage with the 12-week Santhali course on SWAYAM and utilize the primers and textbooks available on the Bharatavani portal.
Developers and ResearchersImpact: They receive public access to massive datasets, including 246 million parallel sentence pairs and 3.7 million monolingual text entries.
Action Required: Access the AIKosh and Medha Bhashika platforms to utilize open-source models and datasets for building language-enabled digital services.
Government Bodies (AICTE, CIIL, NCERT)Impact: These organizations are responsible for creating the foundational content and tools that bridge the gap between traditional education and AI technology.
Action Required: Maintain and update digital repositories and continue the digitization of text and speech data across all scheduled and tribal languages.
Key Entities Referenced
National Education Policy (NEP) 2020: The overarching policy framework that prioritizes multilingualism and the promotion of Indian languages through technology and research.
BHASHINI: An initiative under the Digital India Programme that hosts AI-based language models and translation services for 22 scheduled and several tribal languages.
BharatGen: A multimodal Large Language Model (LLM) project aimed at creating inclusive AI solutions and digital infrastructure for all scheduled Indian languages.
Central Institute of Indian Languages (CIIL): A specialized body under the Ministry of Education that develops foundational primers, linguistic datasets, and digital resources for tribal languages like Santhali and Mundari.
ANUVADINI: An AI-based translation tool developed by the All-India Council of Technical Education (AICTE) to facilitate the translation of educational content into Indian languages.
Ministry of Education
Parliament Question: Mother Tongue Based AI
Learning System for Tribal and Multilingual Area
प्रव तथ: 12 AUG 2026 4:49PM by PIB Delhi
The National Education Policy (NEP) 2020 highlights the importance of multilingualism and places
strong emphasis on the promotion of all Indian languages. In alignment with the objectives of NEP 2020,
the Government of India has undertaken several initiatives promoting education, preservation and research
on Indian languages. These efforts have further been augmented through the adoption of Artificial
Intelligence and Machine Learning (AI/ML) technologies, including the development of Large Language
Models, translation tools and language-enabled digital services in Indian languages.
Government of India has initiated the BharatGen project, which is a multimodal Large Language Model
(LLM) focused on developing efficient and inclusive AI solutions that support all 22 scheduled languages
and enable the creation of a robust digital AI infrastructure. BharatGen is anchored at IIT Bombay, where
the core development of language technologies is undertaken. The initiative is implemented through a
consortium of leading academic institutions including IIT Kanpur, IIT Madras, IIT Hyderabad, IIIT
Hyderabad, IIM Indore, and IIT Mandi.
Further, to facilitate the translation of content into Indian languages, technological advancements include
the development of AI-based translation tools such as ANUVADINI by the All-India Council of Technical
Education (AICTE) and BHASHINI, an initiative under the Digital India Programme.
BHASHINI has developed state-of-the-art AI models for Indian languages through a collaboration of over
70 research partner institutes. BHASHINI platform hosts a repository of over 360 AI-based language
models and provides more than 22 specialized language services. These services include Automatic
Speech Recognition (ASR), Machine Translation (MT), Text-to-Speech (TTS), Optical Character
Recognition (OCR) and Transliteration. The dataset corpus includes 246 million parallel sentence pairs
and 3.7 million monolingual text entries. All datasets and models are publicly accessible via BHASHINI
platform or through the Digital India BHASHINI Division account on the AIKosh platform. Further, the
National Language Translation Mission through the BHASHINI platform is digitising large volumes of
text and speech data across all 22 scheduled languages. The platform also supports tribal languages such
as Bhili and Santhali.
In addition, the Central Institute of Indian Languages (CIIL), a subordinate office under the Ministry of
Education in collaboration with NCERT, has developed foundational primers and linguistic datasets in
various tribal languages including Santhali. Furthermore, it has launched a 12-week Santhali language
course on the SWAYAM platform to promote mother tongue-based education.
Under the Linguistic Data Consortium for Indian Languages (LDC-IL), a project of CIIL, linguistic
resources have been developed for tribal languages, including Santhali and Mundari. The resources
available for Santhali include a speech dataset comprising 50 hours of audio recordings from 60
speakers. In addition, parallel text corpora covering linguistic features and structures have been
developed, comprising 5,200 sentences each in Santhali and Mundari. LDC-IL has also developed web-based applications and tools to support research and development in Indian languages. Several of these
tools, hosted at CIIL data centres, are accessible through the Medha Bhashika website https://medha.ciil.
org
CIIL has also prepared foundational primers in several languages namely Mundari, Kharia, Ho-Hindi,
Kurukh, Korwa (Jharkhand and Chhattisgarh), Bhumij, Santhali-Hindi and Malto. Further, CIIL’s
Bharatavani portal hosts diverse Santhali resources such as Bhashakosha (language learning repository),
Jnanakosha (encyclopaedia), Pathyapustakakosha (textbooks) and Shabdakosha (dictionaries).
This information was given by the Minister of State for Education, Dr Sukanta Majumdar in a written
reply in the Rajya Sabha today.
*****
RRTN
(रलीज़ आईडी: 2298335) आगंतुक पटल : 497
इस वज्ञ को इन भाषाओ ंम पढ़: Urdu , ही , Bengali