**Executive Summary**
This document presents the Ministry of Education's response to Lok Sabha Starred Question No. 34, addressing the promotion of scheduled Indian languages through AI language platforms. The government has undertaken several initiatives aligned with the National Education Policy (NEP) 2020, focusing on AI and Machine Learning technologies to support the 22 scheduled languages. The statement was laid on the table of the House on February 2, 2026.
**Key Points / Main Content**
* **Government Initiatives:**
* Aligned with NEP 2020, focusing on promoting Indian languages through AI/ML.
* Developing Large Language Models, translation tools, and language-enabled digital services.
* **BharatGen Project:**
* Multimodal Large Language Model project.
* Focuses on efficient and inclusive AI solutions supporting all 22 scheduled languages.
* Anchored at IIT Bombay, implemented through a consortium of leading academic institutions (IIT Kanpur, IIT Madras, IIT Hyderabad, IIIT Hyderabad, IIM Indore, IIT Mandi).
* **AI-Based Translation Tools:**
* Development of tools like ANUVADINI by AICTE and BHASHINI under the Digital India Programme.
* **BHASHINI Platform:**
* Developed through collaboration with over 70 research partner institutes.
* Hosts a repository of over 350 AI-based language models.
* Provides over 22 specialized language services (e.g., Automatic Speech Recognition, Machine Translation).
* National Language Translation Mission is digitizing text and speech data across all 22 scheduled languages.
* **IIT Madras Bodhan AI Foundation:**
* Focuses on delivering educational content and services in Indian languages through AI-enabled tools to improve access to quality education in mother tongue.
* **Linguistic Resources:**
* Developed by the Central Institute of Indian Languages (CIIL) under the Linguistic Data Consortium for Indian Languages (LDC-IL) scheme.
* 76 datasets covering all Scheduled languages have been released.
* LDC-IL has developed web-based applications and tools accessible through the Medha Bhashika website (https://medha.ciil.org).
* Steps are being taken for digitizing the corpora of the Tulu language.
**Impact Analysis**
**Ministry of Education**
* **Impact:** Responsible for overall policy and direction, as well as linguistic resources development.
* **Action Required:** Continue supporting and expanding initiatives like Bhashini and BharatGen, and oversee the digitization of additional languages.
**Academic and Research Institutions (e.g., IITs, IIMs, AICTE, CIIL)**
* **Impact:** Involved in the development and deployment of language technologies and research.
* **Action Required:** Continue research and model development, dataset preparation, and baseline AI model development.
**Startups and Technology Companies**
* **Impact:** Opportunity to build applications and solutions on top of the Bhashini language stack.
* **Action Required:** Develop and commercialize applications using the language technologies.
**Citizens/Students**
* **Impact:** Enhanced access to educational content and public services in their mother tongue.
* **Action Required:** Utilize available resources and tools for learning and accessing services.
Key Entities Referenced
National Education Policy (NEP) 2020: A policy highlighting the importance of multilingualism and promoting Indian languages.
BHASHINI: An AI-based language platform initiative under the Digital India Programme for developing AI models and multilingual services for Indian languages.
BharatGen: A multimodal Large Language Model Project focused on developing efficient and inclusive AI solutions that support all 22 Scheduled languages.
Ministry of Education: The ministry responsible for the initiatives promoting scheduled Indian languages through AI language platforms.
Data Consortium for Indian Languages (LDC-IL): A scheme providing resources for developing language technologies for government agencies, researchers, and commercial users.
GOVERNMENT OF INDIA
MINISTRY OF EDUCATION
DEPARTMENT OF HIGHER EDUCATION
LOK SABHA
STARRED QUESTION NO. 34
ANSWERED ON 02.02.2026
PROMOTION OF SCHEDULED INDIAN LANGUAGES THROUGH AI LANGUAGE
PLATFORMS
*34. SHRI SHASHANK MANI:
SHRI MANOJ TIWARI:
Will the Minister of EDUCATION be pleased to state:
(a) the details of initiatives undertaken by the Government to support all 22 scheduled Indian
languages through AI-based language platforms such as Bhashini and BharatGen;
(b) the scope of digitisation of linguistic data, development of multilingual tools and creation of
open language datasets under these initiatives;
(c) the manner in which academic institutions, startups and public sector agencies across the
country have been involved in the development and deployment of these language
technologies; and
(d) whether the Government proposes to expand the coverage of the platforms and include the
Tulu language, spoken predominantly in Dakshina Kannada and Udupi districts of
Karnataka, under these Al-based language platforms for education, governance and public
service delivery, and if so, the details and timelines thereof?
ANSWER
MINISTER OF EDUCATION
(SHRI DHARMENDRA PRADHAN)
(a) to (d): A Statement is laid on the table of the House.
*****STATEMENT REFERRED TO IN REPLY TO PARTS (A) to (D) IN RESPECT OF LOK
SABHA STARRED QUESTION NO. 34 ANSWERED ON 02.02.2026 REGARDING
“PROMOTION OF SCHEDULED INDIAN LANGUAGES THROUGH AI LANGUAGE
PLATFORMS” ASKED BY SHRI SHASHANK MANI & SHRI MANOJ TIWARI”,
HON’BLE MEMBERS OF PARLIAMENT
(a) to (d): The National Education Policy (NEP) 2020 highlights the importance of
multilingualism and places strong emphasis on the promotion of all Indian languages. In
alignment with the objectives of NEP 2020, the Government of India has undertaken several
initiatives promoting education, preservation and research on Indian languages. These efforts
have further been augmented through the adoption of Artificial Intelligence and Machine
Learning (AI/ML) technologies, including the development of Large Language Models,
translation tools and language-enabled digital services in Indian languages.
Government of India has started the BharatGen project, which is a multimodal Large Language
Model Project focused on developing efficient and inclusive AI solutions that support all 22
Scheduled languages and enable the creation of a robust digital AI infrastructure for India’s
unique socio-cultural context and diverse sectors. BharatGen is anchored at IIT Bombay, where
the core development of language technologies is undertaken. The initiative is implemented
through a consortium of leading academic institutions including IIT Kanpur, IIT Madras, IIT
Hyderabad, IIIT Hyderabad, IIM Indore, and IIT Mandi.
Further, to facilitate the translation of content into Indian languages, technological advancements
include the development of AI-based translation tools such as ANUVADINI by the All India
Council of Technical Education (AICTE) and BHASHINI, an initiative under the Digital India
Programme.
Through a collaboration of over 70 research partner institutes, BHASHINI has developed state-
of-the-art AI models for Indian languages. At present, BHASHINI platform hosts a repository of
over three hundred fifty AI-based language models and provides more than twenty-two
specialised language services. These services include Automatic Speech Recognition, Machine
Translation, Text-to-Speech, Optical Character Recognition, Transliteration and several other
language technology capabilities essential for enhancing multilingual digital access in the
country. The National Language Translation Mission through the BHASHINI platform is
digitising large volumes of text and speech data across all 22 scheduled languages. BHASHINI
follows a multi-stakeholder participation model involving academic and research institutions for
fundamental research, dataset preparation and development of baseline AI models. For instance,
IIT Madras has leveraged Bhashini and works closely with other academic partners on research
and model development, while startups and technology companies build applications and
solutions on top of the Bhashini language stack.
The Indian Language Technology Stack for education initiative of the IIT Madras Bodhan AI
Foundation aims to provide technology solutions for making learning more accessible, inclusive,
and effective. These efforts focus on delivering educational content and services in Indian
languages to improve access to quality education in mother tongue through AI enabled tools.
Further, extensive linguistic resources for scheduled Indian Languages have been developed by
the Central Institute of Indian Languages (CIIL) of the Ministry of Education under the LinguisticData Consortium for Indian Languages (LDC-IL) scheme. Since 2019, LDC-IL has provided
resources to government agencies, government-promoted initiatives, researchers, and commercial
and industrial users engaged in developing language technologies.
Under the LDC-IL scheme, 76 datasets covering all Scheduled languages have been released.
Apart from the datasets, LDC-IL has also developed web-based applications and tools to support
research and development in Indian languages. Several of these tools, hosted at CIIL data centres,
are accessible through the Medha Bhashika website https://medha.ciil.org. Expanding the
availability of high-quality digital resources for additional Indian languages, LDC-IL has taken
steps for digitizing the corpora of the Tulu language.
***