Executive Summary:
The Ministry of Science and Technology addresses questions regarding the BharatGen AI Models, a government-supported initiative to develop foundational AI models tailored to Indian languages and societal contexts. The initiative aims to cover 15 Indian languages by December 2025 and all 22 scheduled languages by June 2026. Current applications are in the pilot phase within the agriculture, governance, and defense sectors, with plans for broader deployment across states and districts.
Key Points / Main Content:
* **Overview of BharatGen:**
* BharatGen is a national initiative for developing sovereign foundational AI models.
* It spans multiple modalities: text (Large Language Models), speech (Text-to-Speech and Automatic Speech Recognition), and vision-language systems.
* **Language Coverage and Roadmap:**
* Currently supports 9 Indian languages: Hindi, Marathi, Tamil, Malayalam, Bengali, Punjabi, Gujarati, Telugu, and Kannada.
* Target: 15 Indian languages by December 2025 (including Assamese, Maithili, Nepali, Odia, Sanskrit, and Sindhi).
* Target: All 22 scheduled Indian languages by June 2026.
* **Applications:**
* Pilot applications developed in agriculture, governance, and defense sectors.
* Plans to make applications available across all states and districts upon full deployment.
* **Implementation and Network:**
* Implemented under the National Mission on Interdisciplinary Cyber-Physical Systems (NMICPS) of the Department of Science and Technology (DST).
* Technology Innovation Hubs (TIHs) involved:
* TIH Foundation for IoT and IoE, IIT Bombay (Maharashtra): Central program implementation and coordination hub.
* IITM Pravartak Technologies Foundation, IIT Madras (Tamil Nadu): Implementation partner focusing on solutioning and real-world deployment in governance, security, and media.
* **BharatGen Consortium Institutions and Roles:**
* Indian Institute of Technology, Bombay: Lead institution, guiding research and integration.
* International Institute of Information Technology, Hyderabad: Vision-language document modeling.
* Indian Institute of Technology, Madras: Speech foundation model development and evaluation.
* Indian Institute of Technology, Kanpur: Legal AI research, domain-specific datasets, and tokenization strategies.
* Indian Institute of Technology, Hyderabad: Advanced tokenization and vocabulary optimization.
* Indian Institute of Technology, Mandi: Inclusive multilingual model development and efficient training strategies.
* Indian Institute of Management, Indore: Bharat-centric evaluation and benchmarking, multilingual and multimodal data collection.
* **Partnerships and Deployment:**
* May explore partnerships with research institutions in Karnataka.
* Currently under pilot deployment; not yet released for public and institutional use.
* Plans to extend applicability across all states and districts after full deployment.
Impact Analysis:
* **Government of India (Ministry of Science and Technology, Department of Science and Technology):**
* *Impact:* Responsible for the strategic direction, funding, and oversight of the BharatGen initiative.
* *Action Required:* Continue to monitor progress against milestones, facilitate partnerships, and prepare for wider deployment.
* **Technology Innovation Hubs (IIT Bombay, IIT Madras):**
* *Impact:* Key implementers of the BharatGen project, responsible for technical development, coordination, and real-world deployment.
* *Action Required:* Continue development, testing, and refinement of AI models; coordinate with consortium partners; prepare for broader implementation.
* **Institutions in the BharatGen Consortium (IITs, IIM Indore, IIIT Hyderabad):**
* *Impact:* Contribute specialized expertise in AI research, language modeling, data collection, and evaluation.
* *Action Required:* Fulfill assigned roles in research, development, and testing; collaborate with other consortium members.
* **Academic and Research Institutions (Potential Partners in Karnataka):**
* *Impact:* May have the opportunity to contribute to BharatGen research, language development, or pilot applications.
* *Action Required:* Explore potential partnership opportunities with the BharatGen consortium.
* **Citizens of India (eventually, across all states and districts):**
* *Impact:* Potential beneficiaries of AI applications in healthcare, agriculture, education, and governance tailored to Indian languages and contexts.
* *Action Required:* Await full deployment of BharatGen AI and its integration into public services.
Key Entities Referenced
BharatGen AI Models: A government supported national initiative to develop sovereign foundational AI models tailored to Indian languages and societal contexts.
National Mission on Interdisciplinary Cyber-Physical Systems (NMICPS): A Department of Science and Technology (DST) initiative under which BharatGen is implemented as a project.
Department of Science and Technology: Ministry under which the BharatGen project is being implemented
IIT Bombay, Maharashtra: Host of the Technology Innovation Hub (TIH) for BharatGen, serving as the central program implementation and coordination hub.
IIT Madras, Tamil Nadu: An implementation partner focusing on solutioning and real-world deployment of BharatGen AI technologies.
Hindi: One of the 9 Indian languages currently supported by BharatGen's AI model.
Dewas Shajapur, Madhya Pradesh: A specific district mentioned in the context of potential BharatGen AI applications.
Dr. Jitendra Singh: Minister of State (Independent Charge) of the Ministry of Science and Technology and Earth Sciences
GOVERNMENT OF INDIA
MINISTRY OF SCIENCE AND TECHNOLOGY
DEPARTMENT OF SCIENCE AND TECHNOLOGY
LOK SABHA
STARRED QUESTION No. *257
ANSWERED ON 06/08/2025
BHARATGEN AI MODELS
*257. SHRI JANARDAN SINGH SIGRIWAL:
SHRI JANARDAN MISHRA:
Will the Minister of SCIENCE AND TECHNOLOGY be pleased to state:
(a) the details of the number of Indian languages currently supported by
BharatGen's AI model and the road map for achieving full language coverage;
(b) the specific applications of BharatGen in the healthcare, agriculture,
education, and governance sectors, State and District-wise including Dewas-
Shajapur, Madhya Pradesh;
(c) the number of Technology Innovation Hubs and Translational Research
Parks currently active under the BharatGen network and their thematic focus,
State-wise;
(d) the details of the institutions involved in the BharatGen consortium and
their role in its development and deployment;
(e) whether the Government proposes to partner with academic or research
institutions in Dakshina Kannada to promote BharatGen's research, language
development or pilot applications in the region and if so, the details thereof;
(f) whether the Government has launched or is in the process of launching
BharatGen AI for public and institutional use in the Dewas-Shajapur, Lok Sabha
Constituency of Madhya Pradesh; and
(g) if so, the details thereof including the objectives, features and intended
user groups of BharatGen AI and the benefits to the rural and semi-urban areas
of Dewas-Shajapur?
ANSWER
MINISTER OF STATE (INDEPENDENT CHARGE) OF THE
MINISTRY OF SCIENCE AND TECHNOLOGY AND EARTH SCIENCES
(DR. JITENDRA SINGH)
विज्ञान और प्रौद्योगिकी तथा पथ्ृ िी विज्ञान मंत्रालय के राज्य मंत्री (स्ितंत्र प्रभार)
(डॉ. जितेंद्र स हं )
(a) to (g): A statement is laid on the Table of the House.
Page 1 of 3STATEMENT AS REFERRED IN REPLY TO PARTS (a) to (g) OF LOK
SABHA STARRED QUESTION NO. 257 FOR 06.08.2025 REGARDING
“BHARATGEN AI MODELS”
(a) to (g): BharatGen is the first government supported national
initiative to develop a range of sovereign foundational AI models tailored
to Indian languages and societal contexts. It spans multiple modalities,
including text (via Large Language Models), speech (Text-to-Speech and
Automatic Speech Recognition), and vision-language systems.
Currently, BharatGen models cover 9 Indian languages which include
Hindi, Marathi, Tamil, Malayalam, Bengali, Punjabi, Gujarati, Telugu, and
Kannada.
Roadmap of language coverage across these models includes the
following milestones:
By December 2025, a total of 15 Indian Languages (including
Assamese, Bengali, Gujarati, Hindi, Kannada, Maithili, Malayalam,
Marathi, Nepali, Odia, Punjabi, Sanskrit, Sindhi, Tamil and Telugu)
will be covered.
By June 2026, all 22 scheduled Indian languages will be covered.
BharatGen has developed applications in the sectors of agriculture,
governance and defence, wherein pilots have been carried out. Once
deployed fully, it is planned to make these applications available across
all states and districts.
BharatGen is implemented as a project under the National Mission on
Interdisciplinary Cyber-Physical Systems (NM-ICPS) of Department of
Science and Technology (DST). As part of the BharatGen network,
following Technology Innovation Hubs (TIHs) are currently active:
1. TIH Foundation for IoT and IoE, IIT Bombay (Maharashtra)
This TIH is hosting the BharatGen project and serves as the central
program implementation and coordination hub. It is responsible for
end-to-end execution, including managing the national academic
consortium, overseeing the development of sovereign foundational
AI models across text, speech, and vision, and driving ecosystem
partnerships for compute, data, and talent. The TIH also drives
governance and strategic planning for the BharatGen project,
ensuring cohesive progress across all stakeholders.
Page 2 of 32. IITM Pravartak Technologies Foundation, IIT Madras (Tamil Nadu)
This TIH serves as an implementation partner, focusing on
solutioning and real-world deployment of BharatGen AI
technologies. Thematic focus areas include governance, security,
and media-related use cases.
The following institutions are a part of the BharatGen consortium:
Institution Name Role in BharatGen
Indian Institute of Lead institution, guiding research
Technology, Bombay and integration across
consortium partners
International Institute of Vision-language document
Information Technology, modeling
Hyderabad
Indian Institute of Speech foundation model
Technology, Madras development and evaluation
Indian Institute of Legal AI research, domain-
Technology, Kanpur specific datasets, and developing
tokenization strategies for
multilingual models
Indian Institute of Advanced tokenization and
Technology, Hyderabad vocabulary optimization for large
multilingual LLMs
Indian Institute of Inclusive multilingual model
Technology, Mandi development and research on
efficient training strategies for
LLMs
Indian Institute of Bharat-centric evaluation and
Management, Indore benchmarking of LLMs,
multilingual and multimodal data
collection
While BharatGen is working with the above consortia members, it may
explore partnerships with research institutions in Karnataka.
Since BharatGen AI is currently under pilot deployment phase, it has not
been released for public and institutional use. Once fully deployed, there
is a plan to extend its applicability across all states and districts.
*****
Page 3 of 3