**Executive Summary**
This document provides information on BharatGen AI, India's government-supported initiative to develop foundational AI models tailored for Indian languages and societal contexts. The AI models support multiple modalities, including text, speech, and vision. The document lists the institutions involved in the BharatGen consortium and their roles. The report was posted on February 5, 2026.
**Key Points / Main Content**
* **BharatGen AI Initiative:**
* A national initiative to develop sovereign foundational AI models for Indian languages.
* Supports multiple modalities: text (Large Language Models), speech (Text-to-Speech and Automatic Speech Recognition), and vision-language systems.
* **Language Support:**
* Currently supports 15 Indian languages, including Hindi, Assamese, Bengali, Gujarati, Kannada, Maithili, Malayalam, Marathi, Nepali, Oriya, Punjabi, Sanskrit, Sindhi, Tamil, and Telugu.
* Plans to cover all 22 scheduled Indian languages.
* **Domain-Specific Models:**
* Released fine-tuned models for Ayurveda (Ayur Param), Indian agriculture (Agri Param), and Indian legal domain (Legal Param).
* Models are useful for healthcare, agriculture, education, and governance applications.
* **Technology Innovation Hubs:**
* Two hubs are part of the BharatGen network: TIH Foundation for IoT and IoE, IIT Bombay and IITM Pravartak Technologies Foundation, IIT Madras.
* **Consortium Institutions and Roles:**
* Indian Institute of Technology (IIT), Bombay: Lead institution, guiding research and integration across consortium partners.
* International Institute of Information Technology (IIIT), Hyderabad: Vision-language document modeling.
* IIT Madras: Speech foundation model development and evaluation.
* IIT Kanpur: Legal AI research, domain-specific datasets, and developing tokenization strategies for multilingual models.
* IIT Hyderabad: Advanced tokenization and vocabulary optimization for large multilingual LLMs.
* IIT Mandi: Inclusive multilingual model development and research on efficient training strategies for LLMs.
* Indian Institute of Management (IIM), Indore: Bharat-centric evaluation and benchmarking of LLMs, multilingual and multimodal data collection.
**Impact Analysis**
**Ministry of Science & Technology**
*Impact:* Will gain insights on current projects and impact
*Action Required:* None specified
**Consortium institutions (IIT Bombay, IIIT Hyderabad, IIT Madras, IIT Kanpur, IIT Hyderabad, IIT Mandi, IIM Indore)**
*Impact:* Required to fulfill their designated roles within the BharatGen project.
*Action Required:* Continue fulfilling their stated roles in the BharatGen consortium, contributing to research, development, and evaluation efforts.
**Users of AI Models**
*Impact:* Users in the healthcare, agriculture, education, and governance sectors will benefit from the BharatGen AI models tailored for Indian languages and societal contexts.
*Action Required:* Explore the potential applications of BharatGen models in their respective domains.
Key Entities Referenced
BharatGen: Government-supported national initiative to develop foundational AI models tailored to Indian languages and societal contexts.
Ministry of Science & Technology: The ministry overseeing the BharatGen initiative.
Department of Science and Technology (DST): Department involved in the National Mission on Interdisciplinary Cyber-Physical Systems (NM-ICPS), which supports the Technology Innovation Hubs involved in BharatGen.
Ministry of Science & Technology
PARLIAMENT QUESTION: ROLE OF
BHARATGEN AI
Posted On: 05 FEB 2026 3:23PM by PIB Delhi
BharatGen is the first government supported national initiative to develop a range of sovereign
foundational AI models tailored to Indian languages and societal contexts. It spans multiple modalities,
including text (via Large Language Models), speech (Text-to-Speech and Automatic Speech Recognition),
and vision-language systems.
Currently, BharatGen’s AI models support 15 Indian languages which include Hindi, Assamese, Bengali,
Gujarati, Kannada, Maithili, Malayalam, Marathi, Nepali, Oriya, Punjabi, Sanskrit, Sindhi, Tamil and
Telugu. Soon, all 22 scheduled Indian languages will be covered.
BharatGen has released domain specific fine-tuned models for Ayurveda (Ayur Param), Indian agriculture
(Agri Param) and Indian legal domain (Legal Param). In addition, all BharatGen models (text, speech and
vision) are useful for applications across healthcare, agriculture, education and governance.
Two Technology Innovation Hubs namely TIH Foundation for IoT and IoE, IIT Bombay and IITM
Pravartak Technologies Foundation, IIT Madras under National Mission on Interdisciplinary Cyber-
Physical Systems (NM-ICPS) of Department of Science and Technology (DST) are currently active as part
of BharatGen network.
The following institutions are a part of the BharatGen consortium:
Institution Name Role in BharatGen
Indian Institute of Technology, Lead institution, guiding research and integration across
Bombay consortium partners
International Institute of Vision-language document modeling
Information Technology,
Hyderabad
Indian Institute of Technology, Speech foundation model development and evaluation
Madras
Indian Institute of Technology, Legal AI research, domain-specific datasets, and developing
Kanpur tokenization strategies for multilingual models
Indian Institute of Technology, Advanced tokenization and vocabulary optimization for large
Hyderabad multilingual LLMsIndian Institute of Technology, Inclusive multilingual model development and research on
Mandi efficient training strategies for LLMs
Indian Institute of Management, Bharat-centric evaluation and benchmarking of LLMs,
Indore multilingual and multimodal data collection
*****
NKR/FK
(Release ID: 2223738) Visitor Counter : 187
Read this release in: ही , Tamil , Telugu