Bodhan AI brings four open models to education in Indian languages
Bodhan AI and AI4Bharat released four open-weight models for speech recognition, text to speech, machine translation and optical character recognition. Analytics India Magazine and ETEducation report their language coverage and how interfaces hosted on sovereign infrastructure are being opened to education applications. The sources also identify NVIDIA's Nemotron family as the models' foundation.
Artificial Intelligence··Midday
Four tasks, four language ranges
Bodhan AI and AI4Bharat released four open-weight models for people building education tools in Indian languages. The speech-recognition model supports 27 languages, text to speech supports 23, machine translation supports 22, and optical character recognition supports 23. Analytics India Magazine describes the packages as digital public goods for students and teachers. ETEducation reports the same four tasks and language counts. The release therefore covers speech, translation and document reading through separate models rather than one general-purpose system.[1], [2]
Hosting alongside open weights
Both reports say the models come with hosted interfaces as well as open weights. ETEducation says the interfaces run on sovereign infrastructure and are intended to let organisations build Indian-language applications without recreating foundation capabilities. Analytics India Magazine adds that hosting uses data-anonymisation procedures. The arrangement offers one route for developers that want to run the weights in their own environment and another for education institutions that want a hosted interface. Neither source gives a figure for adoption in real classrooms or products.[1], [2]
From Nemotron to education applications
The models build on NVIDIA's open Nemotron family. Both reports name NeMo, TensorRT-LLM and vLLM in the training and serving layers. ETEducation says the speech-recognition model was post-trained to cover regional dialects and accents. Analytics India Magazine reports AI4Bharat's Mitesh Khapra describing the speech and vision models as public goods made available on digital public infrastructure. The concrete product boundary is the released tasks and language coverage; the reports provide no classroom outcome or independent accuracy measurement.[1], [2]