Description
Sarvam is hiring a Vision-Language Model Researcher to work across the full lifecycle of VLM development, including research, data, training, evaluation, production, and model robustness. The role focuses on vision-language architectures, multilingual training methods, data strategies, Indic multimodal benchmarks, failure modes, and interpretability, with close collaboration with engineers. Candidates should have strong research and experimental skills, a track record of impactful work, and strong PyTorch experience; relevant advanced degrees, publications, multilingual or low-resource experience, document understanding, OCR, and large-scale data curation are bonuses.
