Description
Sarvam is hiring a Data Scientist to own the evaluations function for its sovereign AI vertical. The role designs and maintains domain-specific evaluation frameworks and harnesses for model and system quality, defines metrics with domain experts, runs pre- and post-deployment evaluation cycles, identifies failure modes and distribution shifts, operationalizes eval pipelines with MLOps, manages evaluation datasets, and publishes quality reports. The position requires 3–6 years of data science, machine learning research, or applied AI experience, at least 2 years working with LLMs in production, strong statistics and probability fundamentals, Python proficiency, and experience with evaluation frameworks, prompt engineering, model fine-tuning, or RLHF.
