Description
Sarvam is hiring a senior Performance Engineer, Inference to own the production serving path for large distributed models. The role integrates model and kernel artifacts into a multi-node, multi-tenant serving stack; extends distributed prefill-decode, KV/cache transfer, routing, and scheduling; builds and trains speculative decoders; and owns latency, throughput, GPU utilization, and cost metrics. Candidates need 5+ years in ML systems, including 2+ years at production scale, experience with 100B+ parameter models, source-level knowledge of SGLang, vLLM, Dynamo, or TensorRT-LLM, distributed serving expertise, speculative-decoding competency, and C++/CUDA proficiency.
