Description
Sarvam is hiring an Intel Inference Engineer to own end-to-end deployment of models on Intel NPUs, integrated and discrete GPUs, and x86/AMD64 CPUs. The role covers OpenVINO model conversion and quantization, driver-version compatibility, ONNX Runtime fallback paths, x86 CPU optimization, AVX-512 and AMX intrinsics, and Intel device CI and regression testing. Candidates need at least five years of machine-learning deployment experience, including two years with Intel inference stacks, plus production OpenVINO and x86 profiling expertise.
