Description
Nebius is hiring a Senior Machine Learning Engineer on its Applied AI team to own model and endpoint optimization from artifacts through production deployment. The role focuses on improving latency, throughput, memory efficiency, GPU utilization, and cost per token while maintaining model quality and reliability. Responsibilities include optimizing LLM and VLM endpoints, deploying and extending inference engines, building compression and acceleration workflows, creating reproducible benchmarks, diagnosing bottlenecks, and collaborating with kernel and platform engineers. The position requires strong Python and PyTorch skills, hands-on experience with LLM or VLM inference systems, and knowledge of modern inference stacks.
