Description
DeepMind is hiring an Inference Performance Engineer to improve the speed, cost, and efficiency of large-scale AI model inference. The role analyzes and optimizes inference workloads across application, model, and distributed fleet layers; designs optimization techniques; resolves bottlenecks; models latency-to-cost tradeoffs; and develops metrics and tools to track compute usage. The position requires a bachelor's degree in a technical field or equivalent practical experience, 8 years of software development experience, and experience with Python, C++, serving codebases, AI model execution, throughput-latency tradeoffs, and modern serving architectures. Preferred qualifications include experience with LLM inference serving, open-source inference frameworks, ML profiling, distributed-system observability, and GPU/TPU accelerator performance.
