Skip to main content

Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind at Google

Compensation · Listed in posting

$207,000 – $300,000/yr

Location
Mountain View, California
Type
Full-time
Level
senior
Posted

Description

DeepMind is hiring an Inference Performance Engineer to improve the speed, cost, and efficiency of large-scale AI model inference. The role analyzes and optimizes inference workloads across application, model, and distributed fleet layers; designs optimization techniques; resolves bottlenecks; models latency-to-cost tradeoffs; and develops metrics and tools to track compute usage. The position requires a bachelor's degree in a technical field or equivalent practical experience, 8 years of software development experience, and experience with Python, C++, serving codebases, AI model execution, throughput-latency tradeoffs, and modern serving architectures. Preferred qualifications include experience with LLM inference serving, open-source inference frameworks, ML profiling, distributed-system observability, and GPU/TPU accelerator performance.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation