Description
The company is hiring Senior and Staff Inference Engineers to build and operate production-grade model-serving and inference systems for high-throughput, low-latency AI workloads. The role focuses on GPU utilization, latency and throughput optimization, scalability, reliability, monitoring, and collaboration with AI training, GPU performance, orchestration, platform engineering, and operations teams. The position is hybrid in Bellevue, Washington, with approximately three days per week in the office, and requires U.S. work authorization; visa sponsorship is not available.
