Skip to main content

Senior AI Inference Optimization Engineer at Nuance Labs

Department: Research

Compensation

$250,000 – $350,000/yr

Location
Seattle, Washington
Type
Full-time
Level
senior
Posted

Description

Nuance Labs is hiring a senior AI inference optimization engineer to own end-to-end optimization of LLM, audio, and diffusion-model inference for a real-time, full-duplex multimodal system. The role focuses on KV cache strategies, serving frameworks such as vLLM and SGLang, latency and throughput profiling, quantization, CUDA or Triton kernel optimization, and internal optimization tooling. Candidates should have significant production or high-traffic inference experience, strong Python and PyTorch skills, and familiarity with inference acceleration techniques. The position is in-person in Seattle five days per week and offers $250,000–$350,000 base salary plus equity and visa sponsorship.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation