Description
Nuance Labs is hiring an early-career Inference Engineer to optimize end-to-end inference for its real-time, full-duplex multimodal AI avatar system. The role focuses on LLM, audio, and diffusion-model inference; KV cache and memory-efficient attention strategies; vLLM, SGLang, and TensorRT-LLM serving; latency and throughput profiling; quantization; kernel and batching optimizations; and internal optimization tooling. Candidates should have a BS, MS, or PhD in a relevant field, strong Python and PyTorch skills, and exposure to inference frameworks or ML systems. The position is in-person in Seattle five days per week, offers $200,000–$300,000 base salary plus equity, and provides visa sponsorship and health benefits.
