Skip to main content

Inference Engineer at Nuance Labs

Department: Research

Compensation

$200,000 – $300,000/yr

Location
Seattle, Washington
Type
Full-time
Level
entry
Posted

Description

Nuance Labs is hiring an early-career Inference Engineer to optimize end-to-end inference for its real-time, full-duplex multimodal AI avatar system. The role focuses on LLM, audio, and diffusion-model inference; KV cache and memory-efficient attention strategies; vLLM, SGLang, and TensorRT-LLM serving; latency and throughput profiling; quantization; kernel and batching optimizations; and internal optimization tooling. Candidates should have a BS, MS, or PhD in a relevant field, strong Python and PyTorch skills, and exposure to inference frameworks or ML systems. The position is in-person in Seattle five days per week, offers $200,000–$300,000 base salary plus equity, and provides visa sponsorship and health benefits.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation