Description
Nuance Labs is hiring a Machine Learning Systems Engineer to own distributed training infrastructure for large-scale omni model pretraining. The role builds and operates the training runtime, including job orchestration, distributed execution, checkpointing, recovery, monitoring, debugging, GPU communication, and performance optimization across large GPU clusters. It also develops infrastructure for multimodal data loading, temporal alignment, variable sequence handling, and memory-efficient training. The position requires hands-on experience with large-scale distributed training, major training stacks, and omni or multimodal systems, and offers a Seattle-based in-person role with visa sponsorship.
