Description
The Distributed Training Engineer will optimize large-scale distributed training performance on AWS Trainium by working across PyTorch, JAX, the Neuron compiler and runtime, and the hardware-software boundary. The role owns parallelism strategies, reduced-precision formats, end-to-end profiling, and performance fixes across compute, memory, collectives, and host overhead, while translating gaps into framework requirements and contributing to open-source frameworks. The position requires at least three years of professional software development experience and two years of design or architecture experience, and offers base salaries of $165,200–$223,600 annually in Cupertino and $143,700–$194,400 annually in Seattle.
