Description
The Foundation Model Training Engineer will build and scale distributed pre-training frameworks for large language models, including PyTorch or JAX stacks, DeepSpeed, FSDP, Megatron-LM, and MosaicML Composer. The role develops custom optimizers and attention methods, CUDA/Triton kernels, mixed-precision training, logging, metrics, experiment-tracking tools, and ablation studies; it also involves distributed debugging, productionizing research systems, and mentoring interns and junior engineers. The position requires at least five years of large-scale deep-learning training experience, leadership of a large-scale transformer pre-training run, and expertise with distributed GPU training and NCCL/GLOO debugging. Visa sponsorship is available, and total compensation targets $300,000–$600,000 in base salary plus a target bonus.
