Description
Reflection is hiring a distributed-training engineer to build and scale systems for frontier-model pre-training. The role involves designing and operating large-scale training runs, developing infrastructure across thousands of GPUs, optimizing throughput and GPU utilization, building training pipelines, and debugging distributed-training bottlenecks. Candidates should have experience with large-scale distributed training frameworks, model parallelism, GPU communication libraries, and foundation-model training workflows.
