Description
OpenAI is hiring an SW Engineer to enable production workloads and end-to-end testing on new AI infrastructure platforms. The role involves porting and validating inference and training workloads, building benchmarks and stress tests, analyzing distributed training and inference performance, creating repeatable CI and lab test harnesses, and partnering with systems, fleet, and vendor teams. Candidates need a BS in CS or EE or equivalent practical experience, at least five years of experience in ML systems, performance engineering, distributed systems, or HPC, and strong expertise with PyTorch, distributed training, RDMA, NCCL or RCCL, Python, and performance profiling tools.
