Description
Reducto is hiring an ML Infra Engineer to build and maintain training and inference frameworks, develop benchmarks, explore state-of-the-art advances, and design reliable distributed systems across multi-node, multi-GPU environments. The role focuses on scaling model training and inference workloads, improving GPU utilization and cost efficiency, and creating tooling and observability for ML engineers. It requires strong Python, systems engineering, Kubernetes, and distributed training experience, with at least three years of experience. The position is in person at Reducto’s San Francisco office and includes medical, dental, and vision insurance.

