Description
Atoms is hiring a Staff Machine Learning Infrastructure Engineer to design and scale the large-scale machine learning training infrastructure for autonomous transport models. The role focuses on Kubernetes-based distributed GPU training, high-volume concurrent ML jobs, experiment tracking and MLOps, petabyte-scale data pipelines, and autonomous model validation. It requires 8+ years of software engineering experience, backend systems programming in Go, Python, Java, or similar, Kubernetes expertise, distributed ML frameworks, MLOps pipelines, and high-throughput data engineering. The position is onsite in San Francisco and offers a base salary of $224,000 to $280,000 per year, with equity and other benefits.
