Description
Epsilon Health is hiring an ML infrastructure engineer to design and build distributed training, reinforcement learning, inference, evaluation, and production deployment systems for large foundation models in medical imaging and diagnostics. The role partners with researchers to translate experimentation workflows into production-ready systems, owns GPU-efficient training infrastructure, and contributes to model rollout and monitoring pipelines. Candidates need at least six years of experience with large-scale distributed systems or infrastructure, two years of ML infrastructure experience, strong Python and PyTorch or JAX skills, and deep Kubernetes and cloud infrastructure experience.

