Skip to main content

ML Infrastructure Engineer at Epsilon Health

Department: Engineering

Setup
On-site
Location
San Francisco, California
Level
senior
Posted

Description

Epsilon Health is hiring an ML infrastructure engineer to design and build distributed training, reinforcement learning, inference, evaluation, and production deployment systems for large foundation models in medical imaging and diagnostics. The role partners with researchers to translate experimentation workflows into production-ready systems, owns GPU-efficient training infrastructure, and contributes to model rollout and monitoring pipelines. Candidates need at least six years of experience with large-scale distributed systems or infrastructure, two years of ML infrastructure experience, strong Python and PyTorch or JAX skills, and deep Kubernetes and cloud infrastructure experience.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation