Skip to main content

Infrastructure Engineer at Causal Labs

Department: Infrastructure

Setup
On-site
Location
San Francisco, California
Level
not_specified
Posted

Description

The Infrastructure Engineer will design, deploy, and operate large distributed GPU clusters for training, evaluation, and serving AI workloads. The role extends Kubernetes and Slurm scheduling and orchestration, builds self-serve cluster-management software, manages storage and artifact paths, improves reliability and observability, and partners with researchers to optimize large-scale runs. Candidates should have experience with GPU clusters and container orchestration, strong systems knowledge, cloud platform familiarity, and knowledge of CUDA, NCCL, and distributed-workload profiling.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation