Description
Together AI is hiring an engineer to build and operate GPU fleets for frontier model training and inference. The role focuses on automating GPU cluster provisioning, deployment, upgrades, repairs, and retirement; creating AI infrastructure agents for failure detection and remediation; developing fleet intelligence and validation systems; and improving GPU availability, utilization, performance, and reliability. Candidates need at least three years of experience with distributed systems, infrastructure platforms, or large-scale backend software, along with strong software engineering skills in Python, Go, or Rust and experience with Linux, Kubernetes, Terraform, Ansible, or similar technologies.
