Description
Cerebras Systems is hiring an AI Cluster Operations Engineer to deploy, configure, debug, monitor, and operate large-scale machine learning compute clusters using the Wafer-Scale Engine. The role focuses on container-based services, Python and Go operational platforms, distributed systems, Linux, Docker, Kubernetes, monitoring, automation, incident response, and 24/7 on-call support. Candidates need 6–8 years of relevant experience and are based in the San Francisco Bay Area, Toronto, or Bangalore.
