Description
Cerebras Systems is hiring a Cluster Engineer to build and operate software that manages thousands of Cerebras wafers, servers, and switches as a cloud. The role covers declarative CRD-driven automation, bare-metal networking, OS, application software, cluster installation and security patching, Kubernetes operators for inference workloads, gRPC control-plane services, metrics and log pipelines, failure detection, high availability, and recovery. The position requires at least five years of production distributed systems or infrastructure software experience, strong Go and Python skills, deep Kubernetes knowledge, distributed-systems debugging ability, Prometheus and Grafana expertise, and demonstrated use of AI in engineering workflows.
