Description
Baseten is hiring a Site Reliability Engineer to define and automate day-2 operations for its multi-cloud Kubernetes ML infrastructure platform. The role owns reliability, incident response, observability, runbooks, automated mitigations, SLO/SLI instrumentation, and runtime diagnostics across customer workloads and internal services. Candidates need extensive Kubernetes experience, scalable infrastructure experience, observability and infrastructure-as-code expertise, and familiarity with GitOps and incident management. No prior machine learning experience is required, but knowledge of ML model deployment is beneficial.
