Description
Databricks is hiring a Site Reliability Engineer to bridge software engineering and systems architecture while owning core infrastructure and observability platforms. The role focuses on designing and automating production-grade cloud infrastructure, improving reliability and performance, building CI/CD pipelines, enabling observability, developing automation and AI tooling, managing incidents, and collaborating with Security, Engineering, and Support. Candidates need at least five years of production-level software engineering experience, strong Python and Terraform or Pulumi expertise, cloud and container experience, observability knowledge, distributed systems familiarity, and advanced GitHub Actions knowledge.
