Description
GitLab is hiring an Engineering Manager, Production Engineering - Observability to lead a globally distributed team that builds and operates metrics, logging, alerting, and capacity planning platforms for GitLab.com and GitLab Dedicated. The role owns reliability, scalability, cost, SLO-driven alerting, telemetry gaps, distributed tracing, incident response, and sustainable on-call operations, while guiding technical decisions and using AI tools to support engineering workflows. Candidates should have experience leading observability, platform, or site reliability engineering teams at scale, with knowledge of Prometheus, logging platforms, SLOs, error budgets, capacity forecasting, and production incident coordination.
