Description
Nscale is hiring a Principal/Staff Observability Platform Engineer to own the technical strategy and architecture of its observability platform for GPU clusters, AI workloads, and infrastructure. The role focuses on metrics, logs, traces, alerting, data models, ingestion, retention, cardinality, tooling, standards, incident postmortems, and platform scalability, while partnering with SRE, infrastructure, and AI/ML teams and mentoring the observability team. Candidates need at least 8 years of relevant experience, hands-on knowledge of major observability technologies, strong engineering fundamentals, proficiency in Python or Go, and experience with Kubernetes and infrastructure-as-code. The role offers a base salary of $190,000–$300,000 USD plus potential bonus, equity, and/or commission.
