Skip to main content

Principal Observability Platform Engineer at Nscale

Department: AI Infra Operations

Compensation

$190,000 – $300,000/yr

Type
Full-time
Level
staff+
Posted

Description

Nscale is hiring a Principal/Staff Observability Platform Engineer to own the technical strategy and architecture of its observability platform for GPU clusters, AI workloads, and infrastructure. The role focuses on metrics, logs, traces, alerting, data models, ingestion, retention, cardinality, tooling, standards, incident postmortems, and platform scalability, while partnering with SRE, infrastructure, and AI/ML teams and mentoring the observability team. Candidates need at least 8 years of relevant experience, hands-on knowledge of major observability technologies, strong engineering fundamentals, proficiency in Python or Go, and experience with Kubernetes and infrastructure-as-code. The role offers a base salary of $190,000–$300,000 USD plus potential bonus, equity, and/or commission.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation