Description
OnBoard is hiring a Site Reliability Engineer / Cloud Operations Engineer III for a fully remote United States role. The senior cloud operations engineer will own Datadog observability across the company’s multi-product SaaS platform, including instrumentation, dashboards, SLOs, monitors, alert routing, and migration from a legacy monitoring stack. The role focuses on replacing manual operational work with automation, improving reliability and incident response, and providing technical leadership and mentorship. It requires 5–7 years of cloud operations, site reliability, platform, or DevOps experience, along with expertise in Datadog, PowerShell, Python, Bash, Kubernetes, Azure, infrastructure-as-code, CI/CD, and production incident response.
