Description
Nebius Cloud is hiring a Reliability Engineer to own the reliability, performance, and observability of its GPU inference platform. The role involves building telemetry pipelines, tuning Kubernetes autoscalers, creating Terraform infrastructure modules, improving request routing and retry logic, and developing automation and runbooks for incident detection, isolation, remediation, and post-mortem prevention. The engineer will work with Python or Bash, Prometheus, Grafana, Kubernetes, Terraform, GPU workloads, and MLOps or model-hosting platforms to improve cost, reliability, and self-healing capabilities.
