Skip to main content

Senior Site Reliability Engineer — Token Factory (Inference Platform) at Nebius

Department: SRE

Language
Setup
Remote
Location
Amsterdam, North Holland · Berlin · +2
Level
not_specified
Posted

Description

Nebius Cloud is hiring a Reliability Engineer to own the reliability, performance, and observability of its GPU inference platform. The role involves building telemetry pipelines, tuning Kubernetes autoscalers, creating Terraform infrastructure modules, improving request routing and retry logic, and developing automation and runbooks for incident detection, isolation, remediation, and post-mortem prevention. The engineer will work with Python or Bash, Prometheus, Grafana, Kubernetes, Terraform, GPU workloads, and MLOps or model-hosting platforms to improve cost, reliability, and self-healing capabilities.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation