Description
Nebius is hiring a Senior Site Reliability Engineer (SRE) for its Compute Node team to build and operate the cluster scheduler and node-level services that manage virtual machines across cloud regions. The role focuses on Linux systems engineering, virtualization, containerization, production troubleshooting, observability, incident response, root-cause analysis, and reliability improvements. It requires strong Linux, QEMU/KVM, containerization, debugging, and SRE experience, with optional expertise in Kubernetes, low-level Linux tools, large-scale compute platforms, open-source infrastructure, and hardware or GPU debugging.
