Description
Nebius is hiring a Lead Software Systems Engineer - GPU Performance to build and optimize its hyperscaler platform by analyzing and improving large-scale GPU cluster performance across hardware, system software, networking, virtualization, and distributed communication layers. The role involves troubleshooting GPU cluster performance under training and inference workloads, evaluating hardware and system configurations, supporting escalations, and contributing to hardware and cluster qualification. Candidates need at least five years of system-level software development experience focused on performance optimization and low-level programming, plus three years of Linux systems experience and strong knowledge of server architecture and high-performance computing. The position offers a base compensation range of $170,000 to $300,000 USD and 100% company-paid medical, dental, and vision coverage.
