Description
Nebius is seeking a GPU Cluster Architect to design and optimize large-scale AI infrastructure across compute, networking, storage, and control planes. The role involves architecting GPU cluster topologies, modeling AI/ML workload performance, validating low-latency interconnects, integrating storage, analyzing monitoring signals, and collaborating with reliability, networking, storage, and data center engineering teams. Candidates need at least five years of cluster-design experience, knowledge of modern GPU architectures, experience with InfiniBand and RoCE, systems architecture and hardware reliability expertise, and scripting experience with Python or Go.
