Description
Radical Numerics is hiring a Member of Technical Staff, Distributed Systems to design, build, and operate large-scale GPU clusters and cloud environments for training and inference. The role owns provisioning, capacity planning, scheduling, storage, observability, reliability, and performance optimization across distributed systems, while building a unified compute interface and collaborating with researchers. Candidates should have distributed systems or large GPU-cluster experience, backend software skills in Python or Rust, systems knowledge, and familiarity with modern deep-learning frameworks.
