Description
Reflection is hiring a Compute Platform Engineer to build and maintain tools for automatic remediation, topology-aware scheduling, capacity planning, and hardware debugging across a Kubernetes-based multi-cloud platform. The role also involves designing cluster-management infrastructure for large multi-GPU fleets, implementing monitoring and observability, and preparing for next-generation GPU deployments, multi-cloud storage, petabyte-scale data replication, and GPU-to-GPU networking. Candidates should have systems-level engineering experience, strong coding ability, GPU and Kubernetes expertise, and familiarity with NCCL and high-performance cloud storage.
