Description
This role involves operating, scaling, and optimizing multi-petabyte storage systems tailored for large AI training and inference workloads. Responsibilities include managing high-performance parallel filesystems and object stores, evaluating new technologies like Vast, Weka, Ceph, and Lustre, and solving complex engineering challenges. The individual will also develop Kubernetes-native storage operators and self-service platforms, focusing on automated provisioning, multi-tenancy, performance isolation, and quota enforcement.
