Description
FAR.AI is hiring a Software Engineer, GPU Cluster Infrastructure to lead the technical direction and roadmap of a shared GPU cluster fleet used for large-scale AI research. The role combines hands-on systems engineering with managing a small team of senior engineers, covering architecture, scheduling, storage, networking, security, observability, incident response, and platform reliability. Candidates need at least five years of production systems or infrastructure engineering experience, including GPU, HPC, or large-scale batch platforms, production Kubernetes, infrastructure as code, and strong programming skills. The position is full-time and can be based remotely or in Berkeley, California, or Singapore, with visa sponsorship available for in-person employees.
