Description
The GPU Infrastructure Engineer will design, build, operate, and scale GPU computing clusters for AI/ML workloads across bare metal and public cloud environments. Responsibilities include selecting and validating GPU architectures, configuring operating systems and software stacks, provisioning nodes, designing high-performance networks and storage, monitoring and diagnosing GPU resources, and automating operations. The role requires at least three years of systems engineering experience with large-scale Linux infrastructure, hands-on work in OS, kernel, or drivers, and experience with HPC or AI clusters using GPU servers.
