Description
Lightning AI is hiring an Infrastructure Operations Engineer to scale and operate its GPU infrastructure platform. The role focuses on reliability, automation, incident response, customer provisioning, observability, Linux, AWS, Kubernetes, Terraform, Ansible, storage, networking, and Python or Go automation. It may be fully remote within the United States or hybrid from New York City, San Francisco, Seattle, or London, with occasional offsites. The position requires 8+ years of Linux experience, 5+ years with AWS, and 2+ years each with Kubernetes, Terraform, and Ansible, and offers a base salary of $160,000–$200,000 USD plus equity and benefits.
