Description
Annapurna Labs is seeking a senior engineer to build and maintain infrastructure that monitors and reports on the functionality and performance of large-scale EC2 AI/ML testing workloads. The role involves automating software delivery with internal CI/CD tools, writing Python to spin up large clusters and run ML and HPC benchmarks, using AWS Managed Grafana and Athena to analyze performance data, and creating alerts for functional and performance regressions. The engineer will manage infrastructure across many instance types, software stacks, and Linux operating systems, with strong Linux, networking, performant coding, and HPC/RDMA experience valued.
