Description
LILA is hiring a Staff/Principal DevOps Engineer focused on AI inference infrastructure to design, implement, and optimize GPU-accelerated platform systems for serving machine learning models at scale. The role centers on Kubernetes-based accelerator orchestration, model serving platforms, autoscaling, infrastructure-as-code, observability, CI/CD, and AWS cloud infrastructure, with strong emphasis on reliability, latency, throughput, and cost efficiency for production inference workloads.
