Description
The DevOps/SRE Engineer will deploy, operate, and scale microservices and GPU-based ML inference services across AWS, GCP, on-premises Rancher, RunPod, Scaleway, and Nebius. The role covers Kubernetes, Docker, Helm, CI/CD, Terraform, Ansible, observability, security, compliance, and infrastructure optimization, with a focus on AI product platforms and production reliability. It is fully remote and requires at least five years of DevOps or Site Reliability Engineering experience.
