Description
The role is responsible for architecting, building, and operating end-to-end machine learning pipelines across Google Cloud and AWS, including model training, validation, deployment, logging, monitoring, alerting, and CI/CD. The engineer will collaborate with frontend, backend, research, and infrastructure teams and maintain high availability and performance. Required qualifications include a computer science or related bachelor's degree, at least five years of experience in AI/ML operations, DevOps, or infrastructure engineering, and expertise in Python, TypeScript, Docker, Kubernetes, Terraform, Google Cloud, AWS, machine learning models, and CI/CD pipelines. Computer graphics or physics-based simulation experience, Prometheus/Grafana or ELK monitoring, Vertex AI, and custom domain-specific languages are listed as bonus qualifications.
