Description
Pythian is hiring a Site Reliability Engineer to design, deploy, and operate large-scale distributed systems across compute, storage, networking, and AI/ML environments. The role focuses on Kubernetes, Istio, Linux, automation, monitoring, observability, troubleshooting, infrastructure readiness for AI/ML workloads, and on-call participation. Candidates need Google Cloud, Terraform, microservices, containers, networking, PKI, service mesh, and Linux experience; Golang experience is an asset. The position offers a competitive total rewards package, work-from-home equipment, training, paid vacation and sick days, and a volunteer day off.
