Description
Mozn is hiring a Senior AI Site Reliability Engineer to combine hands-on production SRE responsibilities with the design and deployment of LLM-based agents that automate reliability operations. The role includes on-call incident response, root cause analysis, application debugging and fixes, Kubernetes and cloud operations, observability integrations, agent tool and guardrail design, evaluation against historical incidents, and reporting reliability and toil-reduction impact. It requires at least three years of production LLM software experience, strong Python or similar programming skills, substantial SRE experience, and hands-on Kubernetes, cloud, and observability expertise. The position is associated with Saudi data residency and regulatory requirements.
