Description
The Cortex Real-Time Inference Engineering Lead will design, build, and industrialize low-latency, highly available model-serving platforms for real-time predictive workflows. The role covers online inference architecture, production deployment patterns, autoscaling, performance and latency optimization, monitoring, SLOs, CI/CD, and resilience across cloud and on-premises environments. It requires expertise in Kubernetes, model-serving frameworks, APIs, load testing, observability, and automation, with preferred qualifications including a relevant degree, 5+ years of software or DevOps experience, and programming skills in Python, Go, C++, or Java.
