Description
The company is hiring a Senior ML Systems Engineer to build, operate, and harden production inference systems serving large models at high throughput. The role owns throughput, latency, cost-per-token, reliability, observability, capacity planning, autoscaling, load testing, and production incident resolution, while integrating research-team optimizations such as quantization, custom kernels, scheduling improvements, and memory management. It requires at least three years of experience building production-grade, large-scale serving infrastructure, distributed systems experience, GPU-accelerated inference experience, fluent Python, and systems-level coding ability in C++, CUDA, ROCm, or Triton.
