Description
Together AI is hiring a Distributed Inference Engineer to build and optimize large-scale, fault-tolerant systems that route, load balance, auto-scale, and serve frontier generative AI models across global data centers and model engine pods. The role focuses on low-latency request handling, multi-tenant resource allocation, prefix caching, system profiling, and collaboration with researchers and open-source communities, using technologies such as Rust, Go, Python, Kubernetes, CUDA, Triton, and GPU software stacks.
