Description
The role is a senior machine-learning engineering position focused on building and optimizing production inference systems for language, vision, and speech models. Responsibilities include partnering with research and external partners, designing high-throughput, low-latency serving systems, developing profiling tools and simulators, making technical architecture decisions, and mentoring engineers. The role requires end-to-end machine-learning project leadership, proficiency in PyTorch or JAX, inference-framework experience, and cloud deployment experience with Kubernetes and Docker.
