Description
This role involves scaling and optimizing training systems and core model code, owning critical infrastructure for large-scale training such as GPU/TPU compute and job orchestration. The individual will build reusable JAX training pipelines, collaborate with researchers and model engineers to turn ideas into experiments and production runs, and contribute to core training code evolution. It's a hands-on, high-leverage position intersecting machine learning, software engineering, and scalable infrastructure.
