Description
Luma AI is seeking a Reliability Engineer to join their Infrastructure Engineering team. This role involves architecting and operating large, heterogeneous GPU environments, improving utilization and performance, resolving failures, and building mechanisms to prevent instabilities. The engineer will also focus on defining infrastructure evolution for training and inference scalability, designing scheduling and resource management approaches, and working with research to build necessary systems. Additionally, the role includes hiring and developing exceptional engineers, setting technical depth standards, shaping architecture, and translating reliability constraints into platform strategy.
