Description
Modal is an AI infrastructure company that provides instant GPU access, sub-second container startups, and native storage for AI teams. They are seeking a Staff Reliability Engineer to dramatically improve reliability while scaling their platform and customer base. The role involves identifying architectural changes, fostering a culture of reliability, designing operational processes, participating in on-call rotations, building monitoring systems, and debugging production issues. The ideal candidate will be a deep systems thinker with experience in production-grade code, on-call support for critical services, cloud skills (AWS preferred), and familiarity with large-scale fleet management and capacity planning.
