Summary from listing
Modular is hiring a Performance Engineer to build and operate an automated platform that optimizes LLM inference performance across GPU and ASIC architectures. The role profiles customer workloads, applies optimizations across kernels, inference engines, and distributed systems, develops reusable tooling, and partners with GTM, engineering, and product teams to deliver production performance gains. Candidates need at least five years of experience in distributed systems or performance engineering, with a track record of building durable software tools and libraries. The role is based in Los Altos, California, or can be performed remotely from home, with onboarding conducted in Los Altos.
