Description
Arm is hiring a Principal Software Engineer for its AI Inference Runtime team to set technical direction and lead hands-on development of distributed inference runtime components. The role covers scheduling, batching, KV-cache management, memory allocation, distributed workload execution, kernel development and optimization, performance benchmarking, and production validation. The engineer will partner with AI infrastructure, compute, cloud, framework, compiler, hardware, and research teams, while also mentoring engineers and establishing performance-engineering practices. The position requires 8+ years of experience in relevant systems or AI inference, strong C++, Rust, Python, or comparable programming skills, and expertise in profiling, debugging, and optimizing performance across kernels, runtimes, frameworks, operating systems, and hardware. The salary range is $262,700-$355,400 per year, with hybrid working arrangements determined by team needs.
