Description
Arm is hiring an AI Performance Engineer to optimize production AI models running on Arm technology for customers in the San Jose and Bay Area region. The role focuses on kernel-level and system-level performance and power-efficiency improvements, reference implementations, customer-facing technical communication, and influencing Arm’s IP and software roadmaps. Candidates should have experience with DNN optimization in Triton, CUDA, or similar kernel-level environments, parallel computing, memory hierarchies, Python, C++, modern AI frameworks, and profiling tools. The position is based in San Jose with significant customer-facing work, is hybrid, and requires the right to work in the United States without sponsorship.
