Description
Amazon is hiring a senior software engineer to architect and implement distributed inference support for PyTorch in the Neuron SDK, optimize LLM and other machine learning models for Trainium and Inferentia hardware, and lead performance tuning across the software, hardware, compiler, runtime, and framework stack. The role also involves building model onboarding infrastructure, developing high-performance kernels, profiling bottlenecks, testing, and working with customers and open-source ecosystems. It requires a bachelor's degree, at least five years of professional software development and system design experience, Python and C++ experience, and strong knowledge of machine learning, parallel computing, memory management, and system performance.
