Description
Amazon is hiring an Inference Engineer to architect and implement distributed inference support for PyTorch in the AWS Neuron SDK, optimize large language model families on Trainium and Inferentia hardware, and tune performance for latency and throughput. The role involves system-level programming, kernel development, profiling, testing, customer enablement, and cross-functional collaboration, with requirements including at least three years of professional software development and systems design experience, a bachelor's degree in computer science or equivalent, and experience with C++, Python, machine learning, and parallel computing.
