Description
AWS is hiring an Inference Engineer to architect and implement distributed inference support for PyTorch in the Neuron SDK, optimize large language model families on Trainium and Inferentia hardware, and tune performance for latency and throughput. The role involves kernel and system-level optimization, profiling, testing, customer enablement, and collaboration across compiler, runtime, framework, and hardware teams. It requires a bachelor's degree in computer science or equivalent, at least three years of professional software development and system design experience, and expertise in C++, Python, machine learning, and parallel computing.
