Description
Amazon is hiring a senior software engineer to architect and implement distributed inference support for PyTorch in the AWS Neuron SDK, optimizing large language model families on Trainium and Inferentia hardware. The role covers model enablement, performance profiling, kernel and system-level optimization, testing, production deployment, and customer collaboration, with a focus on latency and throughput across the software, hardware, and machine learning stack.
