Description
AWS is hiring a Distributed Inference Engineer to architect and implement distributed inference support for PyTorch in the Neuron SDK, optimize LLM and other large-model families on Trainium and Inferentia hardware, and tune performance for latency and throughput. The role involves distributed computing architecture, high-performance kernels, system-level profiling, testing, customer model enablement, and cross-functional collaboration. It requires a bachelor's degree, at least three years of professional software development and system design experience, Python and C++ experience, and strong knowledge of machine learning, parallel computing, memory management, and computer architecture. The position is based in Cupertino, California, with a base salary of $165,200 to $223,600 annually.
