Description
AWS is hiring an AI/ML Engineer to build efficient, stable foundation models for long-horizon agentic workloads. The role combines model development with systems and hardware optimization, including training, fine-tuning, evaluation, latency and throughput improvements, distributed-training profiling, and optimization using CUDA, Triton, custom GPU kernels, and specialized accelerators such as AWS Trainium. Candidates should have strong machine-learning and deep-learning experience, familiarity with PyTorch, JAX, or similar frameworks, and experience with large-scale model training, serving, and hardware-aware optimization.
