Description
AWS is hiring a Systems Development Engineer to build automation, diagnostic tooling, and predictive health infrastructure for large-scale AI/ML accelerator and GPU server fleets. The role covers fleet health monitoring, predictive failure detection, test automation, Linux and GPU subsystem debugging, telemetry pipelines, CI/CD, and cross-functional collaboration with hardware, firmware, software, operations, and ODM teams. It requires regular on-site presence at a manufacturing partner site and lists extensive programming, systems design, Linux, and professional software development experience, along with a bachelor's degree or equivalent work experience.
