Description
Oracle is hiring a Principal Systems Engineer for its AI Infra Operations team to lead GPU infrastructure operations in Oracle Cloud Infrastructure. The role focuses on building and maintaining automation and operational tooling across multiple regions, improving observability and reliability, supporting GPU fleet health and incident response, documenting runbooks, and mentoring junior engineers. Candidates need at least eight years of software operations or infrastructure automation experience, strong Python and Bash skills, expert Linux administration, distributed systems knowledge, and experience with GPU, cloud, observability, AI agents, and on-call operations. The US hiring range is $79,900 to $187,000 per annum, with possible bonus and equity.
