Description
The job involves designing, building, and maintaining distributed systems and platforms, focusing on cloud orchestration, API service development, and site reliability engineering (SRE) practices. The role requires a strong understanding of IT architecture, systems design, and integration, aiming to enhance developer efficiency and system reliability through automation and observability. Responsibilities include building a cloud orchestration portal, developing scalable API services for infrastructure capabilities, eliminating repetitive tasks, implementing monitoring for problem diagnosis, defining and tracking SRE metrics, collaborating with engineers to meet infrastructure needs, and adhering to security best practices. The ideal candidate will have a background in Computer Science or a related STEM field, software development experience, proficiency in infrastructure-as-code and scripting tools, and system operations experience. Familiarity with CI/CD processes and metrics/logging systems is a plus.
