Description
The OnCall Site Reliability Engineer will maintain system health and performance, ensure high availability, respond to incidents, troubleshoot issues, collaborate with development and operations teams, document procedures, and participate in 24/7 on-call rotations. The role requires DevOps or system administration experience, cloud and container knowledge, familiarity with monitoring and logging tools, strong problem-solving and communication skills, and native Ukrainian language ability. The technical stack includes AWS, Kubernetes, Terraform, ElasticSearch, Kafka, Grafana, VictoriaMetrics, Vector, Opsgenie, and PostgreSQL or Cassandra. The position offers remote, office, or hybrid work arrangements and includes healthcare coverage for employees in Ukraine.
