Skip to main content

Senior Site Reliability Engineer at CloudFactory

Department: Engineering

Language
Setup
Hybrid
Location
Berlin
Type
Full-time
Level
senior
Posted

Description

CloudFactory is hiring a Site Reliability Engineer to own the reliability, observability, developer tooling, automation, and secure-by-default infrastructure of its AI platform and core services. The role covers model serving and inference infrastructure, GPU-backed endpoints, autoscaling, SLOs, incident response, ML and LLM observability, GitOps, Terraform, and reusable platform components. Candidates need at least five years of infrastructure, DevOps, or SRE experience with Kubernetes, production operational experience, cloud and scripting skills, and strong cross-functional collaboration; AI platform experience is strongly preferred.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation