Description
Boson AI is hiring a Site Reliability Engineer to design, operate, and improve reliable infrastructure supporting large-scale AI training and inference workloads. The role owns operational workflows, monitoring, alerting, incident response, capacity planning, provisioning, configuration management, and deployment automation across networking, GPU clusters, storage, scheduling, and AI platforms. It requires at least four years of production-operations experience and deep expertise in at least one infrastructure area, with Linux administration, scripting, troubleshooting, and distributed-systems skills. The position is based in Toronto or remote.
