Description
Sciforium is hiring a Distributed Training and Inference Engineer to build, optimize, and maintain the software stack powering large-scale AI training and serving workloads. The role spans CUDA/ROCm runtimes, JAX and PyTorch frameworks, distributed system configuration, multi-node GPU cluster integration, profiling, and debugging of hardware–software interactions. Candidates should have at least five years of experience in ML systems, distributed training, or related fields, along with strong Python and C++ skills and familiarity with modern distributed training technologies.
