Description
Nebius is hiring a GPU Inference Software Engineer to develop and optimize low-level kernels and runtime components for AI inference, improve GPU platform performance, profile and debug system- and hardware-level issues, and integrate support for new GPU architectures. The role requires strong C++ or GPU programming expertise, systems-level software experience, profiling and debugging skills, and knowledge of CPU/GPU architecture and memory hierarchy. Preferred qualifications include CUDA, ROCm, CUTLASS, Triton, TensorRT, and other inference-engine technologies.
