Description
Inferact is hiring an AMD GPU performance engineer to build and optimize AMD GPU backends, kernels, runtime paths, and benchmarking infrastructure for vLLM. The role focuses on improving performance-critical inference paths such as attention, GEMM, sampling, KV cache, and communication-heavy operations using ROCm, HIP, Triton, CK, AITER, and related tools. The position is based in Singapore, offers an annual salary of S$200,000 to S$400,000 plus equity, provides medical, dental, and vision coverage, and offers case-by-case visa sponsorship.
