Description
Inferact is hiring an AMD GPU performance engineer to build and optimize AMD GPU backends, kernels, runtime paths, and benchmarking infrastructure for vLLM. The role focuses on improving performance-critical inference paths such as attention, GEMM, sampling, KV cache, and communication-heavy operations using ROCm, HIP, Triton, CK, AITER, and related tools. The position is based in Singapore and offers an annual salary of S$200,000 to S$400,000 plus equity, along with medical, dental, and vision coverage and case-by-case visa sponsorship.
