Description
Inferact is hiring an AMD GPU performance engineer to build and optimize AMD GPU backends, kernels, runtime paths, and benchmarking infrastructure for vLLM. The role focuses on improving performance-critical inference paths such as attention, GEMM, sampling, KV cache, and communication-heavy operations using ROCm, HIP, Triton, CK, AITER, and related tools. The position is based in San Francisco, California, with remote work considered for exceptional candidates, and offers a $200,000–$400,000 USD annual salary plus equity, health, dental, vision, and 401(k) benefits.
