Description
Inferact is hiring a Performance Engineer to write CUDA kernels and low-level optimizations that maximize the performance of vLLM across hundreds of accelerator types, including NVIDIA GPUs and emerging silicon. The role involves GPU architecture optimization, profiling, benchmarking, and close collaboration with hardware vendors. It is based in Singapore and offers an annual salary of S$200,000 to S$400,000 plus equity, along with medical, dental, and vision coverage and case-by-case visa sponsorship.
