Description
Inferact is hiring a TPU performance engineer to build and optimize vLLM's TPU backends, compiler integrations, runtime paths, and benchmarking infrastructure using JAX, XLA, Pallas, and related tooling. The role focuses on improving correctness, latency, and throughput for production model serving on Google TPUs, with responsibilities spanning kernel and inference-path optimization, performance profiling, and benchmarking. The position is based in Singapore and offers an annual salary of S$200,000 to S$400,000 plus equity, along with medical, dental, and vision coverage and case-by-case visa sponsorship.
