Description
Inferact is hiring a TPU performance engineer to build and optimize vLLM's TPU backends, compiler integrations, runtime paths, and benchmarking infrastructure using JAX, XLA, Pallas, and related tooling. The role focuses on improving correctness, latency, and throughput for production model serving on Google TPUs, with work spanning inference systems, kernels, compilers, and hardware architecture. The position is based in San Francisco, California, with remote work considered for exceptional candidates, and offers a $200,000–$400,000 USD annual salary plus equity, health, dental, vision, and 401(k) benefits.
