Bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or a similar field
Required Previous Experiences
Hands-on experience building or optimizing TPU workloads using JAX, XLA, Pallas, or related compiler and runtime tooling
Experience optimizing ML kernels or inference paths, including attention, GEMM, sampling, KV cache, fused kernels, or backend runtime paths
Strong performance profiling and benchmarking skills using measurements, compiler artifacts, correctness tests, and reproducible benchmarks
Preferred Qualifications
Experience with vLLM, SGLang, TensorRT-LLM, XLA-based serving, or other LLM inference systems
Familiarity with batching, KV cache, decoding, serving tradeoffs, and backend performance constraints in production inference systems
Experience with compiler technologies such as XLA, MLIR, LLVM, Pallas, or other kernel DSLs, including lowering, fusion, and backend code generation
Knowledge of quantization methods such as INT8, FP8, mixed precision, or TPU-specific numeric formats, including accuracy and performance tradeoffs
Contribution to vLLM, JAX/XLA, Pallas, PyTorch/XLA, compiler projects, or other open-source ML infrastructure
Building TPU benchmarking infrastructure or automated performance regression detection for accelerator workloads
Work with Google TPU ecosystem stakeholders, accelerator platform teams, or early-access programs to ship backend, compiler, or inference performance improvements
Inferact is a startup founded by creators and core maintainers of vLLM, an open-source LLM inference engine. Its mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster.