Bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or a similar field
Required Previous Experiences
Hands-on experience optimizing AMD GPU workloads using ROCm, HIP, Triton, CK, AITER, or similar AMD ecosystem tools
Experience optimizing ML kernels or inference paths such as attention, GEMM, sampling, KV cache, fused kernels, or communication-heavy runtime paths
Strong performance profiling and benchmarking skills using measurements, hardware counters, correctness tests, and reproducible benchmarks
Preferred Qualifications
Experience with vLLM, SGLang, TensorRT-LLM, ROCm-based serving, or other LLM inference systems
Familiarity with batching, KV cache, decoding, serving tradeoffs, and backend performance constraints in production inference systems
Experience with compiler and kernel technologies such as Triton, MLIR, LLVM, CK, AITER, HIP, or other kernel DSLs and backend libraries
Knowledge of quantization methods such as INT8, FP8, mixed precision, or AMD hardware-specific numeric formats, including accuracy and performance tradeoffs
Contribution to vLLM, ROCm, HIP, Triton, CK, AITER, PyTorch, compiler projects, or other open-source ML infrastructure
Building AMD GPU benchmarking infrastructure or automated performance regression detection for accelerator workloads
Work with AMD, accelerator platform teams, or early-access programs to ship backend, compiler, or inference performance improvements
Inferact is a startup founded by creators and core maintainers of vLLM, an open-source LLM inference engine. Its mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster.