Description
Inferact is hiring an Inference Runtime Engineer to develop and optimize the vLLM inference engine for large language and diffusion models. The role focuses on model execution across diverse hardware and architectures, including transformer variants, KV-cache and prefix-caching systems, hybrid serving, and multimodal inference. Candidates need a bachelor's degree or equivalent experience, strong Python and PyTorch skills, and experience with LLM inference systems. The position is based in Singapore, offers S$200,000 to S$400,000 annually plus equity, and provides medical, dental, and vision coverage.
