Description
Inferact is hiring a Developer Relations Engineer to educate developers about vLLM and related AI inference systems. The role involves writing technical deep dives, building demos, creating tutorials, contributing documentation and examples, hosting workshops, and explaining concepts such as KV cache, continuous batching, prefix caching, quantization, GPU serving, and latency versus throughput. Candidates should have a relevant bachelor's degree or equivalent experience, strong LLM inference and model-serving knowledge, experience with vLLM or adjacent systems, and a public portfolio of technical artifacts. The role is based in San Francisco, with remote work considered for exceptional United States candidates, and offers $200,000–$400,000 USD plus equity, visa sponsorship on a case-by-case basis, and health, dental, vision, and 401(k) benefits.
