Description
Inferact is seeking highly autonomous generalist engineers to contribute to the development and optimization of vLLM, the world's AI inference engine. This globally remote role involves working across various aspects of the vLLM stack, including low-level GPU kernels, distributed systems, and cloud orchestration. Engineers will operate independently, focusing on identifying and solving high-leverage problems to enhance AI inference efficiency.
