Description
Inferact is seeking a Performance Engineer to optimize vLLM, the world's AI inference engine, by developing and optimizing kernels for various accelerators like NVIDIA GPUs. The role involves working directly with hardware vendor teams to maximize performance across generations of hardware.
