Description
Callosum is hiring an Inference Engine Software Engineer to extend and adapt SGLang and vLLM for heterogeneous compute. The role focuses on hardware-aware scheduling, memory management, execution pipelines, parallelism, disaggregation, and caching strategies, while contributing to upstream open-source forks and collaborating with an Accelerator Systems Software engineer. Candidates should have deep knowledge of inference-serving frameworks, high-performance Python and C++/CUDA systems, large-model serving, disaggregated architectures, and open-source development. The position is based in person at Callosum’s London office and offers competitive salary, equity, private healthcare, visa sponsorship, and relocation benefits.
