Description
ElastixAI is seeking a Systems-Minded AI Software Engineer to join their core inference platform team. The successful candidate will design and extend the low-level serving stack, leveraging open-source frameworks like vLLM, SGLang, and TensorRT-LLM, while developing new model sharding and scheduling logic. They will integrate deeply with proprietary AI accelerators, optimizing for throughput, latency, and scalability. This role involves collaborating with ML and hardware engineers, building APIs for flexible model deployment, and profiling/debugging across various layers of the stack.
