Description
Crusoe is hiring an LLM Inference Engineer to own the end-to-end inference stack, bringing modern optimization techniques into production and improving latency, throughput, cost, and reliability for large language models. The role involves designing serving architectures, profiling and tuning frameworks such as vLLM and SGLang, working down to CUDA kernels, adapting optimizations across models, and partnering with customer engineering teams to move workloads from proof of concept through monitored production services. The position requires a degree in a relevant field, production coding experience, familiarity with LLM inference optimization and GPU behavior, and strong communication skills; compensation is up to $250,000-$300,000 plus bonus and equity.

