Skip to main content

LLM Inference Engineer at Crusoe

Department: Cloud Engineering

Compensation

$250,000 – $300,000/yr

Setup
On-site
Location
San Francisco, California
Type
Full-time
Level
senior
Posted

Description

Crusoe is hiring an LLM Inference Engineer to own the end-to-end inference stack in production, improving large language model performance, latency, throughput, and cost. The role involves optimizing serving architectures, profiling and tuning frameworks such as vLLM and SGLang, working down to CUDA kernels, adapting optimization methods across models, and partnering with customer engineering teams to move workloads from proof of concept to monitored production services. The position requires a degree in a relevant field, production coding experience, familiarity with LLM inference optimization and GPU behavior, and strong communication skills; Python is preferred. Compensation is up to $250,000-$300,000 plus bonus, with equity and comprehensive health, dental, and vision benefits.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation