Skip to main content

Senior Machine Learning Engineer, LLM Inference Optimization at Nebius

Department: ML

Compensation

$195,200 – $262,200/yr

Location
Palo Alto, California
Type
Full-time
Level
senior
Posted

Description

Nebius is hiring a Senior Machine Learning Engineer on its Applied AI team to own model and endpoint optimization from artifacts through production deployment. The role focuses on improving latency, throughput, memory efficiency, GPU utilization, and cost per token while maintaining model quality and reliability. Responsibilities include optimizing LLM and VLM endpoints, deploying and extending inference engines, building compression and acceleration workflows, creating reproducible benchmarks, diagnosing bottlenecks, and collaborating with kernel and platform engineers. The position requires strong Python and PyTorch skills, hands-on experience with LLM or VLM inference systems, and knowledge of modern inference stacks.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation