Skip to main content

Senior LLM Inference Optimization Engineer at Intel

Compensation

$195,200 – $361,200/yr

Setup
Hybrid
Location
California · Phoenix, Arizona · +2
Type
Full-time
Level
senior

Description

Intel is seeking an experienced software engineer specializing in optimizing large language model inference engines for local and edge hardware. The role focuses on profiling and improving llama.cpp and vLLM performance across GPU/iGPU and CPU environments, tuning KV cache, batching, scheduling, quantization, startup, model loading, and lifecycle management, while contributing patches to open-source engines. The position requires a BS/MS in a related STEM field, 8+ years of software development, strong C++ or Python, systems-level Linux development, low-level debugging, and demonstrated CPU/GPU performance optimization experience. It is hybrid and based in Santa Clara, California, with additional locations in Phoenix, Arizona; Folsom, California; and Hillsboro, Oregon.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation