Skip to main content

Senior Performance Engineer - LLM Inference Frameworks at NVIDIA

Department: Architect

Language
Setup
On-site
Location
Yoqneam Illit, North District · Raanana, Center District · +2
Type
Full-time
Level
senior
Posted

Description

NVIDIA is hiring a Software Engineer to design, implement, and optimize high-performance inference pipelines for large language models on GPUs. The role focuses on profiling and tuning model execution, memory management, speculative decoding, context caching, quantization, and benchmarking systems. Candidates need a relevant bachelor's, master's, or higher degree, at least five years of software development experience, strong Python and software engineering skills, and experience with PyTorch and HuggingFace. The position is hybrid.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation