Skip to main content

Senior Software Engineer, LLM Inference at NVIDIA

Department: Architect

Language
Setup
On-site
Location
Tel Aviv, Tel-Aviv District
Level
senior
Posted

Description

NVIDIA is hiring a Senior Software Engineer focused on performance optimization and generative AI. The role involves implementing and optimizing inference algorithms for LLM and omnimodal architectures, profiling inference pipelines, writing and tuning CUDA and Triton GPU kernels, solving distributed inference problems, and contributing to production-grade software in open-source inference libraries. Candidates need a degree in computer science or computer engineering, at least five years of experience in performance-critical systems, deep knowledge of deep learning architectures, and experience optimizing deep learning workloads on NVIDIA GPUs.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation