Description
NVIDIA is seeking a Deep Learning Research Engineer to develop and improve LLM inference algorithms, benchmarks, profiling workflows, and experimental frameworks. The role involves prototyping low-latency and high-throughput inference algorithms, profiling performance on NVIDIA hardware, identifying optimization opportunities, and translating research into production software. Candidates need an MSc or equivalent industrial research experience, at least five years of applied research, research engineering, or algorithm engineering experience, strong Python and PyTorch skills, and experience with large-scale GPU clusters.
