Skip to main content

Principal Architect, AI Inference Networking - NIXL and Dynamo at NVIDIA

Department: Architect

Language
Setup
On-site
Location
Huangpu, Shanghai · Shenzhen, Guangdong Province
Type
Full-time
Level
senior
Posted

Description

NVIDIA is hiring a Senior Inference Architect to drive NIXL adoption in China, profile customer inference stacks, fix throughput and time-to-first-token issues, write NIXL backends and plugins, extend Dynamo integrations, define dynamic data-path APIs, and contribute optimizations upstream. The role requires an M.Sc. or Ph.D. in a relevant field, 12+ years of experience building or optimizing large-scale distributed systems, strong C++ and Python systems programming, RDMA and GPU memory expertise, modern LLM serving knowledge, and fluent Mandarin and English. The architect will work with core NIXL and Dynamo engineers in Israel and the United States.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation