Description
The role focuses on optimizing GPU performance for real-time autonomous driving workloads, including sensor processing and neural network inference. Responsibilities include developing parallel computing and GPU-accelerated software, designing onboard GPU architectures, profiling bottlenecks, and debugging embedded GPU software to improve latency, throughput, resource utilization, and runtime stability. The position requires a bachelor’s or master’s degree in a relevant field, strong GPU and parallel computing knowledge, experience with profiling and inference technologies, real-time embedded systems, and C/C++ and Python; preferred qualifications include CUDA, NVIDIA embedded platforms, model quantization, concurrent GPU workload management, and sensor data compression.
