Description
Veeda AI is hiring a Member of Technical Staff focused on machine learning performance for multimodal video world models. The role owns distributed training throughput, inference latency and GPU utilization, numerical stability, CUDA and Triton kernel development, compiler and NCCL optimization, and fault diagnostics for large-scale multi-node training and inference. Candidates need a bachelor's degree or equivalent hands-on experience, deep PyTorch and large-scale parallelism experience, Python and C++/CUDA fluency, and profiling expertise; experience with GPU clusters, kernel libraries, long-sequence parallelism, ROCm or TPU workloads, and fault-tolerant training is preferred.
