Description
Pika is hiring Senior/Staff Inference Engineers to accelerate and optimize inference for its video and language models. The role focuses on inference acceleration, quantization, attention optimization, GPU parallelism, CUDA and NCCL kernel development, distributed workloads, and production deployment of AI models. Candidates need at least five years of engineering experience, strong inference and GPU expertise, and familiarity with video generation and large language models. The position is based in Palo Alto, California, with a three-to-five-day-per-week onsite arrangement.
