Description
OpenAI is hiring an Inference Performance Engineer to model inference performance across application, model, and fleet layers. The role builds cost-to-serve estimates from microbenchmarks, analyzes end-to-end inference workloads, develops tools to identify latency and throughput bottlenecks, and partners with engineering and research teams to translate performance insights into production improvements and future capacity projections.
