Skip to main content

Inference Performance Engineer at OpenAI

Department: Scaling

Compensation

$266,000 – $500,000/yr

Location
San Francisco, California
Level
senior
Posted

Description

OpenAI is hiring an Inference Performance Engineer to model inference performance across application, model, and fleet layers. The role builds cost-to-serve estimates from microbenchmarks, analyzes end-to-end inference workloads, develops tools to identify latency and throughput bottlenecks, and partners with engineering and research teams to translate performance insights into production improvements and future capacity projections.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation