Skip to main content

Inference Runtime Engineer at Inferact

Department: Research & Engineering

Compensation

$200,000 – $400,000/yr

Setup
On-site
Location
San Francisco, California
Type
Full-time
Level
not_specified
Posted

Description

Inferact is hiring an Inference Runtime Engineer to develop and optimize the vLLM inference engine for large language and diffusion models. The role focuses on model execution across diverse hardware and architectures, including transformer variants, KV-cache and prefix-caching systems, hybrid serving, and multimodal inference. Candidates need a bachelor's degree or equivalent experience, strong Python and PyTorch skills, and experience with LLM inference systems. The position is based in San Francisco, California, with remote work considered for exceptional United States candidates, and offers $200,000–$400,000 USD plus equity, visa sponsorship on a case-by-case basis, and health, dental, vision, and 401(k) benefits.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation