Skip to main content

Senior Backend Engineer, Inference Platform at Together AI

Department: EngineeringEducation: education_optional

Compensation

$160,000 – $250,000/yr

Location
San Francisco, California
Type
Full-time
Level
senior
Posted

Description

Together AI is hiring a Distributed Inference Engineer to build and optimize large-scale, fault-tolerant systems that route, load balance, auto-scale, and serve frontier generative AI models across global data centers and model engine pods. The role focuses on low-latency request handling, multi-tenant resource allocation, prefix caching, system profiling, and collaboration with researchers and open-source communities, using technologies such as Rust, Go, Python, Kubernetes, CUDA, Triton, and GPU software stacks.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation