Skip to main content

Site Reliability Engineer, Post Training at ThinkingMachines

Department: Core Engineering

Compensation

$300,000 – $350,000/yr

Setup
On-site
Location
San Francisco, California
Type
Full-time
Level
senior
Posted

Description

Thinking Machines is hiring a Site Reliability Engineer to own the reliability, performance, and uptime of large-scale post-training and reinforcement learning training jobs. The role partners with research teams during active model runs, debugs failures across accelerators, networking, storage, schedulers, and training frameworks, builds monitoring and automated recovery, improves checkpointing and fault tolerance, develops internal tools, and participates in an on-call rotation. The position is based in San Francisco, California, with an expected annual salary of $300,000–$350,000 USD and visa sponsorship.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation