Skip to main content

Software Engineer, GPU Infrastructure Provisioning at Together AI

Location
London, England
Type
Full-time
Level
not_specified
Posted

Description

Together AI is hiring a Software Engineer for its Research and Inference team to build a manifest-driven infrastructure control plane that provisions, operates, scales, heals, and decommissions GPU inference and training clusters. The role involves designing lifecycle state machines, declarative self-service APIs, durable workflows, reconciliation and drift detection, event-driven systems, reliability mechanisms, and production software delivered through testing and CI/CD. Candidates need strong software engineering experience in Go, Python, Rust, or similar, plus workflow orchestration and control-plane or orchestration-system experience; bare-metal, networking, and GPU infrastructure experience are advantageous.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation