Description
The company is hiring multiple STAFF AI TRAINING INFRASTRUCTURE ENGINEERS to build and scale distributed training systems for large AI models across GPU clusters. The role focuses on reliability, efficiency, fault tolerance, checkpointing, recovery, production pipelines, diagnostics, and automation for AI infrastructure. Candidates should have hands-on experience with distributed training systems, large-scale machine learning infrastructure, multi-node GPU training, and complex distributed systems. The position is hybrid in the Bellevue, Washington area, with approximately three days per week in the office, and requires U.S. work authorization; visa sponsorship is not available.
