Description
Telnyx is hiring a senior or staff Inference Infrastructure Architect for its founding China team to operate and expand a bare-metal B300 GPU fleet and build serverless and dedicated LLM inference platforms. The role covers vLLM and SGLang serving, Kubernetes GPU scheduling, model routing, KV caching, weight distribution, autoscaling, observability, and cost optimization, with upstream contributions to open-source infrastructure. Candidates should have production LLM-serving experience, Kubernetes and GPU-fleet operations expertise, and strong performance-engineering skills; the position is remote from mainland China with visa sponsorship available for relocation.
