Skip to main content

Inference Infrastructure Architect (Remote) at Telnyx

Department: Product

Language
Setup
Remote - China
Type
Full-time
Level
staff+
Posted

Description

Telnyx is hiring a senior or staff Inference Infrastructure Architect for its founding China team to operate and expand a bare-metal B300 GPU fleet and build serverless and dedicated LLM inference platforms. The role covers vLLM and SGLang serving, Kubernetes GPU scheduling, model routing, KV caching, weight distribution, autoscaling, observability, and cost optimization, with upstream contributions to open-source infrastructure. Candidates should have production LLM-serving experience, Kubernetes and GPU-fleet operations expertise, and strong performance-engineering skills; the position is remote from mainland China with visa sponsorship available for relocation.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation