Description
Tether is hiring an AI Model Serving Engineer to design, deploy, and optimize model-serving architectures and inference frameworks for advanced AI systems. The role focuses on high-throughput, low-latency, low-memory, and scalable inference across resource-constrained devices, edge platforms, and GPU clusters, including controlled inference testing, performance benchmarking, bottleneck diagnosis, and integration of optimized serving pipelines into production systems. Candidates need a computer science or related degree, strong AI research and low-level kernel optimization experience, Metal Shading Language expertise, and knowledge of modern model architectures and inference techniques.

