Description
Hyperbolic Labs is hiring an Inference Engineer to build and operate model inference capabilities on Forge and Kubernetes across distributed, heterogeneous clusters. The role covers model deployment, serving-engine evaluation, production monitoring, gateways, endpoints, optimization, autoscaling, KV-cache orchestration, and customer inference debugging. Candidates should have broad inference knowledge, deep Kubernetes experience, familiarity with distributed inference concepts and NVIDIA Dynamo, and a record of building products that serve real traffic.
