Description
Hewlett Packard Enterprise is hiring a Senior Software Engineer, Inference to build and evolve the model runtime for HPE AI Essentials, an inference platform for enterprise large language models. The role focuses on engine integration, continuous batching, KV cache management, quantized execution, distributed execution, Kubernetes orchestration, latency and throughput optimization, and customer issue resolution. It is primarily hybrid, requiring work from an HPE office on average two days per week, with the primary work location listed but potentially any other HPE site in the United States; remote work options may be considered. The position requires at least eight years of software engineering experience, including 1–2+ years on LLM inference runtimes or production model serving, and a degree in computer science or a related field.
