Description
Sterling Computers is hiring a Backend ML Engineer in North Sioux City, South Dakota, to take AI/ML systems from prototype to production. The role focuses on building inference APIs, retrieval and orchestration pipelines, integrating large language models, and operating ML infrastructure at scale. Responsibilities include designing scalable RESTful and streaming APIs, integrating and tuning LLMs and embedding models, building document ingestion and chunking pipelines, evaluating retrieval and generation quality, containerizing and deploying workloads with Docker and Kubernetes, optimizing databases and caching, implementing CI/CD and monitoring, and collaborating with frontend, research, and product teams. The position requires Python and FastAPI expertise, production ML experience, familiarity with RAG and vector databases, cloud deployment experience, and willingness to travel 25% to 50%.
