Description
The role is responsible for designing, building, and maintaining scalable model-hosting and inference systems for large-scale AI models, including LLMs, embedding models, STT, and TTS. It covers inference optimization, framework integration, performance monitoring, hardware collaboration, and production reliability, as well as end-to-end model fine-tuning pipelines. The position requires a relevant advanced degree, at least three years of team leadership experience, and five or more years of AI experience in inference optimization, hosting, and fine-tuning pipeline development.
