Description
Deepgram is hiring a Senior Software Engineer - Model Evaluation & AI Systems to build and operate automated evaluation pipelines, harnesses, canaries, and monitoring systems for speech-to-text, text-to-speech, LLM, RAG, agent, and multimodal models. The role defines evaluation methodologies, translates research benchmarks into pass/fail gates, integrates quality checks into CI/CD, and partners with Research, DevOps, and product teams to protect model quality and customer experience. Candidates should have a relevant degree or equivalent experience, at least five years of software or QA engineering experience, backend or scripting experience in Python, Rust, Go, or a similar language, and experience with automated test pipelines or evaluation frameworks.
