Description
Deepgram is hiring an Applied ML Engineer on its Partner Platform Engineering team to port speech models to non-NVIDIA and edge platforms. The role adapts model structure, operators, quantization, precision, and architecture to fit target hardware, validates accuracy and latency, builds repeatable deployment pipelines, and collaborates with embedded AI engineers and silicon vendors. It requires production experience deploying ML models to edge or non-NVIDIA hardware, knowledge of quantization and inference runtimes, Python and PyTorch, and the ability to automate model conversion and deployment.
