Skip to main content

Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco at Plaud

Department: Global Product R&D Center

Compensation

$195,000 – $365,000/yr

Setup
Hybrid
Location
San Francisco, California
Type
Full-time
Level
senior
Posted

Description

Plaud is hiring a SpeechLLM Research Engineer to build and optimize high-throughput, low-latency inference systems for large language models and foundational speech models. The role focuses on real-time streaming, continuous batching, KV cache management, GPU architecture, distributed multi-GPU inference, and model compression or quantization. Candidates should have hands-on experience with LLM or speech-model serving, GPU systems, and real-time conversational AI. The position offers a $195,000–$365,000 base salary plus bonus and equity, healthcare, retirement matching, paid time off, and a hybrid office arrangement requiring at least three days in the office per week.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation