Description
Dolby is hiring a Research Intern to join its Beijing team and conduct research on deep learning algorithms for speech and audio processing. The role focuses on developing controllable multi-modal emotional speech generation systems, applying generative models such as diffusion models and flow matching, building emotion-conditioning modules, and contributing to research, experimentation, analysis, and academic or patent writing. Candidates should be pursuing or have completed a PhD in deep learning for speech and audio processing, with strong experience in expressive speech generation, generative models, deep learning frameworks, Python, and multi-modal learning.
