About this opportunity
Opportunity Overview
Tencent is hiring a Senior Researcher to work on speech synthesis and multimodal LLMs. The role involves developing advanced speech synthesis and generation algorithms, including text-to-speech, voice conversion, and sound/music generation. The Senior Researcher will collaborate with research and engineering teams to optimize speech and audio synthesis systems for online applications. The ideal candidate will have a strong foundation in speech/audio processing and modern generative models. The Senior Researcher will explore and advance full-duplex/streaming multimodal LLM capabilities in speech understanding, generation, and real-time spoken interaction. The role requires a Ph.D. in Computer Science, Electrical Engineering, Signal Processing, or a closely related field. Tencent values diversity and believes that diverse voices fuel innovation, allowing the company to better serve its users and the community. The Senior Researcher will work in a dynamic environment, contributing to the development of cutting-edge speech synthesis and multimodal LLM technologies. The role offers the opportunity to collaborate with cross-functional teams, from prototyping to production, and to publish research at top-tier venues.
Responsibilities
- Develop and optimize speech and audio synthesis systems for online applications
- Explore and advance full-duplex/streaming multimodal LLM capabilities in speech understanding, generation, and real-time spoken interaction
- Collaborate cross-functionally with research and engineering teams from prototyping to production
- Develop advanced speech synthesis and generation algorithms based on LLMs, multimodal/omnimodal LLMs
Requirements & Qualifications
- Ph.D. in Computer Science, Electrical Engineering, Signal Processing, or a closely related field
- Strong foundation in speech/audio processing and modern generative models (e.g., diffusion, flow matching, autoregressive, codec-based approaches)
- Hands-on experience extending LLMs to speech/audio modalities (e.g., speech tokenizers, multimodal adapters, speech-text joint training)
- Experience with full-duplex or streaming spoken dialogue systems and real-time interaction modeling is a plus
- Proficient in Python and deep learning frameworks (e.g., PyTorch)
- Experience with distributed training is a plus
- Track record of publications at top-tier venues (e.g., ICML, NeurIPS, ICLR, ACL, ICASSP, Interspeech)
How to Apply
- Review the job description and requirements to ensure you are a strong fit for the role.
- Prepare your application materials, including your resume and a cover letter.
- Click the Apply button below to submit your application.
Ready to apply?