Jobs
Tencent
Senior Researcher, Speech Synthesis and Multimodal LLM
Tencent is hiring a Senior Researcher to work on speech synthesis and multimodal LLMs. The role involves developing advanced speech synthesis and generation algorithms, including text-to-speech, voice conversion, and sound/music generation. The Senior Researcher will collaborate with research and engineering teams to optimize speech and audio synthesis systems for online applications. The ideal candidate will have a strong foundation in speech/audio processing and modern generative models. The Senior Researcher will explore and advance full-duplex/streaming multimodal LLM capabilities in speech understanding, generation, and real-time spoken interaction. The role requires a Ph.D. in Computer Science, Electrical Engineering, Signal Processing, or a closely related field. Tencent values diversity and believes that diverse voices fuel innovation, allowing the company to better serve its users and the community. The Senior Researcher will work in a dynamic environment, contributing to the development of cutting-edge speech synthesis and multimodal LLM technologies. The role offers the opportunity to collaborate with cross-functional teams, from prototyping to production, and to publish research at top-tier venues. Responsibilities Develop and optimize speech and audio synthesis systems for online applications Explore and advance full-duplex/streaming multimodal LLM capabilities in speech understanding, generation, and real-time spoken interaction Collaborate cross-functionally with research and engineering teams from prototyping to production Develop advanced speech synthesis and generation algorithms based on LLMs, multimodal/omnimodal LLMs Requirements & Qualifications Ph.D. in Computer Science, Electrical Engineering, Signal Processing, or a closely related field Strong foundation in speech/audio processing and modern generative models (e.g., diffusion, flow matching, autoregressive, codec-based approaches) Hands-on experience extending LLMs to speech/audio modalities (e.g., speech tokenizers, multimodal adapters, speech-text joint training) Experience with full-duplex or streaming spoken dialogue systems and real-time interaction modeling is a plus Proficient in Python and deep learning frameworks (e.g., PyTorch) Experience with distributed training is a plus Track record of publications at top-tier venues (e.g., ICML, NeurIPS, ICLR, ACL, ICASSP, Interspeech) 2. Prepare your application materials, including your resume and a cover letter. 3. Click the Apply button below to submit your application.