About this opportunity
Opportunity Overview
Tencent is hiring a Video Generation Foundation Model Researcher to design and iterate on the next-generation video generation foundation architecture. This role involves exploring core technologies for long video generation, including long-context attention, KV Cache compression, and memory mechanisms. The researcher will lead video pre-training at the 10-100 billion token scale and define data mixture and curriculum learning strategies. The ideal candidate has a Ph.D. in AI-related fields with first-author papers at top-tier conferences and experience training video/image generation models from scratch. The researcher will be responsible for advancing unified modeling capabilities for multi-resolution and multi-frame-rate videos and researching high-compression-ratio video tokenizers. They will also track industry state-of-the-art, lead comparative experiments, drive technical roadmap decisions, and publish academic papers. The role requires a deep understanding of the design trade-offs in Video VAE/Tokenizers and experience with PyTorch and large-scale distributed training. Tencent values diversity and believes that diverse voices fuel innovation and allow the company to better serve its users and the community. The company fosters an environment where every employee feels supported and inspired to achieve individual and common goals.
Responsibilities
- Design and iterate on the next-generation video generation foundation architecture
- Explore core technologies for long video generation, including long-context attention, KV Cache compression, and memory mechanisms
- Research high-compression-ratio video tokenizers and advance unified modeling capabilities for multi-resolution and multi-frame-rate videos
- Lead video pre-training at the 10-100 billion token scale, defining data mixture and curriculum learning strategies
- Explore Scaling Laws, define scientific scale-up paths, and continuously improve key model capabilities
- Track industry state-of-the-art, lead comparative experiments, drive technical roadmap decisions, and publish academic papers
Requirements & Qualifications
- Ph.D. in AI-related fields with first-author papers at top-tier conferences
- Proficient in the principles and engineering implementation of diffusion models and autoregressive generation
- Experience training video/image generation models from scratch
- Highly proficient in PyTorch and large-scale distributed training
- Deep understanding of the design trade-offs in Video VAE/Tokenizers
- 3+ years of relevant research experience
- Publications related to video generation or diffusion models
- Experience in core industry product R&D or hands-on experience with the latest technologies is preferred
- Experience leading end-to-end video foundation model projects is preferred
How to Apply
- Review the job description and requirements to ensure you are a good fit for the role.
- Prepare your application materials, including your resume, cover letter, and any required publications or writing samples.
- Click the Apply button below to submit your application.
Ready to apply?