About this opportunity
Opportunity Overview
Tencent's Data and AI Platform Team is hiring an intern to work on the deployment needs of their overseas gaming business in large language model and reinforcement learning scenarios. The role involves developing, optimizing, and implementing high-quality AI computing infrastructure. The team is responsible for ensuring the stability and performance of ultra-large-scale training jobs. This internship is with a team focused on the overseas gaming business. The internship will involve working on complex systems engineering challenges, including the development and optimization of AI job scheduling logic and the addressing of compute bottlenecks in gaming scenarios. The team uses a range of technologies, including PyTorch, DeepSpeed, and Megatron-LM. The ideal candidate will have a strong technical background and experience working with distributed systems and AI computing infrastructure. They will also have excellent collaboration and communication skills, with the ability to work effectively with cross-functional teams.
Responsibilities
- Participate in the implementation of large-scale distributed training solutions
- Own the engineering delivery of data parallelism, model parallelism, and ZeRO techniques
- Continuously tune GPU utilization and ensure the stability of ultra-large-scale training jobs
- Develop and optimize AI job scheduling logic
- Address compute bottlenecks in complex gaming scenarios
- Own the full engineering pipeline from model training to inference serving
- Participate in operator profiling, model quantization, and the construction of high-performance inference pipelines
- Actively embrace AI Coding tools to boost development efficiency
- Drive Harness Engineering practices to ensure extreme reliability of the underlying infrastructure
Benefits
- Opportunity to work with a leading technology company
- Gain experience in AI computing infrastructure development and optimization
- Collaborate with cross-functional teams on complex systems engineering challenges
Requirements & Qualifications
- Bachelor's degree or above in Computer Science, Computer Architecture, High-Performance Computing, or related fields
- Proficient in at least one of Python, C++, or Go
- Deep understanding of the PyTorch framework
- Hands-on experience engineering distributed training with DeepSpeed, Megatron-LM, or equivalent frameworks
- Solid understanding of distributed systems principles
- Familiarity with NCCL, RDMA networking, or high-performance storage
- Working knowledge of containerized infrastructure (Docker / Kubernetes)
- Demonstrable experience with AI Coding tools
- Prior work in Harness Engineering
- Exceptional learning agility, clear logical thinking, and the ability to collaborate effectively with cross-functional teams
- Fluent proficiency in English
How to Apply
- Review the job description and requirements to ensure you are a good fit.
- Prepare your resume and any additional required documents.
- Click the Apply button below to submit your application.
Ready to apply?