Fall 2026

Video Generation Content Understanding and Feedback Research Intern

Posted on 6/22/2026

Tencent

Tencent

10,001+ employees

Global tech conglomerate: social, gaming, cloud

No salary listed

London, UK

In Person

On-site in London, United Kingdom.

Category
AI & Machine Learning (2)
,
Required Skills
Python
PyTorch
Reinforcement Learning

Get referred to Tencent

See people who can refer or advise you

Requirements
  • Currently pursuing a Ph.D. or Master's degree in AI-related fields (video understanding, video prediction, reinforcement learning, or multimodal generation).
  • Good understanding of the internal mechanisms of video generation/VLM models and diffusion model principles.
  • Academic or project experience in interactive video generation, controllable generation, or multimodal understanding.
  • Solid coding and engineering skills; able to assist in building and debugging model training pipelines.
  • Proficient in Python/PyTorch
Responsibilities
  • Assist in researching the model's ability to understand generated content, including parsing semantics, objects, relationships, and spatial structures.
  • Help implement state tracking and evaluate consistency modeling for generated videos.
  • Participate in exploring "unified generation-understanding" model architectures.
  • Assist in evaluating causal consistency control and physical constraints for "action input → video output" pipelines.
  • Contribute to researching hybrid architectures aimed at achieving low-latency feedback for real-time interactive generation.
  • Work with the team to test and refine the end-to-end closed loop of "generation → understanding → control → feedback."
  • Assist in validating interactive capabilities in simulated/game scenarios and help build evaluation metrics.
  • Track the latest industry academic papers and open-source projects related to video understanding and generation.
Desired Qualifications
  • Publications in related fields are a strong plus.
  • Familiarity with Game AI, simulation environments, or reinforcement learning frameworks is preferred.

Tencent is a Chinese technology conglomerate that operates a wide range of consumer platforms and enterprise services. It connects over a billion users through WeChat and QQ, combining messaging, social features, and mobile payments, while Tencent Cloud offers AI, big data, and cloud infrastructure for businesses. It stands out by blending a huge user base with major investments in gaming studios and an integrated ecosystem that spans media, fintech, cloud, and enterprise tools. Its goal is to create a large, connected digital ecosystem for people and businesses in China and worldwide, using AI-powered products and services.

Company Size

10,001+

Company Stage

IPO

Headquarters

Shenzhen, China

Founded

1998

Get referred to Tencent

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Weixin mini-programs deepen merchant monetization without leaving Tencent's ecosystem.
  • Tencent Cloud can benefit from enterprise digital transformation across multiple industries.
  • Hunyuan AI integration expands advertising, search, content, and enterprise tooling.

What critics are saying

  • ByteDance keeps siphoning attention and ad spend from Weixin and Tencent.
  • Alibaba Alipay and Tencent Cloud rivals compress payments and enterprise margins.
  • China regulation can hit gaming, payments, content, and cloud simultaneously.

What makes Tencent unique

  • Weixin and QQ anchor Tencent's one-billion-user consumer ecosystem.
  • Tencent combines gaming, payments, cloud, and AI inside one platform stack.
  • Digital Jingdezhen shows Tencent uses AI and games for cultural preservation.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Professional Development Budget

Growth & Insights and Company News

Headcount

6 month growth

0%

1 year growth

0%

2 year growth

0%
36Kr
Aug 3rd, 2026
China's AI Applications Achieve Full - fledged Product Matrix for the First Time Behind $300M Financing and $2B Valuation

The era of standalone products has ended, and AI applications have stepped into the enterprise - centric era.

Cryptocurrencey.com
Jul 30th, 2026
Tencent open-sources AngelSpec: A unified training framework for MTP and block-parallel speculative decoding on Hy3 models.

Tencent open-sources AngelSpec: A unified training framework for MTP and block-parallel speculative decoding on Hy3 models. July 30, 2026 7 Mins Read Tencent has released AngelSpec, an open-source, torch-native training framework for speculative-decoding draft models. The release covers both autoregressive multi-token prediction (MTP) and the block-parallel DFlash family. Most speculative-decoding work searches for one drafter that scores well on an averaged benchmark mixture. Real serving traffic does not look like that mixture. AngelSpec treats workload heterogeneity as a first-class design constraint, and specializes structure, training data, and verification depth around it. Why one universal drafter underperforms. Speculative decoding is lossless. A lightweight drafter proposes several future tokens, and the target model verifies them together in one forward pass using rejection sampling. Acceleration then depends on two things: how many draft tokens get accepted, and how long the complete draft-verify round takes. Those two quantities move in opposite directions across domains. In high-entropy open-ended conversation, many continuations are semantically valid. The target may pick any one of them, so acceptance decays quickly with proposal depth. Generating and verifying a long block wastes compute. Autoregressive MTP drafting fits this regime because it proposes a shorter candidate sequence. Code and mathematical reasoning behave differently. Programming syntax, repeated identifiers, formal expressions, and step-by-step derivations constrain future tokens more strongly. These workloads create longer predictable spans, which is exactly what block-parallel drafting amortizes well. AngelSpec therefore ships two complementary drafters, not one compromise. The MTP model is trained on rich, diverse conversation-oriented data. The block-diffusion model is strengthened with code- and mathematics-focused samples. The MTP path: Training-Time Test and target-model rollout. The original Hy3 model is trained with a single MTP layer and no recurrent self-conditioned unrolling. At inference the block can be reused recurrently, but the training objective never prepared it for a long self-generated chain. Errors accumulate with depth, so the second and third draft positions accept at substantially lower rates than the first. AngelSpec addresses this train-inference mismatch with a shared-parameter, multi-depth scheme. It retains D logical prediction depths but reuses one physical MTP block. During training that block is autoregressively unrolled for D steps, with each prediction fed into the next invocation. Following the Training-Time Test principle of EAGLE-3, depth k+1 receives the arg max prediction from depth k instead of the ground-truth token. Parameters are shared, but supervision stays depth-specific: the teacher target advances one future position at every depth. Two further choices carry most of the gain. First, the target backbone and target language-model head are frozen, and MTP inputs from the backbone are detached. The drafter improves its proposal distribution without touching the distribution that verification must preserve. Second, training uses target-model rollout: responses are generated by the frozen target rather than taken from the original reference corpus. That produces the exact token choices, hidden-state trajectories, and local uncertainty patterns MTP has to approximate at serving time. The measured effect is concentrated where it should be. At T = 0, mean acceptance moves from 52.8% to 66.4%, and mean accepted length from 2.58 to 2.99. The first-position rate is nearly preserved at 0.799 | 0.814. The deeper positions carry the delta: p3 climbs from 0.290 to 0.706 on GSM8K, and from 0.387 to 0.757 on HumanEval. DFly: hybrid target conditioning and a predecessor-conditioned AR head. DFly is the block-diffusion architecture. It builds on DFlash with two structural changes. Hybrid target-conditioning backbone: DFlash concatenates hidden states from multiple target layers and transforms them with a fully connected layer, producing one shared context feature. That single context is then supplied to every draft layer, which limits layer specialization. DFlare instead learns per-layer fusion weights, giving each draft layer its own view of the target hierarchy, but drops DFlash's learned cross-layer transformation. DFly composes them: the FC branch establishes a common semantic basis, and the per-layer fusion is applied as a residual refinement on top. The extra branch introduces only D x T scalar weights, and its softmax coefficients can be precomputed after training. Predecessor-conditioned autoregressive head: A parallel backbone predicts each block position from the accepted context only. It cannot see which continuation was actually selected at earlier draft positions, which produces suffix acceptance decay. DFly places a small sequential head after the parallel backbone, converting position-wise marginal predictions into prefix-conditioned distributions. The expensive backbone stays fully parallel; only the small head runs left to right. Early performance. On Qwen3-8B, DFly reaches 5.41 average mean accepted length, against 5.32 for DSpark, 4.57 for DFlash, and 3.24 for MTP. It takes the best result on all five math and code benchmarks. DSpark stays slightly ahead on MT-Bench at 3.77 versus 3.67, which is consistent with DFly being positioned for code and math. On Hy3-A21B, the margin is wider. DFly reaches 4.79 against 3.69 for DFlash and 3.00 for MTP - relative gains of 29.8% and 59.7%. It improves every one of the six reported benchmarks. The cumulative ablation on Hy3-A21B under greedy decoding traces where that comes from: DFlash backbone 3.77, DFly backbone 4.40, plus Markov head 4.56, and hidden correction 4.60, and code/math data 4.75. The data expansion adds 700K prompts - 500K code from OpenCodeInstruct and OpenCodeReasoning, 200K math from Big-Math. Any prompt sharing a contiguous 16-token span with an evaluation example is removed before response generation. Inside the framework. AngelSpec is built on TorchSpec and extends it in several places. The foundation is disaggregated: inference engines run the frozen target model and stream hidden states through a Mooncake-backed RDMA store directly to distributed training workers. Hidden states are captured inside vLLM worker processes through public vLLM APIs - a speculative hidden-state extraction hook and a custom KV connector - without forking the engine. The extensions that matter for this work: * TTT rollout unrolled in parallel over the whole sequence, with the causal-prefix-plus-diagonal attention structure reproduced implicitly via compiled FlexAttention and logsumexp merging. Memory stays close to a single causal pass. * Long-context training with Ulysses sequence parallelism, validated at context lengths up to 128k tokens. Each local shard carries a D-token halo so depth-shifted supervision stays rank-local. * Document-aware sequence packing with three isolation mechanisms - an attention document gate, a depth-shift document gate specific to the MTP path, and document-local position encoding. Cross-document isolation is covered by unit tests verifying zero attention leakage. * Evaluation server that periodically runs genuine speculative decoding against the latest checkpoint on dedicated GPUs, reporting mean accepted length and per-position acceptance as measured by the serving engine itself. * Pluggable interfaces at three levels: targets (runtime vLLM plugin entry point, no source patch), objectives (composed over a shared base, selected by config), and optimizers (Muon as a drop-in alternative to AdamW). Backends are tiered: vLLM is first-class, with SGLang and HuggingFace Transformers supported at community tier. Key takeaways. * Seven checkpoints ship on Hugging Face and ModelScope, including no-think and high-think DFly variants. * AngelSpec trains six draft architectures - DFly, DFlash, DFlare, Eagle3, DSpark, MTP - behind one config-driven pipeline. * DFly lifts mean accepted length on Hy3-A21B to 4.79, versus 3.69 for DFlash and 3.00 for MTP. * On HY3-295B-A21B with TP=8, DFly-8 delivers a 1.98-2.40x speedup over autoregressive decoding across concurrency 4 to 64. * D-cut pushes live-traffic throughput to 981 tok/s at c64, +15.7% over DFly, while giving up 2.8% acceptance. Check out the Paper, GitHub Repo, Documentation, Hugging Face Collection and ModelScope Collection. All credit for this research goes to the researchers of this project. Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.

Great Handshake
Jul 23rd, 2026
Tencent: beyond WeChat - The investments and influence shaping global tech.

Tencent: beyond WeChat - The investments and influence shaping global tech. Most Western executives still think of Tencent as "the WeChat company." That framing undersells the reality by an order of magnitude. With a market capitalization that has oscillated between $300 billion and $550 billion over the past three years, Tencent Holdings Limited is one of the largest companies in the world by any measure - and its influence extends far beyond Chinese borders into gaming, cloud infrastructure, fintech, healthcare, and venture capital on six continents. Understanding Tencent is not optional for anyone doing business in or with China. Whether you are a foreign brand seeking Chinese distribution, a Western game developer eyeing global publishers, or a fintech startup wondering who the strategic buyers in your market might be, Tencent is almost certainly a relevant variable. From instant messenger to super-platform. Tencent was founded in Shenzhen in 1998 by Ma Huateng (Pony Ma) and four co-founders, initially as a provider of OICQ, a messaging service. The company listed on the Hong Kong Stock Exchange in 2004 at HK$3.70 per share - a price that appreciated more than 600 times over the following two decades. The pivot that defined Tencent's modern identity came in 2011 with the launch of WeChat (Weixin in China). What began as a mobile messaging app evolved into China's dominant super-app: social networking, mobile payments, e-commerce, government services, healthcare booking, and hundreds of thousands of Mini Programs on a single platform. As of 2025, WeChat reported over 1.3 billion monthly active users. WeChat Pay - alongside rival Alipay - effectively replaced credit cards in China, with the two platforms processing an estimated 90 percent of mobile payment volume. The gaming empire. Tencent's games division generated approximately 170 billion yuan (roughly $23 billion USD) in fiscal year 2024, making it the largest game publisher on the planet by revenue. Domestically, Honor of Kings (王者荣耀) boasts over 100 million daily active users in China alone. Internationally, Tencent's footprint is built through equity stakes rather than organic publishing: * Riot Games (100% owned since 2015) - developer of League of Legends and Valorant * Supercell (84% stake, acquired 2016 for $8.6 billion) - developer of Clash of Clans * Epic Games (approximately 40% stake) - developer of Fortnite and the Unreal Engine * Ubisoft (approximately 9.9% stake as of 2022) These holdings gave Tencent revenue exposure into virtually every major gaming franchise without the regulatory friction of full acquisitions - a capital deployment strategy built for the current geopolitical era. The global investment portfolio. Beyond gaming, Tencent operates one of the world's most active corporate venture arms, with over 800 investments and acquisitions globally spanning healthcare, logistics, electric vehicles, AI, and enterprise software. Notable international stakes include Snap Inc., Sea Limited (Southeast Asia's dominant e-commerce platform), MercadoLibre (Latin America's largest e-commerce player), and an early strategic position in Tesla. The strategic logic is consistent: Tencent invests in founders and platforms that control consumer attention and transaction infrastructure in their markets. Rather than operating businesses directly, Tencent provides capital, technology integration, and ecosystem access in exchange for data, distribution, and optionality. This model has drawn CFIUS review in the US and merger notification requirements in the EU. Understanding the evolving landscape for Chinese outbound investment shapes how deals get structured and where approvals are required. Cloud, AI, and enterprise. Tencent Cloud is the fourth-largest cloud provider globally and second-largest in China, with annual revenue exceeding 100 billion yuan. The platform serves clients in financial services, retail, healthcare, and media and is expanding across Southeast Asia, the Middle East, and Europe. In AI, Tencent has deployed large language models under its Hunyuan brand and integrated AI into WeChat, advertising systems, and enterprise tools. For Western companies operating in China, Tencent Cloud's data residency and cybersecurity certifications make it a relevant option for locally-compliant infrastructure. Navigating China's AI regulations and data localization requirements is a growing compliance task for any company running digital operations in-country. The regulatory reset: 2021-2023. Tencent was not immune to China's tech sector crackdown. In November 2021, regulators forced Tencent Music to relinquish exclusive licensing deals on anti-monopoly grounds. In 2023, the People's Bank of China fined the WeChat Pay subsidiary Tenpay $533 million for AML compliance gaps. The gaming division faced suspended new-title approvals and mandatory limits on minors' gaming hours. The net effect was a structural shift: Tencent accelerated international revenue diversification, with overseas gaming revenue growing to represent roughly 30 percent of the segment by 2024. The regulatory cycle also prompted Tencent to restructure investments, distributing stakes in Meituan and JD.com directly to shareholders. For Western partners, this episode illustrates that strategic value within Chinese tech conglomerates can be redistributed quickly when government priorities shift - a reason for clear-eyed engagement, not avoidance. WeChat as a business tool for foreign companies. For most foreign companies, the operationally relevant parts of Tencent's ecosystem are WeChat Official Accounts and Mini Programs. An Official Account functions like a combination of a website, email newsletter, and social media profile. As of 2025, over 20 million Official Accounts were registered on the platform. WeChat Mini Programs give foreign brands a transaction-capable digital storefront inside the app without requiring a separate download. For B2B operators, WeChat Work (WeCom) integrates with external WeChat contacts and is a standard CRM and communication tool for China-facing sales teams. Using WeChat effectively for B2B lead generation requires understanding both Official Account strategy and WeChat Work's customer management capabilities. Official resources and regulatory context. Foreign companies seeking to work within Tencent's ecosystem or understand the regulatory environment governing Chinese tech platforms should consult primary sources directly. Tencent publishes annual reports and partner program details through its investor relations portal at tencent.com/investors. For cybersecurity and data handling governance, the Cyberspace Administration of China publishes official guidance at cac.gov.cn. On the US side, companies considering technology partnerships with Tencent-affiliated entities should be familiar with CFIUS review processes. The US Treasury's CFIUS resources are available at treasury.gov/cfius. The bottom line. Tencent is simultaneously China's dominant social platform, the world's largest game publisher, a global venture capital force, a cloud infrastructure provider, and the transaction layer beneath hundreds of millions of daily consumer interactions. For Western professionals in cross-border trade, technology, or investment, that scale demands engagement. The companies that understand Tencent's ecosystem, investment patterns, and regulatory environment will find both partnership opportunities and competitive intelligence. Those that treat it as a Chinese domestic story will find themselves repeatedly surprised by whose name appears in their cap table, their competitor's partnership announcement, or their distribution channel across Asia. Tencent is a global company that happens to be headquartered in Shenzhen. Building a complete picture of the US-China business landscape requires treating it as such.

Tencent
Jul 20th, 2026
Tencent unveils full-stack Embodied Intelligence solution; ADP 4.0 Launches globally at WAIC 2026.

Tencent unveils full-stack Embodied Intelligence solution; ADP 4.0 Launches globally at WAIC 2026. July 20, 2026 Tencent announced a series of product and technology upgrades across embodied intelligence and AI agents at the 2026 World Artificial Intelligence Conference (WAIC 2026). The upgraded full-stack embodied intelligence solution spans cloud infrastructure, models, platforms and applications, and is designed to help robots and system developers improve efficiency and performance. For AI agents, Tencent introduced a portfolio of differentiated solutions to enhance productivity for both enterprises and individuals. Tencent Cloud's enterprise-grade Agent Development Platform 4.0, or ADP 4.0, is now available globally, alongside the launch of an ecosystem initiative covering major industries and application scenarios. The platform has already been deployed across more than 30 industries, supporting use cases ranging from smart customer service and knowledge management to media production and more. Tencent WorkBuddy is also now available on iOS, Android and HarmonyOS for individual users. Tencent has continued to accelerate its full-stack AI development in 2026. Total API calls to its Hy3 model increased more than 68-fold compared with the previous-generation Hy2 within a single week of launch, claiming the top position on OpenRouter's global LLM usage leaderboard. Breakthroughs in Tencent's foundation models have accelerated the evolution of applications such as embodied intelligence and WorkBuddy. "AI is evolving from chatbot interactions to task-based collaboration, expanding from personal use to workplaces and enterprises, and from a single assistant to teams and entire organizations," said Songtao Lin, Vice President of Tencent. "We will continue strengthening our models and platform capabilities to make AI more practical, collaborative and accessible to everyone." Tencent Debuts Full-Stack Approach to Embodied Intelligence Tencent presented its full-stack embodied intelligence strategy for the first time at WAIC, demonstrating a comprehensive approach across foundation models, native agents, and development. Together, these capabilities are designed to help robots better understand their surroundings, make decisions and carry out tasks in the physical world. At the model layer, Tencent introduced a portfolio comprising Hy-Embodied-VLA-0.5, Hy-Embodied-VLM-1.0 and Hy-Embodied-RxBrain-1.0. Each model supports a different part of a robot's intelligence. RxBrain-1.0 serves as the reasoning engine, combining language-based reasoning with visual understanding to support decision-making. VLM-1.0 focuses on perception, delivering performance comparable with Tencent's previous flagship model while using only one-tenth of the computing resources. Trained on more than ten thousand hours of high-quality data, VLA-0.5 brings vision, language and physical action together in a single model and can be deployed across different types of robots. The model portfolio has already been tested in real-world scenarios, including retail guidance, visitor assistance and eldercare services. Tencent also introduced the TairosAgent embodied AI agent framework and the Apexio agent, which work together to help robots continuously perceive their environment, make decisions and take action. Tairos, Tencent's embodied intelligence platform launched in 2025, has also been upgraded. The platform will provide a range of open-source capabilities, agents, development tools and services, helping lower development barriers across the full value chain from models and robotic systems to real-world applications and accelerate the adoption of embodied intelligence. At the infrastructure layer, Tencent Cloud has developed a three-tier architecture for embodied intelligence spanning computing infrastructure, model services, perception and interaction. The cloud and AI foundation supports the full development cycle, from research and model training to deployment. Tencent Cloud is also working with industry partners and has launched the cloud-based Embodied-AI-as-a-Service (EaaS) offering, helping bring embodied intelligence from technical development to broader real-world adoption. "True intelligence emerges when language, vision, spatial understanding, physical control and environmental feedback work together," said Zhengyou Zhang, Chief Scientist of Tencent and Director of Tencent Robotics X Lab, and Director of Futian Lab. Zhang added, "These capabilities must be tested in a continuous loop of perception, physical interaction and action. The latest developments reflect Tencent's approach to embodied intelligence, moving beyond isolated digital intelligence towards systems that can understand, interact with and respond to the physical world as an integrated whole." Tencent Cloud Launches Global ADP 4.0; WorkBuddy App Debuts Across Major Mobile Stores As AI enters its next phase, agents are emerging as an important driver of productivity. Simon Wu, Vice President of Tencent Cloud, announced the global launch of Tencent Cloud's enterprise-grade Agent Development Platform, ADP 4.0. It now connects with widely used communication platforms such as LINE and Telegram, supports custom time-zone scheduling and automatic language adaptation, integrates a range of leading foundation models with localized support, and connects with mainstream SaaS platforms including Google Workspace. "The rapid adoption of AI agents is now a clear trend, but identifying the right use cases will be critical," said Wu. "By combining full-stack capabilities with open APIs, Tencent Cloud ADP will work with partners to develop practical, industry-specific solutions and accelerate the adoption of AI agents across a wider range of industries and application scenarios." For AI agents to become part of day-to-day business operations, they must first be able to access and understand an organization's own knowledge. Tencent LearnShare also announced an upgraded enterprise knowledge integration solution, enabling AI agents to more comprehensively retrieve, interpret and apply internal knowledge. For individuals and teams, Tencent officially launched the standalone WorkBuddy app. Since its launch in March, WorkBuddy has evolved through frequent product upgrades and has become a widely used AI productivity tool in China. From executing tasks efficiently in the digital world to acting autonomously in the real world, Tencent is drawing on its full-stack capabilities across models, tools and platforms to equip AI with both the intelligence to understand and the ability to act. Looking ahead, Tencent will continue to invest in foundational technology, strengthen its platform capabilities and accelerate the practical adoption of AI across industries.

Lapaas
Jul 7th, 2026
Tencent raises $1.5B from Kuaishou share sale, stake drops to 9.37%

Tencent Mobility, a wholly owned subsidiary of Tencent Holdings, has raised approximately $1.5 billion through an off-market block trade of shares in Chinese short-video platform Kuaishou Technology. The transaction involved 272.9 million Kuaishou Class B shares at HK$43.25 per share, representing a 6% discount to Monday's closing price. The sale reduced Tencent's stake from 15.68% to 9.37%, meaning it no longer qualifies as a substantial shareholder. The move signals Tencent's strategic shift towards artificial intelligence investments. Days before the sale, Tencent invested $200 million in Kling AI, Kuaishou's text-to-video generation tool. The company has also committed ¥10 billion to Chinese AI firm DeepSeek. Kuaishou cushioned market impact by repurchasing 174.84 million of its own shares worth approximately $1.06 billion. Both companies stated their strategic partnership remains intact despite the ownership change.

INACTIVE