The recently concluded FIFA World Cup in the US, Canada, and Mexico was supported by Tencent Cloud, which powered official match broadcasts in 17 countries and regions, covering two-thirds of the official authorized platforms in Asia-Pacific and domestically. Li Yutao, Vice President of Tencent Cloud and Head of International Product Technology, revealed in a recent interview that AI was first used on a large scale in the production of live broadcasts. Most processes, including image enhancement, intelligent directing, automatic editing, smart horizontal-to-vertical conversion, quality inspection, and traceability, were fully automated by AI. The support for a single World Cup event is just one aspect of AIGC entering large-scale production applications.
Li Yutao refers to what Tencent Cloud is doing as "Harness" in the multimodal field — focusing not on how models generate content, but on the entire supply chain from generation to stable delivery to users, including preprocessing, quality inspection, stitching, encoding, distribution, and format adaptation.
He observed that under the wave of generative AI, multimedia applications are facing multiple challenges, especially in scenarios such as live streaming, on-demand viewing, media creation, and multimedia interaction. These are constrained by network transmission, video quality, user experience, and the叠加 of various features, making it difficult for a single model to solve all problems. Many manufacturers, upon seeing new multimodal large models, immediately try to connect them all — today connecting one company, tomorrow another when a new model emerges. After integration, they inevitably face various issues. If most companies encounter similar problems, a common demand can be identified from customer feedback: someone should handle the underlying tasks, allowing companies to focus more on business innovation and upper-level applications. How to identify user needs, find models that solve those needs, fix inherent model flaws, and deliver the final service to users is exactly the challenge that a video and audio PaaS cloud service provider must tackle.
As early as June 5 at the Tencent Cloud AI Industrial Application Conference, Tencent Cloud's video and audio division launched its AI brand "Tencent Cloud WAND," building six self-developed media-specific models and over 60 AI capabilities covering the entire content generation, understanding, processing, and encoding pipeline. It is specifically trained for vertical scenarios such as e-commerce, short dramas, education, and sports live streaming. Li Yutao said that the team will integrate all models together, not only large language models, but also multimodal models, even small-parameter models for image enhancement or small-parameter application models — they had already started working on the multimodal Harness last year. This coincided with the increasing number of generative video models this year, and the industry's understanding of the concept of Harness has become more aligned.
Notably, video generation is seen by the industry as the second AI transformation scenario to have a commercial closed loop, following AI coding. Since the beginning of this year, competition in the large model video scene has accelerated rapidly, and the market structure has undergone significant changes. Domestic vendors are advancing quickly in commercialization, with Seedance under ByteDance and Kuaishou's KeLing AI forming a "duopoly" in the AI video generation market. Public data shows that by March this year, KeLing AI's ARR approached $500 million, growing four times in a year; other reports indicated that by the first half of the year, Seedance 2.0's ARR reached $2 billion. Although ByteDance denied that the revenue data was "too high and inconsistent with reality," it did confirm that Seedance 2.0 has crossed the "productivity turning point." Meanwhile, in the overseas market, OpenAI announced in March this year that it would stop the independent app, API interface, and ChatGPT built-in video functions of Sora, exiting the consumer-level AI video generation market.
When developing To B products and solutions, ecological cooperation is placed at the top of Li Yutao's strategy. He also admitted that his team has been discussing a question: as the link between generation and interaction continues to shorten, will the position of Harness be ignored or bypassed? His judgment is that at least based on observations over the past six months, it will not be, and the importance of Harness may even increase. If large models are considered backend productivity, the improvement in productivity will significantly stimulate demand for front-end applications, and the breadth of demand growth will always exceed the improvement in productivity — large models can never fully meet the ultimate needs of end users, and there will always be a "supply-demand gap," which is precisely the necessity of Harness.
In his view, Tencent Cloud's biggest opportunity in the multimodal Harness lies in the fact that many video and audio scenarios are worth rethinking. The second opportunity is that in the future, more, stronger, and more comprehensive models will emerge. How to quickly apply these models to enterprise-level production services and let end users experience the latest large model experience is exactly what the multimodal Harness aims to solve. And continuously applying the latest models to the front of customers, the entire production process and tool chain becomes particularly important.
Since last year, the entire AI computing power has begun to shift from training to inference. Many companies now spend more on inference computing than on training. This indicates that a large number of AIGC applications are entering large-scale production. In China, short dramas and animated series are typical scenarios, and they are already widely using multimodal technologies for production. In the Asia-Pacific region, video and audio solutions are the "frontline" of Tencent Cloud's international expansion, leading in revenue, market share, and industry awareness. In the past five years, Tencent Cloud has made internationalization an important strategic direction, investing heavily in business resources and infrastructure abroad. The company disclosed that its overseas business has maintained double-digit growth over the past three years.
Li Zhicheng, General Manager of Tencent Cloud Video and Audio, mentioned that whether it's short dramas, animated series, e-commerce, or tool-based scenarios, the common feature is that clients initially care about effectiveness and influence, but later focus on ROI — whether it can generate profit and create value. The most frequently asked questions now are about the use of short dramas and e-commerce tools, and in the next 1-2 years, they are expected to gradually become widespread. Many users and enterprises will start using them, just like large language models — two or three years ago, they were mainly used in vertical industries, but now they have become ubiquitous and part of daily work.
With the arrival of the Agent era, Tencent Cloud's aforementioned capabilities will be combined with video and audio intelligents, AI hardware, embodied robots, and cloud phones. Relying on global infrastructure and edge inference, they will support AI applications to quickly go global and be deployed locally. Li Yutao introduced that the team has already collaborated with several embodied intelligence companies and autonomous cleaning vehicle companies, such as Zhuji, and developed TRRO remote real-time control (remote digital human solution), to address core issues such as low-latency stable communication and audio-video data transmission during manual intervention in weak network conditions.
Looking ahead, Li Yutao summarized the competition among cloud providers in the multimodal field into three directions. First is multimodal training and computing power. Unlike language model training, a large proportion of the investment in the multimodal field still goes into understanding and generating content. In the future, there will be a lot of end-to-end voice interaction and end-to-end video generation, which require significant computing power and substantial infrastructure investment from cloud providers. Second is storage and data accumulation. Multimedia requires higher storage capacity than text, and the requirements for storage stability, data cleaning, and acquisition are different. A large amount of video data needs to be cleaned and processed before pre-training, and different modalities need to be combined, with processing before training being very different from text-based tasks. Third is the application production side. Cloud service providers offering more end-user-oriented inference services and comprehensive processing capabilities are key factors influencing customer choices — how to effectively meet customer demands generated by innovation trends in operations is equally critical for Tencent Cloud.
