In the summer of 2026, along the Huangpu River, the World Artificial Intelligence Conference was held as scheduled. At this event, HiDream.ai unveiled a major new product on the spotlight — vivago R1, the world's first multimodal intelligent agent with unlimited content creation duration. Alongside this release, the company also launched the Belt and Road Token Export Alliance in collaboration with institutions such as CAS Cerebro and initiated the Physical Intelligence Innovation Consortium with companies like Feijieke. The company aims to push AI from being a tool to becoming a capable intelligent partner through technological innovation and ecological collaboration.
During the conference, Mei Tao, founder and CEO of HiDream.ai, delivered a keynote speech titled "Toward World Models: Native Multimodal Driving Intelligent Agent Capabilities" at a forum hosted by the China Academy of Information and Communication Technology. He systematically outlined the trend of bidirectional advancement between large models and intelligent agents toward world models. His conclusion was clear: top-tier models have entered the realm of human genius, but a significant industrial gap remains between basic large models and real-world applications — challenges including difficulty in deployment, weak adaptability, and lack of control. Intelligent agents inherently possess self-memorization, logical planning, tool scheduling, and multi-terminal collaboration capabilities, which can translate the generation and reasoning advantages of basic models into standardized, verifiable, and deliverable industry capabilities. The combination of basic models and intelligent agents is becoming the core paradigm for AI empowering various industries.

Mei Tao further pointed out a practical dilemma: when facing complex tasks, a single intelligent agent has obvious capability limits, and it is difficult to handle cross-chain, high-precision creative tasks alone. Only by enabling multiple intelligent agents to work together like a professional team can they tackle real industry challenges. To ensure that this intelligent agent cluster operates stably, efficiently, and orderly, an intelligent agent operating system called AgentOS is essential. Therefore, HiDream developed its own HD-AgentOS, which uses a three-layer coupling architecture — resource layer, system layer, and capability layer — to support the entire industrial implementation system: the resource layer provides basic resources and atomic capabilities that intelligent agents can call, forming the execution foundation; the system layer handles full-process operations, risk control, monitoring, and security governance; the capability layer packages models, tools, knowledge, and processes into domain-specific skills. With this operating system, multi-agent collaboration can break through the limitations of solo efforts, forming a professional team that can perceive, reason, execute, and evolve.
vivago R1 is the result of this core technology and is the world's first multimodal creative intelligent agent with unlimited video generation and editing capabilities. Its launch marks a paradigm shift in AI within the creative production field — evolving from merely assisting in generating individual materials to autonomously planning, scheduling, and completing long-chain creative tasks as a smart partner. HiDream summarizes its core advantages as "long, long, and stable," precisely addressing three long-standing industry issues: limited duration, disorganized logic, and unstable performance.
The first "long" refers to unlimited duration and full-scenario adaptability without restrictions. Most mainstream AI video products are locked within 15 to 30 seconds, making them unable to meet long-term commercial creation needs. vivago R1 breaks through this time barrier, supporting continuous and coherent video generation of any length, covering all types of long-form creative scenarios such as short dramas, documentaries, brand promotions, and film productions, filling the gap in long-video creation in the industry. The second "long" refers to long-task thinking, ensuring consistent and unified creative logic. Long videos are not simply a collection of short clips; they require a complete narrative logic, a unified visual style, stable character IP, and scene settings. vivago R1 possesses long-task thinking capabilities, autonomously performing detailed logical planning, shot layout, and style calibration, effectively addressing problems such as sudden scene changes, character inconsistency, narrative gaps, and style fragmentation. The third "stable" refers to high-stability generation, enabling commercial-scale delivery. Traditional large models suffer from probabilistic output deviations, resulting in low usable rates of finished content and high enterprise debugging costs. By leveraging AgentOS's full-process scheduling and governance, vivago R1 turns the entire creation process into a controllable and adjustable workflow, raising the effective usability rate of content to 85%, far exceeding the industry average. This significantly reduces the time and labor required for repeated modifications and adjustments, allowing generated content to directly enter commercial production, brand promotion, and market delivery.
This capability particularly shines in continuous, long-chain scenarios like social media creation and commercial marketing. Even a seemingly simple storyboard generation requires not only following the script to draw scenes but also understanding cinematography language, narrative rhythm, emotional expression, and even differentiating the requirements of short video vlogs, short dramas, and product material videos. This deep industry understanding cannot be achieved by simply calling a large model API. vivago R1 uses a multi-agent collaboration mechanism to atomize and professionally schedule the entire workflow — from understanding tasks, planning storylines, creating storyboard scripts, generating core materials, to producing ultra-long videos. It integrates narrative, shots, characters, style, sound, and aesthetics organically, truly achieving the transition from a single tool to a full-chain creative system.
Beyond the product, HiDream also made two major ecological moves during WAIC. It jointly launched the Physical Intelligence Innovation Consortium Initiative with innovative enterprises and top research institutions, aiming to gather forces from industry, academia, and research to build a new industrial collaborative landscape for native multimodal world models. At the same time, HiDream formed a strategic partnership with institutions such as CAS Cerebro, signing an agreement to establish the Belt and Road Token Export Alliance. Based on the core concept of "Token without borders, AI without islands," the alliance aims to promote Chinese native multimodal AI capabilities across oceans and borders, connect global intelligent infrastructure, and help and co-build the intelligent upgrade and digital ecosystem connectivity of Belt and Road countries and regions.
