Welcome to the "AI Daily" section! This is your guide to exploring the world of artificial intelligence every day. Every day, we present you with the latest content in the AI field, focusing on developers to help you understand technical trends and learn about innovative AI product applications.

Fresh AI products Click to learn more:https://app.aibase.com/zh

1. MiniMax releases the general multimodal model H3, with a price less than one-third of mainstream models for 2K audio-visual generation in 15 seconds

MiniMax released the general multimodal model H3, which achieves unified understanding and generation of text, images, videos, and audio, marking a key step toward multimodal generation models moving from specialized experts to general assistants. H3 demonstrates commercial-level stability in scenarios such as instruction following, brand information presentation, and V2V action transfer, and can already be applied in multiple commercial fields such as advertising, e-commerce, gaming, and UI design. At the same time, during the initial design phase, H3 considered compatibility with multiple domestic chips, and its open source will further reduce the technical barriers for domestic enterprises in multimodal AI applications.

image.png

【AiBase Summary:】

📹 H3 supports native stereo audio-visual output, and can generate up to 15-second 2K resolution content.

🧠 Introduces Contextual Omni Representation technology, making natural language a universal bridge connecting various modalities.

📦 MiniMax announced that it will open-source the H3 model weights under the premise of compliance with relevant laws and regulations, promoting the development of an open-source ecosystem.

2. ByteDance Seedance 2.5 released: 30 seconds of continuous shots, AI video starts telling stories

ByteDance released the Seedance 2.5 video generation model, extending the single video generation duration to 30 seconds, and has long narrative capability, multi-modal reference, and editing capabilities, enabling complete creation.

image.png

【AiBase Summary:】

🎥 Seedance 2.5 extends video generation duration to 30 seconds, supporting multi-shot storytelling and logically connected shot organization.

🔄 Supports multi-round extension capability, maintaining consistency of character subjects, scenes, visual styles, and sound effects.

👥 Supports inputting up to 30 images, 10 video and audio segments, accurately restoring multiple people's appearances and voices.

3. DeepSeek-V4-Flash official version launched, 13 billion active parameters drive the Agent battlefield

The official version of DeepSeek-V4-Flash API was publicly tested, marking a significant progress in the Agent track. The model performed well in multiple benchmark tests, and its active parameters are only 13 billion, offering significant cost advantages compared to the flagship model V4-Pro. At the same time, its self-developed Agent framework DeepSeek Harness made its debut and supports OpenAI's Responses API format, facilitating developers' application migration.

【AiBase Summary:】

🔥 The official version of DeepSeek-V4-Flash is launched, with performance approaching the flagship model V4-Pro.

💡 Active parameters are only 13 billion, with significant cost advantages, suitable for Agent scenarios.

🚀 Supports OpenAI Responses API format, facilitating ecosystem migration.

4. DeepMind releases Gemini Robotics ER 2, unlocking multi-robot collaboration for the first time

The article introduces the Gemini Robotics ER2 model released by DeepMind, which has made major breakthroughs in multi-robot collaboration, real-time decision-making, and task monitoring, marking a solid step forward in the commercialization of embodied intelligence.

image.png

【AiBase Summary:】

🧠 The new embodied reasoning model Gemini Robotics ER2 achieves breakthroughs in multi-robot collaboration.

🚀 Realizes real-time decision-making and continuous task monitoring, improving task execution efficiency.

🤝 Introduces multi-robot collaboration mechanisms, promoting the commercialization of embodied intelligence.

5. Google integrates Nano Banana2 into Google Earth, supporting online AI-generated geographical scene images

Google announced that the latest Nano Banana2 AI image generation model has been integrated into Google Earth, providing AI image generation features for users worldwide, enhancing the visualization capabilities of geographic space content.

image.png

【AiBase Summary:】

🌍 Google integrated the Nano Banana2 AI image generation model into Google Earth, enhancing the visualization of geographic space content.

🎨 Users can generate high-precision custom images by inputting scene descriptions, applicable to multiple scenarios such as education, professional applications, and personal creation.

🚀 Generative AI is further integrated into digital maps and spatial computing platforms, bringing richer visualization application spaces for industries such as education, urban planning, and architectural design.

6. Huawei opensources the 505B parameter openPangu-2.0-Pro model, with weights and inference code simultaneously released

Huawei officially opened the openPangu-2.0-Pro large model under the openPangu AI model brand, releasing model weights, basic inference code, and technical reports simultaneously, further advancing the Ascend AI ecosystem construction, providing best practices for developers and enterprises based on Ascend-native training and inference technologies.

【AiBase Summary:】

🧠 openPangu-2.0-Pro is a large-scale mixture-of-experts (MoE) language model trained on Ascend NPU, with a total parameter size of approximately 505B (505 billion).

🚀 The model supports a context length of 512K, with a training data scale of about 34T Tokens, and improves overall performance through the post-training phase.

🌐 Huawei opensources the openPangu-2.0-Pro model, further promoting the construction of the Ascend AI ecosystem, providing best practices for developers and enterprises.

Details link: https://www.huaweicloud.com/product/modelarts/studio

7. Alibaba Qwen releases Qwen-Audio-3.0-ASR-Flash, overcoming the last mile of professional scenarios in speech recognition

Alibaba Qwen released Qwen-Audio-3.0-ASR-Flash, overcoming the last mile of professional scenarios in speech recognition. The new model achieved upgrades in five dimensions, including long audio context memory, built-in industry-specific dictionaries, hierarchical hot word customization, integrated speech polishing, and multilingual support. The model has been widely applied in multiple fields, such as meeting minutes, live subtitles, educational recordings, and smart customer service.

【AiBase Summary:】

🧠 Long audio context memory, allowing reference to previous text during transcription, ensuring consistency of names and technical concepts throughout several hours of meetings.

💼 Built-in industry-specific dictionaries, achieving a recall rate of 95.36% in medical scenarios and 91.87% in IT programming.

🌐 A single model covers over 30 languages, with an average semantic error rate of 17.09% across seven languages, surpassing models like Azure and Gemini.

8. Tesla China car systems officially integrate the Doubao large model

Tesla China officially released the 2026.14.13 version of the car system software update, with the most notable feature being the introduction of the Doubao large model, which enhances the intelligent interaction experience of the car system. However, this feature requires users to activate the advanced in-car entertainment service to use.

【AiBase Summary:】

✅ Tesla China began rolling out the 2026.14.13 version of the car system software update.

🤖 This update introduced the Doubao large model, enhancing the intelligent interaction experience.

🔑 The AI large model function requires activation of the advanced in-car entertainment service.