Welcome to the "AI Daily" section! This is your guide to exploring the world of artificial intelligence every day. Every day, we present you with the latest content in the AI field, focusing on developers to help you understand technical trends and innovative AI product applications.

Fresh AI products click to learn more:https://app.aibase.com/zh

1. Black Forest Lab releases Flux3: The first native audio-generating multimodal foundation model, capable of producing synchronized audio-video in 20 seconds

The Flux3 multimodal foundation model released by Black Forest Lab is the first to achieve native audio generation and shows excellent performance in audio-video synchronization, image generation, and action control, demonstrating its leadership in the field of artificial intelligence.

image.png

AiBase Summary:

🧠 Flux3 is the first multimodal foundation model that supports native audio generation, capable of generating a 20-second synchronized audio-video segment at once.

📊 In early testing, Flux3 outperformed other top models such as Luma Ray3.2 and Runway Gen-4.5 on multiple benchmarks.

🤖 Flux3 also collaborated with Mimic Robotics to develop the Flux-mimic action model for the robotics industry, which has been tested in production tasks at Audi factories.

2. Claude Opus now supports voice mode: Upgraded from casual conversation to real-time assistant capable of using tools and switching languages

The upgrade of Claude's voice mode significantly enhances its ability to handle complex questions, while adding support for multiple models and tools, allowing users to complete tasks more efficiently through voice interaction.

image.png

AiBase Summary:

🧠 Claude voice mode upgrade supports Opus, Sonnet, and Haiku models, enhancing the ability to handle complex questions.

🛠️ Supports connecting tools like Gmail and Slack, enabling voice commands to perform tasks such as modifying schedules and generating documents.

🌐 Expands multilingual support, allowing users to switch languages during conversations, but requires manual settings.

3. Kuaishou enters the AI interactive content market, opens first batch of creator recruitment

Kuaishou officially launched the 'AI Interactive Content' creation feature, marking the competition in short video platforms extending from traditional content length to interactive experiences. This feature allows users to actively participate in content direction through text input and option selection, breaking the traditional one-way transmission model of short videos. Industry analysts believe this is a key move for Kuaishou in the context of saturated short video competition, aiming to improve user retention and enhance recommendation algorithms.

AiBase Summary:

🤖 Kuaishou launches AI interactive content creation feature, opening up the first batch of partner creator recruitment.

🕹️ AI interactive content is built on large models, allowing users to actively influence content direction, breaking the traditional one-way transmission model of short videos.

📈 Industry analysts believe this move is a key step for Kuaishou in the saturated short video market, aiming to enhance user retention and improve recommendation algorithms.

4. Xiaopeng Humanoid Robot Guangzhou Factory Begins Small-Batch Trial Production, Expected to Achieve Mass Production in 2026

Xiaopeng Humanoid Robot Guangzhou Factory officially begins small-batch trial production, marking the final countdown to mass production. The company plans to achieve mass production in 2026 and gradually apply it to global stores and commercial scenarios, accelerating commercialization through the integration of core group capabilities.

AiBase Summary:

🚀 Xiaopeng Humanoid Robot begins small-batch trial production at the Guangzhou factory, with the mass production line entering the final joint debugging stage.

💼 Chairman He Xiaopeng of the Xiaopeng Group will personally serve as CEO of the robot business, driving the commercialization process.

🌐 Xiaopeng plans to achieve humanoid robot mass production in 2026 and gradually enter global stores and commercial scenarios.

5. Tencent Merges Multimodal and Large Language Model Departments: Yao Shunyu Takes Charge to Explore Full-Modal Capabilities

Tencent announced the merger of the multimodal model department and the large language model department, establishing a basic model department managed by Yao Shunyu. This move aims to improve model development and collaboration efficiency, exploring the intelligent limits of full-modal models.

AiBase Summary:

🧠 Tencent merged the multimodal model department with the large language model department, establishing a basic model department under Yao Shunyu's management.

🚀 The merger aims to improve model development and collaboration efficiency, exploring the intelligent limits of full-modal models.

📊 Yao Shunyu previously served as the head of the large language model department. This merger marks further integration of Tencent's Metaverse technology roadmap.

6. Microsoft Releases Self-Developed Dual Models MAI-Image-2.5-Pro and MAI-Voice-2-Flash: No Third-Party Distillation, Already Applied in Bing and PowerPoint

Microsoft released self-developed AI models MAI-Image-2.5-Pro and MAI-Voice-2-Flash, targeting high-quality image generation and high-concurrency voice interaction scenarios. They emphasize traceable training data and no reliance on third-party model distillation. Both models have been applied in multiple products and significantly improved efficiency and cost-effectiveness.

image.png

AiBase Summary:

🧠 MAI-Image-2.5-Pro is a high-precision image generation model developed by Microsoft, supporting natural language editing instructions and already applied in Bing Image Creator and PowerPoint.

🔊 MAI-Voice-2-Flash is an optimized voice model, with twice the speed and 32% lower cost, already applied to T-Mobile and other customers.

📊 MAI-Transcribe-1.5 has been used in medical solutions, processing 28 million patient consultation records, supporting 58 languages, and reducing error rates by 50%.

7. Tencent Launches WorkBuddy Bench: A Comprehensive Benchmark Suite That Integrates Coding, Web, Office, and Security Into One Intelligent Agent Testing Environment

Tencent launched the WorkBuddy Bench multi-domain evaluation suite, covering four areas: coding, web, office, and security. All tasks are reverse-engineered from real-world scenarios to prevent memorizing answers. The evaluation methods are diverse, emphasizing pollution resistance, and public datasets are provided to enhance transparency.

AiBase Summary:

📌 WorkBuddy Bench covers four fields: coding, web, office, and security, with 260 tasks to prevent memorizing answers.

💡 Each task is generated by reverse engineering from real-world scenarios to ensure that the task prompts cannot be found through web searches.

🔒 Evaluation uses a multi-dimensional scoring mechanism, including rule checks, LLM/VLM evaluation, and agent evaluation, ensuring fairness.

8. Alibaba Open Sources 0.8B Document Parsing Model OvisOCR2, Leading in OmniDocBench

Alibaba's OvisOCR2 model achieved a major breakthrough in document parsing, achieving end-to-end parsing with a 0.8B parameter scale, surpassing traditional pipeline methods, providing efficient support for RAG retrieval, intelligent Q&A, and enterprise knowledge bases.

image.png

AiBase Summary:

✅ OvisOCR2 is the first end-to-end document parsing model to fully surpass traditional pipeline methods.

💡 The model combines real and synthetic data, using various techniques to enhance performance.

🚀 The model is open-sourced and compatible with mainstream inference frameworks, lowering the barrier for high-precision document parsing.