SALMONN Framework: Expanding General Auditory Capabilities of Large Language Models


Google launched Gemini Omni1.1 Flash, an AI video model with up to 4K output, upgrading scene expansion, camera control, and efficiency for developers and creators. It supports direct scene extension, pushing multimodal video generation toward higher resolution and longer duration.....
Tencent Cloud supported live streaming of the US-Canada-Mexico World Cup across 17 countries, covering most authorized platforms in Asia-Pacific and China. AI was adopted at scale for the first time in live production, automating image enhancement, intelligent directing, auto-editing, and horizontal-to-vertical conversion, marking a milestone for AIGC mass application.....
After a six-month ban, ChatGPT returns to WhatsApp in the European Economic Area on July 13, 2026, following EU Commission intervention. Users can access multimodal conversations without registration by calling 1-800-CHATGPT.....
On June 23, 2026, Volcano Engine launched its Seedance2.5 video generation model at the Summer FORCE Conference, set to release in July. It offers three key breakthroughs: direct generation of 30-second native videos, joint generation from up to 50 multimodal assets, and local editing with visual consistency. President Tan Dai noted that video generation is crucial to world models.....
At the Build 2026 conference, Microsoft released its first advanced reasoning model, MAI-Thinking-1, with 35 billion parameters, achieving top performance in software engineering benchmark tests. The model was trained from scratch using clean data and did not use external sources, marking a significant step forward for Microsoft in self-developed AI and building a comprehensive matrix of scenarios.