German artificial intelligence startup Black Forest Labs has officially released the multimodal foundation model Flux3. Built on the Self-Flow architecture, it features dedicated image, video, audio, and motion codecs, integrating a unified understanding and generation of the physical world and digital environment.
The most striking feature is audio-video synchronization. Flux3 is the first multimodal model to support native audio generation, capable of outputting synchronized audio-video clips up to 20 seconds in length. Its capabilities cover text, image, and video-to-video conversion, transition based on keyframes, and multi-character dialogue. For a single model, combining "hearing" and "seeing" into one base means it is no longer just a specialist in images or videos.

In early testing, under settings of 720p resolution and 10-second clips, Flux3 defeated Luma Ray3.2 with a 93% win rate, and had a 77% advantage over Runway Gen-4.5; even when compared to top models like Seedance2.0 and Gemini Omni Flash, it still maintained a slight edge.
Motion has also been integrated into reality. Black Forest Labs, in collaboration with Mimic Robotics, developed the action model Flux-mimic for the robotics field, which is already running production task tests in an Audi factory—multimodal capabilities have moved from the screen to the production line. In terms of release rhythm, BFL has chosen to proceed in stages: Flux3Video has already been launched, and Flux3Image and the open-source weight version "Flux3Dev" will be released in the near future.
