Alibaba officially released the new generation of image generation foundation model Qwen-Image-3.0. As the third-generation foundation model in the Qwen Image Generation and Editing series, the new model supports ultra-long text input of up to 4.5K tokens, enabling the generation of complex content such as formulas, geometric graphics, logical derivations, and multi-layered UI interfaces in one go. It also natively supports 12 languages and more than 20 fonts, further enhancing commercial-grade text and image content generation capabilities.

Compared to the previous generation, Qwen-Image-3.0 has increased the text input length by 4.5 times. In scenarios that rely on long prompts, such as film storyboards, knowledge diagrams, and product explanation pages, users no longer need to compress their requirements. They can describe the visual structure, text content, visual style, and layout details as if writing a design document, resulting in more accurate generated results and reducing the cost of repeated modifications and manual debugging.
In terms of complex content understanding, Qwen-Image-3.0 further enhances semantic parsing and spatial layout capabilities, allowing it to generate a nine-grid knowledge diagram covering multiple topics in one go, ensuring clarity of text and accuracy of content. Additionally, the model supports complex logical nesting, capable of generating combinations of multiple visual elements such as web pages, software interfaces, chat windows, and posters based on a single instruction. For example, in complex scenarios like "generating a hand-poured coffee poster for a chat interface using the Qwen App within a VSCode programming interface," the model can accurately understand multi-layered UI relationships while maintaining consistent styles across all interfaces, even displaying small font text of about 10 pixels clearly.

The official stated that Qwen-Image-3.0 has strong detail restoration capabilities in scenarios such as portrait photography, math tests, ancient stone inscriptions, and live streaming pages, with key optimizations made for complex text and image layouts and high-density information carrying capacity.
