Recently, the official WeChat AI team officially announced that the general multimodal embedding model WeMM-Embedding will be open-sourced globally, and three different parameter versions - 2B, 4B, and 9B - have been released. The model demonstrates strong multimodal compatibility, supporting text, images, and videos comprehensively, and can also process visual documents and any mixed multimodal input natively.

Technical Foundation and Core Advantages of Multimodal Isomorphism

According to the WeChat Vision team, the core logic of WeMM-Embedding lies in enabling the system to deeply understand text, images, and videos simultaneously, mapping and unifying them within the same semantic space for efficient comparison and precise retrieval.

In the internationally authoritative multimodal retrieval and evaluation benchmark MMEB-v2, this model achieved first place among similar solutions with its outstanding overall performance, surpassing all previously submitted open-source and commercial closed-source models.

Large-Scale Business Deployment and Open Source Ecosystem

As a technically refined product developed internally, WeMM-Embedding is not limited to the laboratory stage. Currently, the model has been widely applied in multiple core businesses within the WeChat ecosystem. Its actual application scenarios include Moments search, Video Number recommendations, Official Account recommendations, and WeChat e-commerce, among others. The daily call volume has already exceeded 1 billion times.

To further promote the prosperity of the technical community and cutting-edge innovation, the WeChat team has also publicly released the complete technical ecosystem resources. Developers and researchers can not only conduct in-depth research through relevant code repositories but also access the corresponding model sets on the model hosting platform.