Self-Supervised AI Model GedankenNet: Farewell to Real Data Feeding, Redefining Holographic Microscopy Reconstruction


The robotics research team at Google DeepMind recently released a robotics project called RT-2. This project took 7 months to develop and uses a large model for training. RT-2 has capabilities such as symbol understanding, reasoning, and human recognition, and can think and complete tasks based on human instructions. By combining the large model with the robot's operational capabilities, RT-2 can accomplish tasks that involve logical leaps, such as from 'extinct animals' to 'plastic dinosaurs'. The results of this project performed well in various sub - category tests, with performance up to three times that of the previous generation of robot models. This research result demonstrates the potential of large models in robotics research and is expected to drive the development of robots in the future.
Meta Intelligence OS is a startup founded by Bloomberg. It has developed a series of large models based on the open-source model RWKV and aims to become the Android in the era of large models. The RWKV model has superior performance and low cost in inference tasks, thus attracting customers from industries such as finance, law firms, and smart hardware. The business model of Meta Intelligence OS is model customization based on private data and internal AI Agent development. The company hopes to solve the problems of API call latency and data security by deploying large models on terminal devices. Currently, RWKV versions are available on Windows, Mac, and Linux computers, and Android and iOS versions are also in development. Meta Intelligence OS is raising funds and collaborating with chip companies and computing power platforms to create benchmark customers. Luo Xuan said that the decisive battlefield for large models is on hardware, and both terminal devices and the cloud require dedicated chips.
ElevenLabs has unveiled its new Scribe speech-to-text model, boasting unprecedented accuracy levels, particularly in English where it achieves a remarkable 96.7% accuracy rate.
The rapid advancement of Artificial Intelligence (AI) large language model technology in recent years has led to a price war intensifying market competition. DataBao's latest statistics indicate this downward price trend will continue in 2025. Recently, companies like ByteDance and Alibaba Cloud have announced price reductions for their AI large language models, attracting significant industry attention. For example, ByteDance's Doubao large language model announced a price cut in December last year, reducing the price of its visual understanding model to 0.003 yuan/thousand tokens.
Microsoft has unveiled Phi-4, a new multimodal and mini-model designed to significantly improve processing capabilities across speech, vision, and text. This release represents a notable advancement in AI technology.