Fu Sheng Unveils Orion-14B Large Model with 14 Billion Parameters


Zhipu disclosed important technical signals of its new flagship large models GLM-5.4 and GLM-5.5: the parameter scale of both models has officially entered the trillions, and they will explore two new paths in the underlying technical architecture, seen as a step forward for domestic large models into the deep waters of AGI. The 'dual-track new path' has attracted widespread attention in the industry, marking Zhipu's advancement in the evolution of next-generation large model technology.
Tencent Hunyuan research examines batch size in online LLM RL. As models self-generate data and rollout/training scaling rates diverge, classical critical batch size theory needs revisiting. Covering GRPO and PPO, it gives practical guidance for efficient training on large GPU clusters.....
The Alibaba Cloud Yingying team is developing the AI device QwenBook, positioned as a native intelligent agent computer, with a form similar to the Qwen tablet, still in the early trial phase, focusing on office scenarios, integrating the Qwen3.8-Max and Flash models. The product definition and hardware are self-developed, integrating the Yingying Cloud PC system and end-cloud capabilities, with a team of about 150 people.
Zhipu AI launches GLM-5.3-FlashX, with the API also going live, offering a maximum output of 200 tokens/s. It focuses on intelligence, price, and speed, providing high-throughput, low-latency inference for enterprise developers. The predecessor, GLM-5.3-Flash, was previously introduced overseas under the name Ox Alpha. It gained popularity due to its strong intelligence and cost-effectiveness at the same size, with increasing usage volume. Zhipu is now supporting growing demand.
Ant Group's AI Security Lab opens the internal security guardrail technology called SingProbe. Unlike external review, it reuses the hidden states of the base model for inference, achieving real-time token-level risk assessment with less than 0.5% additional computing power. It can simultaneously perform intent classification, safety detection, and hallucination recognition, and can block risks at the millisecond level before output in high-risk scenarios such as healthcare. It has been adapted to 29 mainstream large models.