The mobile interaction methods of artificial intelligence are undergoing a disruptive transformation. OpenAI has recently officially announced the full introduction of voice-based intelligent agents (Agents) in its ChatGPT mobile application. This means that, in the future, when using a smartphone, users will not only be able to chat with AI as before through voice, but also directly issue oral commands to AI to navigate between various external applications and tools, allowing AI to autonomously perform complex real-world tasks behind the scenes.
From a core functionality perspective, this update fully transfers ChatGPT's previously highly praised Work capabilities on desktops to the mobile platform. For Plus and Pro subscription users, the Work tab on the mobile app can now directly create documents, draft emails, and summarize Slack messages, which are common office tasks; it also supports further expansion into more complex scenarios, such as website building, presentation creation, calling cloud browsers, and even handling various complex financial tasks. For Free and Go users, OpenAI has mainly opened support for various plugins and connected external applications, ensuring that users at all levels can benefit from the efficiency boost provided by smart tools.
In terms of interaction experience and underlying logic, the new mobile update has comprehensively optimized the voice mode. During the conversation, users not only hear smooth voice responses but can also see more complete and clear text output on the screen. For Plus subscribers, the application interface provides a more convenient seamless switching mechanism, allowing users to freely switch between text and voice modes without the hassle of ending the current conversation and starting over. In addition, cross-device continuous work capability makes the entire process extremely smooth: users can completely use fragmented time during commutes or outings to give task instructions to ChatGPT via phone voice, then return to their computer to continue viewing and completing tasks, achieving seamless integration across multiple devices.
Looking back at its technological evolution, this feature upgrade is no accident. In July of this year, OpenAI launched a new GPT-Live voice model, focusing on breaking through technical bottlenecks in natural conversation, interruption handling, and continuous communication. It was quickly deployed to desktops, empowering Codex software development and desktop intelligent agents. The extension to the mobile platform marks that OpenAI is redefining the functional positioning of mobile applications - the "dialogue box" that used to mainly handle Q&A and search has now officially evolved into a "super entrance" capable of scheduling external tools and replacing humans to perform multi-step operations. With the accelerated integration of voice, intelligent agents, and cross-device ecosystems, in the future, users only need to state their goals, and AI will take over the subsequent tedious operations.
Join Now