When Musk promoted the new version of his AI platform Grok on a social media platform, he uploaded an AI-generated video of the late basketball legend Kobe Bryant as a test case. Grok not only fully analyzed the visuals and scenes but also accurately extracted the key dialogue, ultimately determining the video to be a deepfake AI-generated content—showcasing new video interaction capabilities.
From the chat logs released by Musk, it can be seen that Grok first analyzed the footage frame by frame: "Kobe was wearing a black suit, delivering a monologue with his signature 'Mamba spirit,' accompanied by pointing, gestures, smiling, and leaning forward towards the camera, with lighting that had a cinematic quality." The model then accurately extracted key dialogues from the video such as "The strongest intelligent model SpaceX has ever built," "Grok 4.5, greatness comes from perseverance, Mamba, exit," among others.
Comprehensive determination as Deepfake, the AI video of Kobe is exposed
The most critical part was the comprehensive judgment—Grok combined all visual elements, dialogue, and contextual information to finally determine that this footage was a deepfake AI-generated video, and pointed out that Kobe had passed away in 2020, and all the images of the person in the video were generated by AI rather than real footage. This means that Grok not only can "understand" video content, but also has the reasoning ability to distinguish authenticity.
From audio summaries to full-dimensional understanding, video interaction achieves a generation leap
The video interaction feature newly launched by Grok allows users to directly upload local video files or paste links in the conversation window, giving various instructions to the AI. Supported functions include summarizing the entire video, identifying people and objects in the visuals, breaking down action details, or continuously asking questions about any specific detail.
Compared to the previous mode of only extracting audio text to generate summaries, the new Grok has achieved a significant upgrade—capable of linking continuous frames, audio information, and contextual logic to comprehensively and deeply understand the entire video content. When AI evolves from "listening to videos" to "watching videos," and from content summarization to authenticity verification, the ceiling of video interaction has been raised once again.
