Gemini Omni Turns Real Footage Into Surreal Video, Blurring The Line Between Source And Output

Author: Qoo Media

Google is pushing its video AI in a new direction with Gemini Omni, a multimodal model that is less about generating scenes from scratch and more about transforming real footage into something far more surreal. That shift makes the tool stand apart from Veo, which has been known for consistent elements and highly accurate lip-sync in AI-generated video.

Gemini Omni first appeared in the form of Omni Flash, where Google positioned it as a model that can combine text, audio, still images, and video into outputs that may look very different from the original input. Video remains the main focus, but the broader pitch is clear: this is a model built for imaginative edits that can preserve character consistency across frames while still pushing the result into more experimental territory.

A broader multimodal approach

Unlike a simple video generator, Gemini Omni is designed to analyze several inputs at once while shaping the story being built. That means the creative process does not end with a single prompt, because users can keep refining the result with natural-language instructions until the output matches the intended vision.

Google also says the model understands the physical world in a way that should help the generated video follow real-world physics. The company points to gravity, kinetic energy, and fluid dynamics as part of that understanding, which could become a meaningful differentiator if the output holds up in practice.

Built on Veo, but aiming beyond it

The new model still sits on top of the foundation Google built with Veo across its AI video ecosystem. Veo Gen 3 is known for assembling full scenes with consistent elements and very polished lip-sync, but Veo 3 and the newer 3.1 version are described as being limited to fully AI-generated video from text and audio.

Gemini Omni expands that scope by opening the door to turning real-life video into new material that is more experimental and visually detached from the source. In that sense, the model is not just another update in the lineup, but a different class of multimodal system.

Input can come from several directions

Users can start with an image, audio, or video reference, and text can also serve as the only starting point if no source material is available. Google is also introducing the idea of using a digital version of oneself, which adds a more personal layer to the tool.

Through the Avatars feature, users can create videos that show a character that looks and sounds like them. That broadens Gemini Omni’s use beyond visual effects and into character creation, digital identity experiments, and more personalized content production.

Where access is coming from

Google is rolling out Gemini Omni through the Flash model inside the Gemini app, and it is available to paying subscribers on the Google AI Plus, Pro, and Ultra tiers. The company is also bringing it to Flow, Google’s AI filmmaking tool, which suggests a role beyond quick experiments and into a more serious creative workflow.

For users who want to try it without paying, Google is offering another path through YouTube. There, Omni can be used to create Remixes from existing Shorts, and the free access is also available in YouTube Create, extending AI video tools to a wider group of creators.

Early examples and an open question

Google has already shown two early examples to demonstrate what the model can do. One features comedian Adam Waheed, while the other involves YouTuber Happy Kelli, both used to show how Gemini Omni can reshape video material into something more dramatic and unusual.

The demonstrations focus on how far real footage can be altered while keeping the main subject recognizable. One issue remains unresolved, though, because Google has not said whether creators will be able to restrict or block their content from being remixed with AI.

That question is likely to become more important as Shorts Remix tools spread. The easier it becomes for AI to reshape original videos into new works, the more attention will fall on who controls that content in the first place.

Source: www.androidauthority.com
Latest