Google introduces Gemini Omni, a new multimodal AI model from its DeepMind division, unveiled at Google I/O 2026. The model is designed to generate and edit video using a conversational workflow, taking multiple types of user input—including video, images, text, and audio—and producing video outputs aligned with prompts. Several outlets describe Gemini Omni as capable of “transformational” video editing, not just creation, allowing users to modify existing media through instructions rather than traditional editing steps. TechRadar and others highlight a specific initial model variant, Gemini Omni Flash, described as optimized for creating and editing video through conversation. CNET and Engadget frame the system as eventually extending beyond video to other kinds of outputs derived from inputs, while at the same time noting that the current rollout starts with video. The Next Web reports that some features are held back, including speech-editing, and that SynthID watermarking is enabled by default. Decrypt and Business Standard add that Google positions the model to produce realistic results, including physics-aware “cinematic” outputs. Coverage also links related Gemini Omni updates to Google’s broader content-generation tools.