Gemini Omni is a conversational creative studio for generating, directing, and editing videos from multimodal inputs. It is designed to let creators describe a scene or requested change in natural language instead of working through conventional timeline-based editing.
What it does
Generates creative videos from text prompts and reference assets, including images, video, and audio.
- Combines multiple input types into a cohesive video output.
- Lets users modify videos conversationally, such as requesting a background change, camera adjustment, or different character action.
- Creates remixed variations of existing videos from text prompts.
Key features
- Text-to-video generation with prompts for scene descriptions, camera motion, and dynamic actions.
- Multi-image fusion and support for image, text, video, and audio references.
- Chat-based video editing and remixing without requiring complex timeline editing.
- Character and scene consistency across shots, according to the site.
- Native synchronized audio generation, including dialogue, background music, ambience, and action sound effects, according to the site.
- Gemini Omni 1.1 Flash is presented with 10-second deep context, 40-second continuous scene extension, 360p fast drafts, and start and end keyframes. What makes it different Gemini Omni centers the video-creation workflow on natural conversation: users can direct and revise content by describing the desired result. The site also positions its multimodal reference handling as a way to work from combinations of text, images, video, and audio rather than a single prompt type.