Time each story beat
Use ranges such as [0–3s], [3–6s], and [6–10s] when the sequence needs several ordered actions.
Use Gemini Omni 1.1 Flash on Ezier to generate 3–10-second videos from text or images, guide first and last frames, preserve reference subjects, extend scenes, and refine edits with plain-language instructions.
Move from a text brief to reference-led motion, then refine a clip without rebuilding the whole scene. Each prompt makes the subject, timing, camera, sound, and protected details explicit.
Create a 5-second photorealistic 16:9 shot looking straight down a weathered wooden pier that disappears into dense morning fog. Dozens of gray-and-white seagulls line both railings and gather on the damp boards; a few turn their heads, shuffle, flap, and lift briefly while the rest remain naturally spaced. Use a slow centered push-in, a level horizon, cool overcast light, calm gray water, and soft atmospheric depth. Generate distant gull calls, light wingbeats, faint water movement, and damp outdoor ambience. Keep the pier geometry, bird scale, fog density, and camera axis stable. No people, boats, text, logos, duplicate birds, warped railings, sudden weather changes, or cuts.
Choose the model, then confirm the available duration, aspect ratio, resolution, and credit estimate for the shot.
Start from text, add a first or last frame, attach subject references, or provide a source video for editing or extension.
Describe the action, camera path, timing, sound, and final state. Name every face, object, or composition detail that must stay consistent.
Check the result for continuity and instruction following, then request one focused edit or extend the strongest scene.
Use text, images, and video as clearly assigned inputs. Direct the transition, preserve reference subjects, or request one focused change in natural language.

Provide a starting image and an ending image, then describe the path between them. Specify what changes over time and what must remain visually continuous.

Use reference images to define a person, object, or visual style. Repeat the identity details and tell the model whether each image is a reference or a literal frame.

Make edit instructions narrow and explicit. Name the single change, then protect the subject, pose, camera, background, timing, and sound that should stay untouched.
Match the output to the job with a defined duration, format, resolution, and reference plan.
Start with a detailed scene brief or animate a high-resolution image by naming the subject motion, camera movement, and environmental effects.
Use first-and-last-frame interpolation to plan where the sequence begins and ends while the model generates the movement between them.
Continue from a generated result and describe the next edit in plain language, protecting the details that should not change.
Append a new three-to-ten-second continuation and introduce a referenced character or object while keeping the existing scene direction coherent.
Give every input a role, describe motion in sequence, and name the details that must remain unchanged.
Use ranges such as [0–3s], [3–6s], and [6–10s] when the sequence needs several ordered actions.
State whether an image is the first frame, last frame, subject reference, object reference, or style reference.
For edits, name the one change first and follow it with the face, pose, camera path, background, and audio to preserve.
Write constraints such as “no scene cuts,” “no dialogue,” or “no added objects” directly into the generation brief.
Ready to create?
Turn a clear brief into a controllable video sequence.
Define the scene, reference roles, camera path, sound, and protected details before you generate.
Answers about generation, reference media, interpolation, editing, duration, resolution, and prompt structure.
Gemini Omni 1.1 Flash is Google's multimodal model for fast video generation, editing, interpolation, and extension. It accepts text, images, and video and produces high-resolution video with audio.
A generated clip or continuation can last from 3 to 10 seconds. Use timed prompt sections when several actions need to happen in a specific order.
Google lists 360p and 720p generation, plus upscaled 1080p and 4K output. Videos are generated at 24 FPS in 16:9 or 9:16 format.
Yes. Add a high-resolution source image and describe the subject movement, camera path, lighting change, and environmental motion you want to see.
Yes. Provide two images as the starting and ending frames, then describe the transition, pacing, camera movement, and details that should remain consistent between them.
Yes. Give the model a source video and a focused instruction to revise it or append a new 3–10-second continuation. Reference media can guide new subjects or objects in the extension.
Describe the subject, action, setting, shot order, camera movement, lighting, sound, and final frame. Assign every uploaded image or video a clear role and state what must not change.
Next up
Try a nearby video model, open video tools, or prepare a stronger source image.

AI models
Compare models
Choose a better fit before you generate.

AI video tools
Turn images into motion
Animate a still or create a clip.

AI image tools
Create or edit images
Make, clean up, or upscale stills.

Omni Flash
Conversational editing
Use for short video generation and conversational edits.

Veo 3.1
Cinematic control
Use for cinematic shots and camera moves.

Seedance 2.5
Longer reference-led video
Use for longer scenes with richer reference direction.

Wan 3.0
Longer beta video generation
Use for text or start-image video up to 30 seconds.

LTX 2.5
4K audio-video workflow
Use for 1080p or 4K video with Fast, Pro, and generated-audio controls.

Kling 3.0
Action and motion
Use for action, motion, and story beats.

AI models
Compare models
Choose a better fit before you generate.

AI video tools
Turn images into motion
Animate a still or create a clip.

AI image tools
Create or edit images
Make, clean up, or upscale stills.

Omni Flash
Conversational editing
Use for short video generation and conversational edits.