Picking a video model starts with what you’re feeding it: a blank prompt, a single start frame, two keyframes to interpolate between, reference clips, or existing footage to edit. The tables below cover every provider’s models, resolutions, and durations for Video Generation mode — or jump straight to the sub-mode map at the bottom to find every model that supports the workflow you need.
Model coverage matches the in-app Video Generation picker. For pricing, see the Pricing page. Plan differences: Plans & Billing.
Google
Veo
Google video generation with up to 4K on supported models.
| Model | Sub-modes | Notes |
|---|---|---|
| Veo 3.1 Lite | Text to Video, Start Frame, Interpolation | Most affordable Veo; default |
| Veo 3.1 Fast | + References | Quick previews |
| Veo 3.1 | + References | Full quality |
Resolutions: 720p, 1080p, 4K · Duration: 4 / 6 / 8s (1080p and 4K lock to 8s)
Gemini
| Model | Sub-modes | Notes |
|---|---|---|
| Omni 1.1 Flash (preview) | Text to Video, Start Frame, Interpolation, References, Edit, Extend | 360p / 720p / 1080p / 4K; 3–10s; up to 6 mixed references (up to 3 videos, 3s each) |
The previous Omni Flash model is retired from the picker but remains available when restoring older generations.
Alibaba
HappyHorse
| Model | Sub-mode |
|---|---|
| HappyHorse 1.1 T2V | Text to Video |
| HappyHorse 1.1 I2V | Start Frame (default) |
| HappyHorse 1.1 R2V | References (image refs) |
| HappyHorse 1.0 Video Edit | Edit |
HappyHorse 1.0 T2V / I2V / R2V are superseded in the picker but remain for restoring older generations. 1.0 Video Edit stays available.
Resolutions: 720p / 1080p · Duration: 3–15s · seeds, optional watermark
Wan
| Model | Sub-modes |
|---|---|
| Wan 3.0 (beta) | Text to Video, Start Frame, Interpolation, References, Edit |
| Wan 3.0 Prime (beta) | Same, with significantly faster generation at a higher rate |
| Wan 2.7 T2V | Text to Video |
| Wan 2.7 I2V | Start Frame, Interpolation, Continuation |
| Wan 2.7 R2V | References (image or video refs) |
| Wan 2.7 Video Edit | Edit |
| Wan 2.2 KF Flash | Interpolation |
| Wan 2.6 Flash / Wan 2.6 | Start Frame (retired) |
| Wan 2.6 R2V Flash / R2V | References (retired) |
Wan 3.0 is all-in-one: one model covers every sub-mode it lists, so switching sub-modes does not change the model. It renders 2–30s at 480p / 720p / 1080p with audio on by default, and takes mixed references — images and video together in one tray (ten slots), plus an audio reference — cited in the prompt as “Image 1”, “Video 1”, “Audio 1”. Reference video and reference audio are capped at 15s each in total. Video input additionally counts against a 30s combined budget with the render, so a long video reference lowers the maximum output length; audio references do not. Wan 3.0 takes no negative prompt.
Wan 3.0’s Edit regenerates from the source clip as a reference rather than preserving it frame for frame — reach for Wan 2.7 Video Edit when the original footage has to survive the edit, and for Wan 2.7 I2V for Continuation, which Wan 3.0 does not offer.
Wan 2.6 is retired: superseded in the picker, but still restoring older generations.
Resolutions: 480p–1080p · optional audio · negative prompts (2.x only) · seeds
KlingAI
Kling
| Model | Sub-modes | Notes |
|---|---|---|
| Kling 3.0 Turbo | Text to Video, Start Frame | Fast; 720p / 1080p |
| Kling 3.0 | + Interpolation, Motion Control | Core multi-shot + motion ref |
| Kling 3.0 Omni | Text to Video, Start Frame, Interpolation, References | Mixed-media refs; 4K on many paths |
| Kling Avatar | Avatar | Talking avatar from start frame + audio |
Duration: typically 3–15s · motion control uses a motion reference input
BytePlus / ByteDance
Seedance
| Model | Notes |
|---|---|
| Seedance 2.5 | Up to 30s; 480p–1080p; up to 50 refs (image, video, audio); timestamped editing |
| Seedance 2.0 | Multimodal generation; 480p–1080p; up to 15s; image + video refs |
| Seedance 2.0 Fast | Faster; 480p / 720p |
| Seedance 1.5 Pro | Previous gen with audio output |
Seedance 2.5 is the only model that accepts an audio reference on its own, and the only one that honours whole-second timestamps in a prompt. When editing or extending, it keeps the source clip’s aspect ratio and length, so those pills are hidden.
OmniHuman
| Model | Notes |
|---|---|
| OmniHuman 1.5 | Avatar sub-mode: image + required audio → talking portrait; 720p / 1080p |
LTX
| Model | Notes |
|---|---|
| LTX 2.5 Fast | Native multi-shot; 720p–4K; up to 20s at 720p/1080p, 10s at 1440p/4K |
| LTX 2.5 Pro | Native multi-shot at higher fidelity; 720p / 1080p; max 10s |
| LTX 2.3 Fast | 1080p / 1440p / 4K; durations up to 20s |
| LTX 2.3 Pro | Higher fidelity; max 10s; audio-driven & interpolation |
Sub-modes: Text to Video, Start Frame, Interpolation.
LTX 2.5 renders several connected shots from a single prompt, holding character, scene, lighting, style, and voice across the cuts. Reframe remains LTX 2.3 Pro only.
xAI (Grok)
| Model | Sub-modes | Notes |
|---|---|---|
| Imagine Video 1.0 | Text to Video, Start Frame, References, Edit, Extend | Full suite |
| Imagine Video 1.5 | Start Frame only | Image-to-video |
Resolutions: 480p / 720p · Duration: 1–15s
Luma
| Model | Sub-modes | Notes |
|---|---|---|
| Ray 3.2 | Text to Video, Start Frame, Interpolation, Edit, Reformat | HDR, seamless loop; 540p–1080p; 5 or 10s |
Sub-mode map (who supports what)
The first six rows are Video Generation sub-modes; Edit, Extend, and Reformat belong to the separate Video Edit mode.
| Sub-mode | Typical providers |
|---|---|
| Text to Video | Veo, Gemini, Wan 3.0/2.7, HappyHorse, Kling, Seedance, LTX, Grok, Luma |
| Start Frame | Most providers |
| Interpolation | Veo, Gemini Omni 1.1, Wan 3.0/2.7/2.2, Kling, LTX, Luma |
| References | Veo Fast/Full, Gemini, Wan 3.0/2.7, HappyHorse, Kling Omni, Seedance 2.5/2.0, Grok 1.0 |
| Motion Control | Kling 3.0 |
| Avatar | Kling Avatar, OmniHuman |
| Edit | Gemini Omni 1.1 Flash, Wan 3.0/2.7, HappyHorse 1.0 Edit, Seedance 2.5/2.0, Grok 1.0, Luma |
| Extend | Gemini Omni 1.1 Flash, Wan 2.7, Seedance 2.5/2.0, Grok 1.0 |
| Reformat | Luma Ray 3.2 |