Video Providers

All video generation models currently available in Hylope, organized by provider.

Picking a video model starts with what you’re feeding it: a blank prompt, a single start frame, two keyframes to interpolate between, reference clips, or existing footage to edit. The tables below cover every provider’s models, resolutions, and durations for Video Generation mode — or jump straight to the sub-mode map at the bottom to find every model that supports the workflow you need.

Model coverage matches the in-app Video Generation picker. For pricing, see the Pricing page. Plan differences: Plans & Billing.


Google Google

Veo

Google video generation with up to 4K on supported models.

Model Sub-modes Notes
Veo 3.1 Lite Text to Video, Start Frame, Interpolation Most affordable Veo; default
Veo 3.1 Fast + References Quick previews
Veo 3.1 + References Full quality

Resolutions: 720p, 1080p, 4K · Duration: 4 / 6 / 8s (1080p and 4K lock to 8s)

Gemini

Model Sub-modes Notes
Omni 1.1 Flash (preview) Text to Video, Start Frame, Interpolation, References, Edit, Extend 360p / 720p / 1080p / 4K; 3–10s; up to 6 mixed references (up to 3 videos, 3s each)

The previous Omni Flash model is retired from the picker but remains available when restoring older generations.


Alibaba Alibaba

HappyHorse

Model Sub-mode
HappyHorse 1.1 T2V Text to Video
HappyHorse 1.1 I2V Start Frame (default)
HappyHorse 1.1 R2V References (image refs)
HappyHorse 1.0 Video Edit Edit

HappyHorse 1.0 T2V / I2V / R2V are superseded in the picker but remain for restoring older generations. 1.0 Video Edit stays available.

Resolutions: 720p / 1080p · Duration: 3–15s · seeds, optional watermark

Wan

Model Sub-modes
Wan 3.0 (beta) Text to Video, Start Frame, Interpolation, References, Edit
Wan 3.0 Prime (beta) Same, with significantly faster generation at a higher rate
Wan 2.7 T2V Text to Video
Wan 2.7 I2V Start Frame, Interpolation, Continuation
Wan 2.7 R2V References (image or video refs)
Wan 2.7 Video Edit Edit
Wan 2.2 KF Flash Interpolation
Wan 2.6 Flash / Wan 2.6 Start Frame (retired)
Wan 2.6 R2V Flash / R2V References (retired)

Wan 3.0 is all-in-one: one model covers every sub-mode it lists, so switching sub-modes does not change the model. It renders 2–30s at 480p / 720p / 1080p with audio on by default, and takes mixed references — images and video together in one tray (ten slots), plus an audio reference — cited in the prompt as “Image 1”, “Video 1”, “Audio 1”. Reference video and reference audio are capped at 15s each in total. Video input additionally counts against a 30s combined budget with the render, so a long video reference lowers the maximum output length; audio references do not. Wan 3.0 takes no negative prompt.

Wan 3.0’s Edit regenerates from the source clip as a reference rather than preserving it frame for frame — reach for Wan 2.7 Video Edit when the original footage has to survive the edit, and for Wan 2.7 I2V for Continuation, which Wan 3.0 does not offer.

Wan 2.6 is retired: superseded in the picker, but still restoring older generations.

Resolutions: 480p–1080p · optional audio · negative prompts (2.x only) · seeds


KlingAI KlingAI

Kling

Model Sub-modes Notes
Kling 3.0 Turbo Text to Video, Start Frame Fast; 720p / 1080p
Kling 3.0 + Interpolation, Motion Control Core multi-shot + motion ref
Kling 3.0 Omni Text to Video, Start Frame, Interpolation, References Mixed-media refs; 4K on many paths
Kling Avatar Avatar Talking avatar from start frame + audio

Duration: typically 3–15s · motion control uses a motion reference input


BytePlus BytePlus / ByteDance

Seedance

Model Notes
Seedance 2.5 Up to 30s; 480p–1080p; up to 50 refs (image, video, audio); timestamped editing
Seedance 2.0 Multimodal generation; 480p–1080p; up to 15s; image + video refs
Seedance 2.0 Fast Faster; 480p / 720p
Seedance 1.5 Pro Previous gen with audio output

Seedance 2.5 is the only model that accepts an audio reference on its own, and the only one that honours whole-second timestamps in a prompt. When editing or extending, it keeps the source clip’s aspect ratio and length, so those pills are hidden.

OmniHuman

Model Notes
OmniHuman 1.5 Avatar sub-mode: image + required audio → talking portrait; 720p / 1080p

LTX LTX

Model Notes
LTX 2.5 Fast Native multi-shot; 720p–4K; up to 20s at 720p/1080p, 10s at 1440p/4K
LTX 2.5 Pro Native multi-shot at higher fidelity; 720p / 1080p; max 10s
LTX 2.3 Fast 1080p / 1440p / 4K; durations up to 20s
LTX 2.3 Pro Higher fidelity; max 10s; audio-driven & interpolation

Sub-modes: Text to Video, Start Frame, Interpolation.

LTX 2.5 renders several connected shots from a single prompt, holding character, scene, lighting, style, and voice across the cuts. Reframe remains LTX 2.3 Pro only.


xAI xAI (Grok)

Model Sub-modes Notes
Imagine Video 1.0 Text to Video, Start Frame, References, Edit, Extend Full suite
Imagine Video 1.5 Start Frame only Image-to-video

Resolutions: 480p / 720p · Duration: 1–15s


Luma

Model Sub-modes Notes
Ray 3.2 Text to Video, Start Frame, Interpolation, Edit, Reformat HDR, seamless loop; 540p–1080p; 5 or 10s

Sub-mode map (who supports what)

The first six rows are Video Generation sub-modes; Edit, Extend, and Reformat belong to the separate Video Edit mode.

Sub-mode Typical providers
Text to Video Veo, Gemini, Wan 3.0/2.7, HappyHorse, Kling, Seedance, LTX, Grok, Luma
Start Frame Most providers
Interpolation Veo, Gemini Omni 1.1, Wan 3.0/2.7/2.2, Kling, LTX, Luma
References Veo Fast/Full, Gemini, Wan 3.0/2.7, HappyHorse, Kling Omni, Seedance 2.5/2.0, Grok 1.0
Motion Control Kling 3.0
Avatar Kling Avatar, OmniHuman
Edit Gemini Omni 1.1 Flash, Wan 3.0/2.7, HappyHorse 1.0 Edit, Seedance 2.5/2.0, Grok 1.0, Luma
Extend Gemini Omni 1.1 Flash, Wan 2.7, Seedance 2.5/2.0, Grok 1.0
Reformat Luma Ray 3.2
Esc
Type to search the documentation and changelog.