Catalogue vidéo

Modèles vidéo

Comparez les modèles de génération vidéo actuellement disponibles sur APIXO, puis affinez la liste par capacité, récence de lancement et prix.

Type de modèle

Tout60 Image13 Vidéo22 Audio3 Texte22

Flux de travail

22 modèles

10 fournisseurs dans cette famille

Google

Gemini Omni

Gemini Omni is Google's multimodal video generation model for creating videos from text, image references, source video, reusable audio assets, and character asset IDs.

NouveauTexte vers vidéoImage vers vidéo

à partir de $0.1/secVoir

bytedance

Seedance 2.0

Seedance 2.0 is ByteDance's multimodal video model supporting text-to-video, first-and-last-frames, and omni-reference modes. APIXO exclusive: unlimited concurrency, real-person portrait support, and hidden capabilities.

NouveauPopulaireTexte vers vidéoImage vers vidéo

à partir de $0.55/secVoir

bytedance

Seedance 2.0 Fast

Seedance 2.0 Fast is the speed-optimized variant of ByteDance's multimodal video model. It supports text-to-video, first-and-last-frames, and omni-reference modes with lower per-second pricing and the same APIXO-exclusive capabilities.

NouveauPopulaireTexte vers vidéoImage vers vidéo

à partir de $0.45/secVoir

Alibaba

HappyHorse

HappyHorse is Alibaba's video generation and editing model for text-to-video, image-to-video, reference-guided generation, and video-edit workflows with 720p/1080p output.

NouveauPopulaireTexte vers vidéoImage vers vidéo

à partir de $0.375/secVoir

Alibaba

Wan 2.7

Wan 2.7 is Alibaba's video generation and editing model for text-to-video, image-to-video, reference-guided generation, and video-edit workflows with optional audio input and 720p/1080p output.

NouveauTexte vers vidéoImage vers vidéo

à partir de $0.1/secVoir

Alibaba

Wan 2.6

Wan 2.6 is Alibaba's multi-mode video generation model for text, image, flash image, reference, and flash reference workflows, with optional audio input and 720p/1080p output.

NouveauTexte vers vidéoImage vers vidéo

à partir de $0.025/secVoir

OpenAI

Sora 2 Pro

Sora 2 Pro is OpenAI’s premium video generation model with higher quality output, supporting text-to-video and image-to-video at 720p and 1080p resolutions with flexible durations of 10 or 15 seconds.

NouveauTexte vers vidéoImage vers vidéo

à partir de $3/secVoir

bytedance

Seedance 1.5 Pro

Seedance 1.5 Pro is ByteDance's per-second video model for fast text-to-video and image-to-video generation with 480p/720p output, optional sound, aspect ratio control, and fixed-lens camera stability.

NouveauTexte vers vidéoImage vers vidéo

à partir de $0.01/secVoir

hailuo

Hailuo 2.3

Hailuo 2.3 is Miniax's async video model with standard and pro modes for text-to-video and image-to-video generation. Standard mode supports 6s/10s at 768p, while pro mode returns fixed 5s at 1080p.

NouveauTexte vers vidéoImage vers vidéo

à partir de $0.336/secVoir

hailuo

Hailuo 2.3 Fast

Hailuo 2.3 Fast is Miniax's speed-optimized image-to-video model with standard and pro modes. Standard supports 6s/10s at 768p, while pro returns fixed 6s output at 1080p.

NouveauImage vers vidéo

à partir de $0.192/secVoir

xai

Grok Video

Grok Video is xAI's async video generation model for text-to-video and image-to-video workflows, with optional continuation via task_id + index and style control.

NouveauTexte vers vidéoImage vers vidéo

à partir de $0.09/secVoir

Alibaba

Wan 2.2 Animate

Wan 2.2 Animate API is Alibaba's character animation model that combines one source image and one motion video to generate stylized animated outputs with animate/replace behavior.

NouveauEffets vidéo

à partir de $0.04/secVoir

MeiGen

InfiniteTalk

InfiniteTalk converts one photo plus audio into audio-driven talking or singing avatar videos with precise lip synchronization. Supports up to 10 minutes at 480p or 720p resolution.

NouveauImage vers vidéo

à partir de $0.15/secVoir

kling

Kling 3.0 Std

Kling 3.0 Std is Kuaishou's standard-quality video generation model with text-to-video, image-to-video, and motion-control modes. It supports clips up to 15 seconds with optional sound generation and flexible aspect ratios.

NouveauTexte vers vidéoImage vers vidéo

à partir de $0.42/secVoir

vidu

Vidu Q3

Vidu Q3 is a per-second video generation model that combines standard and Turbo text-to-video plus image-to-video workflows in one API. It supports single-image animation, first-and-last-frame transitions, optional sound and BGM, and output up to 1080p.

NouveauTexte vers vidéoImage vers vidéo

à partir de $0.036/secVoir

kling

Kling 2.5 Turbo Pro

Kling 2.5 Turbo Pro is Kuaishou's high-speed video model for text-to-video and image-to-video creation. It supports 5-10 second clips, optional tail-frame images, aspect ratio control for text-to-video, plus negative prompts and CFG scale guidance.

NouveauTexte vers vidéoImage vers vidéo

à partir de $0.3/secVoir

Lightricks

LTX-2 19B

LTX-2 19B is Lightricks' open-source 19B diffusion transformer for cinematic video generation. It supports text-to-video and image-to-video workflows, LoRA conditioning, and high-fidelity outputs up to 1080p in the API.

NouveauTexte vers vidéoImage vers vidéo

à partir de $0.012/secVoir

kling

Kling 2.1

Kling 2.1 is Kuaishou's multi-tier video model with Standard, Pro, and Master modes for image-to-video and text-to-video creation. It supports 5-10 second clips, optional tail images for Pro, and aspect ratio control for Master text-to-video.

NouveauTexte vers vidéoImage vers vidéo

à partir de $0.2/secVoir

kling

Kling 2.6

Kling 2.6 is Kuaishou's native audio-visual video model that generates video, speech, sound effects, and ambience in one pass. It supports text-to-audio-visual and image-to-audio-visual creation with Chinese and English voice generation and up to 10-second clips.

NouveauTexte vers vidéoImage vers vidéo

à partir de $0.3/secVoir

Google

Veo 3.1

Google DeepMind’s upgraded AI video model with lite, fast, and quality routes, 4/6/8 second duration control, 720p/1080p/4k output, and multi-image reference workflows.

Texte vers vidéoImage vers vidéo

à partir de $0.15/secVoir

Alibaba

Wan 2.5

Wan 2.5 is Alibaba's video generation model for text-to-video and image-to-video workflows, with optional audio input, 480p/720p/1080p output, 5 or 10 second clips, and prompt expansion.

Texte vers vidéoImage vers vidéo

à partir de $0.04/secVoir

OpenAI

Sora 2

Sora 2 is OpenAI’s latest AI video generation model, supporting both text-to-video and image-to-video. It delivers realistic motion, physics consistency, with improved control over style, scene, and aspect ratio—ideal for creative apps and social media content.

Texte vers vidéoImage vers vidéo

à partir de $0.2/secVoir