APIXO
    APIXO

    Créez tout le contenu que vous pouvez imaginer. Approuvé par des millions de créateurs.

    Support@apixo.ai

    Modèles

    Qwen Image 3.0Qwen Image 3.0 ProSeedream 5.0 ProNano Banana 2Nano Banana ProGPT-Image-2MiniMax Image 01Wan 2.6 ImageWan 2.5 ImageImage UpscalerImage Watermark RemoverGPT-Image-1Seedream 5.0Seedream 4.5Flux 2Flux KontextFLUX 3MidjourneySora 2Sora 2 ProVeo 3.1
    Gemini OmniWan 2.5Wan 2.6Wan 2.7Seedance 2.0Seedance 2.0 FastSeedance 2.0 MiniKling 3.0 StdKling 3.0 TurboKling 2.5 Turbo ProHappyHorseHailuo 2.3Vidu Q3Grok ImageGrok VideoMiniMax Speech 2.8MiniMax VoiceSuno V5GPT-5.5Claude Sonnet 4.6Gemini 3.1 Pro Preview

    Produit

    À la uneModèles d'imageModèles vidéoModèles audioModèles de texte

    Ressources

    AffiliationInformations sur l'entrepriseConditions d'utilisationPolitique de confidentialitéPolitique de remboursementPolitique de livraison / fourniture numérique

    Autre

    TarifsDocumentation API

    © 2026 APIXO. Tous droits réservés.

    LA DISPONIBILITÉ DES MODÈLES, LA TARIFICATION ET LE ROUTAGE PEUVENT VARIER SELON LE FOURNISSEUR ET LA RÉGION.

    APIXO
    TarifsDocumentation API
    1. Accueil
    2. Modèles
    3. Modèles vidéo

    Catalogue vidéo

    Modèles vidéo

    Comparez les modèles de génération vidéo actuellement disponibles sur APIXO, puis affinez la liste par capacité, récence de lancement et prix.

    Voir la documentation API

    Type de modèle

    Tout76Image20Vidéo31Audio3Texte22

    Flux de travail

    Type de modèle

    Tout76Image20Vidéo31Audio3Texte22

    Flux de travail

    31 modèles

    12 fournisseurs dans cette famille

    Google

    Gemini Omni

    Gemini Omni is Google's multimodal video generation model for creating videos from text, image references, source video, reusable audio assets, and character asset IDs.

    NouveauTexte vers vidéoImage vers vidéo
    à partir de $0.1/sVoir

    bytedance

    Seedance 2.0

    Seedance 2.0 is ByteDance's multimodal video model supporting text-to-video, first-and-last-frames, and omni-reference modes. APIXO exclusive: unlimited concurrency, real-person portrait support, and hidden capabilities.

    NouveauPopulaireTexte vers vidéoImage vers vidéo
    à partir de $0.0573/sVoir

    bytedance

    Seedance 2.0 Fast

    Seedance 2.0 Fast is an APIXO route for ByteDance Seedance 2.0 workflows, exposing text-to-video, first-and-last-frames, and omni-reference modes with optional sound, web search, and 480p/720p output.

    NouveauPopulaireTexte vers vidéoImage vers vidéo
    à partir de $0.044/sVoir

    Alibaba

    HappyHorse

    HappyHorse is Alibaba's video generation and editing model for text-to-video, image-to-video, reference-guided generation, and video-edit workflows with 720p/1080p output.

    NouveauPopulaireTexte vers vidéoImage vers vidéo
    à partir de $0.125/sVoir

    bytedance

    Seedance 2.5

    Seedance 2.5 is ByteDance's audiovisual video model for longer, reference-driven storytelling. It combines timeline-based direction, multimodal creative guidance, synchronized sound, multilingual performance, and selective revision to help production teams develop connected scenes instead of isolated short clips.

    NouveauTexte vers vidéoImage vers vidéo
    à partir de $0.085/sVoir
    MiniMax H3

    MiniMax

    MiniMax H3

    MiniMax H3 is a general-purpose multimodal video model that combines text, image, video, and audio context in one generation workflow. It creates videos up to 2K and 15 seconds with native stereo sound, while supporting first-and-last-frame control, mixed-media references, motion transfer, multi-shot storytelling, and instruction-guided video editing.

    NouveauTexte vers vidéoImage vers vidéo
    à partir de $0.09/sVoir

    bytedance

    Seedance 1.5 Pro

    Seedance 1.5 Pro is ByteDance's per-second video model for fast text-to-video and image-to-video generation with 480p/720p/1080p output, optional sound, aspect ratio control, and fixed-lens camera stability.

    NouveauTexte vers vidéoImage vers vidéo
    à partir de $0.0108/sVoir

    Alibaba

    Wan 2.7

    Wan 2.7 is Alibaba's video generation and editing model for text-to-video, image-to-video, reference-guided generation, and video-edit workflows with optional audio input and 720p/1080p output.

    NouveauTexte vers vidéoImage vers vidéo
    à partir de $0.1/sVoir

    Alibaba

    Wan 2.6

    Wan 2.6 is Alibaba's multi-mode video generation model for text, image, flash image, reference, and flash reference workflows, with optional audio input and 720p/1080p output.

    NouveauTexte vers vidéoImage vers vidéo
    à partir de $0.025/sVoir

    OpenAI

    Sora 2 Pro

    Sora 2 Pro is OpenAI’s premium video generation model with higher quality output, supporting text-to-video and image-to-video at 720p and 1080p resolutions with flexible durations of 4, 8, or 12 seconds.

    NouveauTexte vers vidéoImage vers vidéo
    à partir de $0.3/sVoir

    bytedance

    Seedance 2.0 Mini

    Seedance 2.0 Mini is an APIXO lower-cost route for ByteDance Seedance 2.0 workflows, exposing text-to-video, first-and-last-frames, and omni-reference modes with optional sound, web search, and 480p/720p output.

    NouveauTexte vers vidéoImage vers vidéo
    à partir de $0.028/sVoir

    kling

    Kling 3.0 Turbo

    Kling 3.0 Turbo is Kuaishou's fast video generation model for text-to-video and single-image image-to-video workflows with 720p/1080p output and 3-15 second clips.

    NouveauTexte vers vidéoImage vers vidéo
    à partir de $0.112/sVoir

    hailuo

    Hailuo 2.3

    Hailuo 2.3 is MiniMax's async video model with standard and pro modes for text-to-video and image-to-video generation. Standard mode supports 6s/10s at 768p, while pro mode returns fixed 5s at 1080p.

    NouveauTexte vers vidéoImage vers vidéo
    à partir de $0.056/sVoir

    hailuo

    Hailuo 2.3 Fast

    Hailuo 2.3 Fast is MiniMax's speed-optimized image-to-video model with standard and pro modes. Standard supports 6s/10s at 768p, while pro returns fixed 6s output at 1080p.

    NouveauImage vers vidéo
    à partir de $0.032/sVoir

    xai

    Grok Video

    Grok Video is xAI's async video generation model for text-to-video and image-to-video workflows, with optional continuation via task_id + index and style control.

    NouveauTexte vers vidéoImage vers vidéo
    à partir de $0.02/sVoir

    Alibaba

    Wan 2.2 Animate

    Wan 2.2 Animate API is Alibaba's character animation model that combines one source image and one motion video to generate stylized animated outputs with animate/replace behavior.

    NouveauEffets vidéo
    à partir de $0.04/sVoir

    MeiGen

    InfiniteTalk

    InfiniteTalk converts one photo plus audio into audio-driven talking or singing avatar videos with precise lip synchronization. Supports up to 10 minutes at 480p or 720p resolution.

    NouveauImage vers vidéo
    à partir de $0.03/sVoir

    kling

    Kling 3.0 Std

    Kling 3.0 Std is Kuaishou's standard-quality video generation model with text-to-video, image-to-video, and motion-control modes. It supports clips up to 15 seconds with optional sound generation and flexible aspect ratios.

    NouveauTexte vers vidéoImage vers vidéo
    à partir de $0.084/sVoir

    vidu

    Vidu Q3

    Vidu Q3 is a per-second video generation model that combines standard and Turbo text-to-video plus image-to-video workflows in one API. It supports single-image animation, first-and-last-frame transitions, optional sound and BGM, and output up to 1080p.

    NouveauTexte vers vidéoImage vers vidéo
    à partir de $0.04/sVoir

    kling

    Kling 2.5 Turbo Pro

    Kling 2.5 Turbo Pro is Kuaishou's high-speed video model for text-to-video and image-to-video creation. It supports 5-10 second clips, optional tail-frame images, aspect ratio control for text-to-video, plus negative prompts and CFG scale guidance.

    NouveauTexte vers vidéoImage vers vidéo
    à partir de $0.3/sVoir

    Lightricks

    LTX-2 19B

    LTX-2 19B is Lightricks' open-source 19B diffusion transformer for cinematic video generation. It supports text-to-video and image-to-video workflows, LoRA conditioning, and high-fidelity outputs up to 1080p in the API.

    NouveauTexte vers vidéoImage vers vidéo
    à partir de $0.012/sVoir

    kling

    Kling 2.1

    Kling 2.1 is Kuaishou's multi-tier video model with Standard, Pro, and Master modes for image-to-video and text-to-video creation. It supports 5-10 second clips, optional tail images for Pro, and aspect ratio control for Master text-to-video.

    NouveauTexte vers vidéoImage vers vidéo
    à partir de $0.2/sVoir

    kling

    Kling 2.6

    Kling 2.6 is Kuaishou's native audio-visual video model that generates video, speech, sound effects, and ambience in one pass. It supports text-to-audio-visual and image-to-audio-visual creation with Chinese and English voice generation and up to 10-second clips.

    NouveauTexte vers vidéoImage vers vidéo
    à partir de $0.3/sVoir

    Google

    Veo 3.1

    Google DeepMind’s upgraded AI video model with lite, fast, and quality routes, 4/6/8 second duration control, 720p/1080p/4k output, and multi-image reference workflows.

    Texte vers vidéoImage vers vidéo
    à partir de $0.15/sVoir

    Alibaba

    Wan 2.5

    Wan 2.5 is Alibaba's video generation model for text-to-video and image-to-video workflows, with optional audio input, 480p/720p/1080p output, 5 or 10 second clips, and prompt expansion.

    Texte vers vidéoImage vers vidéo
    à partir de $0.05/sVoir

    OpenAI

    Sora 2

    Sora 2 is OpenAI’s synchronized short-video generation model on APIXO, supporting text-to-video and single-image-to-video with realistic motion, generated audio, landscape/portrait framing, and 4-, 8-, or 12-second outputs.

    Texte vers vidéoImage vers vidéo
    à partir de $0.1/sVoir

    Black Forest Labs

    FLUX 3

    FLUX 3 is Black Forest Labs’ multimodal video model for text-to-video, image-to-video, and video extension with native audio and 720p or 1080p MP4 output.

    NouveauTexte vers vidéoImage vers vidéo
    à partir de $0.17/sVoir

    xai

    Grok Imagine Video 1.5

    Grok Imagine Video 1.5 is xAI's image-to-video model. It accepts 1–7 reference images and generates 1–15 second videos at 480p, 720p, or 1080p.

    NouveauImage vers vidéo
    à partir de $0.02/sVoir

    kling

    Kling 3.0 Omni

    Kling 3.0 Omni exposes nine modes through one endpoint: standard, pro, and 4K tiers across text-to-video, image-to-video, and reference-to-video workflows.

    NouveauTexte vers vidéoImage vers vidéo
    à partir de $0.084/sVoir

    MiniMax

    MiniMax H3 LoRA

    MiniMax H3 LoRA is the LoRA-stylized version of MiniMax H3 with 480p and 768p output. Reference images, audio, and video can condition reference-to-video generation.

    NouveauTexte vers vidéoImage vers vidéo
    à partir de $0.04/sVoir

    Alibaba

    Wan 3.0 Video

    Wan 3.0 Video supports text-to-video, image-to-video, and reference-to-video workflows with smart duration. Reference mode accepts image, video, audio, file, and link inputs.

    NouveauTexte vers vidéoImage vers vidéo
    à partir de $0.07/sVoir