Text-to-video and image-to-video cost $0.10 per second at 720p or $0.15 at 1080p. Reference-to-video costs $0.20 per second at 720p or $0.30 at 1080p. Total cost is the selected duration multiplied by the applicable rate.
Wan 2.6 Video is APIXO’s operational route for five Alibaba video-generation workflows. Build audiovisual clips from text, animate one source image, or guide characters with ordered image and video references. Standard and Flash modes provide 720p or 1080p output, multi-shot control, custom audio, and mode-specific durations and pricing.
Workflows
Text / Image / References
APIXO Price
$0.025–$0.30 / SecBackend-priced Standard and Flash routes
Resolution
720p / 1080p
Duration
2–15s · References ≤10s
Reference Limit
Up to 5 Files
Loading workspace...
Create a video without source imagery from a required Chinese or English prompt describing subjects, action, setting, camera behavior, dialogue, and sound.
Standard and Flash image modes animate exactly one source image. The prompt remains required and directs movement, camera behavior, and scene development from the initial visual state.
Reference modes accept ordered image and video URLs that map to character1, character2, and later roles for appearance, motion, scene, or available voice guidance.
Choose single for one continuous shot or multi for a sequence of changing shots. The explicit setting takes priority over conflicting shot instructions in the prompt.
Submit one optional MP3 or WAV file for custom synchronization, or let sound-enabled generation produce audiovisual output without claiming exact dialogue, pronunciation, or lip synchronization.
Refine output with a negative prompt, prompt expansion, and a numeric seed. Fixed seeds aid controlled iteration but cannot make probabilistic generation deterministic.
Standard routing covers text-to-video, image-to-video, and reference-to-video. Text and image generation support 2–15 seconds, while reference generation supports 2–10 seconds. The sound field is not valid in these modes.
Flash routing is available only for image-to-video-flash and reference-to-video-flash. Its sound control defaults to true; disabling it produces silent output, ignores supplied audio, and applies the lower Flash rate.
Mode-specific inputs and limits exposed through APIXO.
5
Generation Modes
1–1,500
Prompt Characters
1 Image
Image Mode Input
1–5 Files
Reference Input
1 MP3 / WAV
Custom Audio
Polling / Callback
Result Delivery
Use text-to-video and multi-shot control to explore short campaign narratives, camera changes, and scene progression before committing to filming or detailed post-production. Review anatomy, object contact, logos, and embedded text.
Turn one approved product image, illustration, or campaign visual into motion with either Standard or Flash routing. Compare the result against the source for geometry, proportions, composition, and brand fidelity.
Assign ordered image or video references to character roles for short conversations and multi-character scenes. Validate identity, positioning, voice drift, dialogue attribution, pronunciation, and mouth movement in every result.
Supply narration, music, or sound design to guide a short promotional clip. Check timing and synchronization before publication, or disable Flash sound when only a low-cost visual draft is required.
APIXO exposes five product modes but does not publicly identify the precise upstream checkpoint or snapshot selected behind every route.
Every mode requires a non-empty prompt of up to 1,500 characters; the optional negative prompt supports up to 500 characters.
Image modes require one public JPEG, JPG, non-alpha PNG, BMP, or WebP URL, up to 20 MB, with dimensions of 240–8,000 pixels.
Reference modes accept one to five ordered files: up to five images, up to three MP4 or MOV videos, and no more than five files in total.
Custom audio accepts one public MP3 or WAV URL, 3–30 seconds and up to 15 MB; excess audio is truncated, while a short input leaves the remainder silent.
APIXO does not publish a result-URL retention period for this route. Poll text or image modes after 30 seconds and reference modes after 45 seconds, or use callback delivery for terminal results; store required outputs promptly after completion.
Text-to-video and image-to-video cost $0.10 per second at 720p or $0.15 at 1080p. Reference-to-video costs $0.20 per second at 720p or $0.30 at 1080p. Total cost is the selected duration multiplied by the applicable rate.
Image Flash costs $0.025/$0.0375 per second at 720p/1080p without sound and $0.05/$0.075 with sound. Reference Flash costs $0.05/$0.08 without sound and $0.10/$0.16 with sound.
Only image-to-video-flash and reference-to-video-flash accept sound. It defaults to true. When false, the generated video is silent and any supplied audio_urls value is ignored.
Set shot_type to multi and keep prompt_extend enabled so Wan 2.6 can rewrite and optimize the shot description. Use single when the scene should remain one continuous shot.
Tasks use pending, processing, success, and failed states. Successful video URLs appear in resultJson.resultUrls; failures expose failCode and failMsg. Callback delivery requires a reachable public callback endpoint.
Discover more AI models for your next creative workflow