Both modes use the same APIXO rate: $0.05 per generated second at 480p, $0.10 at 720p, and $0.15 at 1080p. Five-second outputs cost $0.25, $0.50, or $0.75; ten-second outputs cost $0.50, $1.00, or $1.50.
Wan 2.5 Video is APIXO’s operational route for Alibaba’s Wan 2.5 video family. Generate 5- or 10-second audiovisual clips from a text prompt or one source image, with optional custom audio, three resolution tiers, prompt expansion, negative prompting, seed-based iteration, and asynchronous delivery.
Workflows
Text / Image to Video
APIXO Price
$0.05–$0.15 / SecBackend-priced 480p, 720p, and 1080p tiers
Resolution
480p / 720p / 1080p
Duration
5 / 10 Seconds
Output
1 MP4 Video
Loading workspace...
Wan 2.5 generates video with audio rather than returning silent footage by default, reducing the need to assemble an initial soundtrack in a separate generation step.
Supply one public MP3 or WAV file as an audio reference. Wan uses it to synchronize the generated visuals and, where applicable, mouth movement with the supplied music or voiceover. Exact wording, pronunciation, speaker identity, and lip synchronization still require review.
Describe subjects, action, environments, camera behavior, dialogue, music, or sound effects in Chinese or English using a prompt of up to 1,500 characters.
Animate one source image while using an optional prompt to guide motion and scene development. The image establishes visual context without guaranteeing exact preservation of every detail.
Prompt rewriting is enabled by default to enrich shorter instructions. Disable it when preserving carefully authored wording matters more than the potential creative benefit.
Use a negative prompt to discourage unwanted traits and a numeric seed to improve repeatability. Identical settings and seeds do not guarantee identical videos.
Submit a required, non-empty prompt of up to 1,500 characters without a source image. Select 5 or 10 seconds, one of three resolutions, and a compatible aspect ratio. Optional audio, negative prompt, prompt expansion, seed, and watermark controls can refine the request.
Provide exactly one public image URL and optionally describe the intended movement or scene. Do not submit an aspect-ratio field in this mode; the source image guides the initial visual state and composition but does not guarantee exact identity, text, logos, geometry, background, or framing.
Current production options for the wan-2-5-video operational route.
1–1,500
Prompt Characters
1 Public URL
Source Image
1 MP3 / WAV
Audio Reference
500 Characters
Negative Prompt
0–2,147,483,647
Seed Range
Polling / Callback
Delivery
Turn a campaign brief into short product reveals, lifestyle scenes, or launch concepts with an initial soundtrack. Review product proportions, object contact, hands, logos, and embedded text before moving an idea into production.
Add motion to product photography, illustrations, posters, or character art using a single source image. Check identity, composition, background, and fine-detail fidelity against the original asset after generation.
Pair a short music excerpt, narration draft, or sound-design reference with generated visuals for promotional and social concepts. Validate spoken wording, pronunciation, mouth movement, and audiovisual timing independently.
Explore camera direction, scene composition, motion, and training or explainer concepts through short drafts. Use lower-resolution, 5-second generations for iteration before committing budget to longer, higher-resolution versions.
APIXO exposes Alibaba’s Wan 2.5 video family under one route but does not publicly document a mode-by-mode mapping to wan2.5-t2v-preview and wan2.5-i2v-preview.
Image-to-video accepts one public JPEG, JPG, non-alpha PNG, BMP, or WebP URL no larger than 20 MB; each dimension must be 240–8,000 pixels.
Text-to-video supports 16:9, 9:16, and 1:1 at every resolution; 4:3 and 3:4 require 720p or 1080p.
Optional audio must use one public MP3 or WAV URL, no larger than 15 MB, with a documented duration range of 3–30 seconds.
Typical processing ranges are 40–120 seconds at 480p, 60–180 seconds at 720p, and 90–250 seconds at 1080p. APIXO recommends the first poll after 40, 60, or 90 seconds respectively, followed by 10-second polling intervals; these timings are estimates rather than an SLA.
Result URLs are temporary, with no fixed public retention period. Download required outputs promptly and review them for safety and source-material permissions.
Both modes use the same APIXO rate: $0.05 per generated second at 480p, $0.10 at 720p, and $0.15 at 1080p. Five-second outputs cost $0.25, $0.50, or $0.75; ten-second outputs cost $0.50, $1.00, or $1.50.
Audio longer than the selected video duration is truncated to the first 5 or 10 seconds. If it is shorter, the remaining video segment is silent. APIXO publishes no additional charge for custom audio, image input, or prompt expansion.
Upstream Wan 2.5 can use prompt-described voice, sound effects, and background music when generating audio. However, APIXO provides no guarantees for exact dialogue, pronunciation, voice identity, speaker attribution, or lip synchronization, and exposes no voice-cloning controls.
Tasks run asynchronously with pending, processing, success, or failed states. Polling returns the final MP4 URL inside resultJson.resultUrls; callback mode sends the terminal payload to a public HTTPS callback URL. Failures include machine-readable code and message fields.
APIXO does not expose arbitrary duration, custom dimensions, configurable FPS, output count, CFG scale, motion strength, camera presets, shot count, reference weights, multi-shot mode, editing, continuation, or a silent-output switch. Enabling `watermark` adds the fixed text `AI 生成` at the bottom right. Alibaba documents its Wan 2.5 Preview outputs as 30 fps H.264 MP4, while APIXO exposes no codec or frame-rate controls.
Discover more AI models for your next creative workflow