APIXO
Video with AudioText / Image Input5 / 10 Seconds

Wan 2.5 Video is APIXO’s operational route for Alibaba’s Wan 2.5 video family. Generate 5- or 10-second audiovisual clips from a text prompt or one source image, with optional custom audio, three resolution tiers, prompt expansion, negative prompting, seed-based iteration, and asynchronous delivery.

Workflows

Text / Image to Video

APIXO Price

$0.05–$0.15 / SecBackend-priced 480p, 720p, and 1080p tiers

Resolution

480p / 720p / 1080p

Duration

5 / 10 Seconds

Output

1 MP4 Video

Create with Wan 2.5

Loading workspace...

Alibaba Wan on APIXO

Generate short video and its soundtrack through one operational route

Wan 2.5 combines visual generation with automatic audio, including content-matched music and sound effects. APIXO also accepts one optional MP3 or WAV reference for custom audiovisual synchronization, supporting both prompt-first creation and animation from a single source image.

Capabilities

Direct scenes, movement, framing, and sound

Integrated Audiovisual Output

Wan 2.5 generates video with audio rather than returning silent footage by default, reducing the need to assemble an initial soundtrack in a separate generation step.

Custom Audio Synchronization

Supply one public MP3 or WAV file as an audio reference. Wan uses it to synchronize the generated visuals and, where applicable, mouth movement with the supplied music or voiceover. Exact wording, pronunciation, speaker identity, and lip synchronization still require review.

Prompt-Guided Scenes

Describe subjects, action, environments, camera behavior, dialogue, music, or sound effects in Chinese or English using a prompt of up to 1,500 characters.

Single-Image Animation

Animate one source image while using an optional prompt to guide motion and scene development. The image establishes visual context without guaranteeing exact preservation of every detail.

Automatic Prompt Expansion

Prompt rewriting is enabled by default to enrich shorter instructions. Disable it when preserving carefully authored wording matters more than the potential creative benefit.

Controlled Iteration

Use a negative prompt to discourage unwanted traits and a numeric seed to improve repeatability. Identical settings and seeds do not guarantee identical videos.

Generation modes

Start with a written scene or one visual anchor

Text-to-Video

Submit a required, non-empty prompt of up to 1,500 characters without a source image. Select 5 or 10 seconds, one of three resolutions, and a compatible aspect ratio. Optional audio, negative prompt, prompt expansion, seed, and watermark controls can refine the request.

Prompt → Video

Image-to-Video

Provide exactly one public image URL and optionally describe the intended movement or scene. Do not submit an aspect-ratio field in this mode; the source image guides the initial visual state and composition but does not guarantee exact identity, text, logos, geometry, background, or framing.

Image → Motion
Route specifications

Inputs and controls exposed by APIXO

Current production options for the wan-2-5-video operational route.

1–1,500

Prompt Characters

1 Public URL

Source Image

1 MP3 / WAV

Audio Reference

500 Characters

Negative Prompt

0–2,147,483,647

Seed Range

Polling / Callback

Delivery

Production uses

Develop short audiovisual concepts for review

Campaign ideation

Prototype Promotional Scenes

Turn a campaign brief into short product reveals, lifestyle scenes, or launch concepts with an initial soundtrack. Review product proportions, object contact, hands, logos, and embedded text before moving an idea into production.

Image activation

Animate Key Visuals

Add motion to product photography, illustrations, posters, or character art using a single source image. Check identity, composition, background, and fine-detail fidelity against the original asset after generation.

Audio-led creative

Build Around Supplied Sound

Pair a short music excerpt, narration draft, or sound-design reference with generated visuals for promotional and social concepts. Validate spoken wording, pronunciation, mouth movement, and audiovisual timing independently.

Previsualization

Test Shots and Explainers

Explore camera direction, scene composition, motion, and training or explainer concepts through short drafts. Use lower-resolution, 5-second generations for iteration before committing budget to longer, higher-resolution versions.

Notes and FAQ

Operational rules and review boundaries

Before you generate

01

APIXO exposes Alibaba’s Wan 2.5 video family under one route but does not publicly document a mode-by-mode mapping to wan2.5-t2v-preview and wan2.5-i2v-preview.

02

Image-to-video accepts one public JPEG, JPG, non-alpha PNG, BMP, or WebP URL no larger than 20 MB; each dimension must be 240–8,000 pixels.

03

Text-to-video supports 16:9, 9:16, and 1:1 at every resolution; 4:3 and 3:4 require 720p or 1080p.

04

Optional audio must use one public MP3 or WAV URL, no larger than 15 MB, with a documented duration range of 3–30 seconds.

05

Typical processing ranges are 40–120 seconds at 480p, 60–180 seconds at 720p, and 90–250 seconds at 1080p. APIXO recommends the first poll after 40, 60, or 90 seconds respectively, followed by 10-second polling intervals; these timings are estimates rather than an SLA.

06

Result URLs are temporary, with no fixed public retention period. Download required outputs promptly and review them for safety and source-material permissions.

Frequently asked questions

Both modes use the same APIXO rate: $0.05 per generated second at 480p, $0.10 at 720p, and $0.15 at 1080p. Five-second outputs cost $0.25, $0.50, or $0.75; ten-second outputs cost $0.50, $1.00, or $1.50.

Audio longer than the selected video duration is truncated to the first 5 or 10 seconds. If it is shorter, the remaining video segment is silent. APIXO publishes no additional charge for custom audio, image input, or prompt expansion.

Upstream Wan 2.5 can use prompt-described voice, sound effects, and background music when generating audio. However, APIXO provides no guarantees for exact dialogue, pronunciation, voice identity, speaker attribution, or lip synchronization, and exposes no voice-cloning controls.

Tasks run asynchronously with pending, processing, success, or failed states. Polling returns the final MP4 URL inside resultJson.resultUrls; callback mode sends the terminal payload to a public HTTPS callback URL. Failures include machine-readable code and message fields.

APIXO does not expose arbitrary duration, custom dimensions, configurable FPS, output count, CFG scale, motion strength, camera presets, shot count, reference weights, multi-shot mode, editing, continuation, or a silent-output switch. Enabling `watermark` adds the fixed text `AI 生成` at the bottom right. Alibaba documents its Wan 2.5 Preview outputs as 30 fps H.264 MP4, while APIXO exposes no codec or frame-rate controls.

Explore Other Models

Discover more AI models for your next creative workflow