It supports text-to-video, image-to-video, and reference-to-video. All 3 modes offer 480p / 720p / 1080p output.
Wan 3.0 Video LoRA
Wan 3.0 Video LoRA generates video from text, from first and last frames, or from a mix of image, video, and audio references. Choose 480p, 720p, 1080p output, set an explicit clip length, or let smart duration decide.
Modes
Text / Image / Reference
APIXO Price
From $0.0525 / second
Duration
Smart or 2–30 Seconds
Reference Types
Image / Video / Audio
Resolution
480p / 720p / 1080p
Create with Wan 3.0 Video LoRA
Loading workspace...
A reference toolkit for LoRA-guided creative briefs
Three Reference Channels
Guide a reference-to-video request with images, videos, or audio instead of compressing every instruction into the prompt alone.
First-and-Last-Frame Images
Image-to-video accepts up to 2 images — the starting frame, plus an optional ending frame. Reference mode accepts up to 10 images.
Video Reference Guidance
Attach up to 5 reference videos of 1–15 seconds each, with at most 15 seconds of reference video in total. Their actual seconds are billed at the selected resolution rate.
Audio References Without a Surcharge
Add up to 5 audio references of 1–15 seconds each in reference mode. Audio references can shape the brief without adding a separate audio-input fee.
Prompt Optional Outside Text Mode
A prompt is required only for text-to-video. Image-to-video and reference-to-video accept an empty prompt when the supplied media already carries the brief.
Smart Duration
Turn on smart duration when the model should determine the output length, or choose any whole number from 2 through 30 seconds for an explicit clip duration.
Published Wan 3.0 Video LoRA API parameters
3 Generation Modes
Modes
480p / 720p / 1080p
Resolution
Smart or 2–30 Seconds
Duration
Up to 10 Images
Image References
Up to 5 Videos
Video References
Enabled by Default
Sound
When richer reference context improves the production brief
Reference-Guided Scene Development
Combine image references with short video clips to communicate subjects, visual direction, and motion cues within one reference-to-video request.
First-to-Last-Frame Motion
Provide a starting image and an optional ending image when the desired clip needs to move between two known visual states.
Audio-Led Reference Context
Extend a visual brief with audio references when pacing, mood, or a spoken track should inform the generated clip.
Smart-Length Creative Exploration
Use smart duration when the appropriate clip length should follow the generated result, then switch to a fixed 2–30 second duration for controlled production passes.
Wan 3.0 Video LoRA input and billing guidance
Notes
A prompt is required only for text-to-video. Image-to-video and reference-to-video accept an empty prompt.
Smart duration uses -1. The backend initially reserves the cost of 30 seconds, then settles against the successful video’s actual duration and refunds the difference.
Reference-video seconds are billed at the same resolution rate as output seconds. Reference audio has no separate input charge.
For a fixed positive duration, output seconds plus total reference-video seconds cannot exceed 30 seconds.
Sound is enabled by default and does not change the price. Set the sound parameter to false when a request should generate a silent clip.
Reference-to-video accepts up to 10 images, 5 videos, and 5 audio files. Each reference video or audio file must be 1–15 seconds, with at most 15 seconds per media type.
Frequently Asked Questions
Output seconds are charged from $0.0525 to $0.21 per second, depending only on the output resolution — the rate table lists every tier. All 3 generation modes share the same rate; reference-to-video additionally bills the actual seconds of any reference videos.
When duration is -1, the backend reserves 30 seconds of cost because the final duration is not known yet. After a successful result, billing is settled using the actual duration and the difference is refunded.
Reference-to-video accepts images, videos, and audio. It supports up to 10 images, 5 videos, and 5 audio references, with each reference video or audio file running 1–15 seconds.
No. Each reference-video second is billed at the selected resolution rate, while reference audio does not add a separate input fee.
Both share the same parameters and modes. This model outputs 480p, 720p, 1080p, while the Pro model raises the ceiling to 1080p, 2K, 4K at its own per-second rates.
Explore Other Models
Discover more AI models for your next creative workflow
