From $0.0525 / SecondThree Reference Input TypesSmart Duration

Wan 3.0 Video LoRA generates video from text, from first and last frames, or from a mix of image, video, and audio references. Choose 480p, 720p, 1080p output, set an explicit clip length, or let smart duration decide.

Modes

Text / Image / Reference

APIXO Price

From $0.0525 / second

Duration

Smart or 2–30 Seconds

Reference Types

Image / Video / Audio

Resolution

480p / 720p / 1080p

Create with Wan 3.0 Video LoRA

Loading workspace...

LoRA-Tuned Video Generation

Direct video with images, clips, audio, and smart duration

Wan 3.0 Video LoRA brings three reference input types into one video API. Start from text or from first and last frames, or combine reference media in reference-to-video mode while choosing output resolution, framing, sound, and duration behavior.

Capabilities

A reference toolkit for LoRA-guided creative briefs

Three Reference Channels

Guide a reference-to-video request with images, videos, or audio instead of compressing every instruction into the prompt alone.

First-and-Last-Frame Images

Image-to-video accepts up to 2 images — the starting frame, plus an optional ending frame. Reference mode accepts up to 10 images.

Video Reference Guidance

Attach up to 5 reference videos of 1–15 seconds each, with at most 15 seconds of reference video in total. Their actual seconds are billed at the selected resolution rate.

Audio References Without a Surcharge

Add up to 5 audio references of 1–15 seconds each in reference mode. Audio references can shape the brief without adding a separate audio-input fee.

Prompt Optional Outside Text Mode

A prompt is required only for text-to-video. Image-to-video and reference-to-video accept an empty prompt when the supplied media already carries the brief.

Smart Duration

Turn on smart duration when the model should determine the output length, or choose any whole number from 2 through 30 seconds for an explicit clip duration.

Model Details

Published Wan 3.0 Video LoRA API parameters

3 Generation Modes

Modes

480p / 720p / 1080p

Resolution

Smart or 2–30 Seconds

Duration

Up to 10 Images

Image References

Up to 5 Videos

Video References

Enabled by Default

Sound

Use Cases

When richer reference context improves the production brief

Visual Continuity

Reference-Guided Scene Development

Combine image references with short video clips to communicate subjects, visual direction, and motion cues within one reference-to-video request.

Timed Transitions

First-to-Last-Frame Motion

Provide a starting image and an optional ending image when the desired clip needs to move between two known visual states.

Multimedia Briefs

Audio-Led Reference Context

Extend a visual brief with audio references when pacing, mood, or a spoken track should inform the generated clip.

Adaptive Timing

Smart-Length Creative Exploration

Use smart duration when the appropriate clip length should follow the generated result, then switch to a fixed 2–30 second duration for controlled production passes.

Notes & FAQ

Wan 3.0 Video LoRA input and billing guidance

Notes

01

A prompt is required only for text-to-video. Image-to-video and reference-to-video accept an empty prompt.

02

Smart duration uses -1. The backend initially reserves the cost of 30 seconds, then settles against the successful video’s actual duration and refunds the difference.

03

Reference-video seconds are billed at the same resolution rate as output seconds. Reference audio has no separate input charge.

04

For a fixed positive duration, output seconds plus total reference-video seconds cannot exceed 30 seconds.

05

Sound is enabled by default and does not change the price. Set the sound parameter to false when a request should generate a silent clip.

06

Reference-to-video accepts up to 10 images, 5 videos, and 5 audio files. Each reference video or audio file must be 1–15 seconds, with at most 15 seconds per media type.

Frequently Asked Questions

It supports text-to-video, image-to-video, and reference-to-video. All 3 modes offer 480p / 720p / 1080p output.

Output seconds are charged from $0.0525 to $0.21 per second, depending only on the output resolution — the rate table lists every tier. All 3 generation modes share the same rate; reference-to-video additionally bills the actual seconds of any reference videos.

When duration is -1, the backend reserves 30 seconds of cost because the final duration is not known yet. After a successful result, billing is settled using the actual duration and the difference is refunded.

Reference-to-video accepts images, videos, and audio. It supports up to 10 images, 5 videos, and 5 audio references, with each reference video or audio file running 1–15 seconds.

No. Each reference-video second is billed at the selected resolution rate, while reference audio does not add a separate input fee.

Both share the same parameters and modes. This model outputs 480p, 720p, 1080p, while the Pro model raises the ceiling to 1080p, 2K, 4K at its own per-second rates.

Explore Other Models

Discover more AI models for your next creative workflow