APIXO
Video GenerationMultimodal ReferencesLower-Cost Route

Seedance 2.0 Mini is APIXO’s lower-cost route for ByteDance Seedance 2.0, not a separately documented upstream ByteDance model. It exposes prompt generation, first-and-last-frame animation, and mixed-media reference workflows with optional generated sound and web search, 480p or 720p output, and clips documented across a 4–15 second range.

APIXO Modes

3 Workflows

APIXO Price

$0.012–$0.041 / Sec

Resolution

480p / 720p

Duration

4–15 Seconds

Seedance 2.0 Launch

February 12, 2026

Create with Seedance 2.0 Mini

Routed by APIXO

Access focused Seedance 2.0 workflows with Mini-specific pricing

ByteDance officially presents Seedance 2.0 as one unified multimodal audio-video model; it does not publish a separate Mini architecture or model ID. APIXO’s Mini name identifies this route’s verified resolution ceiling, workflow controls, reference limits, and lower per-second rates rather than an independently documented upstream checkpoint.

Capabilities

Guide short videos with prompts, keyframes, and mixed media

Multimodal Conditioning

Combine written direction with images, videos, and audio so the output can draw from visual composition, subjects, motion, camera treatment, effects, or sound-related context.

Keyframe-Constrained Motion

Animate one opening image or generate a transition between ordered first and last frames, giving the sequence clearly defined starting and ending frames rather than general reference guidance.

Reference-Led Continuity

Carry recognizable subjects, environments, styling, and movement cues into a new clip. Results remain generative and may diverge when inputs contain conflicting direction.

Joint Audiovisual Output

Use Seedance 2.0’s audio-video generation capability to create sound alongside the visuals. APIXO also exposes silent generation through its sound control.

Complex Motion Interpretation

Follow prompts involving physical action, camera movement, interacting subjects, lighting, and scene changes with the underlying Seedance 2.0 model’s improved controllability.

Web-Informed Context

Use one to three audio references alongside image or video context to guide rhythm, vocal texture, ambience, or event timing. APIXO does not document unchanged soundtrack retention, so uploaded audio should be treated as generation context.

APIXO Configuration

Mini route production limits

Verified output and reference options available through APIXO.

3

Generation Modes

4–15 Seconds

Duration Range

6 Options + Auto

Aspect Ratios

Up to 9

Image References

Up to 3

Video References

Up to 3

Audio References

Production Uses

Develop short-form video with selective reference control

Concept Development

Explore Prompt-Led Scenes

Turn written briefs into compact horizontal, vertical, square, or ultrawide video concepts with optional sound. The Mini route suits storyboarding, campaign ideation, social concepts, and early visual tests where 480p or 720p is sufficient.

Keyframe Animation

Create Videos from First and Last Frames

Animate an approved opening frame or bridge supplied start and end images. Use the workflow for product reveals, storyboard transitions, title treatments, character moments, and design-led motion where the first and last frames need explicit control.

Campaign Creation

Combine Brand References

Mix product imagery, environments, motion footage, and optional audio cues to direct an advertising concept. Review logos, packaging, identity, materials, object geometry, and conflicts among references before production use.

Motion Direction

Transfer Pacing and Camera Ideas

Supply footage and audio as conditioning for action rhythm, camera language, performance timing, or effects in a new generation. APIXO does not document direct source-video editing or preservation of the reference video’s soundtrack.

Reference Planning

Match each source asset to the correct workflow

APIXO validates different media structures for prompt, keyframe, and omni-reference generation.

Choose Prompt or Keyframe Control

Use text-to-video when the prompt defines the complete scene. Choose first-and-last-frames when one image must establish the opening or two ordered images must anchor both ends of the motion.

Assemble Omni References

Combine up to nine images, three videos, and three audio files. Each video or audio file must run for 2–15 seconds, with cumulative duration limited to 15 seconds per media type.

Configure and Retrieve Output

Select resolution, duration, framing, generated sound, and optional generated sound and web search. Submit the task asynchronously, then obtain the resulting MP4 URL through status polling or webhook delivery.

Notes & FAQ

Route identity, duration, and billing boundaries

Implementation Notes

01

ByteDance documents Seedance 2.0, but no separate upstream Seedance 2.0 Mini architecture, checkpoint, or model ID is publicly identified.

02

APIXO’s visible creation selector lists 4, 5, 10, and 15-second options while describing the supported output range as 4–15 seconds.

03

First-and-last-frames accepts one starting image or two ordered images representing the starting and ending frames.

04

Audio-only omni-reference requests are not supported; uploaded audio must accompany at least one image or video reference.

05

Generated sound can be disabled; uploaded audio provides generation context and is not documented as a soundtrack that will be copied unchanged.

06

APIXO recommends waiting approximately 180–300 seconds before the first status poll; actual completion time can vary.

Frequently Asked Questions

No separate Mini model is identified in ByteDance’s public Seedance materials. Mini is an APIXO route label with documented 480p and 720p output, specific pricing, and focused access to three Seedance 2.0 workflows.

Text-to-video, first-and-last-frames, and omni-reference without video input cost $0.019 per output second at 480p or $0.041 per output second at 720p. Image and audio references do not add billable seconds.

The rate is $0.012 per second at 480p or $0.025 at 720p. APIXO applies it to the generated duration plus the combined duration of all reference videos.

APIXO does not document source-audio preservation. Uploaded video and audio provide multimodal conditioning, while the separate sound field controls whether the output contains model-generated audio.

Review anatomy, hands, identity, object contact, rapid motion, embedded text, reference fidelity, transitions, audio interpretation, pronunciation, and lip synchronization. Multi-character scenes or conflicting references may require simplification and regeneration.

Explore Other Models

Discover more AI models for your next creative workflow