APIXO
Video GenerationMultimodal InputAPI Available

Seedance 2.0 is ByteDance’s multimodal audio-video generation model, available on APIXO for prompt-only creation, keyframe animation, and reference-guided workflows. Combine text with image, video, and audio references to direct motion, composition, camera language, sound, and multi-shot storytelling, with APIXO output options ranging from 480p to 4K.

Inputs

Text / Image / Video / Audio

APIXO Price

$0.0573–$1.04/ Sec

Resolution

480p / 720p / 1080p / 4K

Duration

4–15 Seconds

Released

Feb 12, 2026

Create with Seedance 2.0

Loading workspace...

Available on APIXO

Direct Multimodal Audio-Video Generation with Seedance 2.0 API

Build short-form video workflows around prompts, keyframes, or mixed reference media. Seedance 2.0 can draw composition, motion, camera language, visual effects, and sound characteristics from supplied assets, while APIXO exposes selectable duration, aspect ratio, audio, and resolution controls through one asynchronous generation endpoint.

Capabilities

Direct Motion, Story, and Sound with Multimodal References

Mixed-Media Direction

Omni-reference generation can combine a prompt with as many as nine images, three video clips, and three audio clips. Use different assets to guide subject appearance, composition, movement, camera behavior, visual effects, and sound characteristics within one request.

Complex Motion Modeling

Seedance 2.0 improves motion stability and physical plausibility for interactions involving multiple subjects or demanding actions. This makes it useful for sports, performance, and cinematic movement where timing and contact between people or objects matter.

Multi-Shot Storytelling

Generate video sequences with multiple shots while directing character interactions, action progression, camera choices, and narrative pacing through natural-language instructions. Subject and story consistency are improved, though intricate multi-subject scenes may still benefit from iterative refinement.

Unified Audio-Video Output

Seedance 2.0 jointly generates visuals and audio, coordinating dialogue, effects, music, and on-screen action. APIXO includes a sound control, so applications can request an audio-video result or disable generated sound when a silent production asset is required.

Keyframe-Guided Animation

APIXO’s first-and-last-frames mode accepts one starting image or a start-and-end pair. This provides stronger control over visual endpoints for transitions, storyboard animation, product movement, and other shots that need a defined opening or destination frame.

Prompt-Driven Camera Planning

Describe shot design, lighting, performance, and camera movement in the prompt. Seedance 2.0 can plan camera language and visual presentation from those instructions, helping creators move beyond static image animation toward more deliberately staged sequences.

API Details

Seedance 2.0 Options on APIXO

Verified modes, reference limits, framing controls, and delivery behavior exposed by the APIXO endpoint.

Text-to-Video / First-and-Last-Frames / Omni-Reference

Modes

Up to 9 Images

Image References

Up to 3 Video + 3 Audio

Media References

Auto + 6 Fixed Ratios

Aspect Ratios

MP4 URL Array

Output

Polling or Webhook

Result Delivery

Use Cases

Production Workflows for Seedance 2.0

Product Marketing

Reference-Guided Product Reveals

Combine product images with a motion reference and written art direction to create short launch visuals. Guide framing, camera movement, lighting, pacing, and sound while preserving the product’s recognizable appearance across a polished promotional sequence.

Previsualization

Keyframe Animation and Storyboards

Supply a starting frame or defined first-and-last-frame pair to explore transitions between planned compositions. Film, animation, and design teams can use this workflow for storyboard motion tests, sequence planning, camera experiments, and rapid previsualization before full production.

Narrative Production

Multi-Shot Character Scenes

Translate longer scene directions into short, connected shots with coordinated performance, camera language, dialogue, and effects. The model’s improved instruction following and motion stability suit character-driven concepts, action beats, and cinematic scene exploration.

Music and Performance

Audio-Directed Visual Sequences

Use audio references alongside images or video to guide rhythm and sound characteristics, then generate synchronized visuals and audio. This supports performance concepts, music-led edits, branded audiovisual loops, and social clips that need motion shaped around an existing sonic direction.

Quick Start

Move from Seedance 2.0 Testing to Integration

Validate prompts and reference combinations before connecting the asynchronous API to a production workflow.

Test a Mode in the Playground

Open Seedance 2.0 in the APIXO Playground and choose text-to-video, first-and-last-frames, or omni-reference. Add eligible reference assets, write the creative direction, and compare duration, framing, sound, and resolution settings against your intended use case.

Submit an Asynchronous Generation Task

Create an API key and send the prompt, mode, output settings, and any public reference URLs to the Seedance 2.0 generation endpoint. Store the returned task identifier so your application can track the long-running generation job.

Retrieve and Store the Video

Poll the status endpoint or configure an HTTPS callback for final delivery. When processing succeeds, parse the returned result data, retrieve the MP4 URL, and copy important output to persistent storage because APIXO result URLs are temporary.

Notes & FAQ

Seedance 2.0 Integration Notes

Operational Notes

01

First-and-last-frames mode requires one or two reference images representing the opening frame or both endpoints.

02

Omni-reference accepts up to nine images and up to three video and three audio references.

03

Each video or audio reference must be 2–15 seconds; each media type has a 15-second combined-duration limit.

04

Audio-only omni-reference requests are rejected; include at least one image or video whenever audio references are supplied.

05

APIXO accepts any integer output duration from 4 through 15 seconds, while the Create interface presents quick presets.

06

Complex multi-subject consistency, embedded text, intricate editing effects, and occasional audio distortion may require additional generations or prompt refinement.

Frequently Asked Questions

APIXO exposes text-to-video, first-and-last-frames, and omni-reference through the Seedance 2.0 endpoint. Text-to-video needs no reference media, keyframe mode requires one or two images, and omni-reference supports mixed image, video, and audio guidance.

APIXO supports 480p, 720p, 1080p, and 4K output. Available aspect ratios are auto, 16:9, 4:3, 1:1, 3:4, 9:16, and 21:9, covering landscape, square, portrait, and ultrawide delivery formats.

Billing is per second, with the unit rate determined by resolution and whether omni-reference includes video references. Text-to-video and first-and-last-frames use the no-video-reference rate. With a video reference, billed seconds include both reference-video duration and generated output duration.

Video generation is asynchronous and typically takes about 4–8 minutes, with omni-reference jobs sometimes taking longer. Actual time varies with resolution, duration, prompt complexity, reference accessibility, routing, and queue conditions, so use callbacks or measured polling intervals in production.

APIXO exposes an optional web-search control for Seedance 2.0 generation requests. Enable it only when the video concept needs current information; it is separate from the reference-media workflow and does not remove the prompt or mode requirements.

Explore Other Models

Discover more AI models for your next creative workflow