Yes. The model is live on the APIXO asynchronous task endpoint and can be called with the same API key as every other model in the catalog. Generation is also available in the Playground on this page.
MiniMax H3 is a general-purpose multimodal video model that combines text, image, video, and audio context in one generation workflow. It creates videos up to 2K and 15 seconds with native stereo sound, while supporting first-and-last-frame control, mixed-media references, motion transfer, multi-shot storytelling, and instruction-guided video editing.
Inputs
Text / Image / Video / Audio
APIXO Price
$0.09–$0.13 / second
Resolution
768P / 2K
Duration
5-15 Seconds
Audio
Native Stereo Sound
Loading workspace...
Combine written direction with image, video, and audio references in one request. Assign each source a role such as appearance, motion, camera, voice, rhythm, environment, or style.
Define an opening frame, an ending frame, or both to guide transformations, transitions, product reveals, and storyboard beats while preserving a clearer visual trajectory.
Use reference images, videos, and audio together to communicate character identity, product cues, movement, camera language, vocal performance, music, or ambience.
Generate voice, sound effects, ambience, and music together with the video as native stereo audio rather than assembling isolated audio tracks after generation.
Build short sequences with coordinated shot changes, scene progression, action continuity, and audiovisual pacing from one prompt-led or reference-guided workflow.
Transform existing visual material through natural-language direction or transfer motion from video references while reviewing identity, composition, occlusion, and source-detail fidelity.
Combine product imagery, written direction, motion references, and sound intent to develop campaign concepts. Review labels, logos, materials, dimensions, and product geometry before approving generated commercial assets.
Establish a recurring subject with visual references, then guide movement, camera behavior, voice, or scene context through additional media. Evaluate identity continuity, anatomy, expression, and performance timing across the complete result.
Translate a treatment into short scenes that combine action, camera language, ambience, and audiovisual pacing. Use generated footage to evaluate ideas before committing to physical production or detailed post-production.
Explore style, environment, subject, or scene changes through natural-language instructions and optional references. Inspect temporal boundaries, occlusion, unintended modifications, and source-detail preservation frame by frame.
Confirmed against the live APIXO deployment.
Text / Image / Video / Audio
Input
T2V / I2V / Reference
Workflows
768P / 2K
Resolution
5–15 Seconds
Duration
Native Stereo
Audio
6 Ratios + Adaptive
Framing
MiniMax H3 runs on the APIXO asynchronous task endpoint. Submit to generateTask and poll statusTask, or supply a callback URL for webhook delivery.
Three modes are exposed: text-to-video, image-to-video, and reference-to-video. Image-to-video takes 1-2 images as the first and optional last frame; reference-to-video needs at least one image or video reference.
Output is 768p or 2K at 5-15 seconds, whole seconds only. Resolution sets the per-second rate. Prompts are limited to 4,000 characters. Reference video and audio are each capped at 3 files, 2-15 seconds per file and 15 seconds in total.
“MiniMax H3” is the APIXO model ID. Do not carry over parameter values, quality tiers, or billing behavior from Hailuo 2.3 or Hailuo 2.3 Fast — they are separate routes with different contracts.
References guide generation but do not guarantee exact identity, typography, logos, materials, or product geometry.
Review anatomy, hands, object contact, rapid motion, occlusion, camera transitions, speech, and audiovisual synchronization before release.
Yes. The model is live on the APIXO asynchronous task endpoint and can be called with the same API key as every other model in the catalog. Generation is also available in the Playground on this page.
Text-to-video, image-to-video, and reference-to-video. Reference-to-video accepts up to 9 images, 3 videos, and 3 audio files as combined creative context; audio cannot be supplied on its own.
768p or 2K output at 5 to 15 seconds, in whole seconds. 2K is the default and the highest tier the route exposes; 768p bills at a lower per-second rate. Aspect ratio can be set to 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16 for text-to-video and reference-to-video; image-to-video takes its framing from the supplied first frame.
Output video is billed per second in every mode, at the rate that matches your chosen resolution — 768p is cheaper per second than 2K. In reference-to-video, the actual seconds of your reference videos are billed at the same per-second rate, and the first 5 reference images are free with each additional image charged separately. Reference audio is not billed. See the pricing tab for the current rates.
Yes. Audio is produced in the same pass as the picture as native stereo, with speech, sound effects, ambience, and music modelled jointly rather than as separate tracks.
No. Mixed-media references communicate creative intent but do not guarantee exact identity, motion, voice, product structure, or frame-level reproduction. Production use requires visual and audio review.
Discover more AI models for your next creative workflow