Yes. The model is live on the APIXO asynchronous task endpoint and can be called with the same API key as every other model in the catalog. Generation is also available in the Playground on this page.
MiniMax H3 is a general-purpose multimodal video model that combines text, image, video, and audio context in one generation workflow. It creates videos up to 2K and 15 seconds with native stereo sound, while supporting first-and-last-frame control, mixed-media references, motion transfer, multi-shot storytelling, and instruction-guided video editing.
Inputs
Text / Image / Video / Audio
APIXO Price
$0.09–$0.13 / second
Resolution
768P / 2K
Duration
5-15 Seconds
Audio
Native Stereo Sound
Create with MiniMax H3

MiniMax H3 Capabilities
Unified Context
Combine written direction with image, video, and audio references in one request. Assign each source a role such as appearance, motion, camera, voice, rhythm, environment, or style.
First & Last Frames
Define an opening frame, an ending frame, or both to guide transformations, transitions, product reveals, and storyboard beats while preserving a clearer visual trajectory.
Mixed-Media References
Use reference images, videos, and audio together to communicate character identity, product cues, movement, camera language, vocal performance, music, or ambience.
Native Stereo Sound
Generate voice, sound effects, ambience, and music together with the video as native stereo audio rather than assembling isolated audio tracks after generation.
Native Multi-Shot
Build short sequences with coordinated shot changes, scene progression, action continuity, and audiovisual pacing from one prompt-led or reference-guided workflow.
Editing & Motion Transfer
Transform existing visual material through natural-language direction or transfer motion from video references while reviewing identity, composition, occlusion, and source-detail fidelity.
Production concepts suited to MiniMax H3
Audiovisual Product Stories
Combine product imagery, written direction, motion references, and sound intent to develop campaign concepts. Review labels, logos, materials, dimensions, and product geometry before approving generated commercial assets.
Reference-Led Performances
Establish a recurring subject with visual references, then guide movement, camera behavior, voice, or scene context through additional media. Evaluate identity continuity, anatomy, expression, and performance timing across the complete result.
Narrative Previsualization
Translate a treatment into short scenes that combine action, camera language, ambience, and audiovisual pacing. Use generated footage to evaluate ideas before committing to physical production or detailed post-production.
Footage Transformation
Explore style, environment, subject, or scene changes through natural-language instructions and optional references. Inspect temporal boundaries, occlusion, unintended modifications, and source-detail preservation frame by frame.
MiniMax H3 Model Specifications
Confirmed against the live APIXO deployment.
Text / Image / Video / Audio
Input
T2V / I2V / Reference
Workflows
768P / 2K
Resolution
5–15 Seconds
Duration
Native Stereo
Audio
6 Ratios + Adaptive
Framing
MiniMax H3 availability and integration guidance
Notes
MiniMax H3 runs on the APIXO asynchronous task endpoint. Submit to generateTask and poll statusTask, or supply a callback URL for webhook delivery.
Three modes are exposed: text-to-video, image-to-video, and reference-to-video. Image-to-video takes 1-2 images as the first and optional last frame; reference-to-video needs at least one image or video reference.
Output is 768p or 2K at 5-15 seconds, whole seconds only. Resolution sets the per-second rate. Prompts are limited to 4,000 characters. Reference video and audio are each capped at 3 files, 2-15 seconds per file and 15 seconds in total.
“MiniMax H3” is the APIXO model ID. Do not carry over parameter values, quality tiers, or billing behavior from Hailuo 2.3 or Hailuo 2.3 Fast — they are separate routes with different contracts.
References guide generation but do not guarantee exact identity, typography, logos, materials, or product geometry.
Review anatomy, hands, object contact, rapid motion, occlusion, camera transitions, speech, and audiovisual synchronization before release.
Frequently Asked Questions
Text-to-video, image-to-video, and reference-to-video. Reference-to-video accepts up to 9 images, 3 videos, and 3 audio files as combined creative context; audio cannot be supplied on its own.
768p or 2K output at 5 to 15 seconds, in whole seconds. 2K is the default and the highest tier the route exposes; 768p bills at a lower per-second rate. Aspect ratio can be set to 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16 for text-to-video and reference-to-video; image-to-video takes its framing from the supplied first frame.
Output video is billed per second in every mode, at the rate that matches your chosen resolution — 768p is cheaper per second than 2K. In reference-to-video, the actual seconds of your reference videos are billed at the same per-second rate, and the first 5 reference images are free with each additional image charged separately. Reference audio is not billed. See the pricing tab for the current rates.
Yes. Audio is produced in the same pass as the picture as native stereo, with speech, sound effects, ambience, and music modelled jointly rather than as separate tracks.
No. Mixed-media references communicate creative intent but do not guarantee exact identity, motion, voice, product structure, or frame-level reproduction. Production use requires visual and audio review.
Explore Other Models
Discover more AI models for your next creative workflow