Design creates an original custom voice from a written description and preview script. Clone instead derives a voice from one reference recording. The two modes have different required inputs but both return a reusable MiniMax voice ID.
MiniMax Voice on APIXO creates reusable custom voice IDs through two distinct workflows: text-described Voice Design and single-clip Voice Clone. Design includes a required voice preview, while Clone can return preview audio when preview text is supplied. The resulting voice ID can be reused for subsequent text-to-speech generation through the separate MiniMax Speech 2.8 route.
APIXO Modes
Design / Clone
Price
$0.50 / request
Result
Reusable Voice ID
Clone Input
One Audio URL
Availability
Live on APIXO
Create with MiniMax Voice
Loading workspace...
Two Paths to a Reusable MiniMax Voice
Text-Described Voice Design
Describe the desired accent, apparent age, tone, pacing, energy, and intended use without supplying source audio. MiniMax generates a custom voice and reads the provided preview text so the result can be evaluated.
Single-Clip Voice Clone
Clone mode derives a reusable voice identity from exactly one public reference-audio URL. The recording acts as the speaker source rather than as speech content for conversion or dubbing.
Reusable Voice Asset
Both workflows return a MiniMax voice ID instead of only a finished audio file. That identifier can be stored and selected in compatible downstream synthesis requests when a recurring narrator or character is needed.
Preview Before Synthesis
Voice Design requires preview text and returns audio for evaluating the generated voice. Clone mode can also produce preview audio when optional preview text is supplied, allowing teams to review the result before wider synthesis use.
Clone Cleanup Controls
APIXO exposes optional noise reduction and volume normalization for clone references. These controls can improve source preparation when a permitted recording contains inconsistent level or manageable background noise.
Recognition Guidance
Clone mode includes a language hint and an adjustable accuracy control. The hint assists recognition of the reference speech, while accuracy lets applications tune the requested cloning behavior within APIXO’s supported range.
Design an Original Voice or Clone an Authorized Speaker
Voice Design from a Written Brief
Provide a voice description, preview text, and valid identifier prefix. This mode is suited to fictional characters and brand narrators because it creates a voice from requested traits without requiring a real speaker recording.
Voice Clone from One Recording
Supply one public MP3, M4A, or WAV URL from a speaker you are authorized to reproduce. Optional cleanup, recognition, and accuracy controls prepare the source before APIXO returns the reusable cloned voice ID. Supply optional preview text when an audio sample is needed for validation.
MiniMax Voice Inputs on APIXO
Each mode has separate requirements and should be implemented independently.
Description + Preview
Design Input
1–500 Characters
Design Preview
Exactly One URL
Clone Input
MP3 / M4A / WAV
Clone Formats
0–1
Clone Accuracy
Letter-First · 6+ Chars
Voice ID Prefix
Build Reusable Voices for Production Systems
Original Brand Narrators
Design a voice around traits such as warmth, authority, pace, age, and accent, then validate it with a representative preview. The resulting ID can anchor product tours, explainers, and campaign narration without copying a real individual.
Licensed Speaker Workflows
Create a reusable voice from a clean recording supplied under appropriate consent and usage terms. This can extend an approved narrator or performer across localized scripts and recurring content while keeping source authorization explicit.
Recurring Character Voices
Generate distinct voices for game characters, learning companions, digital presenters, or fictional assistants. Store the returned identifiers with character metadata, then use the separate synthesis route whenever new dialogue is required.
Consistent Reading Voices
Establish a familiar custom voice for educational material, interface guidance, or accessible reading experiences. Review pronunciation and emotional delivery in downstream synthesis, particularly when scripts change language or contain specialized terminology.
Voice Creation, Activation, and Rights
Important Notes
MiniMax Voice creates reusable voice assets; it does not expose final text-to-speech narration controls.
Design mode requires a voice description, preview text, and valid voice ID prefix.
Clone mode requires exactly one publicly reachable MP3, M4A, or WAV reference URL.
A newly created voice that is never used for later synthesis may become unavailable after seven days.
APIXO recommends beginning status checks about ten seconds after creation, but actual processing time can vary.
Clone only voices you are authorized to process and reproduce; successful creation does not establish consent or commercial rights.
Frequently Asked Questions
No. The route creates and previews a reusable voice asset. Pass its voice ID to APIXO’s separate MiniMax Speech 2.8 route when you need to synthesize complete narration, dialogue, or application speech.
APIXO charges $0.50 for each accepted Voice Design or Voice Clone creation request. Later text-to-speech generation is a separate operation and follows the character-based pricing of the selected MiniMax Speech 2.8 tier.
Use a clear, authorized recording focused on one speaker. Music, overlapping voices, long silence, strong echo, clipping, and heavy background noise can weaken the usable speaker signal even when the file is technically accepted.
Clone mode provides language hints for supported recognition contexts, and compatible MiniMax synthesis models may reuse custom voice IDs across languages. Accent, pronunciation, identity, and emotional character can drift, so cross-language results require speaker-aware review.
Explore Other Models
Discover more AI models for your next creative workflow
